I-Hsuan(Ethan) Huang

@EthanCornell · User

GitHub profile ↗ · Compare

Ethan is a dedicated software engineer with a Master's in CS from Cornell University, specializing in algorithms, data structures, and systems optimization.

Cornell UniversityAustin, TX0 followers69 repositories

Repositories

EthanCornell/aim-racestudio3-mac

Free one-click installer to run AiM RaceStudio 3 on Apple Silicon Macs: a notarized, drag-to-Applications DMG. No Windows, Parallels, or CrossOver. Community project, not affiliated with AiM.

★ 0Forks 0

EthanCornell/PicoGPT

A hands-on fork of NanoGPT with FlashAttention-2 CUDA kernels, INT8/INT4 GPTQ quantization, paged KV-cache reuse, and continuous batching, turning a tiny Shakespeare model into a full-speed GPU LLM inference demo.

★ 2PythonForks 0

EthanCornell/mini_malloc

A comprehensive memory allocation library implementation featuring multiple levels of sophistication, from basic first-fit allocation to security-enhanced allocators with extensive debugging capabilities.

★ 1CForks 0

EthanCornell/UltraSIMD

Ultra-fast SIMD library delivering 18.7× speedups through AVX-512 optimization, supporting F32/F16/I8/BF16 data types with 100% accuracy across 175 test cases.

★ 2C++Forks 1

EthanCornell/pytorch

Tensors and Dynamic neural networks in Python with strong GPU acceleration

★ 0PythonForks 0

EthanCornell/cy86

CY86 is a simplified, intermediate assembly language created to streamline the translation from high-level programming logic to optimized low-level machine code, enabling faster development, easier debugging, and more efficient performance tuning in compiler and systems projects.

★ 0C++Forks 0

EthanCornell/tvm

Open deep learning compiler stack for cpu, gpu and specialized accelerators

★ 0Forks 0

EthanCornell/iree

A retargetable MLIR-based machine learning compiler and runtime toolkit.

★ 0Forks 0

EthanCornell/onnxruntime

ONNX Runtime: cross-platform, high performance ML inferencing and training accelerator

★ 0C++Forks 0

EthanCornell/ph

A family of header-only, very fast and memory-friendly hashmap and btree containers.

★ 0C++Forks 0

EthanCornell/ppToken

A Standards‑Compliant C/C++ Pre‑Processor Tokeniser with Full Phase 2 & 3 Translation

★ 0C++Forks 0

EthanCornell/ray

Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.

★ 0Forks 0

EthanCornell/RedLockTree

Header-only C++17 red-black tree with per-node locks—parallel look-ups, serialized writers, and a built-in stress test for heavy-load correctness.

★ 0C++Forks 0

EthanCornell/MPI-Wandering-Salesman-Solver

Blazing-fast branch-and-bound TSP solver (≤ 18 cities) in single-file C using the MPI message-passing model; auto-detects triangular inputs and runs locally or via PBS with one command.

★ 0CForks 0

EthanCornell/mini-migration

Mini-Migration — Cross-platform resumable file-transfer tool. C++17 core, Objective-C++ macOS layer; built for Apple Backup & Migration workflows.

★ 1C++Forks 0

EthanCornell/Distrbuted-Filesystem

micro-DFS is a minimalist, log-structured distributed file system designed for educational purposes and edge computing scenarios. Built as a reference implementation, it demonstrates core distributed systems concepts with a focus on simplicity, performance, and reliability.

★ 1CForks 0