MasterJH5574/modern-gpu-programming-for-mlsys
A tutorial on modern GPU programming for machine learning systems
Fourth-year PhD student at CMU / Building MLC / ML Systems / Deep Learning Compilers / @apache TVM PMC
A tutorial on modern GPU programming for machine learning systems
High-performance GPU kernels written in TIRx.
Open deep learning compiler stack for cpu, gpu and specialized accelerators
🖥️A CPU in Verilog that implements the RISC-V 32b integer base user-level real-mode ISA.
TVM FFI
DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scaling
SGLang is a fast serving framework for large language models and vision language models.
Project of MS110 - Computer System (2) - Spring 2020 - SJTU
Enable everyone to develop, optimize and deploy AI models natively on everyone's devices.
Project for SJTU CS158 Data Structure(Honor) Spring 2020
A standalone GEMM kernel for fp16 activation and quantized weight, extracted from FasterTransformer
Project website of FlashInfer project
TensorRT LLM Benchmark Configuration
CSD Blog
🕹️A simple analytical engine emulator used for course "Great Ideas in Computer Science"
Development repository for the Triton language and compiler
🔪Mx-Star Compiler Project
FlashInfer: Kernel Library for LLM Serving
Course Project of CS392, Database Management System, SJTU, 2021 Spring
Bringing large-language models and chat to web browsers. Everything runs inside the browser with no server support.
Bringing stable diffusion models to web browsers. Everything runs inside the browser with no server support.
Minimal BASIC Interpreter
A simple RISC-V simulator
Temp repo for prototyping relax(relay next), the effort will be upstreamed. We use the wiki pages on this repo to host design docs.