zhan4808/cache-barrier
exploring the L2 cache barrier in MLA reconstruction GEMMs and why INT4 quantization fails to reduce decode latency
exploring the L2 cache barrier in MLA reconstruction GEMMs and why INT4 quantization fails to reduce decode latency
hardware counter-grounded kernel optimization for transformer GEMMs
Config files for my GitHub profile.
FlagGems is an operator library for large language models implemented in the Triton Language.
personal portfolio website
CW&T Time Since Launch
decode-focused triton fused int4 gemm and transformer kernels for small-batch LLM inference
SGLang is a high-performance serving framework for large language models and multimodal models.
A high-throughput and memory-efficient inference and serving engine for LLMs
Hosts MLSys Competition Track Problems
NanoGPT (124M) in 2 minutes
FlashInfer Bench @ MLSys 2026: Building AI agents to write high performance GPU kernels
Tile primitives for speedy kernels
The LLVM Project is a collection of modular and reusable compiler and toolchain technologies.
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
A toolkit for SQLite databases, with a focus on application development
real time videos you can talk to
building meaningful relationships with intelligent mesh and automated networking
TensorFlow's Visualization Toolkit
Public Lab Documents for ECE 36200
your own museum curator
Development repository for the Triton language and compiler
GPTQ: Accurate Post-training Quantization of Generative Pretrained Transformers