RubiaCx/Megatron-LM
Ongoing research training transformer language models at scale, including: BERT & GPT-2
Hi~ o(* ̄▽ ̄*)ブ
Ongoing research training transformer language models at scale, including: BERT & GPT-2
A Distributed Attention Towards Linear Scalability for Ultra-Long Context, Heterogeneous Data Training
一个使用 NextJS + Notion API 实现的,部署在 Vercel 上的静态博客系统。
SGLang is a fast serving framework for large language models and vision language models.
An independent Python feature port of Claude Code, entirely rewritten from scratch. Educational Purpose only.
TurboDiffusion: 100–200× Acceleration for Video Diffusion Models
A unified inference and post-training framework for accelerated video generation.
Light Image Video Generation Inference Framework
A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hopper, Ada and Blackwell GPUs, to provide better performance with lower memory utilization in both training and inference.
My learning notes/codes for ML SYS.
2025 个税计算器
A collection of specialized agent skills for AI infrastructure development, enabling Claude Code to write, optimize, and debug high-performance systems.
[ICML 2025] XAttention: Block Sparse Attention with Antidiagonal Scoring
Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels
Enjoy the magic of Diffusion models!
slime is an LLM post-training framework for RL Scaling.
A Quirky Assortment of CuTe Kernels
Scalable and memory-optimized training of diffusion models
text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)
FlagGems is an operator library for large language models implemented in Triton Language.
Tritonbench is a collection of PyTorch custom operators with example inputs to measure their performance.
trition-lang/triton -> facebookexperimental/triton
Quantized Attention that achieves speedups of 2.1-3.1x and 2.7-5.1x compared to FlashAttention2 and xformers, respectively, without lossing end-to-end metrics across various models.
📚Tensor/CUDA Cores, 📖150+ CUDA Kernels, ⚡️⚡️toy-hgemm library with WMMA, MMA and CuTe (98%~100% TFLOPS of cuBLAS 🎉🎉).