sunyi0505/flash-linear-attention
๐ Efficient implementations for emerging model architectures
๐ Efficient implementations for emerging model architectures
Efficient Triton Kernels for LLM Training
LLaMA Factory Document
MCore-Bridge: Providing Megatron-Core model definitions for state-of-the-art large models and making Megatron training as simple as Transformers โ with support for 300+ large language models (Qwen3-Next, GLM-5.2, Deepseek-V4, MiniMax-2.7, ...) and 200+ multimodal large models (Qwen3.5, Qwen3-Omni, Gemma4, ...).
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
veRL: Volcano Engine Reinforcement Learning for LLM
The fastest repo in history to surpass 50K stars โญ, reaching the milestone in just 2 hours after publication. Better Harness Tools, not merely storing the archive of leaked Claude Code but also make real things done. Now rewriting in Rust.
๐ค Transformers: State-of-the-art Machine Learning for Pytorch, TensorFlow, and JAX.
Clean, minimal, accessible reproduction of DeepSeek R1-Zero
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
A high-throughput and memory-efficient inference and serving engine for LLMs
An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.
๐ A simple way to launch, train, and use PyTorch models on almost any device and distributed configuration, automatic mixed precision (including fp8), and easy-to-configure FSDP and DeepSpeed support