zheliuyu/Liger-Kernel-dev
Efficient Triton Kernels for LLM Training
Interested in @flash-linear-attention+NPU @Liger-Kernel+NPU @Ascend/docs @verl/HybridFlow+NPU
Efficient Triton Kernels for LLM Training
MoonEP-Triton-Distributed: MoonEP reimplemented on Triton-distributed with NVSHMEM symmetric memory.
🚀 Efficient implementations for emerging model architectures
A collection of out-of-tree extensions for the Triton language and compiler
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-R1, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, Phi4, ...) (AAAI 2025).
The repository provides docs & api & tutorials & FAQ and all things like that.
DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
LLaMA Factory Document
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
CI system for Ascend
tracking-map is used to track the progress of adapting some popular ecosystem tools for the Ascend NPU.
VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo