li-lizhe/peft
๐ค PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.
๐ค PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.
VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
Ascend NPU open-source ecosystem: torchaudio-free adaptation, tutorials, projects and roadmap
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
Triton language and compiler for Ascend NPU
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
Build voice agents with open-source models
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Wan: Open and Advanced Large-Scale Video Generative Models
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer
xDiT: A Scalable Inference Engine for Diffusion Transformers (DiTs) with Massive Parallelism
๐ Geometric Computer Vision Library for Spatial AI
A framework for few-shot evaluation of language models.
Making large AI models cheaper, faster and more accessible
Deep and online learning with spiking neural networks in Python
End-to-End Speech Processing Toolkit
A Next-Generation Training Engine Built for Ultra-Large MoE Models
SOTA Open Source TTS
An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)
Enjoy the magic of Diffusion models!
The largest collection of PyTorch image encoders / backbones. Including train, eval, inference, export scripts, and pretrained weights -- ResNet, ResNeXT, EfficientNet, NFNet, Vision Transformer (ViT), MobileNetV4, MobileNet-V3 & V2, RegNet, DPN, CSPNet, Swin Transformer, MaxViT, CoAtNet, ConvNeXt, and more
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, Phi4, ...) (AAAI 2025).
MiniMax-M1, the world's first open-weight, large-scale hybrid-attention reasoning model.
Official code for "F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching"
๐ A simple way to launch, train, and use PyTorch models on almost any device and distributed configuration, automatic mixed precision (including fp8), and easy-to-configure FSDP and DeepSpeed support
The agent engineering platform.
SGLang is a high-performance serving framework for large language models and multimodal models.
Machine learning metrics for distributed, scalable PyTorch applications.