z1ying/aibrix
Cost-efficient and pluggable Infrastructure components for GenAI inference
Full-stack engineer exploring AI infra & open source ๐ฑ
Cost-efficient and pluggable Infrastructure components for GenAI inference
A high-throughput and memory-efficient inference and serving engine for LLMs
slime is an LLM post-training framework for RL Scaling.
Development repository for the Triton language and compiler
Modern RL Post-training Infrastructure: Optimized for NVIDIA/AMD GPUs with a focus on vLLM and DeepSpeed integration, CUDA/ROCm/Triton kernels, and transparent hardware-aware scaling.
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
Unsloth Studio is a web UI for training and running open models like Gemma 4, Qwen3.6, DeepSeek, gpt-oss locally.
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
Fast and memory-efficient exact attention
FlashInfer: Kernel Library for LLM Serving
MiniCPM4 & MiniCPM4.1: Ultra-Efficient LLMs on End Devices, achieving 3+ generation speedup on reasoning tasks
Supercharge Your LLM with the Fastest KV Cache Layer
SGLang is a high-performance serving framework for large language models and multimodal models.
A framework for efficient model inference with omni-modality models
Automatic development for retrieval augmented generation system