FlamingoPg/sglang
SGLang is a fast serving framework for large language models and vision language models.
Move Faster
SGLang is a fast serving framework for large language models and vision language models.
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
agent multiplexer that lives in your terminal.
TokenSpeed is a speed-of-light LLM inference engine.
Puzzles for learning Triton
Terminal UI for NVIDIA Nsight Systems profiles — timeline viewer, kernel navigator, NVTX hierarchy
Desktop and web interface for OpenCode AI agent
FlashInfer Bench @ MLSys 2026: Building AI agents to write high performance GPU kernels
High Performance LLM Inference Operator Library
FlagGems is an operator library for large language models implemented in Triton Language.
A high-throughput and memory-efficient inference and serving engine for LLMs
PyTorch native post-training library
Flash Attention in ~100 lines of CUDA (forward pass only)
QuTLASS: CUTLASS-Powered Quantized BLAS for Deep Learning
Matrix-Game 2.0: An Open-Source, Real-Time, and Streaming Interactive World Model
Light Video Generation Inference Framework
sgl-spec official
This is the homepage of a new book entitled "Mathematical Foundations of Reinforcement Learning."
This is a repo to track the latest autoregressive visual generation papers.
Config files for my GitHub profile.
AGE animation official website URL release page(AGE动漫官网网址发布页)
SpargeAttention: A training-free sparse attention that can accelerate any model inference.
📖A curated list of Awesome Diffusion Inference Papers with codes, such as Sampling, Caching, Multi-GPUs, etc. 🎉🎉