Docs | Cookbook | Website | Blog | Slack
SGLang is an open-source inference framework for LLMs and multimodal models, optimized for agentic workloads, RL rollouts, and large-scale serving.
👋 Get started below, or meet the community at SGLang Events, including meetups, developer meetings, workshops, and office hours.
Pull the Docker image, which includes SGLang and its dependencies:
docker pull lmsysorg/sglang:latestAlternatively, install SGLang in an activated Python environment with uv:
uv pip install --prerelease=allow sglangNext, launch your model:
- Quickstart: Run your first model and send a request.
- Cookbook: Choose your model and hardware to get a ready-to-run launch command.
SGLang supports a wide range of GPUs, TPUs, NPUs, CPUs, and Apple Silicon platforms.
| Platform | Representative hardware |
|---|---|
| NVIDIA | A100; H100/H200/H800/H20; B200/B300/GB200/GB300; select RTX 30/40/50 series, RTX 6000 Ada / PRO 6000; DGX Spark, Jetson Orin |
| AMD | Instinct MI300X, MI325X, MI350X, MI355X |
| Google TPU | v6e, v7; SGL-JAX / SGL-torchtpu |
| Intel | Arc / Arc Pro B-Series GPUs, Xeon CPUs |
| Apple Silicon | Macs via Metal / MLX |
| Huawei Ascend | A2, A3, 950PR/DT NPUs |
| Moore Threads | MTT S5000 GPUs |
Integrations in progress: AWS Trainium, Alibaba T-Head PPU, Cambricon MLU, Qualcomm QAIC, MetaX, Hygon HCU/DCU, Iluvatar CoreX, and more.
See the Cookbook and platform guides for model compatibility and setup.
| Area | Projects | Purpose |
|---|---|---|
| Education | Mini-SGLang, zero-to-sglang, DeepLearning.AI course | Learn inference engine design and efficient text and image generation through code and hands-on courses. |
| Diffusion | SGLang Diffusion | Image and video generation with diffusion models. |
| Audio | SGLang Omni | Audio model serving for text-to-speech (TTS) and automatic speech recognition (ASR). |
| RL and Post-Training | Miles, slime, AReaL, Tunix, verl | Training frameworks that integrate SGLang for rollout generation. |
| Speculative Decoding | SpecForge | Train draft models for speculative decoding and deploy them with SGLang. |
| KV Cache | HiCache, Mooncake, LMCache | Hierarchical KV caching across GPU memory, host memory, and external storage, with cache transfer and reuse for distributed inference. |
| Deployment and Orchestration | SMG, RBG, llm-d, Ray Serve, NVIDIA Dynamo | Deploy and scale SGLang inference services with routing, load balancing, and cluster orchestration. |
Contributions are welcome, from bug fixes and documentation to model support and performance improvements.
Start from the lmsysorg/sglang:dev Docker image, which provides development tools and most dependencies. Clone or mount your SGLang checkout inside the container, then install it in editable mode from the repository root so tests use your local Python changes:
pip install -e "python"In an activated virtual environment, you can use uv pip install --prerelease=allow -e "python" instead. See the development guide for container setup and testing.
- Fork the repository and create a branch for your changes. For larger changes, discuss your proposal in a GitHub issue or on Slack.
- Make your changes, run the relevant tests, and add regression coverage for fixes or new behavior. Run
pre-commit run --all-filesbefore submitting. - Open a pull request describing the change and how you tested it. Include benchmarks or accuracy evaluations when relevant.
See the contributor guide for formatting, testing, and pull request instructions. Documentation contributors can start with the docs guide.
SGLang is hosted by LMSYS, a non-profit open-source organization.
- Community discussions: Join Slack for technical questions and development discussions.
- Events: Find meetups, workshops, and office hours on SGLang Events.
- Updates: Follow X and LinkedIn for project updates, and the LMSYS Blog for release announcements and technical articles.
- Project resources: Explore the documentation, Cookbook, roadmap, release notes, issue tracker, and contributor guide.
- Contact Us: For enterprise adoption and deployment, technical consulting, sponsorship, or partnership inquiries, please contact [email protected].
- Contributor sponsorship: Long-term active SGLang contributors are eligible for coding agent sponsorship, including Cursor, Claude Code, or OpenAI Codex. To apply, email [email protected] with links to your key commits or pull requests.
SGLang serves production workloads across AI labs, cloud platforms, enterprises, and universities.
We learned the design and reused code from the following projects: Guidance, vLLM, LightLLM, FlashInfer, Outlines, and LMQL.

