yuan-luo/sglang

SGLang is a fast serving framework for large language models and vision language models.

★ 0Forks 0PythonGitHub ↗Compare

Project website ↗

README

SGLang: Fast inference for LLMs and multimodal models

SGLang

PyPI version License: Apache 2.0 PyPI downloads per month

Docs | Cookbook | Website | Blog | Slack

SGLang is an open-source inference framework for LLMs and multimodal models, optimized for agentic workloads, RL rollouts, and large-scale serving.

👋 Get started below, or meet the community at SGLang Events, including meetups, developer meetings, workshops, and office hours.

Get Started

Pull the Docker image, which includes SGLang and its dependencies:

docker pull lmsysorg/sglang:latest

Alternatively, install SGLang in an activated Python environment with uv:

uv pip install --prerelease=allow sglang

Next, launch your model:

  • Quickstart: Run your first model and send a request.
  • Cookbook: Choose your model and hardware to get a ready-to-run launch command.

Supported Hardware

SGLang supports a wide range of GPUs, TPUs, NPUs, CPUs, and Apple Silicon platforms.

Platform Representative hardware
NVIDIA A100; H100/H200/H800/H20; B200/B300/GB200/GB300; select RTX 30/40/50 series, RTX 6000 Ada / PRO 6000; DGX Spark, Jetson Orin
AMD Instinct MI300X, MI325X, MI350X, MI355X
Google TPU v6e, v7; SGL-JAX / SGL-torchtpu
Intel Arc / Arc Pro B-Series GPUs, Xeon CPUs
Apple Silicon Macs via Metal / MLX
Huawei Ascend A2, A3, 950PR/DT NPUs
Moore Threads MTT S5000 GPUs

Integrations in progress: AWS Trainium, Alibaba T-Head PPU, Cambricon MLU, Qualcomm QAIC, MetaX, Hygon HCU/DCU, Iluvatar CoreX, and more.

See the Cookbook and platform guides for model compatibility and setup.

SGL Ecosystem

Area Projects Purpose
Education Mini-SGLang, zero-to-sglang, DeepLearning.AI course Learn inference engine design and efficient text and image generation through code and hands-on courses.
Diffusion SGLang Diffusion Image and video generation with diffusion models.
Audio SGLang Omni Audio model serving for text-to-speech (TTS) and automatic speech recognition (ASR).
RL and Post-Training Miles, slime, AReaL, Tunix, verl Training frameworks that integrate SGLang for rollout generation.
Speculative Decoding SpecForge Train draft models for speculative decoding and deploy them with SGLang.
KV Cache HiCache, Mooncake, LMCache Hierarchical KV caching across GPU memory, host memory, and external storage, with cache transfer and reuse for distributed inference.
Deployment and Orchestration SMG, RBG, llm-d, Ray Serve, NVIDIA Dynamo Deploy and scale SGLang inference services with routing, load balancing, and cluster orchestration.

Development and Contributing

Contributions are welcome, from bug fixes and documentation to model support and performance improvements.

Development setup

Start from the lmsysorg/sglang:dev Docker image, which provides development tools and most dependencies. Clone or mount your SGLang checkout inside the container, then install it in editable mode from the repository root so tests use your local Python changes:

pip install -e "python"

In an activated virtual environment, you can use uv pip install --prerelease=allow -e "python" instead. See the development guide for container setup and testing.

Contribute

  1. Fork the repository and create a branch for your changes. For larger changes, discuss your proposal in a GitHub issue or on Slack.
  2. Make your changes, run the relevant tests, and add regression coverage for fixes or new behavior. Run pre-commit run --all-files before submitting.
  3. Open a pull request describing the change and how you tested it. Include benchmarks or accuracy evaluations when relevant.

See the contributor guide for formatting, testing, and pull request instructions. Documentation contributors can start with the docs guide.

Community and Sponsorship

SGLang is hosted by LMSYS, a non-profit open-source organization.

  • Community discussions: Join Slack for technical questions and development discussions.
  • Events: Find meetups, workshops, and office hours on SGLang Events.
  • Updates: Follow X and LinkedIn for project updates, and the LMSYS Blog for release announcements and technical articles.
  • Project resources: Explore the documentation, Cookbook, roadmap, release notes, issue tracker, and contributor guide.
  • Contact Us: For enterprise adoption and deployment, technical consulting, sponsorship, or partnership inquiries, please contact [email protected].
  • Contributor sponsorship: Long-term active SGLang contributors are eligible for coding agent sponsorship, including Cursor, Claude Code, or OpenAI Codex. To apply, email [email protected] with links to your key commits or pull requests.

Trusted by Industry and Research

SGLang serves production workloads across AI labs, cloud platforms, enterprises, and universities.

Organizations adopting SGLang

Acknowledgment

We learned the design and reused code from the following projects: Guidance, vLLM, LightLLM, FlashInfer, Outlines, and LMQL.

Contributors

merrymercyhnyls2002fzyzcjyzhyncsmickqianBBufch-wanslin1237Fridge003ispobockalisonshaoKangyan-ZhouJustinTong0323mmangkadShangmingCaiCatherineSueb8zhongYing1123Qiaolin-Yualphabetc1yctseng0211ByronHsuyuan-luomichaelzhang-aiyhyang201YAMY1234kpham-sglhzh0425zijiexiasglang-bot

Issues