zihanlin-ai/sglang
SGLang is a high-performance serving framework for large language models and multimodal models.
SGLang is a high-performance serving framework for large language models and multimodal models.
A high-throughput and memory-efficient inference and serving engine for LLMs
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
LMCache-Ascend is a plugin for running LMCache on the Ascend NPU.
Community maintained hardware plugin for vLLM on Ascend
SGLang kernel library for NPU