kyleliang-nv/dynamo
A Datacenter Scale Distributed Inference Serving Framework
A Datacenter Scale Distributed Inference Serving Framework
A high-throughput and memory-efficient inference and serving engine for LLMs
NVIDIA Inference Benchmarks provide recipes in ready-to-use templates for evaluating platform speed. Validate your platform across specific AI use cases across hardware and software combinations.
Common recipes to run vLLM
SGLang is a fast serving framework for large language models and vision language models.
DeepEP: an efficient expert-parallel communication library