shanjiaz/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
maintainer @ Speculators
A high-throughput and memory-efficient inference and serving engine for LLMs
A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
Common recipes to run vLLM
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.