ZewenShen-Cohere/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
@ZewenShen
A high-throughput and memory-efficient inference and serving engine for LLMs
A safetensors extension to efficiently store sparse quantized tensors on disk
Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM