LopezCastroRoberto/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Senior ML Engineer at RedHat AI
A high-throughput and memory-efficient inference and serving engine for LLMs
Humming is a high-performance, lightweight, and highly flexible JIT (Just-In-Time) compiled GEMM kernel library specifically designed for quantized inference.
Fast and memory-efficient exact attention
FlashInfer: Kernel Library for LLM Serving