WoosukKwon/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
@Inferact | @vllm-project
A high-throughput and memory-efficient inference and serving engine for LLMs
[NeurIPS 2022] A Fast Post-Training Pruning Framework for Transformers
FlashInfer: Kernel Library for LLM Serving
A Datacenter Scale Distributed Inference Serving Framework
Enabling PyTorch on XLA Devices (e.g. Google TPU)