nilig/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
IBM Research
A high-throughput and memory-efficient inference and serving engine for LLMs
Inference scheduler for llm-d
llm-d benchmark scripts and tooling
Website for llm-d: This repository builds the website seen at llm-d.ai
Inference payload processor for llm-d
Achieve state of the art inference performance with modern accelerators on Kubernetes
Let my Claude talk to yours.