robertgshaw2-redhat/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
vllm and llm-d
A high-throughput and memory-efficient inference and serving engine for LLMs
llm-d Router: The intelligent entry point for inference requests
Use Fireworks AI models in Claude Code, Codex, Cursor, VS Code, Copilot, and other coding agents.
Humming is a high-performance, lightweight, and highly flexible JIT (Just-In-Time) compiled GEMM kernel library specifically designed for quantized inference.
llm-d is a Kubernetes-native high-performance distributed LLM inference framework
The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthropic, OpenAI, VertexAI, vLLM, Nvidia NIM]
Automation for running spot checks on quality
Kimi-Vendor-Verifier
Model X HW X llm-d
AIPerf is a comprehensive benchmarking tool that measures the performance of generative AI models served by your preferred inference solution.
Experiments
Experiments related to vllm + rl loops
Midstream Fork
Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs
Demo - Distributed Inference on OpenShift (aka llm-d)
llm-d demo
Justfile harness for llm-d
DeepEP: an efficient expert-parallel communication library
Scripts for installing libraries
Benchmarking vLLM
utils for NIM
Comparison of vLLM and TRT-LLM
Helm charts for benchmarking