Robert Shaw

@robertgshaw2-redhat · User

GitHub profile ↗ · Compare

vllm and llm-d

@RedHatOfficialBoston336 followers92 repositories

Repositories

robertgshaw2-redhat/humming

Humming is a high-performance, lightweight, and highly flexible JIT (Just-In-Time) compiled GEMM kernel library specifically designed for quantized inference.

★ 0Forks 0

robertgshaw2-redhat/litellm

The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthropic, OpenAI, VertexAI, vLLM, Nvidia NIM]

★ 0PythonForks 0

robertgshaw2-redhat/aiperf

AIPerf is a comprehensive benchmarking tool that measures the performance of generative AI models served by your preferred inference solution.

★ 0Forks 0