BenjaminBraunDev/k8s.io
Code and configuration to manage Kubernetes project infrastructure, including various *.k8s.io sites
Code and configuration to manage Kubernetes project infrastructure, including various *.k8s.io sites
repo for CI and infrastructure required to maintain llm-d org member repos
llm-d benchmark scripts and tooling
Website for llm-d: This repository builds the website seen at llm-d.ai
Inference scheduler for llm-d
LeaderWorkerSet: An API for deploying a group of pods as a unit of replication
Latency prediction service for ML-model based scoring with llm-d-inference-scheduler
Asynchronous Processor for Inference Gateway. Orchestrator of queues
llm-d is a Kubernetes-native high-performance distributed LLM inference framework
Gateway API Inference Extension
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
GenAI inference performance benchmarking tool
The Triton TensorRT-LLM Backend
Triton Python, C++ and Java client libraries, and GRPC-generated client examples for go, java and scala.
Development Fork of Triton Inference Server