jthomson04/aiperf
AIPerf is a comprehensive benchmarking tool that measures the performance of generative AI models served by your preferred inference solution.
NVIDIA Dynamo
AIPerf is a comprehensive benchmarking tool that measures the performance of generative AI models served by your preferred inference solution.
Fast Tokens
HTTP routing and request-handling library for Rust that focuses on ergonomics and modularity
NVIDIA Inference Benchmarks provide recipes in ready-to-use templates for evaluating platform speed. Validate your platform across specific AI use cases across hardware and software combinations.
Offline optimization of your disaggregated Dynamo graph
SGLang is a high-performance serving framework for large language models and multimodal models.
Scalable toolkit for efficient model reinforcement
Build RL environments for LLM training
Collection of SLURM deployment scripts for various hardware + trtllm disaggregation variants
Open Source Continuous Inference Benchmarking - GB200 NVL72 vs MI355X vs B200 vs H200 vs MI325X & soon™ TPUv6e/v7/Trainium2/3/GB300 NVL72 - DeepSeek 670B MoE, GPTOSS
A high-throughput and memory-efficient inference and serving engine for LLMs
TensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and support state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in performant way.
NVIDIA Inference Xfer Library (NIXL)