A comprehensive benchmarking lab for comparing vLLM, TensorRT-LLM, and Triton Inference Server with Llama 3.1 8B Instruct on RTX A6000.
This repository provides:
- Standardized benchmark harness for fair comparison across engines
- Docker-based deployment for consistent environments
- Jupyter notebooks for step-by-step deployment and analysis
- Automated result aggregation and visualization
inference-engines-lab/
├── notebooks/ # Step-by-step notebooks
│ ├── 00_env_check.ipynb # Environment validation
│ ├── 01_vllm_deploy_and_bench.ipynb
│ ├── 02_tensorrtllm_deploy_and_bench.ipynb
│ ├── 03_triton_deploy_and_bench.ipynb
│ └── 04_compare.ipynb # Results comparison
├── scripts/
│ ├── bench_openai_api.py # Core benchmark harness
│ └── bench_report.py # Results aggregation
├── configs/
│ └── bench_cases.yaml # Test case definitions
├── docker/ # Docker configurations
├── results/
│ ├── raw/ # Per-request JSONL logs
│ └── summary/ # Aggregated CSV files
└── requirements.txt
- GPU: RTX A6000 (48GB VRAM) or similar NVIDIA GPU
- OS: Linux (tested on Ubuntu 22.04)
- Docker: With NVIDIA Container Toolkit configured
- Python: 3.8+ (for notebooks and benchmark scripts)
- HuggingFace Access: For downloading Llama 3.1 8B Instruct
# Install Python dependencies
pip install -r requirements.txt
# Install PyTorch with CUDA support (required for model verification in notebooks)
# Check your CUDA version first: nvidia-smi
# For CUDA 12.1:
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121
# For CUDA 11.8:
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu118
# Or visit https://pytorch.org/get-started/locally/ for the latest installation commandOpen and run notebooks/00_env_check.ipynb to verify:
- GPU availability
- Docker setup
- Model access
- Benchmark harness
Follow the notebooks in order:
- Notebook 01: Deploy vLLM and run benchmarks
- Notebook 02: Deploy TensorRT-LLM and run benchmarks
- Notebook 03: Deploy Triton and run benchmarks
- Notebook 04: Compare all results
Edit configs/bench_cases.yaml to customize test cases:
prompt_lengths: [128, 512, 2048] # tokens
output_lengths: [64, 256] # tokens
concurrency_levels: [1, 4, 16]
requests_per_case: 50# vLLM
python scripts/bench_openai_api.py \
--base-url http://localhost:8000/v1 \
--engine-name vllm \
--config configs/bench_cases.yaml
# TensorRT-LLM
python scripts/bench_openai_api.py \
--base-url http://localhost:8001/v1 \
--engine-name trtllm \
--config configs/bench_cases.yaml
# Triton (if OpenAI-compatible)
python scripts/bench_openai_api.py \
--base-url http://localhost:8002/v1 \
--engine-name triton \
--config configs/bench_cases.yaml- TTFT (Time to First Token): p50, p95, p99
- TPOT (Time Per Output Token): Mean, p50, p95
- Throughput: Tokens per second
- Total Latency: End-to-end request time
- GPU Memory: Peak usage (if available)
Results are saved to:
results/raw/: Per-request JSONL logsresults/summary/: Aggregated CSV files and comparison plots
docker run --gpus all -p 8000:8000 \
-v ~/.cache/huggingface:/root/.cache/huggingface \
vllm/vllm-openai:latest \
python -m vllm.entrypoints.openai.api_server \
--model meta-llama/Llama-3.1-8B-Instruct \
--port 8000 \
--host 0.0.0.0See notebooks/02_tensorrtllm_deploy_and_bench.ipynb for detailed instructions. Requires engine building first.
docker run --gpus all -p 8002:8000 \
-v $(pwd)/model_repository:/models \
nvcr.io/nvidia/tritonserver:latest \
tritonserver --model-repository=/modelsTo ensure fair comparisons:
- ✅ Same model weights (Llama 3.1 8B Instruct)
- ✅ Same generation settings (temperature, top_p, max_tokens)
- ✅ Same test cases (prompts, output lengths, concurrency)
- ✅ Warmup requests before timing
- ✅ Streaming enabled for accurate TTFT
- ✅ GPU memory tracking
- Ensure HuggingFace token is configured:
huggingface-cli login - Check model access permissions for Llama 3.1 8B Instruct
- Verify NVIDIA Container Toolkit:
docker run --rm --gpus all nvidia/cuda:12.0.0-base-ubuntu22.04 nvidia-smi - Check Docker daemon configuration
- vLLM: 8000
- TensorRT-LLM: 8001
- Triton: 8002 (HTTP), 8003 (gRPC), 8004 (Metrics)
This lab is for educational and benchmarking purposes. Model usage is subject to Llama 3.1 license terms.