deepbuilder/inference-engine-lab

★ 0Forks 0Jupyter NotebookGitHub ↗Compare

README

Inference Engines Benchmark Lab

A comprehensive benchmarking lab for comparing vLLM, TensorRT-LLM, and Triton Inference Server with Llama 3.1 8B Instruct on RTX A6000.

Overview

This repository provides:

  • Standardized benchmark harness for fair comparison across engines
  • Docker-based deployment for consistent environments
  • Jupyter notebooks for step-by-step deployment and analysis
  • Automated result aggregation and visualization

Repository Structure

inference-engines-lab/
├── notebooks/              # Step-by-step notebooks
│   ├── 00_env_check.ipynb           # Environment validation
│   ├── 01_vllm_deploy_and_bench.ipynb
│   ├── 02_tensorrtllm_deploy_and_bench.ipynb
│   ├── 03_triton_deploy_and_bench.ipynb
│   └── 04_compare.ipynb             # Results comparison
├── scripts/
│   ├── bench_openai_api.py          # Core benchmark harness
│   └── bench_report.py               # Results aggregation
├── configs/
│   └── bench_cases.yaml              # Test case definitions
├── docker/                           # Docker configurations
├── results/
│   ├── raw/                          # Per-request JSONL logs
│   └── summary/                      # Aggregated CSV files
└── requirements.txt

Prerequisites

  • GPU: RTX A6000 (48GB VRAM) or similar NVIDIA GPU
  • OS: Linux (tested on Ubuntu 22.04)
  • Docker: With NVIDIA Container Toolkit configured
  • Python: 3.8+ (for notebooks and benchmark scripts)
  • HuggingFace Access: For downloading Llama 3.1 8B Instruct

Quick Start

1. Install Dependencies

# Install Python dependencies
pip install -r requirements.txt

# Install PyTorch with CUDA support (required for model verification in notebooks)
# Check your CUDA version first: nvidia-smi
# For CUDA 12.1:
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121
# For CUDA 11.8:
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu118
# Or visit https://pytorch.org/get-started/locally/ for the latest installation command

2. Verify Environment

Open and run notebooks/00_env_check.ipynb to verify:

  • GPU availability
  • Docker setup
  • Model access
  • Benchmark harness

3. Deploy and Benchmark

Follow the notebooks in order:

  1. Notebook 01: Deploy vLLM and run benchmarks
  2. Notebook 02: Deploy TensorRT-LLM and run benchmarks
  3. Notebook 03: Deploy Triton and run benchmarks
  4. Notebook 04: Compare all results

Benchmark Configuration

Edit configs/bench_cases.yaml to customize test cases:

prompt_lengths: [128, 512, 2048]  # tokens
output_lengths: [64, 256]         # tokens
concurrency_levels: [1, 4, 16]
requests_per_case: 50

Running Benchmarks Manually

# vLLM
python scripts/bench_openai_api.py \
  --base-url http://localhost:8000/v1 \
  --engine-name vllm \
  --config configs/bench_cases.yaml

# TensorRT-LLM
python scripts/bench_openai_api.py \
  --base-url http://localhost:8001/v1 \
  --engine-name trtllm \
  --config configs/bench_cases.yaml

# Triton (if OpenAI-compatible)
python scripts/bench_openai_api.py \
  --base-url http://localhost:8002/v1 \
  --engine-name triton \
  --config configs/bench_cases.yaml

Metrics Collected

  • TTFT (Time to First Token): p50, p95, p99
  • TPOT (Time Per Output Token): Mean, p50, p95
  • Throughput: Tokens per second
  • Total Latency: End-to-end request time
  • GPU Memory: Peak usage (if available)

Results

Results are saved to:

  • results/raw/: Per-request JSONL logs
  • results/summary/: Aggregated CSV files and comparison plots

Docker Deployment

vLLM

docker run --gpus all -p 8000:8000 \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  vllm/vllm-openai:latest \
  python -m vllm.entrypoints.openai.api_server \
  --model meta-llama/Llama-3.1-8B-Instruct \
  --port 8000 \
  --host 0.0.0.0

TensorRT-LLM

See notebooks/02_tensorrtllm_deploy_and_bench.ipynb for detailed instructions. Requires engine building first.

Triton

docker run --gpus all -p 8002:8000 \
  -v $(pwd)/model_repository:/models \
  nvcr.io/nvidia/tritonserver:latest \
  tritonserver --model-repository=/models

Fair Comparison Rules

To ensure fair comparisons:

  1. ✅ Same model weights (Llama 3.1 8B Instruct)
  2. ✅ Same generation settings (temperature, top_p, max_tokens)
  3. ✅ Same test cases (prompts, output lengths, concurrency)
  4. ✅ Warmup requests before timing
  5. ✅ Streaming enabled for accurate TTFT
  6. ✅ GPU memory tracking

Troubleshooting

Model Access

  • Ensure HuggingFace token is configured: huggingface-cli login
  • Check model access permissions for Llama 3.1 8B Instruct

Docker GPU Access

  • Verify NVIDIA Container Toolkit: docker run --rm --gpus all nvidia/cuda:12.0.0-base-ubuntu22.04 nvidia-smi
  • Check Docker daemon configuration

Port Conflicts

  • vLLM: 8000
  • TensorRT-LLM: 8001
  • Triton: 8002 (HTTP), 8003 (gRPC), 8004 (Metrics)

References

License

This lab is for educational and benchmarking purposes. Model usage is subject to Llama 3.1 license terms.

Contributors

deepbuilder

Issues