Teagan42/wyoming-whisper-trt

A project that optimizes Wyoming and Whisper for low latency inference using NVIDIA TensorRT

★ 0Forks 0PythonGitHub ↗Compare

README

WhisperTRT

This project optimizes OpenAI Whisper with NVIDIA TensorRT and implements the Wyoming Protocol for Home Assistant integration..

When executing the base.en model on NVIDIA Jetson Orin Nano, WhisperTRT runs ~3x faster while consuming only ~60% the memory compared with PyTorch.

By default, this uses the base (multilingual) model.

WhisperTRT roughly mimics the API of the original Whisper model, making it easy to use. The Wyoming goodies are based off wyoming-faster-whisper with minimal tweaks to use WhisperTRT instead of faster-whisper.

While WhisperTRT was originally built for and tested on the Jetson Orin Nano, this project was built in Docker on an x86 Ubuntu 24.04 VM with a 4070 Ti.

Check out the performance and usage details below!

Performance

All benchmarks are generated by calling profile_backends.py, processing a 20-second audio clip.

Execution Time

Execution time in seconds to transcribe 20 seconds of speech on Jetson Orin Nano. See profile_backend.py for details.

whisper (Jetson) faster_whisper (Jetson) whisper_trt (Jetson) whisper (4070 Ti) faster_whisper (4070 Ti) whisper_trt (4070 Ti)
tiny.en 1.74 sec 0.85 sec 0.64 sec 0.40 sec 0.35 sec 0.07 sec
base.en 2.55 sec Unavailable 0.86 sec 0.71 sec 0.34 sec 0.10 sec

Memory Consumption

Memory consumption to transcribe 20 seconds of speech on Jetson Orin Nano. See profile_backend.py for details.

whisper (Jetson) faster_whisper (Jetson) whisper_trt (Jetson) whisper (4070 Ti) faster_whisper (4070 Ti) whisper_trt (4070 Ti)
tiny.en 569 MB 404 MB 488 MB 672 MB 522 MB 544 MB
base.en 666 MB Unavailable 439 MB 726 MB 514 MB 548 MB

Usage

NOTE: ARM64 dGPU and iGPU containers may take a while to start on first launch after installation or updates. I do not have ARM64 or Jetson devices so several packages such as torch and torch2trt fail to install properly because CUDA is not detected when using QEMU/buildx. If you know how to get around this please reach out to me.

Supported Models

NOTE: Only the official OpenAI models from HuggingFace are currently supported. Other variants which have been quantized or modified are not.

Multilingual

  • tiny
  • base
  • small
  • medium
  • large
  • large-v2
  • large-v3
  • large-v3-turbo

English only

  • tiny.en
  • base.en
  • small.en

Limitations

Single 30 s window. Whisper's encoder operates on a fixed 30-second window, and WhisperTRT transcribes exactly one window per request: audio longer than 30 s is truncated to the first 30 s (there is no sliding-window chunking). This is well-suited to the short utterances Home Assistant sends, but it is not a general long-form transcription backend.

LANGUAGE=auto is effectively English. Upstream whisper.transcribe runs a dedicated language-detection pass and prepends the detected <|lang|> token to the decoder prompt. WhisperTRT does not run that pass — in auto mode no language token is injected, so decoding is unsteered and strongly biased toward English. For reliable non-English transcription, set LANGUAGE explicitly (e.g. LANGUAGE=es) rather than relying on auto. English-only models (*.en) always transcribe English regardless of this setting.

Low-resource / edge tuning

On memory- or compute-constrained devices (e.g. Jetson Orin Nano) a few knobs trade accuracy for a smaller/faster footprint. All are set via the container environment:

  • MODEL=tiny.en — the smallest English model. Substantially faster and lower VRAM than base.en, at some WER cost. For the short, limited-vocabulary utterances Home Assistant sends this is often more than adequate; use a larger model only if you see it mis-hearing commands.
  • DECODER_MODE=simple — the single-engine decoder. Measured (base, RTX 3050) at 834 MiB vs 1032 MiB for the default kv mode — ~200 MiB less VRAM — for ~14% higher latency. A good trade on tight-VRAM devices. See Decoder modes.
  • PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True — an optional, low-risk allocator setting that can reduce reserved-but-unused VRAM from fragmentation. Effect on a steady inference workload is modest, but it costs nothing to enable on a constrained device.

Most of the runtime footprint is the fixed CUDA/cuDNN context rather than the model weights (the TensorRT engines for base total only tens of MiB), so model choice and decoder mode are the levers that actually move VRAM.

Pre-requisites:

  1. Install and configure Docker
  2. Install and configure the Nvidia Container Toolkit

CUDA / driver requirement

This fork builds against CUDA 12 (cu126 wheels), which run on any r525+ driver — including hosts whose nvidia-smi reports CUDA Version: 12.7.

Upstream tracks CUDA 13 (tensorrt_cu13, plain torch). CUDA 13 binaries require an r580+ driver; on an r5xx driver they import successfully and then fail at the first real CUDA call, which is why upstream does not run on a 12.7 host. The pins live in requirements.txt — note the tensorrt meta-package is deliberately absent, since it hard-pins tensorrt_cu13; tensorrt_cu12 provides the same importable tensorrt module.

The Docker base images are unaffected: torch ships its own CUDA runtime wheels and TensorRT is a pip wheel, so the venv is self-contained and the CUDA major version is decided entirely by requirements.txt.

Docker Compose (recommended)

For discrete GPUs (AMD64 or ARM64):

services:
  wyoming-whisper-trt:
    image: ghcr.io/teagan42/wyoming-whisper-trt:latest
    container_name: wyoming-whisper-trt
    ports:
      - 10300:10300
    restart: unless-stopped
    environment:
      MODEL:      "${MODEL:-base}"
      LANGUAGE:   "${LANGUAGE:-auto}"
      URI:        "${URI:-tcp://0.0.0.0:10300}"
      DATA_DIR:   "${DATA_DIR:-/data}"
      COMPUTE_TYPE: "${COMPUTE_TYPE:-float16}"
      DEVICE:     "${DEVICE:-cuda}"
      BEAM_SIZE:  "${BEAM_SIZE:-1}"
      STREAMING:  "${STREAMING:-false}"
      DEBUG:      "${DEBUG:-false}"
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]

For ARM64 with an iGPU like Jetson devices:

services:
  wyoming-whisper-trt:
    image: ghcr.io/teagan42/wyoming-whisper-trt:latest-igpu
    container_name: wyoming-whisper-trt
    restart: unless-stopped
    environment:
      MODEL:      "${MODEL:-base}"
      LANGUAGE:   "${LANGUAGE:-auto}"
      URI:        "${URI:-tcp://0.0.0.0:10300}"
      DATA_DIR:   "${DATA_DIR:-/data}"
      COMPUTE_TYPE: "${COMPUTE_TYPE:-float16}"
      DEVICE:     "${DEVICE:-cuda}"
      BEAM_SIZE:  "${BEAM_SIZE:-1}"
      STREAMING:  "${STREAMING:-false}"
      DEBUG:      "${DEBUG:-false}"
    network_mode: host
    runtime: nvidia
    environment:
      - NVIDIA_VISIBLE_DEVICES=all
      - NVIDIA_DRIVER_CAPABILITIES=compute,utility

Docker (Latest tag on GHCR)

  1. Clone this repository
  2. Browse to the repository root folder
  3. Run the following command based on your platform:

For discrete GPUs (AMD64 or ARM64):

docker run \
  --gpus all \                                # expose all NVIDIA GPUs
  --name wyoming-whisper-trt \                # give the container a name
  -d \                                        # run in detached mode
  -p 10300:10300 \                            # map port 10300 → 10300
  -e MODEL=base \                             # which model to load (tiny, small, base, etc.)
  -e LANGUAGE=auto \                          # default transcription language (`auto` = detect)
  -e COMPUTE_TYPE=float16 \                   # float16, float32, or int8 (int8 = experimental, encoder-only INT8 + FP16 decoder)
  -e DECODER_MODE=kv \                        # kv (fast, ~1 GB VRAM) or simple (lean, ~200 MB less)
  -e DEVICE=cuda \                            # must be `cuda`; TensorRT engines are GPU-only
  -e BEAM_SIZE=1 \                            # 1 = greedy; >1 = beam search (slower, needs DECODER_MODE=kv)
  ghcr.io/teagan42/wyoming-whisper-trt:latest

For ARM64 with iGPU:

docker run \
  --gpus all \                                # expose all NVIDIA GPUs
  --name wyoming-whisper-trt \                # give the container a name
  -d \                                        # run in detached mode
  -p 10300:10300 \                            # map port 10300 → 10300
  -e MODEL=base \                             # which model to load (tiny, small, base, etc.)
  -e LANGUAGE=auto \                          # default transcription language (`auto` = detect)
  -e COMPUTE_TYPE=float16 \                   # float16, float32, or int8 (int8 = experimental, encoder-only INT8 + FP16 decoder)
  -e DECODER_MODE=kv \                        # kv (fast, ~1 GB VRAM) or simple (lean, ~200 MB less)
  -e DEVICE=cuda \                            # must be `cuda`; TensorRT engines are GPU-only
  -e BEAM_SIZE=1 \                            # 1 = greedy; >1 = beam search (slower, needs DECODER_MODE=kv)
  ghcr.io/teagan42/wyoming-whisper-trt:latest-igpu

Docker (Latest GitHub commit, ARM64 and AMD64 with dGPU)

  1. Clone this repository
  2. Browse to the repository root folder
  3. Run docker compose -f docker-compose-github.yaml up -d

Docker (Latest GitHub commit, ARM64 with iGPU)

  1. Clone this repository
  2. Browse to the repository root folder
  3. Run docker compose -f docker-compose-github-igpu.yaml up -d

Development

For development setup, testing, and contribution guidelines, see CONTRIBUTING.md.

This repository includes comprehensive Copilot instructions to help GitHub Copilot coding agent (and human developers) work effectively with the codebase.

Quick Start for Developers

# Clone with submodules
git clone --recursive https://github.com/Jonah-May-OSS/wyoming-whisper-trt.git
cd wyoming-whisper-trt

# Setup development environment
python script/setup --dev

# Format code
python script/format

# Run linting checks (includes ruff, black, isort, mypy)
python script/lint

# Run tests
python script/test

Benchmarking compute types

To compare float32 / float16 / int8 on real hardware (WER, latency, VRAM, engine sizes), three helper scripts are included. They need a CUDA GPU and the venv created by script/setup:

# One-time: build an eval set from LibriSpeech test-clean (~346 MB download,
# CC BY 4.0). Any folder of .wav/.flac files with sibling .txt transcripts
# also works.
./script/prepare_eval_set --output ./eval

# Compare compute types (default kv decoder). Each runs in its own subprocess
# so VRAM numbers are clean; engines are cached per compute type + decoder
# mode under --data-dir.
./script/benchmark --model base --compute-type float16 int8 \
    --eval-dir ./eval --data-dir ./local --json results.json

# Compare decoder modes (see "Decoder modes" below).
./script/benchmark --model base --compute-type float16 --decoder-mode kv \
    --eval-dir ./eval --data-dir ./local --json kv.json
./script/benchmark --model base --compute-type float16 --decoder-mode simple \
    --eval-dir ./eval --data-dir ./local --json simple.json

# Check which layers actually run INT8. Per-layer precision is only recorded
# when the engine was built with --detailed-build.
./script/benchmark --model base --compute-type int8 \
    --eval-dir ./eval --data-dir ./local --detailed-build --force-rebuild
./script/layer_report --model base --compute-type int8 --data-dir ./local

# benchmark and layer_report re-exec themselves into .venv automatically,
# so they work with the system python on the shebang line.

Decoder modes (--decoder-mode, env DECODER_MODE)

Decoding is autoregressive — the text decoder runs once per output token — so its implementation drives latency. Two are available:

  • kv (default): caches self-attention key/values across steps and projects cross-attention K/V once per utterance, across three TensorRT engines. Fastest.
  • simple: a single engine that recomputes the whole token prefix each step. One engine context instead of three, so it uses less VRAM.

Measured on base, 300 utterances, float16 (RTX 3050, TensorRT 10):

kv (default) simple
WER 5.37 % 5.36 %
Latency p95 0.124 s 0.144 s
Latency mean 0.058 s 0.068 s
Realtime factor 128× 110×
VRAM (nvidia-smi) 1032 MiB 834 MiB
Decoder engine(s) 95.9 MiB (3) 51.0 MiB (1)

kv is ~14 % faster; simple uses ~200 MiB less VRAM. Both run well over 100× realtime, so on tight-VRAM devices (e.g. Jetson) simple is a reasonable trade. Note most of WhisperTRT's decode speedup comes from optimizations that apply to both modes (final-position projection, GPU-side mel, fewer host syncs); the KV cache itself is the last ~14 %.

Beam search (--beam-size, env BEAM_SIZE)

BEAM_SIZE=1 (the default) is greedy decoding: take the highest-probability token at every step. Greedy can commit to a locally attractive token that leads to a worse sentence; beam search keeps the top BEAM_SIZE hypotheses alive and picks the best complete sequence by length-normalized log-probability (matching upstream Whisper's default length_penalty=None).

The cost is latency. The decoder engines are built with the batch dimension pinned to 1, so the beams cannot be batched into one call — each keeps its own KV cache and is stepped separately, making decode roughly BEAM_SIZE× the engine calls per token. The encoder pass and the cross-attention projection are computed once and shared.

Beam search requires DECODER_MODE=kv (the default). Under simple the per-step cost is already O(prefix²), so multiplying it by the beam width is not worthwhile; that combination logs a warning and decodes greedily.

Start at 1 and only raise it if you are losing accuracy on hard audio; for short voice-assistant commands greedy is usually indistinguishable.

A note on int8

COMPUTE_TYPE=int8 requests TensorRT implicit INT8 quantization for the encoder (the decoder always stays FP16). Implicit quantization is per-layer optional — TensorRT only uses INT8 where it times faster than FP16 — and it is deprecated in TensorRT 10. Measured with the tools above (default kv decoder):

base, 300 utterances (RTX 3050, TensorRT 10) float16 int8
WER 5.37 % 5.37 %
Latency p95 0.124 s 0.124 s
VRAM (nvidia-smi) 1032 MiB 1032 MiB
Encoder engine size 40.0 MiB 40.2 MiB
INT8 layers in engine — 0 of 105

In other words: on this hardware the int8 engine is identical to float16 and just takes longer to build (WER differences are within run-to-run noise). Other GPU/TensorRT combinations may behave differently — run script/layer_report on your own hardware before assuming any benefit. Guaranteed quantization would require explicit Q/DQ (e.g. via nvidia-modelopt), which only pays off on medium/large encoders.

Silence hallucination suppression

Fed audio with no real speech, Whisper reliably invents plausible-looking text — www.mooji.org, Thank you for watching, subtitle credits, and the like — pulled from its training data. Greedy decoding never selects the model's <|nospeech|> token (a real token always outscores it), so without a guard those phantom transcripts flow straight through to Home Assistant.

Two gates suppress them:

  • No-speech gate (--no-speech-threshold, env NO_SPEECH_THRESHOLD, default 0.6): drops a window whose <|nospeech|> probability at the first decode position is at or above the threshold, exactly as upstream whisper.transcribe does. This is the accurate check and is enabled by default. Set it above 1.0 to disable.
  • Energy gate (--silence-threshold, env SILENCE_THRESHOLD, default 0.0 = disabled): an optional cheap hard cutoff that skips transcription entirely when the audio's normalized ([-1, 1]) RMS is below the threshold. Useful as belt-and-suspenders, but can clip genuinely quiet speech, so it is off unless you opt in (e.g. 0.005).

Both cause the server to emit an empty transcript, which Home Assistant handles as "nothing was said".

Testing and Linting Tools

This project uses modern Python development tools:

  • Ruff: Fast linting and formatting (10-100x faster than traditional tools)
  • pytest: Testing framework with asyncio support
  • black: Code formatter
  • isort: Import sorting
  • mypy: Static type checking

All configuration is in pyproject.toml. See CONTRIBUTING.md for detailed information.

See also:

  • torch2trt - Used to convert PyTorch model to TensorRT and perform inference.
  • NanoLLM - Large Language Models targeting NVIDIA Jetson. Perfect for combining with ASR!

Contributors

JonahMMayrenovate[bot]dependabot[bot]Copilotglados-arc-runners[bot]jaybdubdusty-nvTeagan42teagan-extrahopgithub-actions[bot]Srafingtoncoderabbitai[bot]mkevenaarmend-bolt-for-github[bot]

Issues