TheTom/tqkit

Unified toolkit for benchmarking and integrating TurboQuant+ KV-cache compression across llama.cpp, vLLM, MLX, and vllm-swift.

โ˜… 1Forks 1PythonGitHub โ†—Compare

README

tqkit

๐Ÿšง Work in progress. Pre-alpha. APIs, layout names, and headline numbers will shift as engine ports and validation runs land. Pin a version (pip install tqkit==0.4.2) if you depend on it. Issues and PRs welcome.

Unified toolkit for benchmarking and integrating TurboQuant+ KV-cache compression across LLM inference engines.

What this is

tqkit is a single CLI and Python package that talks to every inference engine that ships TurboQuant+ KV-cache compression:

You bring the inference engine. tqkit autodetects what's installed, runs the canonical benchmark, and prints a reproducible KV-savings table.

Why this exists

KV cache is the dominant memory cost at long context. TurboQuant+ asymmetric (K=FP8, V=4-bit + metadata) shrinks it ~62% (or ~57% accounting for the 4 boundary layers that stay FP16). The savings replicate across engines and hardware vendors. tqkit is the proof, the tool, and the install path.

For a 14B model at 1M tokens of context:

layout KV cache size (all-quantized) fits on MI300X 192GB after weights?
FP16 192 GB no
TQ+ asym (turboquant_k8v4) 73.5 GB headline / ~83 GB realistic with boundary skip yes
TQ+ sym 4-bit (turboquant_4bit_nc) 50.3 GB yes (more headroom)

You can verify the math yourself:

pip install tqkit
tq report --model qwen2.5-14b-instruct-1m --ctx 1M --layout tq+asym
tq table --model qwen2.5-14b-instruct-1m

Install

pip install tqkit

Usage

tq backends                                            # autodetect installed engines
tq report --model qwen2.5-14b-instruct-1m --ctx 32K    # KV cache size for one config
tq table --model qwen2.5-14b-instruct-1m               # full layout ร— ctx grid
tq integrate <backend>                                 # install + serve recipe
tq bench                                               # canonical benchmark (v0.3.0)

Example output:

$ tq report --model qwen2.5-14b-instruct-1m --ctx 1M --layout tq+asym
[KV cache] model: Qwen/Qwen2.5-14B-Instruct-1M
[KV cache] arch: layers=48 kv_heads=8 head_dim=128
[KV cache] layout: tq+asym
[KV cache] per-token: 72.0 KB (vs 192.0 KB FP16)
[KV cache] total @ 1M ctx: 72.0 GB (vs 192.0 GB FP16, 62.5% savings)

Integration recipes

One-page docs for plugging TurboQuant+ into each supported backend live under docs/integrate/.

Docker (AMD ROCm)

docker pull thetom/vllm-turboquant:rocm-7.2
docker run --rm -it \
    --device=/dev/kfd --device=/dev/dri --group-add video --ipc=host \
    -v "$HOME/.cache/huggingface:/root/.cache/huggingface" \
    -p 8000:8000 thetom/vllm-turboquant:rocm-7.2 \
        --model Qwen/Qwen2.5-14B-Instruct-1M --kv-cache-dtype turboquant_k8v4

See docker/README.md for build details.

Status

v0.4.0 โ€” alpha. Shipping today:

  • KV math + tq report + tq table
  • Pinned canonical_bench.yml + tq config
  • Engine bridges (tq bench) for llama.cpp, vLLM (CUDA + AMD), MLX-Swift, vllm-swift
  • Integration recipes for all 5 backends (docs/integrate/)
  • Docker scaffold for AMD ROCm (docker/Dockerfile.vllm-amd)
  • Models supported in tq report: Qwen2.5 7B/14B/32B, Qwen3-8B, Qwen3.6-27B, Qwen3.6-35B-A3B, Qwen3-Next-80B-A3B, Llama-3.1 8B/70B, Mistral-7B
  • 39 tests, 92% line coverage, โ‰ฅ85% gate enforced

See CHANGELOG.md for the full version history.

License

Apache 2.0.

Contributors

TheTom

Issues