rootfs/terminalbench-router-projection

★ 0Forks 0PythonGitHub ↗Compare

README

Terminal-Bench Router Projection

Price–performance Pareto frontiers and oracle router projections for Terminal-Bench 2.0, computed from real leaderboard scores, published token usage, official API list prices, and per-task HF trial artifacts.

No simulated scores or prices — all numbers come from primary sources cited below.

What this repo contains

Artifact Description
fetch_and_analyze.py Downloads the official leaderboard, joins API pricing, computes Pareto frontiers
oracle_router.py Builds per-task oracle routing from HF result.json trial data
oracle_iso_comparisons.py Iso-price / iso-accuracy comparisons (same budget or same accuracy)
generate_plots.py Renders all PNG charts from JSON outputs
output/ Generated JSON data and plots

Key findings (June 2026)

Single-model Pareto frontier (API list price vs best accuracy)

Blended API $/MTok Accuracy Model
$0.23 21.8% GPT-5-Nano
$0.35 39.6% DeepSeek-V3.2
$5.63 78.4% GPT-5.3-Codex
$15.00 90.2% Claude Opus 4.7

GPT-5.5 (84.7%) is not on the frontier — Opus 4.7 is both cheaper ($15 vs $17.50/MTok blended) and more accurate.

Oracle router (perfect task→model routing, Terminus 2 pool)

Using 3,432 real HF trials across 7 models on 89 tasks:

Metric Full pool (7 models) Big-3 only (OAI + Anthropic + Gemini)
Accuracy 92.1% (82/89) 92.1% (82/89)
Measured cost $20.73 $28.44 (+37%)
Unsolved tasks 7 7 (same)

Full pool routing: GLM-4.7 (45), GLM-5 (16), GPT-5.3-Codex (8), Claude Opus 4.6 (6), Gemini 3.1 Pro (5), DeepSeek (2)

Big-3 routing: GPT-5.3-Codex (53), Claude Opus 4.6 (20), Gemini 3.1 Pro (9)

Key insight: GLM and DeepSeek add zero extra solvable tasks — they only provide cheaper routes to tasks the Big-3 models already solve. Restricting to OpenAI + Anthropic + Gemini keeps the same 92.1% ceiling but costs ~$8 more per full benchmark run.

Best single model in pool (Gemini 3.1 Pro): 74.8% @ $407 measured cost.

Iso-price (same budget → oracle accuracy)

Reference model Budget Single accuracy Oracle @ same budget
DeepSeek-V3.2 $13.67 39.6% 91.0%
Gemini 3.1 Pro $406.63 74.8% 92.1%
Claude Opus 4.6 $1,007 68.7% 92.1%

Iso-accuracy (same target → oracle cost)

Target accuracy Single-model cost Oracle cost
74.8% Gemini 3.1 Pro $407 $0.00†
68.7% Opus 4.6 $1,007 $0.00†

† Oracle reaches moderate accuracy targets using mostly GLM-routed tasks with $0 measured cost_usd in HF trials. Full 92.1% oracle still costs $20.73 (full pool) or $28.44 (Big-3 only).

Plots

All plots live in output/:

File Description
pareto_api_price_vs_accuracy.png API price vs accuracy Pareto + full & Big-3 oracle
pareto_benchmark_cost_vs_accuracy.png Paper Table 2 run cost vs accuracy
pareto_terminus2_cost_vs_accuracy.png Terminus 2 agent cost vs accuracy
oracle_router_vs_single_models.png Full & Big-3 oracle vs single-model scatter
oracle_full_vs_big3.png Accuracy & cost bar chart: full pool vs Big-3
oracle_routing_mix_full_vs_big3.png Task routing mix comparison
oracle_iso_price_measured_curve.png Same budget: single vs oracle accuracy
oracle_iso_accuracy_measured.png Same accuracy: single vs oracle cost
oracle_iso_price_api_constrained.png Same API price tier: constrained oracle
oracle_iso_lines_overlay.png Iso-price & iso-accuracy guide lines

Reproduce

python3 -m venv .venv
source .venv/bin/activate
pip install matplotlib

python fetch_and_analyze.py
python oracle_router.py --cache-only
python generate_plots.py

To download HF trial data from scratch (slow, ~30 min):

python oracle_router.py

Data sources

Source URL
Terminal-Bench 2.0 leaderboard https://www.tbench.ai/leaderboard/terminal-bench/2.0
Paper (Table 2 token usage) https://arxiv.org/abs/2601.11868
HF trial artifacts https://huggingface.co/datasets/harborframework/terminal-bench-2-leaderboard
OpenAI API pricing https://developers.openai.com/api/docs/pricing
Anthropic pricing https://www.anthropic.com/pricing

Methodology

Accuracy = task resolution rate (% of trials with verifier reward > 0, aggregated per task and model).

Benchmark run cost = input_tokens × input_price + output_tokens × output_price (paper Table 2) or sum of cost_usd from Harbor result.json (HF trials).

Oracle router = for each task, pick the cheapest model in the pool with ≥1 successful trial. This is an upper bound — it requires knowing which model solves each task in advance.

Big-3 oracle = same routing rule, but pool restricted to OpenAI, Anthropic, and Google/Gemini models only (GLM, DeepSeek, etc. excluded).

Caveats

  1. Agent matters — same model can differ by 20+ points depending on agent scaffold.
  2. Paper vs leaderboard vintages — Table 2 uses 74 tasks; current leaderboard uses 89 tasks.
  3. HF pool is limited — oracle projection uses 7 Terminus-family models with trial data, not all 144 leaderboard entries.
  4. Open-weight models (GLM, etc.) report $0 API price and often $0 measured cost in trials.
  5. Oracle ≠ learnable router — results show headroom from perfect routing, not achievable by a trained classifier today.

License

Analysis scripts: MIT. Terminal-Bench data and trial artifacts belong to their respective owners (Laude Institute / Stanford, HF submitters).

Issues