Price–performance Pareto frontiers and oracle router projections for Terminal-Bench 2.0, computed from real leaderboard scores, published token usage, official API list prices, and per-task HF trial artifacts.
No simulated scores or prices — all numbers come from primary sources cited below.
| Artifact | Description |
|---|---|
fetch_and_analyze.py |
Downloads the official leaderboard, joins API pricing, computes Pareto frontiers |
oracle_router.py |
Builds per-task oracle routing from HF result.json trial data |
oracle_iso_comparisons.py |
Iso-price / iso-accuracy comparisons (same budget or same accuracy) |
generate_plots.py |
Renders all PNG charts from JSON outputs |
output/ |
Generated JSON data and plots |
| Blended API $/MTok | Accuracy | Model |
|---|---|---|
| $0.23 | 21.8% | GPT-5-Nano |
| $0.35 | 39.6% | DeepSeek-V3.2 |
| $5.63 | 78.4% | GPT-5.3-Codex |
| $15.00 | 90.2% | Claude Opus 4.7 |
GPT-5.5 (84.7%) is not on the frontier — Opus 4.7 is both cheaper ($15 vs $17.50/MTok blended) and more accurate.
Using 3,432 real HF trials across 7 models on 89 tasks:
| Metric | Full pool (7 models) | Big-3 only (OAI + Anthropic + Gemini) |
|---|---|---|
| Accuracy | 92.1% (82/89) | 92.1% (82/89) |
| Measured cost | $20.73 | $28.44 (+37%) |
| Unsolved tasks | 7 | 7 (same) |
Full pool routing: GLM-4.7 (45), GLM-5 (16), GPT-5.3-Codex (8), Claude Opus 4.6 (6), Gemini 3.1 Pro (5), DeepSeek (2)
Big-3 routing: GPT-5.3-Codex (53), Claude Opus 4.6 (20), Gemini 3.1 Pro (9)
Key insight: GLM and DeepSeek add zero extra solvable tasks — they only provide cheaper routes to tasks the Big-3 models already solve. Restricting to OpenAI + Anthropic + Gemini keeps the same 92.1% ceiling but costs ~$8 more per full benchmark run.
Best single model in pool (Gemini 3.1 Pro): 74.8% @ $407 measured cost.
| Reference model | Budget | Single accuracy | Oracle @ same budget |
|---|---|---|---|
| DeepSeek-V3.2 | $13.67 | 39.6% | 91.0% |
| Gemini 3.1 Pro | $406.63 | 74.8% | 92.1% |
| Claude Opus 4.6 | $1,007 | 68.7% | 92.1% |
| Target accuracy | Single-model cost | Oracle cost |
|---|---|---|
| 74.8% | Gemini 3.1 Pro $407 | $0.00† |
| 68.7% | Opus 4.6 $1,007 | $0.00† |
† Oracle reaches moderate accuracy targets using mostly GLM-routed tasks with $0 measured cost_usd in HF trials. Full 92.1% oracle still costs $20.73 (full pool) or $28.44 (Big-3 only).
All plots live in output/:
| File | Description |
|---|---|
pareto_api_price_vs_accuracy.png |
API price vs accuracy Pareto + full & Big-3 oracle |
pareto_benchmark_cost_vs_accuracy.png |
Paper Table 2 run cost vs accuracy |
pareto_terminus2_cost_vs_accuracy.png |
Terminus 2 agent cost vs accuracy |
oracle_router_vs_single_models.png |
Full & Big-3 oracle vs single-model scatter |
oracle_full_vs_big3.png |
Accuracy & cost bar chart: full pool vs Big-3 |
oracle_routing_mix_full_vs_big3.png |
Task routing mix comparison |
oracle_iso_price_measured_curve.png |
Same budget: single vs oracle accuracy |
oracle_iso_accuracy_measured.png |
Same accuracy: single vs oracle cost |
oracle_iso_price_api_constrained.png |
Same API price tier: constrained oracle |
oracle_iso_lines_overlay.png |
Iso-price & iso-accuracy guide lines |
python3 -m venv .venv
source .venv/bin/activate
pip install matplotlib
python fetch_and_analyze.py
python oracle_router.py --cache-only
python generate_plots.pyTo download HF trial data from scratch (slow, ~30 min):
python oracle_router.py| Source | URL |
|---|---|
| Terminal-Bench 2.0 leaderboard | https://www.tbench.ai/leaderboard/terminal-bench/2.0 |
| Paper (Table 2 token usage) | https://arxiv.org/abs/2601.11868 |
| HF trial artifacts | https://huggingface.co/datasets/harborframework/terminal-bench-2-leaderboard |
| OpenAI API pricing | https://developers.openai.com/api/docs/pricing |
| Anthropic pricing | https://www.anthropic.com/pricing |
Accuracy = task resolution rate (% of trials with verifier reward > 0, aggregated per task and model).
Benchmark run cost = input_tokens × input_price + output_tokens × output_price (paper Table 2) or sum of cost_usd from Harbor result.json (HF trials).
Oracle router = for each task, pick the cheapest model in the pool with ≥1 successful trial. This is an upper bound — it requires knowing which model solves each task in advance.
Big-3 oracle = same routing rule, but pool restricted to OpenAI, Anthropic, and Google/Gemini models only (GLM, DeepSeek, etc. excluded).
- Agent matters — same model can differ by 20+ points depending on agent scaffold.
- Paper vs leaderboard vintages — Table 2 uses 74 tasks; current leaderboard uses 89 tasks.
- HF pool is limited — oracle projection uses 7 Terminus-family models with trial data, not all 144 leaderboard entries.
- Open-weight models (GLM, etc.) report $0 API price and often $0 measured cost in trials.
- Oracle ≠ learnable router — results show headroom from perfect routing, not achievable by a trained classifier today.
Analysis scripts: MIT. Terminal-Bench data and trial artifacts belong to their respective owners (Laude Institute / Stanford, HF submitters).