This is the Whamp fork of Turbo Agent. It adds native Pi and Codex execution, local-model routing, configurable verification, and shared endpoint admission. See How this fork differs from upstream for the full comparison, current limits, and upstream changes that still need reconciliation.
Turbo Agent is the Claude Code plugin for LLM-as-a-Verifier. It implements an LLM API proxy that improves response quality through concurrent inference, verification, and refinement. It sits between your client (Claude Code, Codex, etc.) and the LLM provider, sending multiple parallel requests and selecting the best response with a Probabilistic Pivot Tournament (PPT) scored by a fine-grained logprob verifier.
Client request
│
[Context Refinement] (optional) rewrite/augment the system prompt for clarity
│
[Concurrent Inference] send N parallel candidates to the backend model
│
[Verification] pivot tournament over the candidates, pick the best one
│
Best response → Client
Verification uses the pivot tournament from the llm-verifier package to pick the best of N candidates.
pip install turbo-agentOr from source:
pip install -e .For turbo agent to work, you need a turbo-agent.yaml. You can copy the reference file in this repo.
turbo-agent.yaml references keys with $VAR_NAME syntax. The recommended way to provide them is a .env file in the project root (next to turbo-agent.yaml) — the proxy loads it automatically on startup. Copy the committed template and fill in your keys:
cp .env.example .env
# then edit .env# .env
VERTEX_API_KEY=your-vertex-key # preferred for Gemini 2.5 logprobs (verifier)
# GEMINI_API_KEY=your-gemini-key # used by gemini/ models (AI Studio)
# OPENAI_API_KEY=... # only if you route to openai/ models
# ANTHROPIC_API_KEY=... # only if you route to anthropic/ models.env is gitignored; .env.example is committed as the template. Keys already
exported in your shell environment work too and take nothing extra. The verifier
and progress monitor use Gemini logprobs, which are best served by a Vertex
AI key (VERTEX_API_KEY + provider: vertex_ai in the config); a plain
GEMINI_API_KEY also works for the gemini/ backend models.
Verify your keys are valid:
turbo-agent checkIt checks every supported provider (Gemini, Vertex AI, OpenAI, Anthropic) and reports each with ✅ / ❌ /
turbo-agent # default port 8888
turbo-agent -p 9000 # custom portANTHROPIC_BASE_URL=http://localhost:8888 claudeexport OPENAI_API_BASE=http://localhost:8888/v1Pi connects through the
proxy as an OpenAI-compatible client, which works with any backend model
(Gemini, OpenAI, Anthropic, OpenRouter, Zai, Kimi, ...) in your
turbo-agent.yaml.
Pi keeps custom providers in ~/.pi/agent/models.json. Generate the turbo
provider block straight from your config:
python integrations/pi/gen_models_json.py turbo-agent.yaml --merge ~/.pi/agent/models.jsonThat registers one turbo model per backend model (metadata such as context
window and max tokens come from your config, so what Pi shows matches what the
proxy runs). Then, with the proxy running:
turbo-agent # project config, or the global default when none
pi # select the model with /model: turbo/<backend-model>The generated block is only stale when turbo-agent.yaml changes: a new Pi
session just re-reads ~/.pi/agent/models.json when you open /model (no
regeneration needed).
Notes:
- The proxy ignores the model id a client requests — it always runs the
backend models from
turbo-agent.yaml— but it now echoes the requested id back in responses, so Pi displays the model you picked. - The proxy ignores client API keys; the
apiKeyin the generated provider is a placeholder. If Pi hides the models until auth is resolved, save any key with/login turboor pass--api-keywhen selecting the model. - Pi always streams. With a verifier configured, Turbo Agent gathers all
candidates and verifies before replaying the best response as a stream, so
each turn costs
num_candidatesfull responses plus verifier calls. Tunenum_candidates(3 is the reference default) andmajority_voting: trueto control cost/latency. - The verifier judge is configurable — it is not tied to Gemini. Set
verifier.model.nameto any litellm-style model:openrouter/...(defaults tohttps://openrouter.ai/api/v1),deepseek/...(hosted DeepSeek), oropenai/...with abase_urlpointing at a local vLLM/SGLang endpoint (full fine-grained logprob reward needs a server that exposes logprobs; OpenRouter degrades to parsing the judge's written score when the upstream provider does not). Ifverifier.modelis omitted the judge defaults to the backend candidate model. - A named
endpointcan cap candidate and judge traffic to one local server. It defines the base URL, bounded FIFO queue, queue deadline, request deadline, and maximum active calls. The opt-inserver60judge adapter uses x-high comparison generations and non-thinking one-token score probes. See Endpoint admission and server60 judging. - Pi counts tokens locally; the proxy also answers
/v1/messages/count_tokenswith an approximate local count so token-counting clients never leak a request to api.anthropic.com.
Edit turbo-agent.yaml. API keys can reference environment variables with
$VAR_NAME syntax. See the reference turbo-agent.yaml file and the
endpoint admission guide.
Config discovery works like pi's settings files — a project config, then a global default:
--config PATH(explicit, always wins)./turbo-agent.yamlin the current directory (project-level)~/.config/turbo-agent/turbo-agent.yaml(global default, used only when the project file doesn't exist; honors$XDG_CONFIG_HOME)
A project file fully replaces the global one (no merging). The .env file
with API keys is always loaded from the same directory as the config that
was chosen, so a global config reads ~/.config/turbo-agent/.env.
| Prefix | Provider |
|---|---|
gemini/ |
Google Gemini |
openai/ |
OpenAI |
anthropic/ |
Anthropic |
| (none) | OpenAI-compatible endpoint |
| Endpoint | Format |
|---|---|
POST /v1/messages |
Anthropic |
POST /v1/messages/count_tokens |
Anthropic (approximate local count) |
POST /v1/chat/completions |
OpenAI |
GET /v1/models |
OpenAI |
GET /visualizer |
Pipeline visualizer UI |
* |
Upstream passthrough to api.anthropic.com |
A built-in web UI at http://localhost:8888/visualizer shows the pipeline DAG for each request — context refinement, all candidate responses, the pairwise tournament comparisons and scores, and the final selection.
To build the frontend (requires Node.js):
cd frontend
yarn install
yarn buildcd frontend && yarn build && cd ..
pip install build twine
rm -rf dist
python -m build
twine check dist/*
twine upload dist/*