A lightweight Python layer that runs one instruction prompt through either Claude Code or OpenAI Codex, chosen explicitly per task. Runs on your CLI subscription logins by default, or on a third-party LLM provider (DeepSeek V4, GLM 5.2, local Ollama) — no code changes, just a flag. Every run is logged to Langfuse automatically when keys are present.
- claude backend → official
claude-agent-sdk - codex backend → thin async subprocess wrapper over
codex exec --json
~750 lines of source, one runtime dependency (claude-agent-sdk).
- Python ≥ 3.11 (uses
asyncio.timeout) claudeCLI installed and logged in (subscription)codexCLI installed and logged in (codex login, ChatGPT subscription)
pip install -e ".[dev]"from ultraagent import run_sync
result = run_sync("read this source code and review it",
backend="codex", # "claude" | "codex"
model="gpt-5.5", # optional; backend default otherwise
cwd="/path/to/repo", # optional working directory
sandbox="read-only") # "read-only" | "workspace-write" | "full-access"
print(result.text) # final answer
print(result.session_id) # resumable session/thread id
print(result.usage) # tokens (+ cost for claude)Structured output — pass a JSON Schema and get a parsed object back:
plan = run_sync("Plan the implementation of a snake game as 3 steps.",
backend="claude",
output_schema={
"type": "object",
"properties": {
"steps": {"type": "array", "items": {"type": "string"}}
},
"required": ["steps"],
})
print(plan.structured["steps"]) # a real list, no text parsingresult.structured holds the parsed answer (claude: native
structured_output; codex: the schema-conforming final message). It is None
when no schema was given, or when the answer failed to parse — result.text
always keeps the raw answer.
By default both backends use your native subscription. Pass provider= (and
optionally model=) to route through a third-party LLM instead:
from ultraagent import run_sync, list_providers
run_sync("review this code", backend="claude", provider="deepseek") # DeepSeek V4
run_sync("review this code", backend="claude", provider="glm") # GLM 5.2
run_sync("review this code", backend="claude", provider="ollama", # local model
model="qwen3.6:35b-a3b-mxfp8")
list_providers() # discover providers, their models, and readinessCLI:
ultra --providers # list providers, models, readiness
ultra "review this code" -b claude -p deepseek # -p / --provider
ultra "explain this bug" -b codex -p ollama -m qwen3.6:35b-a3b-mxfp8Support matrix (verified live):
| provider | claude backend | codex backend | models |
|---|---|---|---|
| deepseek | ✅ native endpoint | deepseek-v4-flash (default), deepseek-v4-pro |
|
| glm | ✅ native endpoint | glm-5.2 |
|
| ollama | ✅ native endpoint | ✅ native | whatever ollama list shows (default = first) |
Structured output (output_schema) works on all three through the claude
backend — verified returning parsed objects from DeepSeek, GLM, and Ollama.
Keys live in ./.env (or the environment); os.environ wins over .env:
DEEPSEEK_API_KEY=sk-...
GLM_API_KEY=xxxxxxxx.yyyyyyyy # bigmodel.cn platform key
# Ollama needs no key — just a running server (`ollama serve`)
¹ Codex + DeepSeek/GLM needs a translation gateway. OpenAI removed
Chat-Completions support from Codex (Feb 2026); it only speaks the Responses
API, which DeepSeek and GLM don't implement. ultra … -b codex -p deepseek
therefore raises a clear BackendError. Use the claude backend (native, no
gateway) — or run a local LiteLLM proxy that bridges
Responses→Chat and strips Codex's non-function tool types, then point a
custom Codex provider at it. Ollama needs no gateway (it speaks Responses
natively). Recipe:
# gateway.yaml — litellm --config gateway.yaml --port 4000
model_list:
- model_name: deepseek-v4-flash
litellm_params: {model: deepseek/deepseek-v4-flash, api_key: os.environ/DEEPSEEK_API_KEY}
litellm_settings:
callbacks: tool_filter.proxy_handler_instance # strips namespace/web_search/image_generation tools# tool_filter.py
from litellm.integrations.custom_logger import CustomLogger
class ToolFilter(CustomLogger):
async def async_pre_call_hook(self, key, cache, data, call_type):
if data.get("tools"):
data["tools"] = [t for t in data["tools"] if t.get("type") == "function"] or None
return data
proxy_handler_instance = ToolFilter()Then run codex against http://localhost:4000/v1 with wire_api="responses".
Async, with streaming events:
from ultraagent import stream, run
async for event in stream("fix the failing test", backend="claude"):
print(event.type, event.text) # start / thinking / text / tool / error / done
result = await run("summarize the repo", backend="codex",
on_event=lambda e: print(e.type))Every AgentEvent keeps the original SDK message / codex JSON in .raw.
Agent-level failures return result.success == False; infrastructure failures
(CLI missing, crash, timeout) raise BackendError.
Every run() is logged to Langfuse as one trace
(prompt in, answer out, plus backend, provider, model, tokens, success,
duration) — for both backends and all providers. It turns on
automatically when your Langfuse keys are present in ./.env or the
environment; no code changes, no dependency:
LANGFUSE_PUBLIC_KEY=pk-lf-...
LANGFUSE_SECRET_KEY=sk-lf-...
LANGFUSE_BASE_URL=https://us.cloud.langfuse.com # optional; this is the default
Then just run tasks — traces appear in your Langfuse dashboard. Disable with
ULTRAAGENT_TELEMETRY=off. Logging is best-effort: a bad key or network blip is
swallowed and never affects the run.
Why app-level and not Claude Code's built-in OpenTelemetry: Claude Code's OTEL
export works, but codex exec emits no OTEL telemetry (upstream gap), so only
an app-level trace covers every backend uniformly. The trace is emitted from
agent.run() over the Langfuse ingestion REST API using stdlib urllib.
ultra "review this code" -b codex -s read-only # progress → stderr, answer → stdout
ultra "add type hints to utils.py" -b claude -m opus -C /path/to/repo
ultra "summarize the repo" --json | jq . # JSONL events for scripting
ultra "quick question" -q # quiet: answer only
ultra "grade this repo" --schema grade.json -q | jq . # answer conforms to the schemaExit codes: 0 success · 1 run failed · 2 bad usage.
| ultraagent | claude (permission_mode) |
codex (--sandbox) |
|---|---|---|
read-only |
dontAsk + Write/Edit/Bash disallowed |
read-only |
workspace-write (default) |
acceptEdits |
workspace-write |
full-access |
bypassPermissions |
danger-full-access |
pytest --cov=ultraagent --cov-report=term-missing # offline suite (99% coverage)
RUN_LIVE=1 pytest -m live -o addopts="" -v # real e2e (uses quota / local server)Offline tests need no network or login: the codex backend is exercised against
a fake codex stub subprocess, the claude backend against a monkeypatched
query yielding real SDK message objects, and providers against mocked keys /
urllib. The live suite auto-discovers which (backend × provider) pairs are
reachable — keys from .env, Ollama from its local server — and skips the rest,
so it runs the native backends plus every provider you have configured.
- The claude adapter strips
ANTHROPIC_API_KEYfrom the child environment so subscription (OAuth) auth always wins. For a provider it instead setsANTHROPIC_BASE_URL+ANTHROPIC_AUTH_TOKENand pins the threeANTHROPIC_DEFAULT_*_MODELslots so Claude Code's internal small-model calls stay on-provider. - Providers live in one registry (
providers.py); adding one is a singleProvider(...)entry. Native runs (provider=None) are byte-identical to before this feature. result.usage["total_cost_usd"]is Claude Code's own client-side estimate from a model-name pricing table — for Ollama and other providers it does not mean you were billed (Ollama is free). Trust it only for native Anthropic runs.- The codex adapter passes
-c notify=[]so your personal turn-end notification hooks don't fire on programmatic runs, and--skip-git-repo-checkso it works outside git repos. - Strict schemas on codex: OpenAI structured output rejects object nodes
lacking
additionalProperties: false, so the codex adapter injects it recursively (your dict is copied, never mutated). Claude accepts the same schema unchanged — write one schema, run it on both. OpenAI additionally expects every property to be listed inrequired; that is left to the caller since auto-adding it would change the schema's meaning. - Long agent messages are safe: the codex stream reader allows 16 MB per JSONL line (asyncio's default is 64 KB).
- Unknown codex event types are skipped, never fatal — schema drift tolerant.
- v2 candidates (deliberately out of scope): session resume (
session_idis already captured for both backends), parallel fan-out across backends, retry policies for orchestrators.