Benchmark AI agent performance: latency, cost, and reliability.
GuideAgent does for AI agents what GuideLLM does for inference servers. It measures what matters in production -- not just whether an agent can solve a task, but how fast, at what cost, and how reliably.
go install github.com/zanetworker/guideagent/cmd/guideagent@latestOr build from source:
git clone https://github.com/zanetworker/guideagent.git
cd guideagent
make install- Create a task suite (
tasks.yaml):
name: "coding-basics"
description: "Simple coding tasks for baseline measurement"
tasks:
- id: "fizzbuzz"
prompt: "Write a fizzbuzz function in Python"
category: coding
complexity: simple
timeout: 60s
expect:
contains: ["def", "Fizz", "Buzz"]
- id: "parse-json"
prompt: "Write a Python function that reads a JSON config file and returns a dictionary"
category: coding
complexity: medium
timeout: 120s- Run a benchmark:
# Single run
guideagent run --backend exec --command "claude --bare -p" --tasks tasks.yaml
# With cost estimation
guideagent run --backend exec --command "claude --bare -p" --tasks tasks.yaml --model claude-sonnet-4-5
# Repeated runs for reliability measurement
guideagent run --backend exec --command "claude --bare -p" --tasks tasks.yaml --runs 5 --model claude-sonnet-4-5
# Against an HTTP agent API
guideagent run --backend http --endpoint http://localhost:8080/agent --tasks tasks.yaml| Category | Metrics |
|---|---|
| Timing | Wall time, TTFA (time to first action), distribution (min/mean/median/p95/p99/max/stddev) |
| Cost | Token counts (input/output), estimated USD cost per model |
| Reliability | Success rate, wall time CV, consistency score (across repeated runs) |
| Validation | Output checked against expected content patterns |
# Console table (default)
guideagent run ... --format console
# Machine-readable JSON
guideagent run ... --format json --output results.json
# Markdown report
guideagent run ... --format markdown --output report.md
# Convert saved results to another format
guideagent report --input results.json --format markdownguideagent modelsClaude (Sonnet 4.5, Opus 4, Haiku 3.5), GPT (4o, 4o-mini), o3, o4-mini, Gemini (2.5 Pro, 2.5 Flash), Codex.
| Backend | Flag | Description |
|---|---|---|
exec |
--command "agent-cli" |
Run any CLI agent as a subprocess |
http |
--endpoint http://... |
Call agent HTTP APIs (JSON request/response) |
Adding a new backend requires one Go file implementing two methods (Name, Run).
cmd/guideagent/ CLI entry point (cobra)
internal/
task/ Task suite YAML parsing
backend/ Backend interface + registry (exec, http)
runner/ Orchestrator with repeated-run support
tokens/ Token usage extraction from agent output
cost/ Model pricing and cost estimation
validate/ Output validation against expectations
metrics/ Statistical distributions + reliability analysis
report/ Console, JSON, and Markdown reporters
config/ YAML config loading
profile/ Run profile types
make check # go vet + go test + go build
make build # build binary
make test # run tests (93 tests across 11 packages)