zanetworker/guideagent

Benchmark AI agent performance: latency, cost, reliability

★ 0Forks 0GoGitHub ↗Compare

README

GuideAgent

Benchmark AI agent performance: latency, cost, and reliability.

GuideAgent does for AI agents what GuideLLM does for inference servers. It measures what matters in production -- not just whether an agent can solve a task, but how fast, at what cost, and how reliably.

Install

go install github.com/zanetworker/guideagent/cmd/guideagent@latest

Or build from source:

git clone https://github.com/zanetworker/guideagent.git
cd guideagent
make install

Quick Start

  1. Create a task suite (tasks.yaml):
name: "coding-basics"
description: "Simple coding tasks for baseline measurement"
tasks:
  - id: "fizzbuzz"
    prompt: "Write a fizzbuzz function in Python"
    category: coding
    complexity: simple
    timeout: 60s
    expect:
      contains: ["def", "Fizz", "Buzz"]

  - id: "parse-json"
    prompt: "Write a Python function that reads a JSON config file and returns a dictionary"
    category: coding
    complexity: medium
    timeout: 120s
  1. Run a benchmark:
# Single run
guideagent run --backend exec --command "claude --bare -p" --tasks tasks.yaml

# With cost estimation
guideagent run --backend exec --command "claude --bare -p" --tasks tasks.yaml --model claude-sonnet-4-5

# Repeated runs for reliability measurement
guideagent run --backend exec --command "claude --bare -p" --tasks tasks.yaml --runs 5 --model claude-sonnet-4-5

# Against an HTTP agent API
guideagent run --backend http --endpoint http://localhost:8080/agent --tasks tasks.yaml

What It Measures

Category Metrics
Timing Wall time, TTFA (time to first action), distribution (min/mean/median/p95/p99/max/stddev)
Cost Token counts (input/output), estimated USD cost per model
Reliability Success rate, wall time CV, consistency score (across repeated runs)
Validation Output checked against expected content patterns

Output Formats

# Console table (default)
guideagent run ... --format console

# Machine-readable JSON
guideagent run ... --format json --output results.json

# Markdown report
guideagent run ... --format markdown --output report.md

# Convert saved results to another format
guideagent report --input results.json --format markdown

Supported Models (for cost estimation)

guideagent models

Claude (Sonnet 4.5, Opus 4, Haiku 3.5), GPT (4o, 4o-mini), o3, o4-mini, Gemini (2.5 Pro, 2.5 Flash), Codex.

Backends

Backend Flag Description
exec --command "agent-cli" Run any CLI agent as a subprocess
http --endpoint http://... Call agent HTTP APIs (JSON request/response)

Adding a new backend requires one Go file implementing two methods (Name, Run).

Architecture

cmd/guideagent/         CLI entry point (cobra)
internal/
  task/                 Task suite YAML parsing
  backend/              Backend interface + registry (exec, http)
  runner/               Orchestrator with repeated-run support
  tokens/               Token usage extraction from agent output
  cost/                 Model pricing and cost estimation
  validate/             Output validation against expectations
  metrics/              Statistical distributions + reliability analysis
  report/               Console, JSON, and Markdown reporters
  config/               YAML config loading
  profile/              Run profile types

Development

make check    # go vet + go test + go build
make build    # build binary
make test     # run tests (93 tests across 11 packages)

Contributors

zanetworker

Issues