Daino is an evaluation harness for AI coding-agent context changes. It helps you compare baseline vs candidate setups with reproducible runs, structured event logs, and statistical verdicts.
- Did this context change improve verifier pass rate?
- Did efficiency improve (tokens, turns, cost, wall time)?
- Did agent behavior improve (repeat reads, hook blocks, tool errors)?
flowchart LR
A[Bundle Manifest\ncontext surface] -->|bundle_id hash| B[Run Execution]
T[Task Fixture + Verifier] --> B
B --> E[NDJSON Event Log\nappend-only]
E --> R[Reduction Layer\nDuckDB + SQL]
R --> C[Bundle Comparison\nA/B verdict + CI]
C --> D{Decision}
D -->|Promote| P[Champion Bundle]
D -->|Hold/Reject| N[Next Candidate]
bun install
bun run reduce
bun run check
bun run compare -- demo00000001 demo00000002Useful commands:
bun run autopsy -- <run_id>
bun run new-run -- --bundle <bundle_id> --task <task_id> --replicate 0- README.md: quick intro, architecture, and essential commands
- DESIGN.md: statistical model, schema, and design rationale
- docs/WORKFLOW.md: day-to-day operating playbook
- docs/CONTRIBUTING.md: contribution expectations and PR hygiene
- docs/TASK-CONTRACT.md: template for new task fixtures and verifiers
src/: core logic (bundle,reduce,compare,stats,autopsy,new-run)sql/: reduction and analysis SQLbundles/: bundle manifests by bundle_idtasks/: fixtures, verifiers, and metadata by task_idruns/: run event streams (curated sample fixtures committed)scratch-runs/: local scratch run outputresults/: generated analysis outputs
- Two demo bundles in
bundles/ - One smoke task in
tasks/T00-smoke/ - Two sample run streams in
runs/
These exist so reduce/check/compare work out of the box.