Azayzel/Daino

Daino is an evaluation harness for AI coding-agent configurations. It lets you test changes to instruction/context surfaces (like prompt files, tool allowlists, hooks, and folder structure) using repeatable A/B runs, structured event logs, and statistical comparison.

★ 2Forks 0TypeScriptGitHub ↗Compare

README

Daino

Daino is an evaluation harness for AI coding-agent context changes. It helps you compare baseline vs candidate setups with reproducible runs, structured event logs, and statistical verdicts.

What Daino Answers

  • Did this context change improve verifier pass rate?
  • Did efficiency improve (tokens, turns, cost, wall time)?
  • Did agent behavior improve (repeat reads, hook blocks, tool errors)?

Harness Overview

flowchart LR
    A[Bundle Manifest\ncontext surface] -->|bundle_id hash| B[Run Execution]
    T[Task Fixture + Verifier] --> B
    B --> E[NDJSON Event Log\nappend-only]
    E --> R[Reduction Layer\nDuckDB + SQL]
    R --> C[Bundle Comparison\nA/B verdict + CI]
    C --> D{Decision}
    D -->|Promote| P[Champion Bundle]
    D -->|Hold/Reject| N[Next Candidate]
Loading

Quick Start

bun install
bun run reduce
bun run check
bun run compare -- demo00000001 demo00000002

Useful commands:

bun run autopsy -- <run_id>
bun run new-run -- --bundle <bundle_id> --task <task_id> --replicate 0

Documentation Map

Repository Layout

  • src/: core logic (bundle, reduce, compare, stats, autopsy, new-run)
  • sql/: reduction and analysis SQL
  • bundles/: bundle manifests by bundle_id
  • tasks/: fixtures, verifiers, and metadata by task_id
  • runs/: run event streams (curated sample fixtures committed)
  • scratch-runs/: local scratch run output
  • results/: generated analysis outputs

Included Demo Scaffold

  • Two demo bundles in bundles/
  • One smoke task in tasks/T00-smoke/
  • Two sample run streams in runs/

These exist so reduce/check/compare work out of the box.

Contributors

lavelyio

Issues