Deva-1903/AutoEval

★ 0Forks 0PythonGitHub ↗Compare

README

AutoEval

AutoEval is a minimal MVP where an agent improves an evaluation suite for retrieval systems.

The target is not model training. The target is better evals: queries and relevance labels that better separate good retrievers from bad retrievers.

Why evolve evals?

Static benchmarks get saturated. When weak systems and strong systems get similar scores, the benchmark loses discriminative power.

AutoEval iteratively proposes new eval items, measures whether they improve separation, and keeps only improvements.

What is in this repo?

  • Deterministic retrieval benchmark with 5 reference systems
  • Synthetic corpus (~300 docs) committed to git
  • Editable eval set (~120 items) and hidden holdout (~80 items)
  • Fitness objective focused on discriminative power, with guardrails
  • Rule-based agent that evolves seed_eval.jsonl
  • Git integration for accepted suite updates with report cards

The checked-in editable suite is intentionally under-discriminative: it starts with many valid but easy items where strong and weak lexical systems tie, so the agent has real room to harden the benchmark.

Quickstart

pip install -r requirements.txt
python run_eval.py --suite editable
python agent.py --minutes 10

Fitness

For each system:

  • suite_score(system) = mean nDCG@10

Discriminative power uses pairwise AUC between good and bad systems:

  • dp_editable = AUC(editable)
  • dp_holdout = AUC(holdout)
  • DP = 0.2 * dp_editable + 0.8 * dp_holdout

Penalties:

  • cost proxy (deterministic runtime estimate)
  • redundancy (query similarity)
  • flakiness (validation failures)

Final:

  • fitness = DP - 0.05*cost_seconds - 0.20*redundancy - 0.50*flakiness

Guardrails against Goodharting

  • Holdout is committed and treated as read-only by the agent
  • Schema checks reject invalid, trivial, or leaky queries
  • Near duplicates are rejected via Jaccard threshold (>0.85)
  • Fitness is holdout-dominated (80%)

Add a new system

  1. Implement a retriever in autoeval/systems/ with:
    • name
    • retrieve(query, k) -> list[doc_id]
  2. Register it in autoeval/systems/__init__.py
  3. Add it to GOOD_SYSTEMS, MID_SYSTEMS, or BAD_SYSTEMS in autoeval/config.py

Add a new generator template

  1. Add a new template function in autoeval/evals/generators.py
  2. Return a valid eval item dict (query, relevant_doc_ids, tags)
  3. Register the template in generate_candidates()
  4. Keep it deterministic through the provided RNG

Example output

python run_eval.py --suite editable

Per-system mean nDCG@10 (editable):
  bm25     1.0000
  hybrid   1.0000
  keyword  1.0000
  random   0.0289
  tfidf    1.0000
DP_editable: 0.7500
DP_holdout : 1.0000
DP_total   : 0.9500
Penalties  : cost=2.0000, redundancy=0.9053, flakiness=0.0000
Final fitness: 0.6689

python agent.py --minutes 1 --no-git

iter=1 accepted delta=0.0500 fitness=0.7189 +3 -3 sha=nogit
agent finished: iterations=73 accepted=1 best_fitness=0.7189

Structure

  • autoeval/: core package
  • tests/: pytest checks for determinism, schema, holdout protection, and fitness sanity
  • reports/: report cards for accepted agent changes
  • leaderboard.md: last accepted suite improvements

Contributors

Deva-1903

Issues