AutoEval is a minimal MVP where an agent improves an evaluation suite for retrieval systems.
The target is not model training. The target is better evals: queries and relevance labels that better separate good retrievers from bad retrievers.
Static benchmarks get saturated. When weak systems and strong systems get similar scores, the benchmark loses discriminative power.
AutoEval iteratively proposes new eval items, measures whether they improve separation, and keeps only improvements.
- Deterministic retrieval benchmark with 5 reference systems
- Synthetic corpus (
~300docs) committed to git - Editable eval set (
~120items) and hidden holdout (~80items) - Fitness objective focused on discriminative power, with guardrails
- Rule-based agent that evolves
seed_eval.jsonl - Git integration for accepted suite updates with report cards
The checked-in editable suite is intentionally under-discriminative: it starts with many valid but easy items where strong and weak lexical systems tie, so the agent has real room to harden the benchmark.
pip install -r requirements.txt
python run_eval.py --suite editable
python agent.py --minutes 10For each system:
suite_score(system) = mean nDCG@10
Discriminative power uses pairwise AUC between good and bad systems:
dp_editable = AUC(editable)dp_holdout = AUC(holdout)DP = 0.2 * dp_editable + 0.8 * dp_holdout
Penalties:
- cost proxy (deterministic runtime estimate)
- redundancy (query similarity)
- flakiness (validation failures)
Final:
fitness = DP - 0.05*cost_seconds - 0.20*redundancy - 0.50*flakiness
- Holdout is committed and treated as read-only by the agent
- Schema checks reject invalid, trivial, or leaky queries
- Near duplicates are rejected via Jaccard threshold (
>0.85) - Fitness is holdout-dominated (
80%)
- Implement a retriever in
autoeval/systems/with:nameretrieve(query, k) -> list[doc_id]
- Register it in
autoeval/systems/__init__.py - Add it to
GOOD_SYSTEMS,MID_SYSTEMS, orBAD_SYSTEMSinautoeval/config.py
- Add a new template function in
autoeval/evals/generators.py - Return a valid eval item dict (
query,relevant_doc_ids,tags) - Register the template in
generate_candidates() - Keep it deterministic through the provided RNG
python run_eval.py --suite editable
Per-system mean nDCG@10 (editable):
bm25 1.0000
hybrid 1.0000
keyword 1.0000
random 0.0289
tfidf 1.0000
DP_editable: 0.7500
DP_holdout : 1.0000
DP_total : 0.9500
Penalties : cost=2.0000, redundancy=0.9053, flakiness=0.0000
Final fitness: 0.6689
python agent.py --minutes 1 --no-git
iter=1 accepted delta=0.0500 fitness=0.7189 +3 -3 sha=nogit
agent finished: iterations=73 accepted=1 best_fitness=0.7189
autoeval/: core packagetests/: pytest checks for determinism, schema, holdout protection, and fitness sanityreports/: report cards for accepted agent changesleaderboard.md: last accepted suite improvements