Which parts of an agent should survive a model upgrade?
An agent can accumulate instructions, tools, and checks as it learns from failures. When its underlying model changes, some of those adaptations may remain useful. Others may become unnecessary or get in the way. Reforge studies how to tell the difference, and whether using that history helps an agent continue improving.
The proposed experiment migrates the same evolved harness to a target model using different strategies: retain it, rebuild it, optimize it normally, apply existing maintenance, or selectively reorganize it using its development history. We will compare immediate performance, subsequent improvement, and total cost on held-out tasks. The evaluator and permission boundaries stay fixed across strategies.
Status: research setup. No empirical model-upgrade results yet. The executable code is a small, fixed-model consolidation pilot. Cross-model migration, the substantive benchmark, and the main study are planned, not implemented. This repository does not claim a novel or superior algorithm before those comparisons.
- Research plan: question, hypotheses, comparisons, measurements, and decision gates.
- Execution prompt: instructions for continuing the project in a fresh agent session.
- Research rationale and closest work: evidence, overlap, and limitations.
- Experiment status: what exists and what still needs to be done.
- Paper outline: the intended report and evidence needed for each section.
- Evaluation boundaries: what the agent may change and what it may not.
Python 3.11 or newer; no third-party runtime dependencies are needed. Run from the repository root:
python -m unittest discover -s tests -v
python scripts/check_repo.py
python -m harness_lab run --provider fixture --output runs/fixture-first --budget-usd 1Choose a new output directory on subsequent fixture runs. Fixtures use deterministic arithmetic answers and invented accounting units. Their scores are not research results or API spending.
The existing pilot implements a solver, a fixed improver, paired candidate evaluation, three inherited treatments, spending reservations, and a barrier that freezes every branch before final test evaluation. See the v0 protocol and pilot instructions. It does not execute model-generated code or expose the host filesystem to the model.
The interesting outcome is not merely a shorter prompt. It is a repeatable change in how well an agent adapts after a model transition, with the cost of diagnosis and consolidation included. A result that ordinary maintenance works equally well, or that removing old instructions harms performance, would change the proposed method and should be reported.
The planned public artifacts are a versioned protocol, code and configurations, permissible data manifests, reproducible figures, a record of failed runs, and a paper that links its claims to evidence. Until measurements exist, there is no results dashboard or completed paper. Publication of benchmark material and raw traces requires a license and privacy review.
Prompt migration, context maintenance, harness evolution, and AI control all have substantial prior work. Reforge's candidate contribution concerns history-informed migration of an evolved harness and its subsequent adaptation, rather than any of those ideas individually. See the source ledger for primary references and precise limits on this claim.
Original repository code and prose are available under the MIT license. Third-party datasets retain their own licenses and are not bundled here.