A benchmark for profiling cooperative and strategic behavior in LLM agents through repeated games
AgentCoOpBench evaluates language models as persistent decision-making agents rather than as one-shot answer generators. Each model plays repeated Iterated Prisoner's Dilemma (IPD) sessions against a fixed suite of algorithmic opponents. The released trajectories support behavioral profiles spanning cooperation, niceness, retaliation, forgiveness, exploitability, consistency, lag-one agreement, and payoff.
This repository is the reproducibility artifact for the paper:
AgentCoOpBench: A Benchmark for Profiling Cooperative and Strategic Behavior in LLM Agents through Repeated Games
The artifact contains the complete evaluation record used in the paper:
- six model-specific experiment notebooks;
- 24,000 round-level decisions and payoffs;
- per-model behavioral profiles and per-opponent summaries;
- the exploratory analysis notebook;
- the definitive scripts used to recompute transition-conditioned metrics and generate the publication figures; and
- PNG, PDF, and SVG figure exports.
No model weights or API credentials are included.
| Component | Setting |
|---|---|
| Models | 6 open-weight LLMs |
| Opponents | 8 deterministic or seeded algorithmic strategies |
| Sessions | 5 sessions per model–opponent pair |
| Rounds | 100 rounds per session |
| Total | 6 × 8 × 5 × 100 = 24,000 rounds |
| Temperature | 0.0 |
| Actions | Cooperate or Defect |
| Stage payoffs | (T=5,\ R=3,\ P=1,\ S=0) |
The opponent suite comprises All Cooperate, All Defect, Copycat, Copykitten, Detective, Grudger, Random, and Simpleton. The game engine, opponent policies, action parser, payoff calculation, and metric implementation are contained in each experiment notebook.
AgentCoOpBench/
├── experiments/
│ ├── run_gemma_4_31b.ipynb
│ ├── run_qwen_3_6_35b.ipynb
│ ├── run_gpt_oss_120b.ipynb
│ ├── run_phi_4.ipynb
│ ├── run_mistral_small_22b.ipynb
│ ├── run_command_r_35b.ipynb
│ ├── analyze_results.ipynb
│ └── results/
├── figures/
├── scripts/
│ ├── generate_evaluation_figures.py
│ ├── generate_forgiveness_retaliation.py
│ └── verify_artifact.py
├── .env.example
└── requirements.txt
The included verifier uses only the Python standard library. It checks the number of models, opponents, sessions, and rounds; validates every action and stage payoff; and confirms the expected result files.
python3 scripts/verify_artifact.pyExpected final line:
Artifact verified: 6 models, 48 model-opponent pairs, 240 sessions, 24,000 rounds.
Create an environment and install the analysis dependencies:
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txtThen regenerate the definitive paper figures from the released CSV files:
python scripts/generate_evaluation_figures.py
python scripts/generate_forgiveness_retaliation.pyThe scripts write vector PDF/SVG and high-resolution PNG files to figures/. They do not call an LLM API.
For the broader exploratory analysis:
cd experiments
jupyter lab analyze_results.ipynbThe released runs used the University of Idaho MindRouter service through its OpenAI-compatible endpoint. Full reruns therefore require access to that service and to the evaluated model identifiers.
cp .env.example .env
# Edit .env and set OPENAI_API_KEY.
cd experiments
jupyter labOpen the notebook for the desired model and run all cells. Each notebook fixes the temperature at zero, executes the same eight opponents and session protocol, and writes three timestamped CSV files under experiments/results/.
To use another OpenAI-compatible provider, update BASE_URL, API_HOST, and the model identifier in the relevant notebook. Backend or model-version differences may change the trajectories even when the protocol is held fixed.
For each model, the artifact includes:
*_history_*.csv: all 4,000 round-level actions, payoffs, cumulative scores, raw model responses, and API-use flags;*_opponent_summary_*.csv: payoff and cooperation aggregates for each of the eight opponents; and*_profile_*.csv: the original model-level aggregate profile emitted by the run notebook.
The publication scripts recompute transition-conditioned forgiveness, retaliation, and realized exploitability directly from *_history_*.csv. In particular, exploitability is the fraction of rounds against All Defect and Detective on which the LLM cooperates while the opponent defects. This trajectory-level recomputation is the source of the values reported in the paper.
Randomis seeded independently for each session by the experiment code.- The game judge and every algorithmic opponent are deterministic given the session seed and prior trajectory.
- LLM decoding uses temperature 0.0; exact replay can still depend on the serving backend and model revision.
- Raw responses are retained so action parsing can be audited.
- API keys are loaded from
.envand are never printed or committed.
During anonymous review, please cite the accompanying manuscript as:
@inproceedings{anonymous2026agentcoopbench,
title = {AgentCoOpBench: A Benchmark for Profiling Cooperative and
Strategic Behavior in LLM Agents through Repeated Games},
author = {Anonymous},
year = {2026}
}No license is asserted in this review artifact. A reuse license and final author metadata can be added with the camera-ready release.
