YakinRubaiat/AgentCoOpBench

Reproducibility artifact for profiling cooperative and strategic behavior in LLM agents through repeated games.

★ 0Forks 0Jupyter NotebookGitHub ↗Compare
benchmarkgame-theoryiterated-prisoners-dilemmallm-agentsmulti-agent-systemsreproducibility

README

AgentCoOpBench

A benchmark for profiling cooperative and strategic behavior in LLM agents through repeated games

AgentCoOpBench evaluates language models as persistent decision-making agents rather than as one-shot answer generators. Each model plays repeated Iterated Prisoner's Dilemma (IPD) sessions against a fixed suite of algorithmic opponents. The released trajectories support behavioral profiles spanning cooperation, niceness, retaliation, forgiveness, exploitability, consistency, lag-one agreement, and payoff.

This repository is the reproducibility artifact for the paper:

AgentCoOpBench: A Benchmark for Profiling Cooperative and Strategic Behavior in LLM Agents through Repeated Games

Forgiveness–retaliation profiles for the six evaluated models

What is included

The artifact contains the complete evaluation record used in the paper:

  • six model-specific experiment notebooks;
  • 24,000 round-level decisions and payoffs;
  • per-model behavioral profiles and per-opponent summaries;
  • the exploratory analysis notebook;
  • the definitive scripts used to recompute transition-conditioned metrics and generate the publication figures; and
  • PNG, PDF, and SVG figure exports.

No model weights or API credentials are included.

Evaluation protocol

Component Setting
Models 6 open-weight LLMs
Opponents 8 deterministic or seeded algorithmic strategies
Sessions 5 sessions per model–opponent pair
Rounds 100 rounds per session
Total 6 × 8 × 5 × 100 = 24,000 rounds
Temperature 0.0
Actions Cooperate or Defect
Stage payoffs (T=5,\ R=3,\ P=1,\ S=0)

The opponent suite comprises All Cooperate, All Defect, Copycat, Copykitten, Detective, Grudger, Random, and Simpleton. The game engine, opponent policies, action parser, payoff calculation, and metric implementation are contained in each experiment notebook.

Repository structure

AgentCoOpBench/
├── experiments/
│   ├── run_gemma_4_31b.ipynb
│   ├── run_qwen_3_6_35b.ipynb
│   ├── run_gpt_oss_120b.ipynb
│   ├── run_phi_4.ipynb
│   ├── run_mistral_small_22b.ipynb
│   ├── run_command_r_35b.ipynb
│   ├── analyze_results.ipynb
│   └── results/
├── figures/
├── scripts/
│   ├── generate_evaluation_figures.py
│   ├── generate_forgiveness_retaliation.py
│   └── verify_artifact.py
├── .env.example
└── requirements.txt

Quick verification

The included verifier uses only the Python standard library. It checks the number of models, opponents, sessions, and rounds; validates every action and stage payoff; and confirms the expected result files.

python3 scripts/verify_artifact.py

Expected final line:

Artifact verified: 6 models, 48 model-opponent pairs, 240 sessions, 24,000 rounds.

Reproduce the paper figures

Create an environment and install the analysis dependencies:

python3 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txt

Then regenerate the definitive paper figures from the released CSV files:

python scripts/generate_evaluation_figures.py
python scripts/generate_forgiveness_retaliation.py

The scripts write vector PDF/SVG and high-resolution PNG files to figures/. They do not call an LLM API.

For the broader exploratory analysis:

cd experiments
jupyter lab analyze_results.ipynb

Rerun the LLM experiments

The released runs used the University of Idaho MindRouter service through its OpenAI-compatible endpoint. Full reruns therefore require access to that service and to the evaluated model identifiers.

cp .env.example .env
# Edit .env and set OPENAI_API_KEY.
cd experiments
jupyter lab

Open the notebook for the desired model and run all cells. Each notebook fixes the temperature at zero, executes the same eight opponents and session protocol, and writes three timestamped CSV files under experiments/results/.

To use another OpenAI-compatible provider, update BASE_URL, API_HOST, and the model identifier in the relevant notebook. Backend or model-version differences may change the trajectories even when the protocol is held fixed.

Result files

For each model, the artifact includes:

  • *_history_*.csv: all 4,000 round-level actions, payoffs, cumulative scores, raw model responses, and API-use flags;
  • *_opponent_summary_*.csv: payoff and cooperation aggregates for each of the eight opponents; and
  • *_profile_*.csv: the original model-level aggregate profile emitted by the run notebook.

The publication scripts recompute transition-conditioned forgiveness, retaliation, and realized exploitability directly from *_history_*.csv. In particular, exploitability is the fraction of rounds against All Defect and Detective on which the LLM cooperates while the opponent defects. This trajectory-level recomputation is the source of the values reported in the paper.

Reproducibility notes

  • Random is seeded independently for each session by the experiment code.
  • The game judge and every algorithmic opponent are deterministic given the session seed and prior trajectory.
  • LLM decoding uses temperature 0.0; exact replay can still depend on the serving backend and model revision.
  • Raw responses are retained so action parsing can be audited.
  • API keys are loaded from .env and are never printed or committed.

Citation

During anonymous review, please cite the accompanying manuscript as:

@inproceedings{anonymous2026agentcoopbench,
  title     = {AgentCoOpBench: A Benchmark for Profiling Cooperative and
               Strategic Behavior in LLM Agents through Repeated Games},
  author    = {Anonymous},
  year      = {2026}
}

License

No license is asserted in this review artifact. A reuse license and final author metadata can be added with the camera-ready release.

Contributors

YakinRubaiat

Issues