This benchmark evaluates large language models (LLMs) using the sixty "parlor puzzles" from the game Blue Prince. These puzzles require using logical reasoning to deduce which of three boxes contains a prize. Each puzzle is repeated three times to reduce the impact of randomness in model responses.
Puzzle texts are excerpted from Blue Prince and remain the property of their original copyright holder.
| Rank | Model | Score | Cost (USD) |
|---|---|---|---|
| 1 | qwen3-next-80b-a3b-thinking |
97% (175/180) | $0.37 |
| 2 | deepseek-r1-0528 |
97% (174/180) | $3.99 |
| 3 | o4-mini |
95% (170/179) | $1.04 |
| 4 | gpt-5-nano |
94% (170/180) | $0.26 |
| 5 | gpt-oss-20b |
94% (169/180) | $0.08 |
| 6 | qwen3-14b |
93% (166/179) | $0.22 |
| 7 | qwen3-30b-a3b |
91% (163/179) | $0.27 |
| 8 | qwen3-32b |
91% (163/180) | $0.24 |
| 9 | gpt-5-mini |
89% (161/180) | $0.52 |
| 10 | gpt-4.1-mini |
89% (160/180) | $0.29 |
| 11 | deepseek-v3-0324 |
89% (160/180) | $1.53 |
| 12 | gpt-4.1 |
87% (157/180) | $1.41 |
| 13 | gemini-2.5-flash |
79% (142/180) | $1.44 |
| 14 | gpt-4.1-nano |
66% (119/180) | $0.11 |
Let's play a logic puzzle.
The rules are:
- You are in the parlor.
- The room contains a blue box, a white box, and a black box.
- The blue box is on the left, the white box is in the middle, and the black box is on the right.
- Each box displays a statement on its lid.
- There will always be at least one box which displays a true statement.
- There will always be at least one box which displays a false statement.
- Exactly one box contains gems; the other 2 are empty.
- The room also contains a wind-up key.
The three boxes display these statements:
- blue box: Only one box is true
- white box: Only one box contains the gems
- black box: The gems are in the white box
Which box contains the gems?
Think step-by-step before writing your answer in exactly this format: <solution>white</solution>.
Python and uv are required.
Define the environment variables required by model providers:
export OPENROUTER_API_KEY=...
export AZURE_API_BASE=...
export AZURE_API_KEY=...
Benchmark a model using its name defined in models.yaml:
uv run bench.py gpt-4.1-nano
Render updated README.md plot and leaderboard:
uv run render.py