A headless Plants vs Zombies: Fusion simulator, and a reinforcement learning agent that learns to play it. Both written from scratch in C++ with no machine learning framework of any kind.
No PyTorch. No TensorFlow. No JAX. No ONNX. No cuBLAS. No cuDNN. Not wrapped, not vendored,
not optional. The tensors, the matrix multiply, the backward pass, the optimiser, the
reinforcement learning algorithm and every CUDA kernel are in this repository. The entire
build is one g++ invocation over three translation units.
- Results
- Build and run
- Architecture at a glance
- Part 1: the simulator
- Part 2: the reinforcement learning wrapper
- Part 3: the neural network
- Part 4: backpropagation and the optimiser
- Part 5: PPO
- Part 6: imagination, a search tree inside the forward pass
- Part 7: the free neuron block
- Part 8: the GPU backend
- Part 9: performance
- Part 10: correctness and testing
- Case study: the reward was exploitable
- What is and is not implemented
- File layout
- Status and roadmap
Real Adventure Day level 1, 500 evaluation games per figure.
| Agent | Win rate | Return | Wave | Kills | Fusions | Steps |
|---|---|---|---|---|---|---|
| Hand-written scripted bot | 27.5% | |||||
| Neural network alone | 70.8% | 7.13 | 8.9 / 10 | 35.2 | 23.2 | 54 |
| Network + imagination search tree | 95.4% | 12.90 | 9.9 / 10 | 43.0 | 30.5 | 60 |
Level 1 hands you exactly two seed packets: Peashooter and Sunflower. Winning requires climbing the fusion ladder, and every rung costs sun you had to grow in advance:
Peashooter + Peashooter = Repeater
Repeater + Peashooter = Split Pea
Split Pea + Peashooter = Gatling Pea
That is a sequential planning problem, which is exactly what search buys and what a purely reactive policy cannot do. On an earlier, easier, synthesised version of level 1 that handed the agent ten seed packets, the search was worth only +9.7 points. On the real level it is worth +24.6. Making the game faithful is what let the search show what it was actually worth.
Honest caveats. The +24.6 figure is a single seed. Seed variance in this setup is about 3.2 points of standard deviation, so the direction is certain but the exact number can move. The earlier +9.7 figure is a properly measured mean over 4 seeds per arm (74.3 sd 3.2 versus 84.0 sd 3.5, Welch t = 4.07, p ~ 0.007, with complete separation between arms).
| Question | Result |
|---|---|
| Does the compressed action head cost accuracy? | No. 73.6% flat vs 73.0% embedded, z = 0.21, p = 0.83, for a head 4.1x smaller |
| Does the free neuron block help? | No. Ablating it moved seeds by -2.4, -0.8, +2.4 points. Averages to zero |
| Does an async actor-learner help? | No. 1.2x faster, win rate collapsed 81.4% to 22.0% |
| Is pure potential-based reward shaping better than bounded-dense? | No. 58.6% vs 73.8% |
Requires a C++23 compiler. CUDA is optional; without it everything runs on the CPU.
.\build.ps1pvzf train train a policy
pvzf eval evaluate a checkpoint over N games
pvzf bot run the hand-written scripted baseline
pvzf gametest 40 simulation correctness tests
pvzf selftest 15 CPU tests, including finite-difference gradient checks
pvzf gputest 13 tests, every kernel against its CPU reference
pvzf simbench simulation throughput measurement
pvzf gpuinfo device capabilities
| Flag | Default | Meaning |
|---|---|---|
--envs N |
64 | parallel games per rollout. 256 is the sweet spot |
--groups N |
2 | double-buffered rollout. 1 disables the overlap |
--hidden A,B |
512,512 | trunk widths |
--head-dim K |
70 | action embedding size. 0 selects the flat head |
--imagine N |
0 | shortlist size for the search tree. 0 disables it |
--imag-branch R |
1 | children per node. 2 makes it a tree rather than a line |
--capped-reward |
off | bounds every dense reward term. Recommended |
--free N |
0 | free neuron block size. Off by default, see part 7 |
--lr F |
3e-4 | learning rate |
--gamma F |
0.997 | discount |
--ent F |
0.01 | entropy bonus |
--epochs N |
4 | passes over each rollout |
pvzf train --envs 256 --head-dim 70 --imagine 16 --imag-branch 2 --capped-reward
pvzf eval --games 500 --load bin/day_tree.bin
bin/day_base.bin and bin/day_tree.bin are the two trained agents from the results table.
flowchart TB
subgraph SIM["Simulator (src/game)"]
B["Board: 91 plants, 39 zombies<br/>53 fusion recipes<br/>100 ticks per second"]
end
subgraph RL["RL wrapper (src/rl)"]
E["Encoder: 1234 floats"]
M["Legality mask: 649 bits"]
R["Reward shaping"]
end
subgraph NET["Network (src/ai)"]
T["Trunk 512, 512"]
I["Imagination search tree"]
P["Policy head"]
V["Value head"]
end
subgraph TRAIN["PPO (src/rl/ppo.h)"]
G["GAE advantages"]
O["AdamW"]
end
B --> E --> T
B --> M --> P
T --> I --> P
T --> V
P --> A["action"] --> B
B --> R --> G --> O --> T
Every box above is hand-written code in this repository.
Deterministic, fixed timestep, 100 ticks per second, matching the original game's logic rate. Headless, it runs orders of magnitude faster than real time.
Board::step(action) is a pure function of (state, action, rng). There is no renderer, no
audio, no frame pacing, no wall clock. That is the entire reason training is fast.
flowchart LR
A["apply action"] --> S["tickSun"] --> W["tickWaves"] --> P["tickPlants"]
P --> J["tickProjectiles"] --> Z["tickZombies"] --> BL["tickBlasts"] --> C["tickCleanup"]
The order is fixed and matters. Plants fire before projectiles move, so a pea spawned this tick does not travel until the next one. Blasts resolve after zombies move, so an explosion catches a zombie at the position it actually reached.
A plant is a row in a table. So is a zombie. So is a fusion recipe.
{ .name = "Peashooter", .cost = 100, .recharge = RC_FAST, .health = 300,
.flags = PF_LAND_ONLY, .attack = Attack::Shoot,
.shootInterval = SEC(1.5f), .proj = ProjKind::Pea, .damage = 20 },{ .name = "Pole Vaulting Zombie", .bodyHealth = 720, .secPerGrid = 2.5f,
.ragedSecPerGrid = 4.7f, .flags = ZF_JUMPS,
.pointCost = 2, .firstLevel = 4, .pickWeight = 2000 },{ P_PEASHOOTER, P_PEASHOOTER, P_REPEATER, "confirmed in-mod" },
{ P_WALLNUT, P_WALLNUT, P_GIANT_WALLNUT, "rolls the lane for 600" },kPlants[] and kZombies[] are positional arrays: the row carries no id, its index in the
table is its id. Enum order and table order must match exactly, which is checked at startup.
Fusion needs no special action, no menu, no mode and no fee. Planting seed packet B onto a tile that already holds plant A is the fusion, exactly as the mod does it. The cost is simply both packets: a Peashooter at 100 plus a Wall-nut at 50 gives you a Peanut, a 4000-health plant that shoots, for 150 sun total.
That single design decision is what makes the two-seed level 1 survivable, and it is why the fusion tech tree is reachable through ordinary play rather than needing a bespoke action type.
Adventure levels 1 to 9, verified line by line against the wiki:
- The real seed unlock order. Level 1 really is Peashooter and Sunflower and nothing else
- The real zombie introduction order: Flag 1, Conehead 2, Buckethead 3, Pole Vaulting 4, Peashooter Zombie 5, the box zombies 6, Cherryshooter Newspaper 7, Screen Door and Buckshooter and Buck-nut 8, Cherry-nut 9
- Wave timing: first wave at 15 seconds, 30 second intervals
- Zombotany, zombies that shoot your plants and can be ducked under by short plants
- The nested box chain: Diamond-box unpacks into Gold-box unpacks into Silver-box unpacks into a Conehead, on death, 14,570 points of nesting
- Pole vaulting blocked by tall plants only, one vault each, fast before and normal after
- Screen doors that eat straight shots but not lobbed, fume or spikeweed damage
- Ladders that bridge a wall and stay bridged for every zombie behind
- Bungees that fall on a tile, steal the plant and leave
- Chompers that swallow whole, except gargantuars, which they merely bite for 40
- Upgrade plants that refuse bare dirt. A Tall-nut grows out of a Wall-nut, at 8000 health
The game runs at 100 ticks per second. The agent decides every 50 ticks, half a second of game time.
Early on the agent almost never has enough sun to plant anything, so nearly every decision is a forced no-op. Asking a network to choose between 649 actions when 648 are illegal teaches it nothing and wastes the entire rollout.
Now the environment only wakes the network when at least one placement is actually legal.
| Metric | Before | After |
|---|---|---|
| Legal actions per decision | 2.5 | 60 |
| Policy entropy | 0.14 | 3.98 |
| Explained variance | 0.99 (critic predicting a constant) | real |
| Return curve | flat | learning |
| Block | Offset | Size | Contents |
|---|---|---|---|
| Tiles | 0 | 864 | 54 tiles x 16 features |
| Zombie bins | 864 | 300 | 6 lanes x 10 distance bins x 5 features |
| Rows | 1164 | 18 | 6 rows x 3: lane active, is water, mower alive |
| Seed slots | 1182 | 40 | 10 slots x 4: plant id, affordable, recharge left, cost |
| Global | 1222 | 12 | sun, wave, tick, plants alive, zombies alive, and so on |
The 16 features per tile are: occupancy, health fraction, cooldown fraction, age, a nine-way
one-hot over the Attack enum, armor, underlay, and crater.
Zombies are binned, not listed. A variable-length set of entities cannot be fed to a dense network. Instead each lane is chopped into 10 distance buckets, 9 on-screen columns plus one off-screen, and each bucket carries count, total health, nearest distance, armor and a flying flag. That turns the threat picture into a fixed-size image with spatial structure the network can learn over.
0 no-op
1..540 plant seed slot S on tile T (10 slots x 54 tiles)
541..594 shovel tile T (54 tiles)
595..648 fire the cannon at tile T (54 tiles)
A 649-bit legality mask accompanies every observation. Illegal logits are set to negative infinity before the softmax, so an illegal action has probability exactly zero rather than merely small.
perKillPoint = 0.020 // scaled by the zombie's wave-budget cost
perFusion = 0.100 // the mechanic we actually want it to use
perSun = 0.0015 // only sun the agent GREW
perPlantLost = 0.010 // per 50 sun of plant destroyed
perMowerLost = 1.000
win = +5.0
lose = -5.0
illegal = -0.010Two design notes that took real experiments to learn:
Only sun the agent grew counts. Sun falling from the sky arrives on a fixed timer no matter what the agent does. Paying for it adds a large, perfectly predictable term to every return, which the critic learns exactly and which therefore contributes nothing but noise to the advantage.
Every dense term must be bounded in time. See the case study
below. --capped-reward sets sunRewardCap = 4000 and fusionRewardCap = 25.
flowchart TD
O["Observation<br/>1234"] --> L1["Linear 1234 to 512"]
L1 --> R1["ReLU"]
R1 --> L2["Linear 512 to 512"]
L2 --> R2["ReLU"]
R2 --> PH["Policy head<br/>Linear 512 to 70"]
R2 --> VH["Value head<br/>Linear 512 to 1"]
PH --> Q["query vector q, 70-dim"]
Q --> DOT["logit[a] = q · emb[a]"]
EMB["actEmb: 649 x 70<br/>learned action embeddings"] --> DOT
DOT --> MSK["apply legality mask"]
MSK --> SM["softmax, sample"]
VH --> VAL["V(s), scalar"]
| Layer | Shape | Parameters |
|---|---|---|
trunk[0] |
1234 to 512 | 632,320 |
trunk[1] |
512 to 512 | 262,656 |
policy |
512 to 70 | 35,910 |
actEmb |
649 x 70 | 45,430 |
value |
512 to 1 | 513 |
| Total | 976,829 |
Weights are initialised orthogonally, with gain sqrt(2) on ReLU layers, gain 0.1 on the
policy head so the initial policy is close to uniform, and gain 1.0 on the value head.
Orthogonal initialisation keeps the singular values of each weight matrix at 1, so gradients
neither explode nor vanish through depth at the start of training.
The obvious design gives the last layer 649 outputs: 512 x 649 = 332,288 weights, over a
quarter of the entire network spent on the output alone. Worse, every action becomes a private
column that learns nothing from any other action, so "plant a Sunflower on row 3 column 2"
and "plant a Sunflower on row 3 column 3" are as unrelated as "no-op" and "fire the cannon".
Instead the trunk emits a 70-dimensional query, and every action owns a learned
70-dimensional embedding. The score of action a is their dot product:
logit[a] = q · actEmb[a]
This is exactly how a language model's output layer works: a hidden state dotted against a per-token embedding matrix. It makes the head 4.1x smaller and, more importantly, it makes similar actions neighbours in embedding space, so learning about one generalises to the others.
We tested whether that compression costs anything. Over 500 games: flat head 73.6% win,
embedded head 73.0%. z = 0.21, p = 0.83. Statistically identical.
There is no autograd. Every operation has an explicit hand-written backward pass.
The backward pass through ReLU needs to know which units were active in the forward pass. Rather than storing a separate boolean mask, the code reads the sign of the stored activation, since a post-ReLU activation is positive exactly when the unit fired.
That is cheap and correct, and it caused the single nastiest bug in this project. The free
neuron block originally merged its output into act[trunk.size()] in place, which
corrupted the very signs the trunk's backward pass depended on. Gradients silently went wrong
in a way no assertion caught. The fix is a separate topMix buffer, so act[trunk.size()]
stays the pure trunk activation. The same fix exists in both the CPU and GPU paths.
Decoupled weight decay, so the decay term is applied directly to the weights rather than folded into the gradient and then rescaled by the adaptive term:
m = b1 * m + (1 - b1) * g
v = b2 * v + (1 - b2) * g^2
w -= lr * ( mhat / (sqrt(vhat) + eps) + wd * w )
Plus global gradient norm clipping: the L2 norm is taken across every parameter tensor at
once, and if it exceeds maxGradNorm all gradients are scaled by the same factor. Clipping
each tensor separately would change the direction of the update, not just its length.
The reported gradient norm is measured before clipping, because a post-clip norm just reads back the clip threshold and tells you nothing.
flowchart LR
ROLL["rollout<br/>256 games in parallel<br/>T steps each"] --> GAE["GAE advantages<br/>gamma 0.997, lambda 0.95"]
GAE --> NORM["normalise advantages"]
NORM --> EP["4 epochs x 4 minibatches<br/>= 16 gradient steps"]
EP --> LOSS["clipped surrogate<br/>+ clipped value loss<br/>+ 0.01 entropy"]
LOSS --> ADAM["AdamW, lr 3e-4<br/>global grad-norm clip"]
ADAM --> ROLL
delta_t = r_t + gamma * V(s_{t+1}) * (1 - done) - V(s_t)
A_t = delta_t + gamma * lambda * (1 - done) * A_{t+1}
gamma = 0.997 is deliberately high. At half a second per decision a game runs several hundred
decisions, and planting a Sunflower only pays off dozens of steps later. The bootstrap term is
cleared on both true termination and truncation.
ratio = exp(logp_new - logp_old)
L = -min( ratio * A, clip(ratio, 1 - 0.2, 1 + 0.2) * A )
Because the update runs 16 gradient steps on data collected by an older policy, the ratio
drifts away from 1. The clip is what makes that safe. approxKL and clipFrac are logged
every update so the drift is visible rather than assumed.
The value prediction is clipped to a trust region around the old prediction, and the loss is the max of the clipped and unclipped squared errors, which prevents a single large critic update from destabilising the advantages the actor is using.
This is the module that produced the 95.4%, and it is the most unusual part of the project.
A plain policy is reactive: it sees a board and reaches for a move. But the fusion ladder needs a plan. So the network imagines before it commits: it forks off the trunk halfway down, rolls a learned world model forward over its own imagined futures, and merges what it finds back into the policy logits.
flowchart TD
H["trunk features h<br/>forked after layer 1"] --> SL["shortlist:<br/>top 16 actions from the<br/>un-imagined logits"]
SL --> ENC["encode into a<br/>128-wide latent"]
ENC --> D1["depth 1: 16 nodes<br/>dyn: latent to latent<br/>rHead: predicted reward<br/>vHead: predicted value"]
D1 --> PR["proposer:<br/>latent to 2 continuations"]
PR --> D2["depth 2: 32 nodes"]
D2 --> LF["leaf: Q = r + gamma * v"]
LF --> BK["max backup<br/>Q(node) = r + gamma * max_k Q(child)"]
BK --> SC["score per shortlisted action"]
SC --> GT["gate(h), initialised to 0"]
GT --> MG["logits[a] += gate(h) * score[a]"]
- A first forward pass produces un-imagined logits. The host takes the top 16 as a shortlist. Searching all 649 actions would be pointless; most are illegal or absurd.
- Each shortlisted action is embedded, reusing the same 70-dimensional
actEmbthe policy head uses, so the search and the policy speak the same language about actions. - A learned dynamics function
dynmaps a 128-wide latent to the next latent. Two heads read off a predicted immediate reward and a predicted value. - At each depth a
proposermaps the current latent tobranchcontinuations, so the imagined actions are state-dependent rather than a fixed learned set. - Values are backed up to the root, and the score of each shortlisted action is added to its logit through a gate.
Node count is shortlist * branch^(depth-1). Children of parent i land at index i * R + k,
which keeps every array batch-major with zero index arithmetic.
The gate starts at exactly zero. On step one the imagination contributes literally nothing. It can only gain influence if the gradient says it helps. This is why bolting it on cannot make things worse by construction.
Max backup, not average. Q(node) = r + gamma * max_k Q(child). This differentiates like
max-pooling: the winning child receives the entire gradient and the losers receive none. The
argmax index is stored per node during the forward pass and replayed in the backward pass.
Both the CPU and the GPU use a strict > on ties, so the two backends can never pick
different winners and diverge.
dQ is dRew. The reward at a node enters its own Q with coefficient exactly 1, so the reward gradient and the Q gradient are the same array. The world model's separate reward supervision must be added after the descent, or it wrongly propagates deeper into the tree than it should.
It is gradient-checked. The whole module, argmax backup included, passes finite-difference
checks at 2.2e-04. That test has a subtlety worth knowing: a max is piecewise linear, so if the
winning branch flips between the +h and -h probes, the central difference measures the
average of two different slopes and the check fails spuriously. The test detects the flip and
resamples that coordinate.
--imag-branch 1 gives a single adaptive line rather than a tree. It is cheaper, and it is
worse in a specific and interesting way, documented in the case study below.
An experiment in architecture without layers, kept in the repository along with the measurement that says it did not help.
flowchart LR
IN["input projection<br/>inDim to 96"] --> C012["clusters 0,1,2<br/>96 neurons, injected"]
C012 -.->|sparse| C34["clusters 3,4<br/>reachable only<br/>through the graph"]
C34 -.->|sparse| C567["clusters 5,6,7<br/>96 neurons, read out"]
C012 -.->|sparse| C567
C567 --> OUT["output projection<br/>96 to outDim"]
OUT --> GATE["gate, starts at 0"]
GATE --> TOP["added to trunk top"]
A pool of 256 neurons wired to itself by a fixed clustered random graph: 8 clusters of 32, 50% connection probability inside a cluster, 5% between, giving 10.3% density and 6,742 edges. There are no layers. The block relaxes over 4 micro-steps:
z[0] = relu(inj)
z[t] = relu(Wm · z[t-1] + b + inj) // the input is re-injected every step
Input enters clusters 0 to 2, output is read from clusters 5 to 7, and clusters 3 and 4 are reachable only through the graph. The projections are 96-wide on each side, so that disjointness is structural rather than enforced by masking.
Wm = W * mask is rebuilt every forward pass, so absent edges have exactly zero derivative,
not merely a small one. That is asserted by a GPU test which reads 0.000e+00.
The recurrent matrix is rescaled to a spectral radius of 0.95, measured by power iteration plus a Gelfand-style geometric mean of the tail. This has to be measured, not assumed, because masking an orthogonal matrix destroys its orthogonality.
Its gate rose from 0 to about 0.60 to 0.72 on every seed. It clearly wanted a voice.
Then we ablated it. eval --no-free loads the trained weights, sets useFree = false, and
replays the identical 500 games. The win rate moved by -2.4, -0.8 and +2.4 points across
three seeds. That averages to nothing.
A large gate does not mean a useful block. Whatever it computes is close to redundant with what the trunk already produces. The ablation is the only measurement that settles this, which is exactly why it was built before the experiment was run. The block is off by default.
CUDA is reached through the driver API and NVRTC directly. Kernels are C++ source
strings inside gpu_kernels.h, compiled at runtime to a sm_86 cubin. There is no nvcc step
in the build and no math library linked.
| Group | Kernels |
|---|---|
| Linear algebra | sgemm_nn, sgemm_nt, sgemm_tn, sgemm_tn_split |
| Activations | relu_fwd, relu_bwd |
| Policy | log_softmax_masked, sample_actions, gather_logp |
| PPO | ppo_policy_grad, ppo_value_grad |
| Optimiser | adamw, sq_norm, col_sum, add_inplace, fill |
| Imagination | imag_child, imag_child_bwd, imag_branch, imag_branch_bwd, imag_backup_leaf, imag_backup_step, imag_backup_bwd, imag_leaf_bwd, imag_merge, imag_merge_bwd |
| Free block | free_mask, free_slice, free_merge, free_merge_bwd, free_gate_bwd |
sgemm_tn tiles only its output. For a weight gradient the output is the weight matrix, so
a [1 x 128] reward head got 2 thread blocks and a [128 x 128] dynamics matrix got 4, on a
56-SM device, to reduce roughly a million rows. M, the only large dimension, received no
parallelism at all.
sgemm_tn_split adds a gridDim.z over slices of M. Each block reduces its own slice
privately, then atomicAdds the partial result. It is only used when beta == 1, because an
overwrite needs a single writer.
| Configuration | Before | After | Speedup |
|---|---|---|---|
--imag-branch 1 update |
2.77 s | 0.78 s | 3.6x |
--imag-branch 2 update |
5.75 s | 1.24 s | 4.6x |
| no imagination | 0.39 s | 0.38 s | unchanged |
Imagination overhead versus a plain network fell from 6.3x to 2.2x.
cuMemcpyDtoDreturns to the host before the copy lands, and a non-blocking stream never waits on the legacy stream. Missing that sync means reading torn weights.CU_STREAM_NON_BLOCKINGis mandatory, or a newly created stream still synchronises with stream 0 and the whole point of the second stream is lost.BNis a#definefor the SGEMM tile width, so no kernel parameter may be namedBN.- The optimiser originally did one blocking download per parameter tensor per step, about 368 round trips per update to retrieve a single scalar. Each tensor now reduces into its own slice of one buffer, downloaded once.
| Metric | Value |
|---|---|
| Simulation throughput | orders of magnitude faster than real time |
| Agent steps/sec, 256 envs | 37,200 |
| Agent steps/sec, 768 envs, groups 2 | 40,500 |
Entity counts are tiny: 5 live zombies out of 6.5 slots, 8.3 plants, 1.5 projectiles. There is no O(n^2) problem here; lane-bucketing would cost more than it saves.
The batch is split in half. A persistent worker thread simulates group A on the CPU while the main thread runs group B's forward pass on the GPU.
flowchart LR
subgraph T0["time 0"]
S0["CPU: simulate group A"]
N0["GPU: forward group B"]
end
subgraph T1["time 1"]
S1["CPU: simulate group B"]
N1["GPU: forward group A"]
end
T0 --> T1
The alternation order is chosen so the RNG draw sequence is identical to serial, which makes
the output bit-for-bit the same as running with --groups 1. That is verified, not assumed.
A full asynchronous actor-learner was also built: a separate actor thread with its own CUDA
context binding, its own non-blocking stream, its own frozen weight copy and its own device
scratch, playing rollout u+1 while the learner updates on rollout u.
It works, and it is 1.2x faster. It also learns far worse:
| Epochs | Sync win rate | Async win rate |
|---|---|---|
| 4 | 81.4% | 22.0% |
| 1 | 67.0% | 37.2% |
More epochs makes sync better and async worse, which is the signature of off-policy drift: PPO already takes 16 gradient steps per update, and async data ends up roughly twice as stale. Making it viable would need V-trace or proper importance-sampling correction on the advantages. Speed that costs correctness is not speed. It is off by default.
pvzf selftest 15 passed CPU
pvzf gputest 13 passed every kernel against its CPU reference
pvzf gametest 40 passed simulation rules and the RL wrapper
Gradients. The full network, including the imagination tree and the free block together, is finite-difference checked. That combined test is the one that catches a rejoin which overwrites instead of accumulating.
GPU equals CPU. Forward at 6.9e-07, backward at 1.5e-06, and absent graph edges have exactly 0.000e+00 gradient.
Determinism. Two runs from the same seed agree on a board fingerprint at every tick.
Game rules, a sample:
- A Tall-nut refuses bare dirt, is legal on a Wall-nut, and grows at 8000 health
- A pole vaulter clears a Wall-nut and reaches
x = -21, the lawnmower, but is stopped dead atx = 415by a Tall-nut. Same zombie, same lane, same seed, only the blocker differs - A Silver-box unpacks into a Conehead when killed, but into nothing when devoured by a Chomper
- Potential-based reward shaping telescopes to zero over a whole episode, measured at 2.8e-07, while a companion assertion proves the shaping was non-trivial so the test cannot pass vacuously
Worth reading, because the search found the hole and the plain policy never did.
perSun was the only reward term unbounded in time. Kills and fusions are limited by what the
level contains, but a Sunflower pays out for as long as the game runs.
The --imag-branch 1 search discovered this and learned to kill slowly, letting the wave
timer idle while its economy kept earning:
| Baseline | Branch 1, uncapped | |
|---|---|---|
| Win rate | 81.3% | 59.3% |
| Return | 8.60 | 11.06 |
| Wave reached | 9.8 | 9.8 |
| Kills | 40.8 | 40.8 |
| Steps | 67 | 123 |
| Sun collected | 3,419 | 5,902 |
Same wave, same kills, twice the length. Its win rate peaked at 84.4% then decayed to 60.6% while its return kept climbing. That divergence between the objective and the goal is the signature to watch for.
Two fixes were built and both are in the repository:
--capped-reward bounds every dense term (sunRewardCap = 4000, fusionRewardCap = 25).
It keeps the curriculum the terms were added to teach and stops paying once the lesson is over.
Branch 1 went from 123 steps and 59.3% to 66 steps and 77.8%. This is the one to use.
--potential-reward replaces the dense terms with potential shaping,
gamma * Phi(s') - Phi(s). This is provably policy-invariant and therefore exploit-proof. It
also learns considerably worse, 58.6% versus 73.8%, and that is expected rather than a bug: the
telescoping property that makes it safe also makes its net incentive exactly zero, so it deletes
the curriculum along with the exploit.
The tree at --imag-branch 2 never drifted in the first place, and its extra steps are not
stalling: its wave reached and its kills both went up.
All seven Day plants at wiki-exact stats. All ten Day zombies with their real behaviours. 53 fusion recipes covering the reachable Day tech tree, including both the pea ladder and the parallel cherry ladder. Status effects: chill, freeze, butter, burn, hypnosis. Armour types and what strips them. Lawnmowers. Craters. Lily pads and flower pots as underlays. Pumpkins as armour. All nine attack kinds.
Known gaps, listed rather than hidden
- The mod has 511 fusions across six tiers; this implements 53
- Damage types are not modelled, so the Cherry-nut Zombie's 150-damage cap against Cherry Bombs and the Cremator type are absent
- The Giant Wall-nut's roll is modelled as an instant whole-lane blast rather than a travelling object
- The Gardening Glove, Fertilizer, Surprise-box and Bucket items are absent, which also means the Odyssey fusion tier is unreachable
- Zombie movement speeds are our own calibration. The wiki's zombie infobox has no speed field at all, so these cannot be verified against it
- Only the Day map is implemented. Night, Pool, Fog and Roof are stubs
build.ps1 the entire build, one g++ invocation
src/game/
defs.h every enum, flag, struct and tunable constant
data.cpp the tables: 91 plants, 39 zombies, 53 recipes
board.h board.cpp the simulation, ~1400 lines
gametest.h 40 correctness tests and the scripted baseline
src/ai/
tensor.h Mat, the flat float matrix everything is built on
nn.h Linear, Net, Imagination, FreeBlock, AdamW, backward passes
gpu.h device mirrors of the above
gpu_kernels.h every CUDA kernel, as NVRTC source
cuda_api.h driver API bindings, loaded dynamically
selftest.h 15 CPU tests including gradient checks
gpu_selftest.h 13 GPU-versus-CPU tests
src/rl/
env.h observation encoding, action masking, reward shaping
ppo.h rollout, GAE, the PPO update, checkpointing
src/core/
rng.h xoshiro256++
worker.h persistent worker thread for the double-buffered rollout
src/main.cpp CLI
Checkpoints are versioned; v5 stores the free block's graph mask verbatim, because the RNG state that generated it is not recoverable.
Done. The simulator, the learning engine, and a trained agent that wins 95.4% of real Day level 1 games.
Next. The rest of the fusion catalogue: 511 recipes across Common, Upgraded, Advanced, Titan, Odyssey and Infusible tiers, and the mechanics that go with them. The rolling nuts, the Cremator damage type, the exploding nuts, the Gardening Glove and the two-tile Titan fusions.
After that. The harder Day levels. The scripted baseline currently scores 27.5% on level 1 and 0% on level 9, so there is a real curriculum waiting. Then the remaining four maps.