zigazaga4/pvzf

★ 0Forks 0C++GitHub ↗Compare

README

pvzf

A headless Plants vs Zombies: Fusion simulator, and a reinforcement learning agent that learns to play it. Both written from scratch in C++ with no machine learning framework of any kind.

No PyTorch. No TensorFlow. No JAX. No ONNX. No cuBLAS. No cuDNN. Not wrapped, not vendored, not optional. The tensors, the matrix multiply, the backward pass, the optimiser, the reinforcement learning algorithm and every CUDA kernel are in this repository. The entire build is one g++ invocation over three translation units.


Contents


Results

Real Adventure Day level 1, 500 evaluation games per figure.

Agent Win rate Return Wave Kills Fusions Steps
Hand-written scripted bot 27.5%
Neural network alone 70.8% 7.13 8.9 / 10 35.2 23.2 54
Network + imagination search tree 95.4% 12.90 9.9 / 10 43.0 30.5 60

Level 1 hands you exactly two seed packets: Peashooter and Sunflower. Winning requires climbing the fusion ladder, and every rung costs sun you had to grow in advance:

Peashooter + Peashooter = Repeater
Repeater   + Peashooter = Split Pea
Split Pea  + Peashooter = Gatling Pea

That is a sequential planning problem, which is exactly what search buys and what a purely reactive policy cannot do. On an earlier, easier, synthesised version of level 1 that handed the agent ten seed packets, the search was worth only +9.7 points. On the real level it is worth +24.6. Making the game faithful is what let the search show what it was actually worth.

Honest caveats. The +24.6 figure is a single seed. Seed variance in this setup is about 3.2 points of standard deviation, so the direction is certain but the exact number can move. The earlier +9.7 figure is a properly measured mean over 4 seeds per arm (74.3 sd 3.2 versus 84.0 sd 3.5, Welch t = 4.07, p ~ 0.007, with complete separation between arms).

Ablations that came out negative, reported anyway

Question Result
Does the compressed action head cost accuracy? No. 73.6% flat vs 73.0% embedded, z = 0.21, p = 0.83, for a head 4.1x smaller
Does the free neuron block help? No. Ablating it moved seeds by -2.4, -0.8, +2.4 points. Averages to zero
Does an async actor-learner help? No. 1.2x faster, win rate collapsed 81.4% to 22.0%
Is pure potential-based reward shaping better than bounded-dense? No. 58.6% vs 73.8%

Build and run

Requires a C++23 compiler. CUDA is optional; without it everything runs on the CPU.

.\build.ps1

Commands

pvzf train      train a policy
pvzf eval       evaluate a checkpoint over N games
pvzf bot        run the hand-written scripted baseline
pvzf gametest   40 simulation correctness tests
pvzf selftest   15 CPU tests, including finite-difference gradient checks
pvzf gputest    13 tests, every kernel against its CPU reference
pvzf simbench   simulation throughput measurement
pvzf gpuinfo    device capabilities

Flags that matter

Flag Default Meaning
--envs N 64 parallel games per rollout. 256 is the sweet spot
--groups N 2 double-buffered rollout. 1 disables the overlap
--hidden A,B 512,512 trunk widths
--head-dim K 70 action embedding size. 0 selects the flat head
--imagine N 0 shortlist size for the search tree. 0 disables it
--imag-branch R 1 children per node. 2 makes it a tree rather than a line
--capped-reward off bounds every dense reward term. Recommended
--free N 0 free neuron block size. Off by default, see part 7
--lr F 3e-4 learning rate
--gamma F 0.997 discount
--ent F 0.01 entropy bonus
--epochs N 4 passes over each rollout

Reproduce the headline number

pvzf train --envs 256 --head-dim 70 --imagine 16 --imag-branch 2 --capped-reward
pvzf eval  --games 500 --load bin/day_tree.bin

bin/day_base.bin and bin/day_tree.bin are the two trained agents from the results table.


Architecture at a glance

flowchart TB
  subgraph SIM["Simulator (src/game)"]
    B["Board: 91 plants, 39 zombies<br/>53 fusion recipes<br/>100 ticks per second"]
  end
  subgraph RL["RL wrapper (src/rl)"]
    E["Encoder: 1234 floats"]
    M["Legality mask: 649 bits"]
    R["Reward shaping"]
  end
  subgraph NET["Network (src/ai)"]
    T["Trunk 512, 512"]
    I["Imagination search tree"]
    P["Policy head"]
    V["Value head"]
  end
  subgraph TRAIN["PPO (src/rl/ppo.h)"]
    G["GAE advantages"]
    O["AdamW"]
  end
  B --> E --> T
  B --> M --> P
  T --> I --> P
  T --> V
  P --> A["action"] --> B
  B --> R --> G --> O --> T
Loading

Every box above is hand-written code in this repository.


Part 1: the simulator

Deterministic, fixed timestep, 100 ticks per second, matching the original game's logic rate. Headless, it runs orders of magnitude faster than real time.

Board::step(action) is a pure function of (state, action, rng). There is no renderer, no audio, no frame pacing, no wall clock. That is the entire reason training is fast.

Tick order

flowchart LR
  A["apply action"] --> S["tickSun"] --> W["tickWaves"] --> P["tickPlants"]
  P --> J["tickProjectiles"] --> Z["tickZombies"] --> BL["tickBlasts"] --> C["tickCleanup"]
Loading

The order is fixed and matters. Plants fire before projectiles move, so a pea spawned this tick does not travel until the next one. Blasts resolve after zombies move, so an explosion catches a zombie at the position it actually reached.

Data-driven design

A plant is a row in a table. So is a zombie. So is a fusion recipe.

{ .name = "Peashooter", .cost = 100, .recharge = RC_FAST, .health = 300,
  .flags = PF_LAND_ONLY, .attack = Attack::Shoot,
  .shootInterval = SEC(1.5f), .proj = ProjKind::Pea, .damage = 20 },
{ .name = "Pole Vaulting Zombie", .bodyHealth = 720, .secPerGrid = 2.5f,
  .ragedSecPerGrid = 4.7f, .flags = ZF_JUMPS,
  .pointCost = 2, .firstLevel = 4, .pickWeight = 2000 },
{ P_PEASHOOTER, P_PEASHOOTER, P_REPEATER,   "confirmed in-mod" },
{ P_WALLNUT,    P_WALLNUT,    P_GIANT_WALLNUT, "rolls the lane for 600" },

kPlants[] and kZombies[] are positional arrays: the row carries no id, its index in the table is its id. Enum order and table order must match exactly, which is checked at startup.

Fusion

Fusion needs no special action, no menu, no mode and no fee. Planting seed packet B onto a tile that already holds plant A is the fusion, exactly as the mod does it. The cost is simply both packets: a Peashooter at 100 plus a Wall-nut at 50 gives you a Peanut, a 4000-health plant that shoots, for 150 sun total.

That single design decision is what makes the two-seed level 1 survivable, and it is why the fusion tech tree is reachable through ordinary play rather than needing a bespoke action type.

What the Day area models

Adventure levels 1 to 9, verified line by line against the wiki:

  • The real seed unlock order. Level 1 really is Peashooter and Sunflower and nothing else
  • The real zombie introduction order: Flag 1, Conehead 2, Buckethead 3, Pole Vaulting 4, Peashooter Zombie 5, the box zombies 6, Cherryshooter Newspaper 7, Screen Door and Buckshooter and Buck-nut 8, Cherry-nut 9
  • Wave timing: first wave at 15 seconds, 30 second intervals
  • Zombotany, zombies that shoot your plants and can be ducked under by short plants
  • The nested box chain: Diamond-box unpacks into Gold-box unpacks into Silver-box unpacks into a Conehead, on death, 14,570 points of nesting
  • Pole vaulting blocked by tall plants only, one vault each, fast before and normal after
  • Screen doors that eat straight shots but not lobbed, fume or spikeweed damage
  • Ladders that bridge a wall and stay bridged for every zombie behind
  • Bungees that fall on a tile, steal the plant and leave
  • Chompers that swallow whole, except gargantuars, which they merely bite for 40
  • Upgrade plants that refuse bare dirt. A Tall-nut grows out of a Wall-nut, at 8000 health

Part 2: the reinforcement learning wrapper

Timing

The game runs at 100 ticks per second. The agent decides every 50 ticks, half a second of game time.

Idle-skip, the change that made learning work at all

Early on the agent almost never has enough sun to plant anything, so nearly every decision is a forced no-op. Asking a network to choose between 649 actions when 648 are illegal teaches it nothing and wastes the entire rollout.

Now the environment only wakes the network when at least one placement is actually legal.

Metric Before After
Legal actions per decision 2.5 60
Policy entropy 0.14 3.98
Explained variance 0.99 (critic predicting a constant) real
Return curve flat learning

Observation: 1234 floats

Block Offset Size Contents
Tiles 0 864 54 tiles x 16 features
Zombie bins 864 300 6 lanes x 10 distance bins x 5 features
Rows 1164 18 6 rows x 3: lane active, is water, mower alive
Seed slots 1182 40 10 slots x 4: plant id, affordable, recharge left, cost
Global 1222 12 sun, wave, tick, plants alive, zombies alive, and so on

The 16 features per tile are: occupancy, health fraction, cooldown fraction, age, a nine-way one-hot over the Attack enum, armor, underlay, and crater.

Zombies are binned, not listed. A variable-length set of entities cannot be fed to a dense network. Instead each lane is chopped into 10 distance buckets, 9 on-screen columns plus one off-screen, and each bucket carries count, total health, nearest distance, armor and a flying flag. That turns the threat picture into a fixed-size image with spatial structure the network can learn over.

Action space: 649

    0        no-op
    1..540   plant seed slot S on tile T   (10 slots x 54 tiles)
  541..594   shovel tile T                 (54 tiles)
  595..648   fire the cannon at tile T     (54 tiles)

A 649-bit legality mask accompanies every observation. Illegal logits are set to negative infinity before the softmax, so an illegal action has probability exactly zero rather than merely small.

Reward

perKillPoint  = 0.020   // scaled by the zombie's wave-budget cost
perFusion     = 0.100   // the mechanic we actually want it to use
perSun        = 0.0015  // only sun the agent GREW
perPlantLost  = 0.010   // per 50 sun of plant destroyed
perMowerLost  = 1.000
win           = +5.0
lose          = -5.0
illegal       = -0.010

Two design notes that took real experiments to learn:

Only sun the agent grew counts. Sun falling from the sky arrives on a fixed timer no matter what the agent does. Paying for it adds a large, perfectly predictable term to every return, which the critic learns exactly and which therefore contributes nothing but noise to the advantage.

Every dense term must be bounded in time. See the case study below. --capped-reward sets sunRewardCap = 4000 and fusionRewardCap = 25.


Part 3: the neural network

flowchart TD
  O["Observation<br/>1234"] --> L1["Linear 1234 to 512"]
  L1 --> R1["ReLU"]
  R1 --> L2["Linear 512 to 512"]
  L2 --> R2["ReLU"]
  R2 --> PH["Policy head<br/>Linear 512 to 70"]
  R2 --> VH["Value head<br/>Linear 512 to 1"]
  PH --> Q["query vector q, 70-dim"]
  Q --> DOT["logit[a] = q · emb[a]"]
  EMB["actEmb: 649 x 70<br/>learned action embeddings"] --> DOT
  DOT --> MSK["apply legality mask"]
  MSK --> SM["softmax, sample"]
  VH --> VAL["V(s), scalar"]
Loading

Layers

Layer Shape Parameters
trunk[0] 1234 to 512 632,320
trunk[1] 512 to 512 262,656
policy 512 to 70 35,910
actEmb 649 x 70 45,430
value 512 to 1 513
Total 976,829

Weights are initialised orthogonally, with gain sqrt(2) on ReLU layers, gain 0.1 on the policy head so the initial policy is close to uniform, and gain 1.0 on the value head. Orthogonal initialisation keeps the singular values of each weight matrix at 1, so gradients neither explode nor vanish through depth at the start of training.

Why the action head is embedded rather than flat

The obvious design gives the last layer 649 outputs: 512 x 649 = 332,288 weights, over a quarter of the entire network spent on the output alone. Worse, every action becomes a private column that learns nothing from any other action, so "plant a Sunflower on row 3 column 2" and "plant a Sunflower on row 3 column 3" are as unrelated as "no-op" and "fire the cannon".

Instead the trunk emits a 70-dimensional query, and every action owns a learned 70-dimensional embedding. The score of action a is their dot product:

logit[a] = q · actEmb[a]

This is exactly how a language model's output layer works: a hidden state dotted against a per-token embedding matrix. It makes the head 4.1x smaller and, more importantly, it makes similar actions neighbours in embedding space, so learning about one generalises to the others.

We tested whether that compression costs anything. Over 500 games: flat head 73.6% win, embedded head 73.0%. z = 0.21, p = 0.83. Statistically identical.


Part 4: backpropagation and the optimiser

There is no autograd. Every operation has an explicit hand-written backward pass.

The ReLU trick, and the bug it caused

The backward pass through ReLU needs to know which units were active in the forward pass. Rather than storing a separate boolean mask, the code reads the sign of the stored activation, since a post-ReLU activation is positive exactly when the unit fired.

That is cheap and correct, and it caused the single nastiest bug in this project. The free neuron block originally merged its output into act[trunk.size()] in place, which corrupted the very signs the trunk's backward pass depended on. Gradients silently went wrong in a way no assertion caught. The fix is a separate topMix buffer, so act[trunk.size()] stays the pure trunk activation. The same fix exists in both the CPU and GPU paths.

AdamW

Decoupled weight decay, so the decay term is applied directly to the weights rather than folded into the gradient and then rescaled by the adaptive term:

m = b1 * m + (1 - b1) * g
v = b2 * v + (1 - b2) * g^2
w -= lr * ( mhat / (sqrt(vhat) + eps) + wd * w )

Plus global gradient norm clipping: the L2 norm is taken across every parameter tensor at once, and if it exceeds maxGradNorm all gradients are scaled by the same factor. Clipping each tensor separately would change the direction of the update, not just its length.

The reported gradient norm is measured before clipping, because a post-clip norm just reads back the clip threshold and tells you nothing.


Part 5: PPO

flowchart LR
  ROLL["rollout<br/>256 games in parallel<br/>T steps each"] --> GAE["GAE advantages<br/>gamma 0.997, lambda 0.95"]
  GAE --> NORM["normalise advantages"]
  NORM --> EP["4 epochs x 4 minibatches<br/>= 16 gradient steps"]
  EP --> LOSS["clipped surrogate<br/>+ clipped value loss<br/>+ 0.01 entropy"]
  LOSS --> ADAM["AdamW, lr 3e-4<br/>global grad-norm clip"]
  ADAM --> ROLL
Loading

Generalised Advantage Estimation

delta_t = r_t + gamma * V(s_{t+1}) * (1 - done) - V(s_t)
A_t     = delta_t + gamma * lambda * (1 - done) * A_{t+1}

gamma = 0.997 is deliberately high. At half a second per decision a game runs several hundred decisions, and planting a Sunflower only pays off dozens of steps later. The bootstrap term is cleared on both true termination and truncation.

The clipped surrogate

ratio = exp(logp_new - logp_old)
L = -min( ratio * A,  clip(ratio, 1 - 0.2, 1 + 0.2) * A )

Because the update runs 16 gradient steps on data collected by an older policy, the ratio drifts away from 1. The clip is what makes that safe. approxKL and clipFrac are logged every update so the drift is visible rather than assumed.

Value loss, also clipped

The value prediction is clipped to a trust region around the old prediction, and the loss is the max of the clipped and unclipped squared errors, which prevents a single large critic update from destabilising the advantages the actor is using.


Part 6: imagination, a search tree inside the forward pass

This is the module that produced the 95.4%, and it is the most unusual part of the project.

A plain policy is reactive: it sees a board and reaches for a move. But the fusion ladder needs a plan. So the network imagines before it commits: it forks off the trunk halfway down, rolls a learned world model forward over its own imagined futures, and merges what it finds back into the policy logits.

flowchart TD
  H["trunk features h<br/>forked after layer 1"] --> SL["shortlist:<br/>top 16 actions from the<br/>un-imagined logits"]
  SL --> ENC["encode into a<br/>128-wide latent"]
  ENC --> D1["depth 1: 16 nodes<br/>dyn: latent to latent<br/>rHead: predicted reward<br/>vHead: predicted value"]
  D1 --> PR["proposer:<br/>latent to 2 continuations"]
  PR --> D2["depth 2: 32 nodes"]
  D2 --> LF["leaf: Q = r + gamma * v"]
  LF --> BK["max backup<br/>Q(node) = r + gamma * max_k Q(child)"]
  BK --> SC["score per shortlisted action"]
  SC --> GT["gate(h), initialised to 0"]
  GT --> MG["logits[a] += gate(h) * score[a]"]
Loading

How it runs

  1. A first forward pass produces un-imagined logits. The host takes the top 16 as a shortlist. Searching all 649 actions would be pointless; most are illegal or absurd.
  2. Each shortlisted action is embedded, reusing the same 70-dimensional actEmb the policy head uses, so the search and the policy speak the same language about actions.
  3. A learned dynamics function dyn maps a 128-wide latent to the next latent. Two heads read off a predicted immediate reward and a predicted value.
  4. At each depth a proposer maps the current latent to branch continuations, so the imagined actions are state-dependent rather than a fixed learned set.
  5. Values are backed up to the root, and the score of each shortlisted action is added to its logit through a gate.

Node count is shortlist * branch^(depth-1). Children of parent i land at index i * R + k, which keeps every array batch-major with zero index arithmetic.

Four properties that make this work

The gate starts at exactly zero. On step one the imagination contributes literally nothing. It can only gain influence if the gradient says it helps. This is why bolting it on cannot make things worse by construction.

Max backup, not average. Q(node) = r + gamma * max_k Q(child). This differentiates like max-pooling: the winning child receives the entire gradient and the losers receive none. The argmax index is stored per node during the forward pass and replayed in the backward pass. Both the CPU and the GPU use a strict > on ties, so the two backends can never pick different winners and diverge.

dQ is dRew. The reward at a node enters its own Q with coefficient exactly 1, so the reward gradient and the Q gradient are the same array. The world model's separate reward supervision must be added after the descent, or it wrongly propagates deeper into the tree than it should.

It is gradient-checked. The whole module, argmax backup included, passes finite-difference checks at 2.2e-04. That test has a subtlety worth knowing: a max is piecewise linear, so if the winning branch flips between the +h and -h probes, the central difference measures the average of two different slopes and the check fails spuriously. The test detects the flip and resamples that coordinate.

Line versus tree

--imag-branch 1 gives a single adaptive line rather than a tree. It is cheaper, and it is worse in a specific and interesting way, documented in the case study below.


Part 7: the free neuron block

An experiment in architecture without layers, kept in the repository along with the measurement that says it did not help.

flowchart LR
  IN["input projection<br/>inDim to 96"] --> C012["clusters 0,1,2<br/>96 neurons, injected"]
  C012 -.->|sparse| C34["clusters 3,4<br/>reachable only<br/>through the graph"]
  C34 -.->|sparse| C567["clusters 5,6,7<br/>96 neurons, read out"]
  C012 -.->|sparse| C567
  C567 --> OUT["output projection<br/>96 to outDim"]
  OUT --> GATE["gate, starts at 0"]
  GATE --> TOP["added to trunk top"]
Loading

A pool of 256 neurons wired to itself by a fixed clustered random graph: 8 clusters of 32, 50% connection probability inside a cluster, 5% between, giving 10.3% density and 6,742 edges. There are no layers. The block relaxes over 4 micro-steps:

z[0] = relu(inj)
z[t] = relu(Wm · z[t-1] + b + inj)      // the input is re-injected every step

Input enters clusters 0 to 2, output is read from clusters 5 to 7, and clusters 3 and 4 are reachable only through the graph. The projections are 96-wide on each side, so that disjointness is structural rather than enforced by masking.

Wm = W * mask is rebuilt every forward pass, so absent edges have exactly zero derivative, not merely a small one. That is asserted by a GPU test which reads 0.000e+00.

The recurrent matrix is rescaled to a spectral radius of 0.95, measured by power iteration plus a Gelfand-style geometric mean of the tail. This has to be measured, not assumed, because masking an orthogonal matrix destroys its orthogonality.

The result

Its gate rose from 0 to about 0.60 to 0.72 on every seed. It clearly wanted a voice.

Then we ablated it. eval --no-free loads the trained weights, sets useFree = false, and replays the identical 500 games. The win rate moved by -2.4, -0.8 and +2.4 points across three seeds. That averages to nothing.

A large gate does not mean a useful block. Whatever it computes is close to redundant with what the trunk already produces. The ablation is the only measurement that settles this, which is exactly why it was built before the experiment was run. The block is off by default.


Part 8: the GPU backend

CUDA is reached through the driver API and NVRTC directly. Kernels are C++ source strings inside gpu_kernels.h, compiled at runtime to a sm_86 cubin. There is no nvcc step in the build and no math library linked.

The kernels

Group Kernels
Linear algebra sgemm_nn, sgemm_nt, sgemm_tn, sgemm_tn_split
Activations relu_fwd, relu_bwd
Policy log_softmax_masked, sample_actions, gather_logp
PPO ppo_policy_grad, ppo_value_grad
Optimiser adamw, sq_norm, col_sum, add_inplace, fill
Imagination imag_child, imag_child_bwd, imag_branch, imag_branch_bwd, imag_backup_leaf, imag_backup_step, imag_backup_bwd, imag_leaf_bwd, imag_merge, imag_merge_bwd
Free block free_mask, free_slice, free_merge, free_merge_bwd, free_gate_bwd

A real optimisation: sgemm_tn was starved

sgemm_tn tiles only its output. For a weight gradient the output is the weight matrix, so a [1 x 128] reward head got 2 thread blocks and a [128 x 128] dynamics matrix got 4, on a 56-SM device, to reduce roughly a million rows. M, the only large dimension, received no parallelism at all.

sgemm_tn_split adds a gridDim.z over slices of M. Each block reduces its own slice privately, then atomicAdds the partial result. It is only used when beta == 1, because an overwrite needs a single writer.

Configuration Before After Speedup
--imag-branch 1 update 2.77 s 0.78 s 3.6x
--imag-branch 2 update 5.75 s 1.24 s 4.6x
no imagination 0.39 s 0.38 s unchanged

Imagination overhead versus a plain network fell from 6.3x to 2.2x.

Other things learned the hard way

  • cuMemcpyDtoD returns to the host before the copy lands, and a non-blocking stream never waits on the legacy stream. Missing that sync means reading torn weights.
  • CU_STREAM_NON_BLOCKING is mandatory, or a newly created stream still synchronises with stream 0 and the whole point of the second stream is lost.
  • BN is a #define for the SGEMM tile width, so no kernel parameter may be named BN.
  • The optimiser originally did one blocking download per parameter tensor per step, about 368 round trips per update to retrieve a single scalar. Each tensor now reduces into its own slice of one buffer, downloaded once.

Part 9: performance

Simulation

Metric Value
Simulation throughput orders of magnitude faster than real time
Agent steps/sec, 256 envs 37,200
Agent steps/sec, 768 envs, groups 2 40,500

Entity counts are tiny: 5 live zombies out of 6.5 slots, 8.3 plants, 1.5 projectiles. There is no O(n^2) problem here; lane-bucketing would cost more than it saves.

Double-buffered rollout

The batch is split in half. A persistent worker thread simulates group A on the CPU while the main thread runs group B's forward pass on the GPU.

flowchart LR
  subgraph T0["time 0"]
    S0["CPU: simulate group A"]
    N0["GPU: forward group B"]
  end
  subgraph T1["time 1"]
    S1["CPU: simulate group B"]
    N1["GPU: forward group A"]
  end
  T0 --> T1
Loading

The alternation order is chosen so the RNG draw sequence is identical to serial, which makes the output bit-for-bit the same as running with --groups 1. That is verified, not assumed.

The async experiment that was rejected

A full asynchronous actor-learner was also built: a separate actor thread with its own CUDA context binding, its own non-blocking stream, its own frozen weight copy and its own device scratch, playing rollout u+1 while the learner updates on rollout u.

It works, and it is 1.2x faster. It also learns far worse:

Epochs Sync win rate Async win rate
4 81.4% 22.0%
1 67.0% 37.2%

More epochs makes sync better and async worse, which is the signature of off-policy drift: PPO already takes 16 gradient steps per update, and async data ends up roughly twice as stale. Making it viable would need V-trace or proper importance-sampling correction on the advantages. Speed that costs correctness is not speed. It is off by default.


Part 10: correctness and testing

pvzf selftest    15 passed    CPU
pvzf gputest     13 passed    every kernel against its CPU reference
pvzf gametest    40 passed    simulation rules and the RL wrapper

What is actually asserted

Gradients. The full network, including the imagination tree and the free block together, is finite-difference checked. That combined test is the one that catches a rejoin which overwrites instead of accumulating.

GPU equals CPU. Forward at 6.9e-07, backward at 1.5e-06, and absent graph edges have exactly 0.000e+00 gradient.

Determinism. Two runs from the same seed agree on a board fingerprint at every tick.

Game rules, a sample:

  • A Tall-nut refuses bare dirt, is legal on a Wall-nut, and grows at 8000 health
  • A pole vaulter clears a Wall-nut and reaches x = -21, the lawnmower, but is stopped dead at x = 415 by a Tall-nut. Same zombie, same lane, same seed, only the blocker differs
  • A Silver-box unpacks into a Conehead when killed, but into nothing when devoured by a Chomper
  • Potential-based reward shaping telescopes to zero over a whole episode, measured at 2.8e-07, while a companion assertion proves the shaping was non-trivial so the test cannot pass vacuously

Case study: the reward was exploitable

Worth reading, because the search found the hole and the plain policy never did.

perSun was the only reward term unbounded in time. Kills and fusions are limited by what the level contains, but a Sunflower pays out for as long as the game runs.

The --imag-branch 1 search discovered this and learned to kill slowly, letting the wave timer idle while its economy kept earning:

Baseline Branch 1, uncapped
Win rate 81.3% 59.3%
Return 8.60 11.06
Wave reached 9.8 9.8
Kills 40.8 40.8
Steps 67 123
Sun collected 3,419 5,902

Same wave, same kills, twice the length. Its win rate peaked at 84.4% then decayed to 60.6% while its return kept climbing. That divergence between the objective and the goal is the signature to watch for.

Two fixes were built and both are in the repository:

--capped-reward bounds every dense term (sunRewardCap = 4000, fusionRewardCap = 25). It keeps the curriculum the terms were added to teach and stops paying once the lesson is over. Branch 1 went from 123 steps and 59.3% to 66 steps and 77.8%. This is the one to use.

--potential-reward replaces the dense terms with potential shaping, gamma * Phi(s') - Phi(s). This is provably policy-invariant and therefore exploit-proof. It also learns considerably worse, 58.6% versus 73.8%, and that is expected rather than a bug: the telescoping property that makes it safe also makes its net incentive exactly zero, so it deletes the curriculum along with the exploit.

The tree at --imag-branch 2 never drifted in the first place, and its extra steps are not stalling: its wave reached and its kills both went up.


What is and is not implemented

Implemented

All seven Day plants at wiki-exact stats. All ten Day zombies with their real behaviours. 53 fusion recipes covering the reachable Day tech tree, including both the pea ladder and the parallel cherry ladder. Status effects: chill, freeze, butter, burn, hypnosis. Armour types and what strips them. Lawnmowers. Craters. Lily pads and flower pots as underlays. Pumpkins as armour. All nine attack kinds.

Known gaps, listed rather than hidden

  • The mod has 511 fusions across six tiers; this implements 53
  • Damage types are not modelled, so the Cherry-nut Zombie's 150-damage cap against Cherry Bombs and the Cremator type are absent
  • The Giant Wall-nut's roll is modelled as an instant whole-lane blast rather than a travelling object
  • The Gardening Glove, Fertilizer, Surprise-box and Bucket items are absent, which also means the Odyssey fusion tier is unreachable
  • Zombie movement speeds are our own calibration. The wiki's zombie infobox has no speed field at all, so these cannot be verified against it
  • Only the Day map is implemented. Night, Pool, Fog and Roof are stubs

File layout

build.ps1                 the entire build, one g++ invocation

src/game/
  defs.h                  every enum, flag, struct and tunable constant
  data.cpp                the tables: 91 plants, 39 zombies, 53 recipes
  board.h  board.cpp      the simulation, ~1400 lines
  gametest.h              40 correctness tests and the scripted baseline

src/ai/
  tensor.h                Mat, the flat float matrix everything is built on
  nn.h                    Linear, Net, Imagination, FreeBlock, AdamW, backward passes
  gpu.h                   device mirrors of the above
  gpu_kernels.h           every CUDA kernel, as NVRTC source
  cuda_api.h              driver API bindings, loaded dynamically
  selftest.h              15 CPU tests including gradient checks
  gpu_selftest.h          13 GPU-versus-CPU tests

src/rl/
  env.h                   observation encoding, action masking, reward shaping
  ppo.h                   rollout, GAE, the PPO update, checkpointing

src/core/
  rng.h                   xoshiro256++
  worker.h                persistent worker thread for the double-buffered rollout

src/main.cpp              CLI

Checkpoints are versioned; v5 stores the free block's graph mask verbatim, because the RNG state that generated it is not recoverable.


Status and roadmap

Done. The simulator, the learning engine, and a trained agent that wins 95.4% of real Day level 1 games.

Next. The rest of the fusion catalogue: 511 recipes across Common, Upgraded, Advanced, Titan, Odyssey and Infusible tiers, and the mechanics that go with them. The rolling nuts, the Cremator damage type, the exploding nuts, the Gardening Glove and the two-tile Titan fusions.

After that. The harder Day levels. The scripted baseline currently scores 27.5% on level 1 and 0% on level 9, so there is a real curriculum waiting. Then the remaining four maps.

Contributors

zigazaga4

Issues