ObservedObserver/q-learning-practice

★ 0Forks 0TypeScriptGitHub ↗Compare

README

Q-Learning Snake

Practice project: tabular Q-learning agent for Snake, with a headless Node train/eval path (no browser/UI required).

Setup

yarn install
cd packages/brain
yarn build

Headless CLI (packages/brain)

# Train and evaluate (writes a Q-table JSON)
yarn train-eval -- --width 10 --height 10 --train-episodes 50000 --eval-episodes 50 --out ./models/model-10x10.json

# Train only
yarn train -- --width 10 --height 10 --episodes 50000 --out ./models/model-10x10.json

# Greedy eval of a saved model
yarn eval -- --model ./models/model-10x10.json --episodes 50

# Random-policy baseline
yarn baseline -- --width 10 --height 10 --episodes 50

Or after build:

node ./build/cjs/cli.js train-eval --width 10 --height 10 --train-episodes 50000 --eval-episodes 50

Tests

cd packages/brain
yarn test

Covers snake rules (eat/grow, wall death, self death), utility fixes, and a real Bellman Q-update against the shipped formula.

Algorithm (canonical: packages/brain)

Piece Design
State Compact relative features (~288 keys): danger straight/left/right, food direction signs, current heading
Actions 3 relative turns: straight / right / left (never reverse)
Update Classic Q-learning: (Q(s,a) \leftarrow Q(s,a) + \alpha[r + \gamma \max_{a'}Q(s',a') - Q(s,a)]); terminal uses (r) only
Exploration ε-greedy with exponential decay
Rewards +10 food, −10 death/starvation, small step cost, light closer/farther shaping

Bugs fixed vs the original implementation

  • bbox used Math.max for minY (corrupt body features)
  • manhattanDis used the comma operator (wrong return)
  • snakeSelfPos indexed the wrong array segment
  • Sparse absolute (x,y,fx,fy,cbbox) state → unlearnable on large grids; replaced with relative state
  • No true ε schedule; opposite-move handling was ad hoc
  • Missing episode step / starvation limits (infinite loops)

Packages

  • packages/brain — game kernel + Q-learning + headless CLI (source of truth)
  • packages/snake — optional React UI (not required for training)
  • packages/move-to-point — separate toy demo

Expected performance

On a 10×10 board after ~50k training episodes, greedy eval over 50 games typically scores mean food ≫ 8 (often ~15–25), far above a random baseline (~0).

Contributors

ObservedObserver

Issues