This repository provides a modular Python implementation of a single-player 14.1 continuous pool environment built on PyBullet, with human-playable UI, episode recording/export, offline replay, and a Gymnasium wrapper for RL training.
lib/engine/engine.py: PyBullet simulation core with table/ball setup, shot execution, foul handling, scoring, undo snapshots, and reward hooks.lib/engine/types.py: Shared dataclasses for configuration, state, actions, and outcomes.lib/data/recorder.py: Episode recorder exporting/importing context-free JSONL timesteps for pretraining and replay.lib/render/renderer.py: Pygame visualization with shot preview and HUD.lib/render/play.py: Interactive human loop for testing and data collection.lib/replay/viewer.py: Offline replay viewer that steps through recorded episodes with the same renderer.lib/envs/pool.py: GymnasiumEnvexposing the simulator for fast RL training with optional episode recording.
Activate the existing pool conda env and install Python deps if missing:
conda activate pool
pip install -r requirements.txtNo additional assets are required.
Run the interactive loop (headless physics, Pygame rendering):
python -m lib.render.human_sessionControls:
- Mouse aim (hold/drag); scroll to adjust power.
[/]cycle called ball (1–15),1-6set called pocket id (0–5),Cclears the call.A/Dhorizontal spin,W/Svertical spin.Spaceshoot,Rreset rack,Uundo last shot,Eexport recording toexports/episode-<timestamp>.jsonl.Escquits. Each shot is appended to a recorder as a stand-alone timestep with observation, action, call, outcome, reward, and done flag.
Load a recording and step through frames:
python -m lib.replay.viewer exports/episode-<timestamp>.jsonlControls: Left/Right step, Space toggle autoplay, Esc quit. The viewer reconstructs each frame directly from the stored observation, so episodes are context-independent.
Use lib.envs.pool.PoolEnv for training:
import gymnasium as gym
from lib.envs.pool import PoolEnv
env = PoolEnv(record=True)
obs, info = env.reset()
action = env.action_space.sample()
obs, reward, terminated, truncated, info = env.step(action)- Observation: flat array
[x, y]per ball (cue + 15 objects) ordered byball_id. - Action space: Dict with
angle(rad),power[0,1],english_horizontal/english_vertical[-1,1],call_ball(0=none, 1–15),call_pocket(6=none, 0–5 pockets). - Rewards: configurable via
RewardWeightson the engine (w_ball,w_complete,w_shot,w_foul,w_illegal). - Set
record=Trueto mirror each timestep into a JSONL recorder for replay/debugging.
- Top-down PyBullet model with simplified English: spin is injected as initial angular velocity on the cue ball; precise throw/curve effects are approximated.
- Pockets are detected geometrically (ball center within pocket radius)
- Illegal or uncalled pocketing resets table to last state, scratches additionally reset cue-ball to starting default position.
- Engine supports fast headless stepping for RL and GUI mode for debugging; undo history stores full PyBullet state snapshots.
Recordings are JSONL, one context-free timestep per line:
{
"observation": [...],
"action": {"angle": 0.7, "power": 0.4, "english_horizontal": 0.0, "english_vertical": -0.2},
"call": {"ball_id": 5, "pocket_id": 2},
"outcome": {"pocketed_balls": [5], "pocketed_pockets": {"5": 2}, "hit_obj_ball": true, "scratched": false, ...},
"next_observation": [...],
"reward": 1.0,
"done": false,
"info": {"foul_reason": null}
}Episodes can be shuffled for batch training because each row contains the full pre/post state.