Shad0wSeven/scratch

★ 0Forks 0Jupyter NotebookGitHub ↗Compare

README

Single-Player 14.1 Pool RL Environment

This repository provides a modular Python implementation of a single-player 14.1 continuous pool environment built on PyBullet, with human-playable UI, episode recording/export, offline replay, and a Gymnasium wrapper for RL training.

Components

  • lib/engine/engine.py: PyBullet simulation core with table/ball setup, shot execution, foul handling, scoring, undo snapshots, and reward hooks.
  • lib/engine/types.py: Shared dataclasses for configuration, state, actions, and outcomes.
  • lib/data/recorder.py: Episode recorder exporting/importing context-free JSONL timesteps for pretraining and replay.
  • lib/render/renderer.py: Pygame visualization with shot preview and HUD.
  • lib/render/play.py: Interactive human loop for testing and data collection.
  • lib/replay/viewer.py: Offline replay viewer that steps through recorded episodes with the same renderer.
  • lib/envs/pool.py: Gymnasium Env exposing the simulator for fast RL training with optional episode recording.

Setup

Activate the existing pool conda env and install Python deps if missing:

conda activate pool
pip install -r requirements.txt

No additional assets are required.

Human Play & Data Capture

Run the interactive loop (headless physics, Pygame rendering):

python -m lib.render.human_session

Controls:

  • Mouse aim (hold/drag); scroll to adjust power.
  • [/] cycle called ball (1–15), 1-6 set called pocket id (0–5), C clears the call.
  • A/D horizontal spin, W/S vertical spin.
  • Space shoot, R reset rack, U undo last shot, E export recording to exports/episode-<timestamp>.jsonl.
  • Esc quits. Each shot is appended to a recorder as a stand-alone timestep with observation, action, call, outcome, reward, and done flag.

Replay Viewer

Load a recording and step through frames:

python -m lib.replay.viewer exports/episode-<timestamp>.jsonl

Controls: Left/Right step, Space toggle autoplay, Esc quit. The viewer reconstructs each frame directly from the stored observation, so episodes are context-independent.

Gymnasium Environment

Use lib.envs.pool.PoolEnv for training:

import gymnasium as gym
from lib.envs.pool import PoolEnv

env = PoolEnv(record=True)
obs, info = env.reset()
action = env.action_space.sample()
obs, reward, terminated, truncated, info = env.step(action)
  • Observation: flat array [x, y] per ball (cue + 15 objects) ordered by ball_id.
  • Action space: Dict with angle (rad), power [0,1], english_horizontal/english_vertical [-1,1], call_ball (0=none, 1–15), call_pocket (6=none, 0–5 pockets).
  • Rewards: configurable via RewardWeights on the engine (w_ball, w_complete, w_shot, w_foul, w_illegal).
  • Set record=True to mirror each timestep into a JSONL recorder for replay/debugging.

Simulation Notes & Assumptions

  • Top-down PyBullet model with simplified English: spin is injected as initial angular velocity on the cue ball; precise throw/curve effects are approximated.
  • Pockets are detected geometrically (ball center within pocket radius)
  • Illegal or uncalled pocketing resets table to last state, scratches additionally reset cue-ball to starting default position.
  • Engine supports fast headless stepping for RL and GUI mode for debugging; undo history stores full PyBullet state snapshots.

Data Format

Recordings are JSONL, one context-free timestep per line:

{
  "observation": [...],
  "action": {"angle": 0.7, "power": 0.4, "english_horizontal": 0.0, "english_vertical": -0.2},
  "call": {"ball_id": 5, "pocket_id": 2},
  "outcome": {"pocketed_balls": [5], "pocketed_pockets": {"5": 2}, "hit_obj_ball": true, "scratched": false, ...},
  "next_observation": [...],
  "reward": 1.0,
  "done": false,
  "info": {"foul_reason": null}
}

Episodes can be shuffled for batch training because each row contains the full pre/post state.

Contributors

alvinzhengqShad0wSeven

Issues