Weather-market intelligence for Polymarket.
Thermocline scans weather contracts, turns forecasts into calibrated probabilities, simulates execution quality on the CLOB, and observes PolyDekos-style adjacent-bucket ladders before any capital is put at risk.
It is built to answer one practical question:
Is this weather market actually mispriced after forecast uncertainty, liquidity, calibration, and event exposure are accounted for?
The current release is a production observation build: cron-ready, tested, snapshotting live market conditions, but deliberately not live trading.
| Area | Status |
|---|---|
| Runtime | System cron every 30 min from /home/builder/weather-edge |
| Latest validation | PYTHONPATH=src pytest tests/ -q โ 205 passed on 2026-06-01 |
| Trading mode | No live trading |
| Paper opening | Tiny observation experiment only: cron may open max 5 SHADOW/PAPER positions at 1 USDC under explicit env flags |
| Calibration gate | Currently blocking true PASS/PAPER readiness; paper/live should stay constrained |
| Empirical forecast result | Poor so far: 31 paper trades, 26 closed, 4 wins / 22 losses, about -20.03 USDC closed PnL as of 2026-06-01 |
| Candidate observation | candidate_observations table records SHADOW/PAPER/PASS/REJECT outcomes for future Gamma resolution |
| PolyDekos ladder | Implemented for read-only observation, fill simulation, and reporting |
| Ladder readiness | Not tradable yet: missing historical fill-level replay, realized PnL/ROI/drawdown, and ladder-level calibration |
Bottom line: this is an observation/calibration build. The forecast edge is not proven. Keep live trading disabled and keep any paper experiment tiny until calibration and ladder readiness gates are explicitly green and the operator authorizes it.
For a complete project handoff, read docs/weather-forecast-handoff-2026-06-01.md.
-
Live trading remains disabled. Paper opening, if enabled, must stay tiny and explicitly bounded by environment flags.
-
Safest default when in doubt:
export WEATHER_EDGE_DISABLE_PAPER_OPEN=1 -
Calibration gate blocks risk-taking when Brier score / bucket calibration are outside thresholds.
-
Ladder fill simulation is read-only: it calls order-book simulation, writes snapshots, and places no orders.
-
No secrets belong in the repo or README. Keep credentials in the runtime environment only.
-
Do not commit runtime artifacts: DBs, logs, reports, snapshots, backups, locks, and generated datasets are ignored.
Thermocline targets weather binary markets such as:
โWill the high in Tokyo be 22ยฐC or higher on May 3?โ
The pipeline:
- discovers Polymarket weather markets;
- parses temperature buckets / thresholds;
- fetches weather forecasts and context;
- computes probabilities with uncertainty;
- rejects unsafe or poorly calibrated opportunities;
- tracks paper accounting and settlement;
- records order-book snapshots and fill simulations;
- produces reports for calibration, risk, and ladder readiness.
Polymarket Gamma/CLOB
โ
โผ
Market discovery + parsing
โ
โผ
Forecast/context layer
- Open-Meteo forecasts
- ensemble / horizon features
- Weather.com / METAR settlement sources
- NASA GISTEMP baseline with cache/fallback for global-temperature markets
โ
โผ
Scanner + probability model
- Gaussian bucket probability
- uncertainty / horizon / regime adjustments
- candidate scoring
โ
โโโโโโโโโโโโโโโโโ
โผ โผ
Single-bucket PolyDekos-style ladder observation
candidate flow - adjacent buckets
- deterministic ladder_id
- parent_ladder_id per leg
- token_id per leg
- read-only fill simulation
- ladder order-book snapshots
โ โ
โผ โผ
Risk and gates Ladder backtest report
- calibration gate - qualitative hit-rate only for now
- event exposure - no historical fill/PnL yet
- sizing caps
โ
โผ
Reports + paper accounting + settlement audit
โ
โผ
Cron heartbeat + DB backup + runtime snapshots
src/weather_edge/
โโโ main.py # CLI entry point and paper-cycle orchestration
โโโ scanner.py # Market scan and probability computation
โโโ candidates.py # PASS/PAPER/REJECT scoring
โโโ calibration.py # Calibration reports/gates
โโโ risk.py # Risk sizing logic
โโโ event_exposure.py # Event-level exposure caps
โโโ uncertainty.py # Horizon/regime uncertainty adjustments
โโโ weather_features.py # Forecast/weather feature extraction
โโโ ladder.py # PolyDekos-style adjacent-bucket ladders
โโโ ladder_fill.py # Read-only per-leg fill simulation + snapshots
โโโ ladder_backtest.py # Qualitative ladder-vs-single report
โโโ closed_trades_audit.py # Settlement/accounting audit
โโโ research_dataset.py # Export research dataset artifacts
โโโ settlement.py # Resolution via weather sources
โโโ db.py # SQLite schema and persistence
โโโ clients/
โโโ clob.py # CLOB order-book / fill simulation helpers
โโโ polymarket.py # Market discovery
โโโ weathercom.py # Weather.com / Wunderground observations
โโโ aviationweather.py # METAR observations
โโโ openmeteo.py # Forecast API
โโโ nasa_gistemp.py # GISTEMP baseline with timeout/cache/fallback
Docs:
docs/weather-forecast-handoff-2026-06-01.md # Complete forecast handoff / restart guide
docs/polydekos-ladder-roadmap.md # Ladder strategy roadmap
docs/ladder-paper-readiness-guide.md # Go/no-go checklist before paper ladder
docs/v1-spec.md # Earlier system spec
Runtime artifacts are under data/, reports/, and logs/ and are intentionally ignored by Git.
pip install -e .
PYTHONPATH=src python3 -m weather_edge.main init-dbPYTHONPATH=src pytest tests/ -qTargeted safety/ladder checks:
PYTHONPATH=src pytest \
tests/test_ladder.py \
tests/test_ladder_fill.py \
tests/test_ladder_backtest.py \
tests/test_nasa_gistemp.py \
tests/test_scanner_global.py \
-qPYTHONPATH=src \
WEATHER_EDGE_DISABLE_PAPER_OPEN=1 \
python3 -m weather_edge.main verify-candidatesOutputs include reports/verified_candidates.json and policy flags such as:
{
"ladder_fill_simulation_read_only": true,
"ladder_fill_simulation_places_orders": false
}PYTHONPATH=src \
WEATHER_EDGE_DISABLE_PAPER_OPEN=1 \
python3 -m weather_edge.main ladder-backtest-report \
--output /tmp/weather_edge_ladder_backtest_report.jsonCurrent interpretation: useful for qualitative comparison, not sufficient for trading because historical fill-level replay and realized ladder PnL are not available yet.
bash scripts/paper_cycle.shProduction cron currently runs:
*/30 * * * * builder cd /home/builder/weather-edge && bash scripts/paper_cycle.shThe wrapper uses a project-local lock:
data/run/weather_edge_paper_cycle.lock
and writes heartbeat state to:
reports/paper_cycle_heartbeat.json
Current CLI commands include:
init-db
fetch-markets
scan
verify-candidates
paper-open
paper-report
paper-settle
reconcile-sources
resolve-candidate-observations
paper-cycle
run-once
calibration-report
calibration-snapshot
audit-closed-trades
risk-sizing-report
export-research-dataset
ladder-backtest-report
recalibrate-sigma
backtest
paper-open exists for controlled experiments, but should not be used automatically while WEATHER_EDGE_DISABLE_PAPER_OPEN=1 and calibration gates are blocking.
Single-bucket probabilities use Gaussian bucket integration:
P(bucket) = ฮฆ((upper - ฮผ) / ฯ) - ฮฆ((lower - ฮผ) / ฯ)
where:
ฮผ= forecast temperature estimate;ฯ= forecast uncertainty;- uncertainty is adjusted by horizon/regime/calibration context.
Risk controls include:
- calibration gate;
- event exposure caps;
- conservative sizing / caps;
- CLOB fill simulation;
- rejection of candidates with insufficient liquidity or unstable inputs.
The ladder path is designed around a range/ladder thesis rather than betting only the exact most likely bucket.
Implemented now:
- deterministic
ladder_id; parent_ladder_idon every leg;token_idpropagation for CLOB simulation;- adjacent buckets / narrow ranges;
- read-only per-leg fill simulation via
ladder_fill.py; - gzip snapshots for ladder books;
- qualitative report comparing:
single_best_bucket,ladder_pm_1c,ladder_pm_2c.
Not ready yet:
- historical fill-level replay by
ladder_id; - realized ladder cost / payout / PnL / ROI / max drawdown;
- ladder-level calibration gate;
- enough resolved ladder observations to justify paper/live.
Do not enable automatic paper/live openings until all are true:
- calibration gate allowed;
- candidate reports show stable non-zero accepted candidates;
- ladder fill snapshots have accumulated across many cycles;
- historical ladder replay works by
event_key + ladder_id; - ladder settlement maps each leg to realized outcomes;
- PnL / ROI / drawdown are computed from real historical snapshots;
- ladder-level calibration metrics are acceptable;
- cron heartbeat remains stable;
- explicit operator approval is given.
Useful code/docs/tests should be committed intentionally. Generated runtime artifacts should not.
Ignored by design:
data/backups/data/cache/data/run/data/snapshots/data/*.db*data/research_dataset*.jsonldata/sigma_calibration.jsonreports/*except placeholderslogs/*except placeholders.hermes/
Before any commit:
git status --short
PYTHONPATH=src pytest tests/ -qCommit/push only after explicit operator approval.
This is a research and observation system for prediction-market strategy development. It is not financial advice. Live trading should remain disabled until the system has proven calibration, execution quality, and risk behavior on resolved historical observations.