SuperMarioYL/blastevict

Context-eviction CLI for coding agents — traces how far each tool output propagated across turns, ranks what is silently eating your context budget, and replay-proves which drops are safe before you cut them (temp=0 ablation replay).

★ 1Forks 0PythonGitHub ↗Compare

Project website ↗

ai-radarblast-radius-proven-evictioncontext-budgetobservability-to-controlpropagation-tracingreplay-proven-safety

README

English | 简体中文

version python license ci replay

blastevict typing

blastevict brand hero — particle convergence of a tool-output blast radius

blastevict is the context-eviction CLI that replay-proves which tool outputs SREs can safely drop. When your coding agent blows through context mid-incident, blastevict traces which persisted tool outputs are silently eating your budget — and replay-proves which ones are safe to evict without breaking the next decision.


Why

Coding-agent sessions (Claude Code, Cursor-class tools) run long and multi-turn. Every tool call — a file read, a log dump, a grep — leaves its output persisted in context for every later turn. Wider context windows didn't fix this; they enabled longer runs, so persisted output accumulates past the point any prevention strategy can keep flat. Mid-run, the agent hits its limit and you must either abort (losing the chain of thought) or blindly trim context the agent may still need.

Existing tools each miss the decision blastevict makes:

Approach What it measures The gap
Output compression (headroom) Reduces each call's footprint before the LLM sees it Prevention only — cannot help once context is already overflowing mid-run
Call-envelope profilers (envelcost) Token cost of a tool call at call time Call-time cost ≠ persistence cost: a cheap call referenced every turn is far more expensive cumulatively
Context editors (excise) Cuts context mid-run Cuts blindly — no cross-turn blast-radius attribution, no eviction-safety check

blastevict introduces a new flow primitive — observability-to-control: it traces how far each output's footprint propagated across turns, then replay-proves the cut is safe by re-running the agent under ablation. The verdict is the product.

blastevict architecture — ingest to trace to replay to report

How it works

A single Python process, four stages:

  1. Ingest — parses a Claude Code session JSONL into an ordered turn model: events, tool calls, and which outputs persisted into later context windows.
  2. Trace — builds the blast-radius footprint for every persisted tool output: how many subsequent turns it propagated across and its cumulative token cost while it sat there.
  3. Replay — for the top-footprint candidates, re-runs the agent's next decision without that output in context (temperature=0, N=3 samples) and compares the action distribution to the original. A compaction event truncates the visible tail — outputs before it stop propagating.
  4. Report — ranks every output by footprint and stamps a replay-proven verdict.
blastevict core process — trace blast radius, replay under ablation, diff the decision

The two new structures

ToolOutputFootprint
  tool_call_id          # which persisted tool output
  origin_turn           # turn where it was produced
  propagated_turns      # turns where it still sat in context
  token_footprint       # cumulative tokens occupied across turns
  propagation_depth     # max turns it persisted

EvictionVerdict
  tool_call_id
  footprint
  replay_samples        # N=3 ablation runs
  action_overlap        # 0.0–1.0, action-distribution similarity
  verdict               # safe-to-evict | load-bearing | uncertain
  determinism_note      # temp=0, N=3, action-distribution comparison basis

The determinism mechanism (stated, not asserted)

Replay calls use temperature=0 with N=3 samples per candidate. Comparison is action-distribution similarity (does the agent still call the same tools / reach the same decision), not exact-token match — because even temperature=0 has residual non-determinism on frontier models. High cross-sample variance forces an uncertain verdict and never produces a false safe-to-evict. This follows the leave-one-out replay-ablation precedent (fixed-seed action-distribution comparison).

Eviction removes the tool_use + tool_result pair together: the messages API requires every tool call to be answered, so the candidate's call and output are ablated as a unit.

Quickstart

# 1. Install (Python 3.12+)
uv tool install blastevict            # or: pip install blastevict
# from source:
git clone https://github.com/SuperMarioYL/blastevict && cd blastevict
uv pip install -e .

# 2. (Optional) enable replay-proven verdicts
export ANTHROPIC_API_KEY=sk-ant-...

# 3. Analyze a Claude Code session
blastevict analyze examples/sample_session.jsonl
# replay is skipped without an API key — you still get the full footprint report;
# set ANTHROPIC_API_KEY to replay-prove the top candidates.

Replay is bounded by cost: only the top-K footprint candidates (default --top-k 5) are replayed. Lower-footprint outputs are reported as not-replayed.

Demo

A real blastevict analyze run on a Claude Code session JSONL (replay skipped, no API key — the footprint report is fully reproducible without secrets):

blastevict analyze output — ranked eviction report
$ blastevict analyze examples/sample_session.jsonl

 ─ blastevict v0.2.0 ─
 examples/sample_session.jsonl
 12 events · 6 persisted tool outputs

 Blast-radius eviction report
 #  tool    origin  depth  tok/turn  footprint  overlap  verdict
 1  Bash       6       3       224       672     —     pending
 2  Read       4       4        89       356     —     pending
 3  Read       2       5        49       245     —     pending
 4  Grep       2       5        33       165     —     pending
 5  Bash      10       1        39        39     —     pending
 6  Edit       8       2        17        34     —     pending

 Cumulative footprint: 1511 token-turns across 6 outputs.
 Verdicts: 6 pending.

 replay not run (ANTHROPIC_API_KEY not set): footprints are exact;
 verdicts are pending, not replay-proven.

With ANTHROPIC_API_KEY set, the top-K rows carry replay-proven safe-to-evict / load-bearing / uncertain verdicts and an overlap column showing the action-distribution similarity. An asciinema cast of a full run is shipped at assets/demo.cast.

Usage

blastevict analyze SESSION.jsonl [OPTIONS]

  SESSION.jsonl              Path to a Claude Code session JSONL.

  --top-k INTEGER             Replay only the top-K footprint candidates.   [default: 5]
  --n-samples INTEGER          Replay samples per candidate (temp=0).       [default: 3]
  --no-replay                  Skip replay; footprint report only.
  --json                       Emit machine-readable JSON.
  --help                       Show this message and exit.

Verdict reference:

Verdict Meaning
safe-to-evict Replay matched the original decision under ablation — removing the output did not change the next move.
load-bearing Replay diverged — removing the output changed the decision. Keep it.
uncertain Samples disagreed (high variance), or overlap sat between thresholds. Never a false safe-to-evict.
pending Replay not run (no API key or below top-K). Footprint is exact; verdict is not replay-proven.

Machine-readable output (--json) mirrors docs/demo-results.json.

Capabilities & integration

blastevict capabilities — Claude Code JSONL in, replay verdicts out
  • Input: Claude Code session JSONL (~/.claude/projects/<session>.jsonl). v0.1 supports Claude Code transcripts only — Cursor / Aider / generic agent-transcript adapters are post-v0.1.
  • Replay backend: the Anthropic messages API, at the model that produced the original session.
  • Output: a ranked rich-text report to stdout, or JSON via --json.
  • Mode: read-only post-mortem analysis. It tells you what to cut; it does not auto-apply evictions to a running session.

Roadmap

Milestone Status What it proves
m1 — trace + replay done Ingest a session, build the blast-radius footprint, run one replay ablation, print the decision-diff. The determinism prototype.
m2 — evict report done (this release) Full blastevict analyze ranked report — every output with footprint, propagation depth, and a replay-proven verdict. The shippable MVP.
m3 — ship + Show HN stub README + bilingual sibling + demo cast; Show HN post.

Planned post-v0.1: Cursor / Aider adapters; live in-session eviction once the safety claim holds on real transcripts; richer action-distribution metrics beyond tool-name Jaccard.

Kill criteria (stated up front)

  • Determinism kill (pre-launch): if no stable decision-diff emerges from replay ablation on 3+ real Claude Code sessions at temp=0 / N=3, the central safety claim ("eviction is replay-proven safe") is unprovable — kill regardless of other signal.
  • Star kill (post-launch): after 30 days on GitHub with 3 launch activities, fewer than 50 stars AND no organic user files an issue describing the eviction decision → kill.
  • Demand kill: if no interviewee or issue-filer describes the eviction decision unprompted within 60 days → the primitive solves a problem users haven't articulated; kill or re-scope.

License

MIT — Copyright (c) 2026 SuperMarioYL. See LICENSE.

Contributors

SuperMarioYL

Issues