Replace the brittle Fixed/Fractional OOG-chain heuristic with the gas-limit multiplier analysis (EIP-8037-aware replay)

#11 · open · 0 comments

View on GitHub ↗

misilva73

## Summary The wallet-fixable vs `contract_broken` decision currently leans on a static gas-forwarding chain-walk in `crates/research/src/oog_chain.rs` (`classify_chain` → `proportional` / `Stipend2300` / `FixedGas` / `FractionalGas`). That heuristic is brittle. In parallel, the producer already records an empirical fixability signal — `would_fit_in_original_limit` + `min_multiplier_to_succeed` — by replaying the tx at a higher gas limit and observing whether it succeeds. This issue proposes two changes: 1. Make the gas-limit multiplier analysis the **primary fixability classifier** (for every repricing schedule, not just 8037), and demote the chain-walk to a **diagnostic** role. 2. Make the fixability replay's gas-limit headroom **schedule-conditional**: keep the configurable max multiplier per run, but for the 8037 schedule let `tx.gas` exceed the EIP-7825 16.7M cap (which activates 8037's already-implemented state reservoir), while every other schedule stays clamped at 16.7M. ## 1. The Fixed/Fractional chain-walk heuristic is brittle In `classify_chain` (`oog_chain.rs`): ```rust const PROPORTIONAL_TOLERANCE: u64 = 100; const FIXED_GAS_THRESHOLD: u64 = 100_000; // ... let cap = parent_gas.saturating_mul(63) / 64; if stack_gas >= cap.saturating_sub(PROPORTIONAL_TOLERANCE) { continue; } // proportional let kind = if stack_gas == 2300 { Stipend2300 } else if stack_gas < FIXED_GAS_THRESHOLD { FixedGas } else { FractionalGas }; ``` Problems: - **Arbitrary constants.** The `FixedGas (< 100_000)` / `FractionalGas (≥ 100_000)` split is a guess; nothing about EVM semantics anchors 100k. - **Non-scaling tolerance → real misclassification.** The proportional test uses a *fixed absolute* tolerance (100 gas). A relative-reserve forward like `target.call{gas: gasleft() - K}(...)` reads as `proportional` (wallet-fixable) when the parent frame is gas-rich, but flips to `FixedGas`/`FractionalGas` (→ `contract_broken`) once `parent_gas` drops below ~`64·K` — even though the pattern is genuinely wallet-fixable. The decision boundary is `parent_gas ≥ 64·(K − tol)`, i.e. it tracks the frame's budget rather than the contract's intent. - **Fractional forwards are mislabeled.** `call{gas: gasleft() / 2}` is tagged `contract_broken` even though it scales with outer gas (so it is at least partially wallet-fixable). ### Proposal Use the **gas-limit multiplier analysis as the fixability oracle**. Because it actually re-runs the tx at a higher limit and observes the outcome, it has none of the static blind spots above. The chain-walk's `proportional`/throttled flag stops *deciding* the bucket and instead only *diagnoses* the `contract_broken` cohort (which frame throttled, and how). ### What this does NOT replace (keep these) - **Diagnosis, not just a verdict.** A failed replay is a single `NULL`, which lumps {structural throttle, non-gas revert, pre-execution rejection} together. The chain-walk + failing-leaf frame analysis is what tells them apart, and it powers the `bottleneck-kinds` / `failure-flow` forensics views. Keep it as the explainer for the broken cohort. - **`Stipend2300` is reliable.** It's an exact `== 2300` match on a protocol constant, not a tunable threshold — a trustworthy diagnostic worth keeping. The brittleness is specifically the `FixedGas`/`FractionalGas` split and the proportional tolerance. - **Non-gas divergences.** The multiplier method only speaks to OOG/gas-limit failures. Status flips from refund accounting, EIP-7623 floor gas, control-flow reverts, and event-log changes still need the outcome/trace ladder. ## 2. Make the fixability replay's headroom schedule-conditional The replay that produces `min_multiplier_to_succeed` already sweeps inflated gas limits (via `--research.gas-limit-multipliers`). What it must change is **how high `tx.gas` is allowed to go, conditional on the schedule under analysis** — so each schedule is judged under its own gas model rather than a single shared one. The reason this matters: EIP-7825 caps `tx.gas` at `2^24 = 16,777,216`. EIP-8037 sizes a **per-tx state reservoir** as a function of `tx.gas`: ```text initial_reservoir = max(0, tx.gas − intrinsic_state_gas − TX_MAX_GAS_LIMIT) // TX_MAX_GAS_LIMIT = 16,777,216 ``` That `max(0, …)` floor is why the reservoir reads as `0` for every historical tx today — 7825 pins `tx.gas` at or below the cap, so the subtraction never goes positive. **The reservoir mechanism is already implemented in the schedule's execution path; it is dormant, not missing.** Feeding the 8037 path a `tx.gas` above the cap is what activates it, and the path sizes the reservoir itself. Proposed mechanism: - **Keep the configurable max multiplier per run** (e.g. `10×` / `20×` of the original gas limit), exposed as a run parameter rather than hardcoded. - **Set the replay's `tx.gas` conditional on the schedule:** - **8037 schedule:** `tx.gas = original_limit × multiplier`, and **relax the EIP-7825 tx-gas-limit validation in the replay** so the inflated tx executes instead of being rejected pre-execution. The 8037 path then derives `initial_reservoir` from `tx.gas` per the formula above — execution gas stays bounded by 7825 internally, the surplus above the cap funds the reservoir, and state-creation charges (CPSB/byte) draw from it; demand beyond it is the spillover / OOG. - **Every other schedule (7904, …):** `tx.gas = min(original_limit × multiplier, 16_777_216)` — stay clamped at the 7825 cap. No reservoir; a single execution pool. > **Verify in this repo before implementing.** Confirm whether the existing sweep already lifts `tx.gas` above 16.7M for 8037, or whether it is silently 7825-capped today. If the sweep is currently capped, the 8037 reservoir never activates even in the sweep, and `min_multiplier_to_succeed` for state-bound txs is being computed against a 16.7M execution wall — i.e. wrong. Relaxing 7825 for the 8037 sweep is then the substantive fix, not a refinement. ### Consequences for classification - **Execution-bound and state-bound OOGs separate cleanly under 8037.** Once execution gas is pinned at the 16.7M cap, a higher multiplier only grows the reservoir. So an execution-OOG is broken regardless of multiplier, while a state-OOG stays fixable by a larger reservoir — a distinction a single-pool limit can't express. - **`inconclusive_needs_higher_sweep` becomes 8037-specific.** For 8037, a tx still OOG at the configured max multiplier is "inconclusive at this ceiling" — raising the multiplier is the operator's lever, and the bucket records txs that exhausted it. For a clamped schedule, the per-tx ceiling is structural (16.7M, reached at `multiplier ≈ 16.7M / original_limit`); a tx still OOG there cannot be rescued by a higher multiplier, so it is `contract_broken` (execution-bound), **not** `inconclusive`. Don't raise `inconclusive_needs_higher_sweep` for clamped schedules. - **The `aa_gas_reestimation` carve-out must not regress.** That bucket is currently split out of `contract_broken` using the chain-walk's `FixedGas` verdict on ERC-4337 EntryPoint OOGs. Once `classify_bucket` stops consuming the chain-walk (§1), AA detection needs its own signal (e.g. recipient == EntryPoint plus "would fit at a higher budget, but the UserOp gas is signed and fixed"). Don't let the demotion silently re-merge AA into `contract_broken`. ### Caveat on the multiplier value `min_multiplier ≈ schedule_gas_used / original_limit` is a *lower-bound* approximation: the 63/64 forwarding reservation (and any `gasleft()`-dependent execution) means the limit needed to avoid a deep-leaf OOG can exceed total gas used. Under 8037 the shortfall can be in execution or in the reservoir, which scale differently with the multiplier. An exact minimum needs a search; approximate bucketing is cheap and likely sufficient. Whichever is chosen should be documented. ## Net change - Lead with the **multiplier analysis** as the fixability classifier for every repricing schedule. The columns (`would_fit_in_original_limit`, `min_multiplier_to_succeed`) are already schedule-agnostic on `divergences`. - Make the replay's `tx.gas` headroom **schedule-conditional**: for 8037, inflate `tx.gas` to `original × multiplier` and relax the 7825 validation so the schedule's own (already-implemented) reservoir activates; for every other schedule, clamp `tx.gas` at the 16.7M 7825 cap. Keep the max multiplier configurable. - Keep the chain-walk (especially `Stipend2300`) as **diagnostic annotation** on the broken cohort, not as the classifier — without regressing the `aa_gas_reestimation` carve-out that currently rides on it. ## Affected files - `crates/research/src/oog_chain.rs` — `classify_chain` thresholds: demote from classifier to diagnostic (keep computing it for annotation / forensics views). - `crates/research/src/divergence.rs` — `classify_bucket`: route the fixability decision through the multiplier signal first; apply the per-schedule ceiling rule (`inconclusive_needs_higher_sweep` only for 8037, structural-`contract_broken` for clamped schedules); give `aa_gas_reestimation` a chain-walk-independent signal. - `crates/research/bin/reth-research/src/main.rs` — set the replay's `tx.gas` per schedule (8037: `original × multiplier` with 7825 validation relaxed so the reservoir activates; others: `min(original × multiplier, 16_777_216)`), and populate `would_fit_in_original_limit` / `min_multiplier_to_succeed` for all schedules. Do **not** compute a `gas_left`/`reservoir` split here — the 8037 path sizes the reservoir from `tx.gas` itself. ## Downstream changes in the forensics (consumer) repo These are **not** part of this issue — they land separately in the consumer once the producer change ships — but flagged so the producer change is reviewed with them in mind: - **Presentation only; no schema change.** `would_fit_in_original_limit` / `min_multiplier_to_succeed` and the reservoir columns already exist on `divergences`, and the consumer only *reads* `bucket` (it never reclassifies). The fixability/multiplier surfacing is 8037-scoped today (the `eip8037_*` API endpoints and `eip8037.html`); generalizing it to 7904 is a UI change, not a data-model one. - **Per-schedule ceiling copy.** The "reservoir activation (dormant today)" framing flips live only on the 8037 view. For clamped schedules the story is "still OOG at the 16.7M execution cap → broken," with no reservoir narrative — different copy, same plumbing. - **Optional sweep-level reservoir telemetry.** The existing reservoir columns reflect the *normal* replay (`= 0` historically). If the consumer wants to show "at multiplier `m`, reservoir was X / spillover Y" for the activated 8037 sweep, the producer would need to emit *per-sweep* reservoir telemetry — a new-column decision to make explicitly. If the scalar `min_multiplier_to_succeed` suffices, no addition is needed.

Comments