EIP-8037 multiplier search appears capped too low — 99% of gas-needy txs report 'unresolved'

#4 · closed · 3 comments

View on GitHub ↗

CarlBeek

## Summary Of all transactions that don't fit their original gas limit under the EIP-8037 schedule, roughly **99% have `min_multiplier_to_succeed = NULL`** in the replay output. Combined with the fact that no resolved multiplier exceeds **~1.81x** anywhere in the corpus, this strongly suggests the replay's multiplier search is capped at or near 2.00x and gives up rather than finding the true minimum for the long tail. ## Observed Live response from `https://repricing-forensics.carlbeek.com/api/eip8037/multiplier-histogram`: | Bucket | Txs | Status-changed txs | Max multiplier in bucket | |------------------|----------:|-------------------:|-------------------------:| | `fits original` | 1,387,313 | 0 | 1.00 | | `1.00–1.25x` | 1,063 | 1,063 | 1.2495 | | `1.25–1.50x` | 841 | 841 | 1.4999 | | `1.50–2.00x` | 414 | 414 | 1.8129 | | `unresolved` | 358,242 | 300,762 | `null` | So among the 360,560 txs that need more gas under EIP-8037: - **2,318 (0.6%)** got a resolved multiplier — all under 2.00x - **358,242 (99.4%)** report `unresolved` The same disparity is even sharper among status-changed txs (baseline succeeded, schedule failed): **2,318 resolved vs. 300,762 unresolved.** A `1.50–2.00x` bucket that exists (414 txs) but no `2.00–4.00x` / `4.00–8.00x` / `>8.00x` buckets, paired with a `max_multiplier` of 1.8129 across the entire corpus, looks like the search topped out around 2x rather than the data genuinely lacking higher multipliers. ## Hypothesis The replay tries multipliers up to some cap (looks to be ~2.00x) and records `min_multiplier_to_succeed = NULL` if no multiplier in that range works, instead of widening the search. Two cases are worth distinguishing: 1. **Search cap is the cause.** Bumping the cap to e.g. 8x or 16x would resolve a large share of the current `unresolved` cohort. This is the case I expect dominates under EIP-8037, since most state-gas shortfalls are linear in the gas limit (bigger tx gas limit → bigger reservoir → enough headroom). 2. **Genuine non-gas failure.** A subset will be contract-broken txs (hardcoded gas in inner `CALL`s, etc.) that no tx-level multiplier can fix. These should remain `unresolved` — that's the correct outcome. The current implementation conflates both cases under the same `unresolved` label, so the downstream analysis can't tell whether \"give it more gas\" works or not for the bulk of affected contracts. ## Why this matters Audience: contract teams using https://repricing-forensics.carlbeek.com to assess whether their contracts will work post-Glamsterdam, and client devs evaluating the proposal. For 99% of affected txs, the per-contract diagnose page and the multiplier-histogram chart can't answer the basic question \"how much do I need to bump my gas limit by?\". The site has been pushing that as the headline framing for EIP-8037 impact, but the data is largely empty. ## Suggested next steps 1. Confirm the upper bound of the multiplier search in the replay code path that fills `min_multiplier_to_succeed` / `extra_gas_needed` (likely a constant or a loop bound in the EIP-8037 replay routine). 2. Either widen the cap (e.g. 16x or until total gas exceeds the block gas limit), or split `unresolved` into two distinct outputs: - `search_exhausted` — search hit its cap without finding a working multiplier - `non_gas_failure` — replay confirmed the tx fails for a reason no gas bump can fix 3. If widening the cap is prohibitively slow, consider exponential-then-binary-search (try 2x, 4x, 8x, 16x, then narrow) rather than a fine-grained linear sweep. ## Notes - The same data column (`min_multiplier_to_succeed`) is what powers the per-contract `p95 multiplier` and `max multiplier` stats on the affected page; with this much null data those percentile aggregates are nearly meaningless. - Downstream tracking: https://repricing-forensics.carlbeek.com/eip8037 — section 3 (multiplier histogram) and the per-contract diagnose page severity numbers. - Related: #3 (forensics writer placeholder values).

Comments

CarlBeek

Addressed in c14e6719f (`feat(research): add replay_halt_oog to distinguish search-exhausted txs`). Confirmed your hypothesis: the cap is the `--research.gas-limit-multiplier` flag (default `1`, you'd been running `2`). When the inflated replay still OOGs, `schedule_replay_success = false` and `min_multiplier_to_succeed` falls through to `NULL` — collapsing the "search exhausted" and "non-gas failure" cases together. Rather than picking a higher fixed cap or building an exponential search (both are expensive in re-execution time), the fix splits the unresolved bucket via a new column on `schedule_divergences` and the parquet hot export: | `min_multiplier_to_succeed` | `replay_halt_oog` | meaning | | --- | --- | --- | | `Some(x)` | `None` | needs `x * tx_gas_limit` (replay succeeded under inflation) | | `None` | `Some(true)` | **search exhausted** — replay OOG'd at the inflated budget; the true minimum exceeds `--research.gas-limit-multiplier` | | `None` | `Some(false)` | **non-gas failure** — replay halted for a non-gas reason or reverted; no extra gas resolves it | So your pipeline can now do `WHERE replay_halt_oog = 1` to recover the "needs more gas than we tried" cohort and re-run the corpus with a higher multiplier (16x, 32x) only for those rows. The `non_gas_failure` rows can stay where they are. A schema migration (`ALTER TABLE schedule_divergences ADD COLUMN replay_halt_oog BOOLEAN`) auto-applies on existing databases. New parquet exports include a nullable `replay_halt_oog` Boolean field. Implementation note: the EVM result's halt reason is generic (`EvmFactory::HaltReason`), so we can't pattern-match the concrete `revm::HaltReason` enum directly. The check uses `Debug` formatting and string-matches the `OutOfGas` prefix, which is stable across all spec-conformant `HaltReason` impls (they share a `From<HaltReason>` bound).

CarlBeek

Addressed in c14e6719f via the new replay_halt_oog column.

CarlBeek

Confirmed by a comprehensive data audit (`/api/_debug/data-audit`): the multiplier-cap symptom shows up in three downstream fields, all uniformly null/zero: | Table | Field | Verdict | Populated | |---|---|---|---:| | `eip8037_tx_impact` | `extra_gas_needed` | UNPOPULATED | 0 / 1,899,013 | | `eip8037_contract_impact` | `max_extra_gas_needed` | UNPOPULATED | 0 / 56,073 contracts | | `eip8037_contract_impact` | `fixable_with_more_outer_gas` | CONSTANT 0 | all 56,073 contracts | `extra_gas_needed` being uniformly NULL (rather than holding per-tx values) confirms the search exits without recording the field at all — there's no "search reached cap" sentinel, just no value. Combined with the `max_multiplier` of 1.81x ceiling already noted in the issue body, this means the consumer can't: - show how much extra gas a tx needs (per-tx number is missing) - aggregate to a per-contract "max needed" (always NULL) - count contracts that are fixable by upping the gas limit (always 0, even though some clearly are — there are 2,318 txs with a known sub-2x multiplier) Splitting `unresolved` into `search_exhausted` vs `non_gas_failure` (per the original suggested next step) plus recording the search bound that was reached would let the consumer surface "search-cap reached at multiplier=N, give us X extra gas to retry" rather than the current "we don't know."