Inference recursion limiter matches an inner `ComposedFunction` against an unrelated outer one and poisons the whole subtree: 10x compile time, 40x allocations for a composed objective over nested ForwardDiff

#63516 · open · 1 comments

View on GitHub ↗

andreasnoack

## Summary With a `ComposedFunction` at the root of a ForwardDiff gradient, any `float ∘ norm` reached below it (LinearAlgebra's `generic_norm1`) with a nested `Dual` argument makes the recursion limiter in `abstract_call_method` treat the two unrelated `call_composed` frames as a growing recursion. It widens the inner argument to `Dual{T,V,N} where {T,V,N}`, and `poison_callstack!` marks every frame up to the root `LimitedAccuracy`, so nothing in between is cached and every call site re-infers the subtree. The cost is the product of call-site counts along the chain. The same computation with a closure at the root is unaffected. Self-contained reproducer (ForwardDiff + LinearAlgebra, below), first gradient, K = 4 levels of 4 call sites each: | Julia | closure root | `ComposedFunction` root | |---|---|---| | 1.12.7 | 2.2 s, 0.62 GiB | 56.1 s, 28.9 GiB | | 1.13.0 | 6.0 s, 0.63 GiB | 61.8 s, 26.1 GiB | | 1.13.1 | 6.1 s, 0.63 GiB | 49.7 s, 25.2 GiB | | 1.14.0-DEV.3428 | 6.4 s, 0.62 GiB | 132.7 s, 41.3 GiB | K = 2 on 1.13.1: 5.9 s vs 8.8 s; the gap grows with the depth of the call chain below the widened call. With a print in `abstract_call_method` where `newsig !== sig` and in `finishinfer!` where a limited result is discarded, the K = 4 composed run logs 380,288 widenings of `call_composed` (all the same `(float, norm)` site against the root) and 1,246,886 discarded frames before I stopped it. ## Mechanism (from the instrumented v1.13.1 Compiler, loaded as the `Compiler` package) Every widening event has this shape: - `sig = call_composed((float, norm), (Dual{Tag{EtaTag}, Dual{Tag{ThetaTag}, Dual{Tag{root}, Float64, 11}, 1}, 1},), kw)` - `cmp = call_composed((sum, g), (Vector{Dual{Tag{root}, Float64, 11}},), kw)` — the user's root composition - `new = call_composed((float, norm), (Dual{T,V,N} where {T,V,N},), kw)` - stack, root first: `gradient → vector_mode_dual_eval! → ComposedFunction → #_#NN → call_composed → g → jacobian → … → jacobian → model → norm → norm1 → generic_norm1 → mapreduce → … → ComposedFunction → #_#NN` The limiter walks the stack for a frame of the same method (`call_composed`); the root frame qualifies, and `edge_matches_sv`'s soft-limit test passes because the root frame's parent (`#_#NN`, the `ComposedFunction` kwcall body) is the same method as the inner call's caller. `type_more_complex` then compares `Tuple{Dual{Dual{Dual}}}` with `Tuple{Vector{Dual}}`, sees the deeper nesting, and `limit_type_size` returns the widened signature. Neither call is recursive in any real sense: different `fs` tuples, finite depth. ## Reproducer ```julia using ForwardDiff, LinearAlgebra # ARGS: composed | closure ; K variant = ARGS[1]; const K = length(ARGS) >= 2 ? parse(Int, ARGS[2]) : 4 struct ThetaTag end; struct EtaTag end const Tag = ForwardDiff.Tag ForwardDiff.:≺(::Type{Tag{F1,V1}}, ::Type{Tag{ThetaTag,V2}}) where {F1,V1,V2} = true ForwardDiff.:≺(::Type{Tag{ThetaTag,V2}}, ::Type{Tag{F1,V1}}) where {F1,V1,V2} = false ForwardDiff.:≺(::Type{Tag{F1,V1}}, ::Type{Tag{EtaTag,V2}}) where {F1,V1,V2} = true ForwardDiff.:≺(::Type{Tag{EtaTag,V2}}, ::Type{Tag{F1,V1}}) where {F1,V1,V2} = false ForwardDiff.:≺(::Type{Tag{ThetaTag,V1}}, ::Type{Tag{EtaTag,V2}}) where {V1,V2} = true ForwardDiff.:≺(::Type{Tag{EtaTag,V1}}, ::Type{Tag{ThetaTag,V2}}) where {V1,V2} = false ForwardDiff.:≺(::Type{Tag{ThetaTag,V1}}, ::Type{Tag{ThetaTag,V2}}) where {V1,V2} = false # same tag: resolve the latent ambiguity of the two rules above ForwardDiff.:≺(::Type{Tag{EtaTag,V1}}, ::Type{Tag{EtaTag,V2}}) where {V1,V2} = false jac(f, x, tag) = ForwardDiff.jacobian(f, x, ForwardDiff.JacobianConfig(tag, x, ForwardDiff.Chunk{1}()), Val{false}()) leaf(A) = norm(vec(A), 1) # norm1 == mapreduce(float ∘ norm, +, x): the inner ComposedFunction work(A, s) = (A * A) .* s .+ transpose(A) ./ (1 + s * s) level(::Val{0}, A, s, i) = leaf(work(A, s)) function level(::Val{k}, A, s, i) where {k} # K levels, 4 distinct call sites each B = work(A, s) r = i % 4 if r == 0; return level(Val(k - 1), B, s + 1, i + 1) elseif r == 1; return level(Val(k - 1), B .* 2, s + 2, i + 1) elseif r == 2; return level(Val(k - 1), B .+ 1, s + 3, i + 1) else; return level(Val(k - 1), B .- 1, s + 4, i + 1) end end function model(θ, η, t) A = [-θ[1]*exp(η[1]) 0.0 0.0; θ[1] -θ[2]*exp(η[2])/θ[3] 0.0; 0.0 θ[2] -θ[3]] return [level(Val(K), A .* t[i], t[i], i) for i in eachindex(t)] end const θ0 = [1.5, 1.0, 30.0] g(t) = (J = jac(θ -> vec(jac(η -> model(θ, η, t), zeros(2), EtaTag())), θ0, ThetaTag()); J' * J) obj = variant == "composed" ? ComposedFunction(sum, g) : (t -> sum(g(t))) t = collect(range(0.5, 24.0, length = 21)) obj(t) @time ForwardDiff.gradient(obj, t) ``` Notes on the ingredients: - The explicit tags with a fixed ordering are what libraries doing nested AD typically define. With ForwardDiff's default tags the inner closures capture the outer `Vector{Dual}`, the inner tags embed the outer `Dual` type, and the limiter instead widens `ForwardDiff.jacobian` itself, which cuts the stack and hides the pattern in this small example (0 `call_composed` candidates). In the larger application below, default tags are worse on both roots (closure 39 s vs 18 s; composed 241 s vs 57-90 s). - The `init`-free `mapreduce` inside `norm1` is the only place a second `ComposedFunction` is on the stack. Any user code with `f ∘ g` below a composed root does the same. - The amplification is ordinary code structure: a function called from several sites inside an uncacheable region is re-inferred per site. - Type stability of the subtree is irrelevant: in the larger application the same composed objective costs 299 s / 194 GiB with every intermediate type concrete and 341 s / 194 GiB with the objective constructed as a concrete type inside a function, versus 12-17 s for a closure root. ## A larger application and the 1.13.1 angle I hit this in a larger application with a composed objective over a nested ForwardDiff Jacobian of an analytical ODE solution; its matrix exponential calls `opnorm`, hence `norm1`. First gradient: 370 s / 194 GiB with the composed root vs 60 s / 10 GiB with a closure root. There it appeared only with 1.13.1 and bisects (both directions, by runtime monkeypatch) to #62993. Before that change, `mapreduce(…; init)` went through `mapfoldl_impl`'s `identity ∘ Fix1(MappingRF, f)`, and the limiter happened to match that composition first, early in the tree, where the poisoned region was small; after it, the first match is the `norm1` site deep inside the solver. The reproducer above has no `init` reductions and shows the pattern on every version, so the heuristic issue is independent of #62993. ## Questions 1. Should the soft-limit test in `edge_matches_sv` match two `call_composed` frames just because both are called from the `ComposedFunction` kwcall body? The recursion variable of `call_composed` is `fs`, which strictly shrinks, so the recursion is well-founded; the limiter instead compares the unrelated argument `x` and reads a deeper `Dual` as growth. This looks like a false positive: inferring the inner call precisely terminates, and every frame in between would then be cacheable. 2. Could `limited_src` be narrowed to frames whose return type or optimization actually depends on the limited value, so that the "large multiplier" the comment in `finishinfer!` accepts stays small? **Disclosure:** the investigation (instrumented Compiler runs, bisection, reproducer) and this text were produced with [Claude Code](https://claude.com/claude-code) (Claude Fable 5.1), on behalf of @andreasnoack. I have reviewed the text.

Comments

andreasnoack

Follow-up: the table in the issue shows the heuristic's cost (composed vs closure root) but not the 1.13.1 change, because the reproducer has no `init` reduction. Adding one above the deep subtree reproduces the regression without any package beyond ForwardDiff. The only change to `g` is: ```julia function g(t) s = t[end] tmax = mapreduce(x -> x / s, max, t; init = zero(eltype(t))) # `init` reduction with a non-identity closure, above the deep subtree ts = t .* tmax J = jac(θ -> vec(jac(η -> model(θ, η, ts), zeros(2), EtaTag())), θ0, ThetaTag()) return J' * J end ``` First gradient, K = 4, same machine: | Julia | closure root | `ComposedFunction` root | |---|---|---| | 1.12.7 | 2.0 s, 0.64 GiB | 1.5 s, 0.55 GiB | | 1.13.0 | 5.9 s, 0.64 GiB | 5.9 s, 0.56 GiB | | 1.13.1 | 5.7 s, 0.64 GiB | 45.8 s, 25.2 GiB | | 1.14.0-DEV.3428 | 6.8 s, 0.63 GiB | 158.5 s, 41.4 GiB | Mechanism, from the instrumented Compiler: on 1.13.0 the `mapreduce(…; init)` goes through `mapfoldl_impl`, whose `identity ∘ Fix1(MappingRF, f)` is the first `call_composed` the limiter meets below the root. Its `ComposedFunction{identity, Fix1{…, closure{Dual}}}` argument is judged more complex than the root's `ComposedFunction{typeof(sum), typeof(g)}` and is widened to `ComposedFunction{O, I} where {O, I}`, so the reduction's result is imprecise, `ts` and everything computed from it is `Any`, and the deep subtree is reached by dynamic dispatch from a fresh root with no `call_composed` on its stack (9 widenings in the whole gradient, none of them `float ∘ norm`, 39 discarded frames). After #62993 that `mapreduce` is an inline loop, inferred precisely, the whole subtree is inferred under the composed root, and `float ∘ norm` inside `norm1` is widened against the root at every call site (384,000 `float ∘ norm` widenings against the root, 1,259,097 discarded frames). So the shield before 1.13.1 was an unrelated widening that happened to cut the inference tree early; #62993 removed it and exposed the `norm1` match. **Disclosure:** runs and text produced with [Claude Code](https://claude.com/claude-code) (Claude Fable 5.1) on behalf of @andreasnoack. I have reviewed the text.