andreasnoack
## Summary With a `ComposedFunction` at the root of a ForwardDiff gradient, any `float ∘ norm` reached below it (LinearAlgebra's `generic_norm1`) with a nested `Dual` argument makes the recursion limiter in `abstract_call_method` treat the two unrelated `call_composed` frames as a growing recursion. It widens the inner argument to `Dual{T,V,N} where {T,V,N}`, and `poison_callstack!` marks every frame up to the root `LimitedAccuracy`, so nothing in between is cached and every call site re-infers the subtree. The cost is the product of call-site counts along the chain. The same computation with a closure at the root is unaffected. Self-contained reproducer (ForwardDiff + LinearAlgebra, below), first gradient, K = 4 levels of 4 call sites each: | Julia | closure root | `ComposedFunction` root | |---|---|---| | 1.12.7 | 2.2 s, 0.62 GiB | 56.1 s, 28.9 GiB | | 1.13.0 | 6.0 s, 0.63 GiB | 61.8 s, 26.1 GiB | | 1.13.1 | 6.1 s, 0.63 GiB | 49.7 s, 25.2 GiB | | 1.14.0-DEV.3428 | 6.4 s, 0.62 GiB | 132.7 s, 41.3 GiB | K = 2 on 1.13.1: 5.9 s vs 8.8 s; the gap grows with the depth of the call chain below the widened call. With a print in `abstract_call_method` where `newsig !== sig` and in `finishinfer!` where a limited result is discarded, the K = 4 composed run logs 380,288 widenings of `call_composed` (all the same `(float, norm)` site against the root) and 1,246,886 discarded frames before I stopped it. ## Mechanism (from the instrumented v1.13.1 Compiler, loaded as the `Compiler` package) Every widening event has this shape: - `sig = call_composed((float, norm), (Dual{Tag{EtaTag}, Dual{Tag{ThetaTag}, Dual{Tag{root}, Float64, 11}, 1}, 1},), kw)` - `cmp = call_composed((sum, g), (Vector{Dual{Tag{root}, Float64, 11}},), kw)` — the user's root composition - `new = call_composed((float, norm), (Dual{T,V,N} where {T,V,N},), kw)` - stack, root first: `gradient → vector_mode_dual_eval! → ComposedFunction → #_#NN → call_composed → g → jacobian → … → jacobian → model → norm → norm1 → generic_norm1 → mapreduce → … → ComposedFunction → #_#NN` The limiter walks the stack for a frame of the same method (`call_composed`); the root frame qualifies, and `edge_matches_sv`'s soft-limit test passes because the root frame's parent (`#_#NN`, the `ComposedFunction` kwcall body) is the same method as the inner call's caller. `type_more_complex` then compares `Tuple{Dual{Dual{Dual}}}` with `Tuple{Vector{Dual}}`, sees the deeper nesting, and `limit_type_size` returns the widened signature. Neither call is recursive in any real sense: different `fs` tuples, finite depth. ## Reproducer ```julia using ForwardDiff, LinearAlgebra # ARGS: composed | closure ; K variant = ARGS[1]; const K = length(ARGS) >= 2 ? parse(Int, ARGS[2]) : 4 struct ThetaTag end; struct EtaTag end const Tag = ForwardDiff.Tag ForwardDiff.:≺(::Type{Tag{F1,V1}}, ::Type{Tag{ThetaTag,V2}}) where {F1,V1,V2} = true ForwardDiff.:≺(::Type{Tag{ThetaTag,V2}}, ::Type{Tag{F1,V1}}) where {F1,V1,V2} = false ForwardDiff.:≺(::Type{Tag{F1,V1}}, ::Type{Tag{EtaTag,V2}}) where {F1,V1,V2} = true ForwardDiff.:≺(::Type{Tag{EtaTag,V2}}, ::Type{Tag{F1,V1}}) where {F1,V1,V2} = false ForwardDiff.:≺(::Type{Tag{ThetaTag,V1}}, ::Type{Tag{EtaTag,V2}}) where {V1,V2} = true ForwardDiff.:≺(::Type{Tag{EtaTag,V1}}, ::Type{Tag{ThetaTag,V2}}) where {V1,V2} = false ForwardDiff.:≺(::Type{Tag{ThetaTag,V1}}, ::Type{Tag{ThetaTag,V2}}) where {V1,V2} = false # same tag: resolve the latent ambiguity of the two rules above ForwardDiff.:≺(::Type{Tag{EtaTag,V1}}, ::Type{Tag{EtaTag,V2}}) where {V1,V2} = false jac(f, x, tag) = ForwardDiff.jacobian(f, x, ForwardDiff.JacobianConfig(tag, x, ForwardDiff.Chunk{1}()), Val{false}()) leaf(A) = norm(vec(A), 1) # norm1 == mapreduce(float ∘ norm, +, x): the inner ComposedFunction work(A, s) = (A * A) .* s .+ transpose(A) ./ (1 + s * s) level(::Val{0}, A, s, i) = leaf(work(A, s)) function level(::Val{k}, A, s, i) where {k} # K levels, 4 distinct call sites each B = work(A, s) r = i % 4 if r == 0; return level(Val(k - 1), B, s + 1, i + 1) elseif r == 1; return level(Val(k - 1), B .* 2, s + 2, i + 1) elseif r == 2; return level(Val(k - 1), B .+ 1, s + 3, i + 1) else; return level(Val(k - 1), B .- 1, s + 4, i + 1) end end function model(θ, η, t) A = [-θ[1]*exp(η[1]) 0.0 0.0; θ[1] -θ[2]*exp(η[2])/θ[3] 0.0; 0.0 θ[2] -θ[3]] return [level(Val(K), A .* t[i], t[i], i) for i in eachindex(t)] end const θ0 = [1.5, 1.0, 30.0] g(t) = (J = jac(θ -> vec(jac(η -> model(θ, η, t), zeros(2), EtaTag())), θ0, ThetaTag()); J' * J) obj = variant == "composed" ? ComposedFunction(sum, g) : (t -> sum(g(t))) t = collect(range(0.5, 24.0, length = 21)) obj(t) @time ForwardDiff.gradient(obj, t) ``` Notes on the ingredients: - The explicit tags with a fixed ordering are what libraries doing nested AD typically define. With ForwardDiff's default tags the inner closures capture the outer `Vector{Dual}`, the inner tags embed the outer `Dual` type, and the limiter instead widens `ForwardDiff.jacobian` itself, which cuts the stack and hides the pattern in this small example (0 `call_composed` candidates). In the larger application below, default tags are worse on both roots (closure 39 s vs 18 s; composed 241 s vs 57-90 s). - The `init`-free `mapreduce` inside `norm1` is the only place a second `ComposedFunction` is on the stack. Any user code with `f ∘ g` below a composed root does the same. - The amplification is ordinary code structure: a function called from several sites inside an uncacheable region is re-inferred per site. - Type stability of the subtree is irrelevant: in the larger application the same composed objective costs 299 s / 194 GiB with every intermediate type concrete and 341 s / 194 GiB with the objective constructed as a concrete type inside a function, versus 12-17 s for a closure root. ## A larger application and the 1.13.1 angle I hit this in a larger application with a composed objective over a nested ForwardDiff Jacobian of an analytical ODE solution; its matrix exponential calls `opnorm`, hence `norm1`. First gradient: 370 s / 194 GiB with the composed root vs 60 s / 10 GiB with a closure root. There it appeared only with 1.13.1 and bisects (both directions, by runtime monkeypatch) to #62993. Before that change, `mapreduce(…; init)` went through `mapfoldl_impl`'s `identity ∘ Fix1(MappingRF, f)`, and the limiter happened to match that composition first, early in the tree, where the poisoned region was small; after it, the first match is the `norm1` site deep inside the solver. The reproducer above has no `init` reductions and shows the pattern on every version, so the heuristic issue is independent of #62993. ## Questions 1. Should the soft-limit test in `edge_matches_sv` match two `call_composed` frames just because both are called from the `ComposedFunction` kwcall body? The recursion variable of `call_composed` is `fs`, which strictly shrinks, so the recursion is well-founded; the limiter instead compares the unrelated argument `x` and reads a deeper `Dual` as growth. This looks like a false positive: inferring the inner call precisely terminates, and every frame in between would then be cacheable. 2. Could `limited_src` be narrowed to frames whose return type or optimization actually depends on the limited value, so that the "large multiplier" the comment in `finishinfer!` accepts stays small? **Disclosure:** the investigation (instrumented Compiler runs, bisection, reproducer) and this text were produced with [Claude Code](https://claude.com/claude-code) (Claude Fable 5.1), on behalf of @andreasnoack. I have reviewed the text.