CUDA `tile_assign()` loses gradients from broadcast source tiles

#2001 · open · 1 comments

View on GitHub ↗

shi-eric

### Bug Description `wp.tile_assign()` produces correct forward values when its source is a `wp.tile_broadcast()` view, but the input gradient is incorrect on CUDA with `block_dim=32`. ### Reproducer ```python import numpy as np import warp as wp @wp.kernel def direct_broadcast(x: wp.array2d[float], y: wp.array2d[float]): src = wp.tile_load(x, shape=(1, 3), storage="shared") broadcast = wp.tile_broadcast(src, shape=(2, 3)) wp.tile_store(y, broadcast) @wp.kernel def assigned_broadcast(x: wp.array2d[float], y: wp.array2d[float]): src = wp.tile_load(x, shape=(1, 3), storage="shared") broadcast = wp.tile_broadcast(src, shape=(2, 3)) dest = wp.tile_zeros(shape=(2, 3), dtype=float, storage="shared") wp.tile_assign(dest, broadcast) wp.tile_store(y, dest) seed = np.array([[1, 2, 3], [4, 5, 6]], dtype=np.float32) for device in ("cpu", "cuda:0"): for kernel in (direct_broadcast, assigned_broadcast): x = wp.array([[1.0, 2.0, 3.0]], dtype=float, device=device, requires_grad=True) y = wp.zeros((2, 3), dtype=float, device=device, requires_grad=True) with wp.Tape() as tape: wp.launch_tiled(kernel, dim=1, inputs=[x], outputs=[y], block_dim=32, device=device) y.grad = wp.array(seed, dtype=float, device=device) tape.backward() print(device, kernel.key, "output:", y.numpy(), "input gradient:", x.grad.numpy()) ``` ### Expected and Actual Results Both kernels produce the expected forward output, `[[1, 2, 3], [1, 2, 3]]`. Each input element is used by both output rows, so the expected input gradient is `[[5, 7, 9]]`. ```text cpu direct_broadcast: [[5. 7. 9.]] cpu assigned_broadcast: [[5. 7. 9.]] cuda:0 direct_broadcast: [[5. 7. 9.]] cuda:0 assigned_broadcast: [[1. 2. 3.]] ``` The CUDA results were consistent across three runs. The assigned kernel also produced the expected gradient with a one-thread CUDA block. ### System Information Warp 1.19.0.dev0, source commit `6678147e6e916f1f4bae097ad3c60bd292fe9c41`; Python 3.12.13; Linux x86_64; CUDA Toolkit 13.0; NVIDIA driver 595.58.03; NVIDIA RTX PRO 6000 Blackwell Server Edition MIG (`sm_120`).

Comments

kvnloo

I can validate this independently on Ampere without competing with the assigned fix. The useful matrix looks like: - direct `tile_broadcast` -> `tile_store`; - `tile_broadcast` -> `tile_assign` -> `tile_store`; - `block_dim=1`; - `block_dim=32`; - one larger block size if supported. For a source row used by two destination rows, the backward oracle is the sum of both rows' upstream gradients. So with: `[[1,2,3],[4,5,6]]` the source gradient must be: `[[5,7,9]]` for both execution forms. If only the assigned form loses the second contribution, that isolates the bug to gradient accumulation through `tile_assign`, not broadcast itself. I can run that CPU/CUDA matrix on the implementation branch.