TomAugspurger
Spilling data from device memory is relatively slow, and slows down subsequent operations that need to load the data back to device memory. We'd like to have better observability into spill (and unspill) events into cudf-polars. With this observability, we should be able to answer 1. How much data did each worker spill to complete this query? 2. How much spilled data did each worker need to unspill to complete this query? 3. Which operators were responsible for spilling? 4. What throughput did I achieve during spill / unspill operations? We'll need to ensure our Quent schema accurately reflects all the memory tiers (device, host, and soon disk). There's lots of details to work out: - We'll need to see how this interacts with https://github.com/NVIDIA/cudf/issues/24401; during a Shuffle, does a transfer to a peer's host memory count as spilling? - We'll need to coordinate with rapidsmpf to ensure we have the data necessary for tracking. We might need to track additional data on spilled buffers so that we can track which operator a buffer belongs to. - We'll need to figure out how this interacts with rapidsmpf's various spill triggers (memory reservations triggering spilling, vs. the periodic background spilling).