TomAugspurger
rapidsmpf's memory reservation system hands out an allocation budget that cudf-polars draws against. We'd like observability into that process, to answer questions like: 1. At any given time, what is the current total reservation on this Worker? 2 For any given reservation, what is its size and memory tier? What operator made this reservation? How long did it take to acquire the reservation? (maybe) how much of it have we drawn against? We'll model reservations as a new FSM in our Quent model. The states are queued, reserved, freed (or freeing?), and exit. The individual records will vary by the state the FSM is in. But we should somehow include - Operator ID - Reservation size - Reservation tier - Wait duration (or is this just computed from the duration between queued -> reserved)? along with the usual FSM attributes giving the time of each transition. We should figure out if there's any information about spilling that ought to be captured here. Does it make sense to ask "how many bytes were spilled as a result of this reservation?" If so, do we have any hope of acquiring that information, given rapidsmpf architecture? xref https://github.com/NVIDIA/cudf/pull/24038, though I might throw that out.