TomAugspurger
Data movement, especially between ranks / workers, can be an expensive component of executing a cudf-polars query. To better understand the execution of queries, we'd like observability of when and how data moves between ranks of a cudf-polars `StreamingEngine`. In particular, we'd like - Potentially some indication on the Quent timeline view about when transfers are happening and how much bandwidth they're consuming. - Summary statistics on transfers (count of transfers, average size, average duration, etc.) This will let us answer questions like 1. How much data (in terms of bytes, or perhaps rows) did I transfer between workers? Unexpectedly high bytes transferred might indicate failure to push down some predicate, or poor physical plan choice. 2. What throughput did I achieve when transferring data (relevant for deciding whether to compression before sending). In our Quent schema, Workers currently have an `instance_name` (https://github.com/NVIDIA/cudf/blob/dfc5fa903079492a7931372494c1f1bed900d22c/python/cudf_polars/quent/model.yaml#L245). We'll likely need to include a numeric `rank` as well. We'll use `DataChannel` to represent the channel between ranks. The individual transfer records will need to come from rapidsmpf. They will include details on - rapidsmpf Collective Operation ID ("collective ID", to avoid confusion with cudf-polars' Operat**or** ID) - collective kind - source rank - destination rank - message ID - metadata_bytes (size of metadata message) - payload_bytes (size of payload message) - MemoryType (the tier) - completion_timestamp See https://github.com/rapidsai/rapidsmpf/issues/996 for the rapidsmpf side. We'll aggregate those by Operator, using the mapping `ReserveOpIDs.collective_id_map` from `IR -> list[int]` (rapidsmpf operation IDs are reused, so we need to include the query ID in there too).