ThomasWaldmann
While looking for other code with the same kind of problem as #10450, I found that `RobustUnpacker` scales very badly when it has to resync. Code (master, same in 1.4-maint and 1.2-maint): https://github.com/borgbackup/borg/blob/09b2969d3f0a4bb295d670ee4fe7088f30f5ed5a/src/borg/archive.py#L2190-L2234 ## What happens `borg check` iterates over the item metadata stream of every archive with `RobustUnpacker` (`robust_iterator`). After a missing item metadata chunk, or when msgpack fails to unpack, it calls `resync()` and then searches the following data byte by byte for the start of a valid item dict: - For every byte offset it tries, `data = data[1:]` copies the whole remaining buffer. Searching B bytes costs O(B²). - If no item starts in the data seen so far, `__next__` stops. The next `feed()` appends one more chunk to `_buffered_data`, and the next `__next__` joins all buffered data again and searches it from byte 0 again. Over K item metadata chunks this is roughly O(K³). This gets really bad when the resync point is inside a big item. The item of a large file contains its whole chunk list, and no item can start inside it. For example, a 1.22 TB stdin item with 12.3 M chunks (as in #10450) has a chunk list of about 500 MB, spread over thousands of item metadata chunks. A single missing item metadata chunk there would make `borg check` practically never finish, and its memory use would grow by the size of that list. ## Measurements Measured on a Ryzen 5 8500GE (Zen 4), Python 3.13, borg master b7015f32d, pinned to one core. The benchmark calls `RobustUnpacker.resync()`, then feeds pieces of a msgpacked chunk list (`[[id32, size], ...]`, i.e. the bulk of a huge item) and iterates after each `feed()`, as `robust_iterator` does. No valid item starts in this data. | after `resync()` | time | |---|---:| | one 32 KiB piece | 0.01 s | | one 64 KiB piece | 0.04 s | | one 128 KiB piece | 0.13 s | | one 256 KiB piece | 0.49 s | | 2 × 64 KiB | 0.17 s | | 4 × 64 KiB | 0.94 s | | 8 × 64 KiB | 6.31 s | A single piece grows about 4× per doubling. Several pieces grow about 6–7× per doubling. Extrapolating, a file with 100k chunks (about 4 MB of chunk list) already needs on the order of an hour of resyncing, and the 12.3 M-chunk item from #10450 would never finish. ## Possible fix - Search with an offset into a `memoryview` of the buffered data instead of slicing `bytes`. - Remember how far the search got. Drop the data that was already searched instead of joining and searching everything again after the next `feed()`. That makes the resync linear in the amount of data it has to skip.