borg check: RobustUnpacker resync is quadratic/cubic in the amount of data it has to skip

#10459 · open · 1 comments

View on GitHub ↗

ThomasWaldmann

While looking for other code with the same kind of problem as #10450, I found that `RobustUnpacker` scales very badly when it has to resync. Code (master, same in 1.4-maint and 1.2-maint): https://github.com/borgbackup/borg/blob/09b2969d3f0a4bb295d670ee4fe7088f30f5ed5a/src/borg/archive.py#L2190-L2234 ## What happens `borg check` iterates over the item metadata stream of every archive with `RobustUnpacker` (`robust_iterator`). After a missing item metadata chunk, or when msgpack fails to unpack, it calls `resync()` and then searches the following data byte by byte for the start of a valid item dict: - For every byte offset it tries, `data = data[1:]` copies the whole remaining buffer. Searching B bytes costs O(B²). - If no item starts in the data seen so far, `__next__` stops. The next `feed()` appends one more chunk to `_buffered_data`, and the next `__next__` joins all buffered data again and searches it from byte 0 again. Over K item metadata chunks this is roughly O(K³). This gets really bad when the resync point is inside a big item. The item of a large file contains its whole chunk list, and no item can start inside it. For example, a 1.22 TB stdin item with 12.3 M chunks (as in #10450) has a chunk list of about 500 MB, spread over thousands of item metadata chunks. A single missing item metadata chunk there would make `borg check` practically never finish, and its memory use would grow by the size of that list. ## Measurements Measured on a Ryzen 5 8500GE (Zen 4), Python 3.13, borg master b7015f32d, pinned to one core. The benchmark calls `RobustUnpacker.resync()`, then feeds pieces of a msgpacked chunk list (`[[id32, size], ...]`, i.e. the bulk of a huge item) and iterates after each `feed()`, as `robust_iterator` does. No valid item starts in this data. | after `resync()` | time | |---|---:| | one 32 KiB piece | 0.01 s | | one 64 KiB piece | 0.04 s | | one 128 KiB piece | 0.13 s | | one 256 KiB piece | 0.49 s | | 2 × 64 KiB | 0.17 s | | 4 × 64 KiB | 0.94 s | | 8 × 64 KiB | 6.31 s | A single piece grows about 4× per doubling. Several pieces grow about 6–7× per doubling. Extrapolating, a file with 100k chunks (about 4 MB of chunk list) already needs on the order of an hour of resyncing, and the 12.3 M-chunk item from #10450 would never finish. ## Possible fix - Search with an offset into a `memoryview` of the buffered data instead of slicing `bytes`. - Remember how far the search got. Drop the data that was already searched instead of joining and searching everything again after the next `feed()`. That makes the resync linear in the amount of data it has to skip.

Comments

mr-raj12

I'd like to take this one. Two things the fix must preserve, otherwise resync loses items (a naive "drop what was searched" version fails the existing `test_missing_chunk` / `test_corrupt_chunk`): - An item start where msgpack runs out of data (the item continues in the next chunk) must be tried again after the next `feed()`, so that data is not "already searched". - `valid_msgpacked_dict` returns False when it has too few bytes to decide, so the last 260 bytes (map16 header + str8 header + 255 byte key) must be searched again after the next `feed()`. A third thing I found while prototyping: inside a chunk list, random ids regularly look like an item start followed by a 32-bit length. `msgpack.Unpacker` preallocates the list for a bogus array32/map32 (~2.5 s each here, the current code pays that too), and such a candidate never completes. Kept pending, it pins the buffer and is tried again after every `feed()`, which is quadratic again. Probing candidates with `unpackb` on a memoryview avoids the copy and the preallocation (its length limits default to the data length). To make it linear, I would forget pending item starts after 4 MiB. Side effect: an item > 4 MiB (~100k chunks) that starts right at a resync point is not recovered. The current code does not finish in that case anyway. OK with that? Prototype numbers (resync over a msgpacked chunk list fed in 128 KiB pieces, one core): 32 MiB 6.7 s, 128 MiB 28.6 s, 512 MiB 105 s, the buffer stays at ~4.3 MB. Output is identical to the current code on 9000 randomized feed/resync/garbage cases. PR against master follows. 1.4-maint / 1.2-maint after review: there, `valid_msgpacked_dict` uses `.startswith()`, which needs a small change for a memoryview.