git's detached auto-maintenance after fetch races with the cache's own rm_rf

#5 · closed · 1 comments

View on GitHub ↗

dazuma

`GitCache#get` fetches via: ```ruby git(repo_dir, ["fetch", "--depth=1", "--force", "origin", "#{commit}:#{local_commit}"], error_message: "Unable to fetch commit: #{commit}") ``` https://github.com/dazuma/git_cache/blob/main/lib/git_cache.rb#L284 Since git ~2.43, `git fetch` ends by spawning post-command maintenance **detached**, so `fetch` returns while that process is still starting up. Verified with `GIT_TRACE=1` on git 2.55, running exactly the fetch above: ``` trace: run_command: git maintenance run --auto --no-quiet --detach ``` That process then writes into the cache repo *after* `fetch` has returned. I measured it creating `.git/objects/maintenance.lock` ~0.4ms after the spawn, and it can go on to write pack files. ### Why this matters here Every path that deletes cache directories uses `FileUtils.rm_rf`, which swallows every error, per entry. If the background process creates a file in a directory that `rm_rf` has already emptied but not yet `rmdir`ed, the `rmdir` fails with `ENOTEMPTY` — and so does every parent, all the way up to the repo's base dir. The whole chain silently survives: - [`remove_repos`](https://github.com/dazuma/git_cache/blob/main/lib/git_cache.rb#L144) — reports a remote as removed while its base dir remains, after which `repo_info` keeps reporting a repo that is no longer really there - [`ensure_repo`](https://github.com/dazuma/git_cache/blob/main/lib/git_cache.rb#L269) — the rebuild path taken when the remote URL doesn't match; can leave the stale repo dir in place - [`remove_sources`](https://github.com/dazuma/git_cache/blob/main/lib/git_cache.rb#L208) Evidence: deleting a cache dir immediately after a fetch left 9 of 10 trees behind, each containing exactly one file that the background repack had written after `rm_rf` walked that directory: ``` cache/467ed9278.../repo/.git/objects/pack/tmp_pack_J4Cfk8 cache/e9a7a63b3.../repo/.git/objects/pack/pack-fe3eb92....pack cache/db900b178.../repo/.git/objects/pack/pack-6508f5d....idx ``` Everything else in each tree was deleted. I found this chasing an intermittent CI failure in toys, which vendors this library: `toys system git-cache show <remote>` succeeded for a remote that was never cached, because the base dir had survived the test's cleanup. ### Suggested fix Disable auto maintenance on this library's own git invocations. Either knob suppresses the spawn — verified with `GIT_TRACE=1` on git 2.55, 3 maintenance trace lines dropping to 0: ``` git -c maintenance.auto=false fetch ... # modern knob git -c gc.auto=0 fetch ... # also works, and predates the above ``` Probably cleanest in the [`git`](https://github.com/dazuma/git_cache/blob/main/lib/git_cache.rb#L231) helper, so every invocation is covered. These are managed cache repos that the library recreates at will, so background repacking buys them nothing, and a detached process mutating them is exactly what the caller is not expecting. Worth considering independently: retrying `rm_rf` until the path is actually gone. `rm_rf`'s silence is what turns a failed deletion into a wrong answer from `repo_info` rather than an error.

Comments

dazuma

### A second failure mode of the same race: `chmod_R` raises The report above covers the quiet half — `rm_rf` swallowing `ENOTEMPTY` and leaving the tree. There is a loud half too, and the "retry `rm_rf` until the path is gone" suggestion at the end doesn't cover it, because the raise happens *before* any removal is attempted. Every deletion path here calls `FileUtils.chmod_R("u+w", ..., force: true)` before `rm_rf`. `force` does less than it looks like (fileutils 1.8.0): ```ruby def chmod_R(mode, list, noop: nil, verbose: nil, force: nil) ... list.each do |root| Entry_.new(root).traverse do |ent| begin ent.chmod(fu_mode(mode, ent.path)) rescue raise unless force # <- guards only the chmod of each entry end end end end ``` The `begin`/`rescue` wraps the chmod of each entry. The **traversal** — `Entry_#traverse` → `preorder_traverse` → `entries` → `Dir.children` — sits outside it. So when the detached maintenance process packs loose objects and removes an emptied `.git/objects/<xx>` fanout directory between the moment the walk lists it and the moment it descends into it, `Errno::ENOENT` escapes `chmod_R` regardless of `force`. Seen on CI in [toys](https://github.com/dazuma/toys), which vendors this library, as an error in a git-cache test teardown: ``` Errno::ENOENT: No such file or directory @ dir_initialize - /tmp/toys_git_cache_test.../cache/f83d.../repo/.git/objects/16 fileutils-1.8.0/lib/fileutils.rb:2163:in 'Dir.children' fileutils-1.8.0/lib/fileutils.rb:2163:in 'FileUtils::Entry_#entries' fileutils-1.8.0/lib/fileutils.rb:2357:in 'FileUtils::Entry_#preorder_traverse' fileutils-1.8.0/lib/fileutils.rb:1820:in 'block in FileUtils.chmod_R' ``` Reproduced without git, to isolate the traversal from the fetch: build a 256-entry `objects/<xx>` tree, then run `chmod_R("u+w", root, force: true)` while another thread removes those entries — standing in for the repack. Being a race, the rate moves around, but the script below **raised on 14, 21, 22, and 23 of 50 runs across four invocations** on my machine (macOS, Ruby 4.0.5, fileutils 1.8.0), always the same `Errno::ENOENT ... @ dir_initialize` on an `objects/<hex>` path. Wrapping the `chmod_R` in `rescue SystemCallError` takes it to **0 of 50**. ```ruby require "fileutils" require "tmpdir" def run_once ::Dir.mktmpdir("race_repro") do |tmp_dir| 256.times do |i| dir = ::File.join(tmp_dir, "objects", format("%02x", i)) ::FileUtils.mkdir_p(dir) ::File.write(::File.join(dir, "loose"), "x" * 64) end packer = ::Thread.new do sleep(0.002) 256.times { |i| ::FileUtils.rm_rf(::File.join(tmp_dir, "objects", format("%02x", i))) } end begin ::FileUtils.chmod_R("u+w", tmp_dir, force: true) nil rescue ::SystemCallError => e e ensure packer.join end end end failures = 50.times.map { run_once }.compact puts "#{failures.size}/50 raised; sample: #{failures.first&.message}" ``` ### Where this bites the library The two sites that walk a live repo tree: - [`remove_repos` L143](https://github.com/dazuma/git_cache/blob/main/lib/git_cache.rb#L143) — walks `repo_base_dir_for(remote)`, which contains `repo/.git` - [`ensure_repo` L268](https://github.com/dazuma/git_cache/blob/main/lib/git_cache.rb#L268) — walks `repo_dir` directly (The other `chmod_R` calls — [L207](https://github.com/dazuma/git_cache/blob/main/lib/git_cache.rb#L207), [L307, L311](https://github.com/dazuma/git_cache/blob/main/lib/git_cache.rb#L307) — walk source checkouts, which the maintenance process doesn't write to, so they look unexposed.) So the same race that can make `remove_repos` silently under-delete can instead make it raise `Errno::ENOENT` at the caller, depending on which process reaches the fanout directory first. Both outcomes are surprising for a method whose contract is "remove this cache entry". ### Bearing on the proposed fix This strengthens the case for `-c maintenance.auto=false` in the [`git` helper](https://github.com/dazuma/git_cache/blob/main/lib/git_cache.rb#L231) over hardening call sites. Every walk of a live repo tree is exposed, `force:` doesn't cover it, and each new walker would have to remember — whereas suppressing the spawn fixes both halves at the source. Until then, callers hardening their own teardowns want `rescue SystemCallError` around the `chmod_R` and a retry, rather than trusting `force:`. That is what toys now does in the three suites that build git caches.