dazuma
`remove_repos` is the only mutating entry point that never takes the `repo.lock` flock. [`remove_refs`](https://github.com/dazuma/git_cache/blob/main/lib/git_cache.rb#L163) and [`remove_sources`](https://github.com/dazuma/git_cache/blob/main/lib/git_cache.rb#L193) both run inside `lock_repo`, and `ensure_repo` runs inside the lock held by `GitCache#get` — but [`remove_repos`](https://github.com/dazuma/git_cache/blob/main/lib/git_cache.rb#L138) deletes the entire base directory, `repo.lock` included, without ever acquiring it. So a `remove_repos` in one process can pull the tree out from under an in-flight `get` in another. The `get` holds the lock, believes it has exclusive access, and its next `git` invocation chdirs into a directory that no longer exists — surfacing as a confusing `GitCache::Error` ("Unable to fetch commit", "Path … does not exist at SHA …") in a process that did nothing wrong. Noticed while fixing #5. The rename-first removal from that fix narrows the damage — a `get` that starts *after* the rename now works on a path disjoint from the one being deleted, where previously an in-place `rm_rf` would delete files out from under a concurrent fetch — but it does not address an already-in-flight `get`. ### Why simply taking the flock in `remove_repos` doesn't fix it The lock file lives *inside* the directory being removed, so removing it hands out a fresh lock to everyone who arrives later. `lock_repo` opens with `File::CREAT`, so the next process to come along recreates `repo.lock` at the canonical path — a different inode — and acquires it immediately, while the remover still holds the old one. flock is attached to the inode rather than the path, so the removal itself is legal while holding the lock; it just stops excluding anyone. ```ruby require "fileutils" require "tmpdir" def contend(path) fork do file = ::File.open(path, ::File::RDWR | ::File::CREAT) exit(file.flock(::File::LOCK_EX | ::File::LOCK_NB) ? 0 : 1) end ::Process.wait2.last.success? ? "ACQUIRED" : "blocked" end ::Dir.mktmpdir do |tmp| base = ::File.join(tmp, "abc123") ::FileUtils.mkdir_p(base) lock_path = ::File.join(base, "repo.lock") holder = ::File.open(lock_path, ::File::RDWR | ::File::CREAT) holder.flock(::File::LOCK_EX) puts "contender, lock file in place: #{contend(lock_path)}" # Exactly what remove_repos does, while the lock is held. ::FileUtils.mv(base, ::File.join(tmp, ".trash-xyz")) puts "holder's fd still usable: #{holder.write('x').positive?}" # Exactly what the next GitCache#get does. ::FileUtils.mkdir_p(base) puts "contender, after removal: #{contend(lock_path)}" end ``` On macOS, Ruby 4.0.5: ``` contender, lock file in place: blocked holder's fd still usable: true contender, after removal: ACQUIRED ``` The same holds for a plain `unlink` + recreate, which is what the pre-#5 in-place `rm_rf` path did. So a flock in `remove_repos` would be a one-directional barrier — it makes the remover wait for an in-flight `get`, which is the bug above and is worth having — but it is not mutual exclusion, and it comes with two further problems: **Windows.** Windows refuses to rename a directory that has an open handle underneath it, and the flock handle would be exactly that. "Hold the lock across the rename" fails there with `EACCES`, falls through to the in-place delete, which fails for the same reason, and `remove_repos` raises. Portably you would have to release before renaming, which reduces it to "wait for quiescence" with a window still open. **Orphaned state.** A `get` blocked on the old inode wakes up when the remover releases, holding a lock on a file that now sits in the trash directory. It rebuilds the repo at the canonical path and writes its state JSON into an inode nobody will ever read again. Not corruption, but silent state loss — and it happens precisely because the lock and the state are the same file inside the removed tree. ### Proposal: move the lock out of the tree it protects Put the lock at `<cache_dir>/<md5>.lock` (or `<cache_dir>/locks/<md5>`) instead of `<cache_dir>/<md5>/repo.lock`. It is then never destroyed by the operation it guards, and all three problems go away together: - `remove_repos` can hold a real lock across the entire removal. - Newcomers block on the same inode instead of minting a new one. - The open handle is outside the renamed directory, so Windows is fine. - No state gets orphaned, because the state file is no longer the lock file. Costs, both manageable: - **Cross-version safety.** A 0.1.x process locking `<base>/repo.lock` and a newer one locking `<md5>.lock` would not exclude each other at all, which is worse than today for anyone sharing a cache directory between gem versions. Bumping `FORMAT_VERSION` from `v1` to `v2` sidesteps this by giving the new layout its own tree; the price is one re-clone, which is cheap for a cache. - **Enumeration and a stray-file wart.** `remotes` would enumerate lock files rather than directories. Worth fixing in the same pass: `repo_info` currently fabricates a `repo.lock` for any remote it is merely *asked about* (`lock_repo` opens `CREAT`, and `RepoLock` fills in `remote` even when nothing was cached), which is half of what made the symptom in #5 visible. Under the new layout that side effect becomes a stray top-level `<md5>.lock` file. ### Interim option If the layout change is not happening soon, an acquire-then-release barrier in `remove_repos` — take the flock, release it, then rename — is a handful of lines and closes the in-flight-`get` case on every platform, at the cost of a small window between release and rename. Not worth doing if the relocation is imminent.