fix(supervision): bound lock steal recursion on progress, not depth - #9
Merged
zeeshaanahmad merged 1 commit intoAug 13, 2026
Merged
Conversation
fm_lock_try_acquire recurses into "<lockdir>.steal" to displace an abandoned holder. Reaching that recursion with the lock path absent means the create failed for a reason a deeper steal cannot fix - a state directory removed out from under a running watcher, an unwritable or full filesystem - so every deeper level failed identically while the name grew by ".steal", past the filesystem's name limit and on without bound. Gate the recursion on there being something to displace, and apply the gate by attempting progress rather than inferring it: the same absence is also the benign race where the holder released between the create attempt and the check, and there the lock is simply free. Retry the create; take it if it succeeds, report contention if the path reappeared, and refuse loudly if it is still absent and still uncreatable. The walk itself stays unbounded by depth. A fully abandoned chain is legitimate recovery that reclaims and cleans up every level, and capping it turns slow, noisy recovery into a permanent refusal that fm_lock_acquire_wait then spins on forever - the reason the depth bound written alongside the signal-deferral fix was reverted rather than repaired. Both properties are pinned by regressions.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Lock steal recursion is bounded on progress, not on depth
Starting from the disproof, not from the depth bound
A depth bound on this same recursion was written during the signal-deferral work (#7) and reverted before it landed, because its own red/green proof showed it would deadlock legitimate recovery: an unbounded walk of a fully abandoned steal chain genuinely reclaims the lock, and capping the walk turns slow, noisy recovery into a permanent refusal that
fm_lock_acquire_waitthen spins on forever. That conclusion holds and is not re-litigated here. This change does not re-land the depth bound in any form.1. The legitimate recursion, characterised first
fm_lock_try_acquirecalls itself in exactly one place: to acquire"<lockdir>.steal", the mutex that serialises displacing an abandoned holder.A level is entered only because the level above it exists on disk with a dead holder. So the recursion depth equals the length of the abandoned chain, and a chain grows one level per crash that happens while the steal mutex is held. Measured against the real code path (a process that acquires
L,L.steal,L.steal.steal,L.steal.steal.stealand is thenkill -9ed):L, every.steallevel is gone, verified by listing the directory;NAME_MAXand each level adds six bytes. Nothing else bounds it, and nothing else should.That is why no depth number is safe to calibrate: the depth that must be tolerated is "however many crashes happened since the last successful reclaim", and the cost of guessing low is a lock nobody can ever take again.
2. What actually runs away, and the guard chosen
The runaway is the case with no abandoned holder at all.
fm_lock_try_createfails for two very different reasons: the lock exists (contention - the case the steal path is for), or it could not be created at all. In the second case the lock path is absent, the pid read is empty,fm_pid_aliveis false andfm_lock_mid_acquire_is_freshis false, so control falls into the steal path anyway - where the deeper level fails for the identical reason, while the name grows by.stealeach time. PastNAME_MAXevery level keeps failing and the recursion never terminates.Reproduced deterministically, with a cause that is not exotic: a state directory removed or moved out from under a running watcher.
fm_lock_try_acquire "$state/gone/.contend.lock"was still recursing after 15s and had to be killed; its stderr showed names grown past the OS limit.The guard is progress, not depth, and it is applied by attempting progress rather than inferring it. Before recursing, if the lock path is absent the create is retried once:
Recursion is therefore untouched whenever there is something to displace, at any depth, which is precisely the property the disproof requires.
Why this over the alternatives the task listed
The progress test also needs no calibration and no new tunable, and its failure mode points the safe way: if it is ever wrong it refuses a lock nobody holds, which the caller retries, instead of refusing a lock it could have recovered.
A benign race the first version got wrong
The first version tested progress by inspection alone - "path absent -> refuse". Running the existing suite surfaced the counterexample immediately:
test_cycle_exit_ledger_links_successor_and_stays_boundedprinted the new refusal for.watcher-down.lockin a state directory that plainly existed. The holder had released between the create attempt and the check, so the lock was simply free, and the old code reached it (wastefully) through the steal path. Inferring the fault from absence turned a benign, non-rare race into a spurious failure and a scary operator warning.Retrying the create fixed it, and the warning no longer appears anywhere in the suite. Recorded because the lesson generalises: absence of a lock is ambiguous evidence, and the cheap way to disambiguate is to attempt the thing rather than to reason about it.
3. Red-then-green
Both cases are pinned in
tests/fm-watcher-lock.test.sh, alongside the existing steal-primitive cases. The runaway case is bounded in the test (150 x 0.1s, thenkill -9and a named failure) because the regression is an unbounded loop and a hanging suite reports nothing.85643b8ok(control)oknot ok - acquiring an uncreatable lock never returned (unbounded steal recursion)okokThe first row is the row that judges the design: it passes on both sides. The guard did not buy termination by breaking the recovery the reverted depth bound broke - that is the case a depth cap of 4 or less would have turned into a permanent refusal.
Upstream check
Fetched
upstream(kunchenguid/firstmate) and searched its open issues forlock,recursion,deadlockandstealbefore implementing. No open upstream issue covers this defect.recursionreturns zero. The nearest neighbours are different faults in different code: kunchenguid#1972 (remote job worker wedged on an orphaned lock temp file), kunchenguid#1508 (Windows/MSYSkill -0reporting every pid alive, sofm_lock_acquire_waitspins), kunchenguid#2251 and kunchenguid#2270 (the watcher/daemon symptoms already addressed by #7). Nothing was adopted from them.Not harness-dependent
The verdict comes from filesystem state and lock ownership, not from anything a vendor emits, so per the coding guidelines this is pinned by a portable regression with real processes and no live-harness guard is owed. No per-harness verification record changes.
Local verification
macOS 26.4.1 (arm64),
/bin/bash3.2.57. Measured against the pushed head.Red evidence was produced in an isolated copy running
bin/restored to85643b8withtests/from this branch, so the only difference is the fix itself.tests/fm-watch-triage.test.sh: pre-existing load flakiness, confirmed on both sidesThis suite could not be brought to a clean run on this machine because it was saturated by unrelated work throughout (load averages 9.7 to 55.6). It failed on both sides, at a different case every run - the exact signature #7 recorded for it, from absorb gates built on fixed wall-clock slices rather than artifacts:
85643b8.seen-*suppressor85643b8working:note with an idle paneBase failed 2/2 and this branch 4/4, never twice at the same case. Converting that suite's gates to artifact-based waits remains its own task, as #7 concluded.
Not run, and why
bin/fm-test-run.shsuite: left to the pipeline, which ran against the pushed headcddb7b7and completed its review, test, document and lint steps with zero findings.live-harness-optinguards: nothing here is harness-dependent (see above), so none is owed.bin/fm-ci-probe.shreportsnonefor this fork, so no check run can ever register and the ci step was skipped deliberately rather than left to time out. The PR shows no checks for that reason, not because any check failed.Merge reasoning (firstmate)