Skip to content

fix(supervision): make signal handling deadlock-safe in watcher/daemon - #7

Merged
zeeshaanahmad merged 6 commits into
mainfrom
fm/watcher-signal-deadlock-redesign
Aug 13, 2026
Merged

zeeshaanahmad merged 6 commits into
mainfrom
fm/watcher-signal-deadlock-redesign

Conversation

@zeeshaanahmad

@zeeshaanahmad zeeshaanahmad commented Aug 13, 2026 •

Copy link
Copy Markdown
Owner

Intent

Redesign watcher signal handling so a signal landing mid-critical-section can never self-deadlock the watcher.

This is a firstmate-repo task on supervision-core code, so the firstmate-coding-guidelines apply throughout, including their deterministic-enforcement and harness-dependent-check sections in full.

ROOT CAUSE. Fleet locks are non-reentrant, and bin/fm-watch.sh traps HUP/INT/TERM as exit 1, which unwinds into watcher_cleanup, which re-acquires locks the interrupted flow still holds. Reachable on EVERY poll through _fm_recovery_marker_arm_check (resurface_after_downtime), not only through fm_wake_append. Observed in the wild as a watcher whose own pid held .wake-queue.lock and .watcher-down.lock with a frozen beacon for 64 minutes. Upstream kunchenguid#2251 reports the same symptom with no fix; kunchenguid#2270 (afk daemon SIGTERM shutdown hang) is the same root cause on the daemon path, so the fix had to cover both surfaces.

REQUIRED DELIVERABLES. (1) The redesign, applied to fm-watch.sh AND every other surface sharing the trap-unwind pattern, explicitly auditing the afk daemon bin/fm-supervise-daemon.sh and any other loop that traps into cleanup while holding locks - no tunnel vision on the watcher alone. (2) A portable regression in tests/ that drives a real signal into the critical section deterministically and proves no deadlock, plus TERM-still-stops-the-watcher coverage, extending existing runner patterns with .test.sh naming. (3) Removal of the quarantine skip the earlier stale-reads change placed on the hanging watch-triage case (tracking name watcher-signal-deadlock-redesign), with its full assertion restored and proven passing. (4) bin/fm-lint.sh clean.

DESIGN CHOICE AND WHY. Two options were pre-scoped and neither was to be inherited blindly: (1) a global deferred-signal trap in fm-watch.sh with per-poll safe points, correct in shape but risking a TERM-immune watcher if a safe point is missed; (2) an ownership-aware fm_lock_acquire_wait, small but requiring an audit of all ~65 callers because ~20 ignore its return value and would silently unlock spawn/teardown/drain critical sections. A hybrid or third design was acceptable if shown to dominate both.

I chose a hybrid and it must be judged as one: make the critical SECTION uninterruptible rather than make the unwind survivable, plus a separate liveness bound so shutdown can never block.

  • Signal deferral (fm_signal_defer_begin/end, fm_lock_section_enter/leave) lives in bin/fm-wake-lib.sh, the library that owns these locks, so every consuming script is covered without enumerating call sites - this is what dominates option 1, because a missed section degrades to today's behaviour for that section rather than to a TERM-immune watcher.
  • Deferral is opt-in per section and deliberately NOT a process-wide trap, because fm-watch.sh holds .watch.lock for its entire life and a blanket deferring trap over a long-held ownership lock would produce the mirror-image defect: a watcher TERM can no longer stop.
  • fm_lock_acquire_wait's return contract is untouched for every existing caller, and the new bound is caller-scoped and deliberately NOT seeded from the environment, so the ~20 callers that ignore the return value cannot begin silently proceeding without a lock they believe they hold. This is what dominates option 2 and is why no 65-caller audit was required.
  • Inside a deferred section, a nested acquisition becomes bounded once a signal is already pending, so a held signal waits on a constant rather than on another process's lifetime; those nested sites already check the return value and their failure paths release the outer lock.
  • The partial patch in the earlier lane's task records is explicitly SUPERSEDED, not inherited: it wrapped call sites inside fm-watch.sh, which its own comments admit misses paths reachable elsewhere.

DELIBERATE DECISIONS A REVIEWER CANNOT SEE IN THE DIFF.

  • fm_lock_try_acquire's self-held reclaim (upstream fix(bin): collapse duplicate supervision wakes kunchenguid/firstmate#2287) is already in the base. I verified empirically rather than assuming, and it genuinely rescues the same-pid case. It is deliberately left in place. What it does not cover, and what this change fixes, is unbounded shutdown against a DIFFERENT live holder, the torn invariant a half-finished section leaves behind, and the fact that its soundness rests on an assumption ("every lock-taking trap path exits rather than resuming") that nothing enforces.
  • procevent_surface_queued's cross-function wake-queue hold is deliberately NOT converted to a deferred section. It hands its lock to wake()'s post-output action, and deferring across wake()'s potentially blocking stdout write is exactly the TERM-immunity trap this design exists to avoid. It is not a deadlock source because cleanup never re-acquires the queue lock.
  • Rather than trust a reading pass for "no tunnel vision", I mechanically extracted every trap-registered handler in bin/ and flagged those that can block on a lock or a child. That found two surfaces beyond the watcher and the daemon: bin/fm-pr-check-migrate.sh (identical shape - traps into a cleanup re-entering the recovery-marker lock while holding the watch lock) and bin/fm-watch-arm.sh's handle_arm_signal (an unbounded kill-and-wait on its watcher child, so the handler whose job is to stop the arm could itself hang). Both are fixed. The bounded child stop has ONE owner in fm-wake-lib.sh rather than a copy per caller. Everything else the audit listed only releases locks in cleanup, which never blocks.
  • watcher_cleanup keeps its existing "retaining stale lock evidence" outcome when the bound is hit; it deliberately does not release the watch lock in that case, because the next arm reclaims it through the ordinary dead-holder path and the retained lock is meaningful evidence.
  • The quarantine is not merely deleted. FM_SIGNAL_DEADLOCK_SKIP and both skip sites are replaced by a new assert_reaped_on_term helper that turns the same condition (reap having to escalate past TERM) into a NAMED FAILURE, so it is enforced rather than tolerated. reap's KILL escalation stays only so a future regression fails loudly instead of hanging CI, and tests/fm-liveness-source.test.sh now proves that reporter fires rather than proving a skip string exists.
  • A recursion depth bound I wrote for an observed .steal.steal.steal... blowup in fm_lock_try_acquire was REVERTED after its own red/green proof showed it was harmful: an unbounded walk of a fully abandoned steal chain legitimately reclaims the lock, and capping it converts slow recovery into a permanent refusal that fm_lock_acquire_wait would spin on forever. Deliberately not shipped; it needs a fix that cleans up the chain, as its own task.
  • tests/fm-wake-queue.test.sh's "a subshell must not reclaim its parent's live hold" half is now gated on BASHPID being available. This is a deliberate capability gate, not a masked failure: lock ownership is keyed on ${BASHPID:-$$}, BASHPID is Bash 4.0+, and on stock macOS Bash 3.2 a subshell reports its parent's pid so the property cannot be expressed. Confirmed failing identically at origin/main.
  • docs/watcher-continuity.md is updated as the maintainer-architecture owner of the watcher recovery contract, with a signal-safety section and a regression pointer; the mechanism itself stays owned by the library header per the one-owner rule. No AGENTS.md change: this is internal mechanism, not an always-loaded operational contract.
  • Per the harness-dependent-check rule, this is NOT harness-dependent - the verdict comes from signal disposition and lock ownership, not from anything a vendor emits - so it is pinned by a portable regression with real processes and no live-harness guard or per-harness verification record is owed.

KNOWN PRE-EXISTING ISSUES DELIBERATELY LEFT TO FOLLOW-UP, NOT REGRESSIONS FROM THIS CHANGE: tests/fm-watch-triage.test.sh's absorb gates are fixed 3-second wall-clock slices (wait_live), which the harness's own wait_watcher_beat comment warns against; under heavy load they flake, measured alternating in isolated checkouts at load ~20 as 2/2 failures on this branch AND 2/2 on origin/main. Converting ~26 gates changes what those cases assert and belongs in its own task rather than inside a supervision-core signal fix.

DELIVERY. PRs go to origin zeeshaanahmad/firstmate (a fork); kunchenguid/firstmate is fetch-only upstream and must never be pushed to. The ci step is skipped because the fork has zero Actions runs and checks never register: bin/fm-ci-probe.sh with the fork named explicitly returns "none". Note its no-argument path returns "present" here because gh repo view resolves to the upstream parent rather than origin - a separate defect being reported, not a reason to let ci poll for a check that can never arrive. Local verification stands in for CI and is recorded in the PR body.

The PR body must include a Local verification section with exact commands, the deterministic red-then-green deadlock-reproduction evidence, the TERM-stops proof, the restored-quarantine proof, all measured against the pushed head, and an explicit list of what was not run and why.

What Changed

  • Added a signal-deferral mechanism to bin/fm-wake-lib.sh (fm_signal_defer_begin/end, fm_lock_section_enter/leave) that holds HUP/INT/TERM during a recovery-marker or wake-queue critical section and re-raises the signal against the caller's own disposition once the section closes, so a signal can no longer unwind into watcher_cleanup mid-section and re-acquire a lock the interrupted flow still holds. _fm_recovery_marker_arm_check (reachable on every watcher poll) and the other recovery-marker helpers in fm-wake-lib.sh now use these deferred sections instead of raw fm_lock_acquire_wait/fm_lock_release.
  • Added a caller-scoped, non-environment-seeded bound (FM_LOCK_ACQUIRE_WAIT_TICKS) on fm_lock_acquire_wait, and a shared fm_child_stop_bounded/_fm_child_is_live helper for bounded TERM-then-KILL child shutdown. Applied the bound to watcher_cleanup in bin/fm-watch.sh and to the recovery-marker release in bin/fm-pr-check-migrate.sh's migration_cleanup (both trap-unwind-into-cleanup surfaces), and applied the bounded child stop to bin/fm-supervise-daemon.sh's new stop_watcher_child and to bin/fm-watch-arm.sh's handle_arm_signal, replacing their unbounded kill; wait shutdown of the watcher child.
  • Replaced the FM_SIGNAL_DEADLOCK_SKIP quarantine in tests/wake-helpers.sh with an assert_reaped_on_term helper that turns a TERM-ignoring reap into a named test failure instead of a silent skip, updated tests/fm-liveness-source.test.sh and tests/fm-watch-triage.test.sh (restoring its previously-skipped assertions) to use it, gated the subshell-reclaim half of tests/fm-wake-queue.test.sh on BASHPID availability, and added tests/fm-watcher-signal-safety.test.sh as a new regression suite (registered in bin/fm-test-run.sh) covering deferred-section signal handling, bounded watcher/daemon shutdown, and TERM-still-stops-the-watcher behavior with real processes and signals.
  • Documented the new signal-safety contract and its regression coverage in docs/watcher-continuity.md.

Risk Assessment

✅ Low: Both prior-round findings (zombie/kill-0 liveness race, and stop_watcher_child's return code being masked by log()) are correctly and completely fixed on independent re-inspection; the core signal-deferral mechanism (fm_signal_defer_begin/end, fm_lock_section_enter/leave, the caller-scoped fm_lock_acquire_wait bound) is internally consistent across every call site with no depth-counter leaks or unpaired enter/leave, all four required deliverables from the user intent are present and verifiable in source (redesign applied to all four trap-unwind surfaces, a real-process/real-signal regression suite, the quarantine skip fully replaced by named assert_reaped_on_term failures, and a clean bin/fm-lint.sh run), and no forbidden behavior was introduced.

Testing

Ran the new tests/fm-watcher-signal-safety.test.sh (5/5 pass) which drives real signals into a real watcher pinned inside its critical section by an external lock holder, and separately reproduced the pre-fix deadlock (not ok) against the base commit's code in an isolated scratch checkout to establish red-then-green evidence; also ran fm-liveness-source, fm-watch-triage (including both restored-quarantine cases), fm-wake-queue, fm-watch-arm, and fm-daemon suites, all passing with no failures, orphaned processes, or worktree artifacts.

Evidence: Red-then-green deadlock reproduction transcript
Deterministic red-then-green reproduction of the watcher signal/lock self-deadlock
=====================================================================================

Method: the new regression test tests/fm-watcher-signal-safety.test.sh (added at
target commit 12484d739da833dc1a18c117a176a0eb353e3c8c) pins a watcher inside its
recovery-marker critical section using a real external lock holder, then sends it
a real TERM and asserts it exits within a bound and abandons no lock.

RED - same test executed against base commit 6d074eca5c7dc5b148db1db9df595a5115d2c78b's
bin/fm-watch.sh and bin/fm-wake-lib.sh (git archive of the base commit, only the new
test's single critical-section case run, no other files modified):

  $ bash tests/fm-watcher-signal-safety.test.sh
  not ok - a watcher signalled inside a contended critical section never exited (self-deadlock)

This reproduces the reported wild symptom: a watcher signalled while holding
.wake-queue.lock and contending .watcher-down.lock does not exit on TERM.

GREEN - same test executed against target commit 12484d739da833dc1a18c117a176a0eb353e3c8c
(the actual worktree under test):

  $ bash tests/fm-watcher-signal-safety.test.sh
  ok - a signal inside a deferred section is held, lets the section finish, then fires through the caller's own handler
  ok - a watcher signalled inside a contended critical section exits, reports why, and abandons no queue lock
  ok - TERM stops the watcher and frees its singleton lock from every position in the poll cycle
  ok - watcher shutdown stays bounded and loud against a marker lock another live process never releases
  ok - daemon shutdown bounds a TERM-ignoring watcher child and still stops an ordinary one on TERM alone

No orphan processes were left behind by either run (verified via `ps aux` after
each); the base-commit scratch checkout was created under /tmp (outside the
worktree and outside git) and removed after use.
Evidence: Full fm-watch-triage.test.sh run showing both restored-quarantine assertions passing
ok - signal_reason_is_actionable: benign absorbed, captain verbs and coalesced batches surfaced
ok - stale_is_terminal: terminal status surfaces, non-terminal and no-status are benign
ok - scan_captain_relevant_statuses lists only captain-relevant statuses
ok - classifier primitives: keyed decisions and activity phases, captain relevance, window-to-task, and overrides
ok - crew_is_provably_working: only working+run-step/pane is provable; idle/finished/parked/failed/unknown surface
ok - status_is_paused: only the leading paused verb matches, and paused is not captain-relevant
ok - crew_absorb_class: working/paused/none from one read; crew_is_paused and crew_is_provably_working agree
ok - signal_crew_provably_working: benign only when every referenced crew is provably working
ok - a secondmate's status signal is never absorbed as provably working; crewmates are unaffected
ok - a no-verb signal whose crew is provably working is absorbed (no exit, no queue, suppressor advanced, beacon present)
ok - a bare turn-end whose crew is provably working (busy pane) is absorbed
ok - a bare turn-end whose crew is not provably working is surfaced (the swallowed-finish fix)
ok - a no-verb working: note whose crew is idle with no running pipeline is surfaced
ok - a secondmate's status note surfaces even while its own agent is busy
ok - a self-announced close never wakes its own home, and the next real note still does
ok - captain-relevant signal is surfaced (queue + exit) and marked surfaced
ok - a stale pane sitting on a terminal status is surfaced (queue + exit)
ok - a stale terminal-looking status is overridden and absorbed while a run is actively working, then wedge-escalated
ok - provably-working non-terminal stale is absorbed on first sight, then wedge-escalated past the threshold
ok - consecutive wedge escalations on the same pane accumulate and demand deep inspection at the threshold
ok - a pane becoming active again resets the consecutive wedge-escalation counter
ok - a busy worker below the turn-age bound remains working with no escalation
ok - a busy worker with a stable pane hash still escalates once its completed-turn age reaches the bound
ok - a busy worker whose pane hash changes every poll still escalates once its completed-turn age reaches the bound
ok - touching a busy worker's completed-turn marker resets the age and prevents an old-age escalation
ok - repeated busy turn-age escalations reuse the existing escalation counter and demand deep inspection at the threshold
ok - the production default busy-turn-age bound is 3600s (5min under does not wedge, 66min over does)
ok - a not-provably-working non-terminal stale is surfaced immediately (never left to wait out the timer)
ok - a declared pause is absorbed on first sight, then re-surfaced as a recheck past the threshold, never wedge-escalated
ok - exited declared-pause and captain-held panes use bounded pause cadence while a live decision gate still surfaces once
ok - a declared paused secondmate re-surfaces on the bounded normal-mode cadence
ok - a non-paused secondmate retains normal stale suppression
ok - a resumed secondmate clears pause and stale tracking before stale exemption
ok - unchanged stale hashes reclassify when a crew enters or leaves pause
ok - a declared pause is periodically rechecked against authoritative active-run state
ok - a paused status overridden by authoritative working preserves its wedge timer and escalates
ok - matching non-terminal stale suppressors repair missing or corrupt stale-since timers
ok - triage log capping handles wc byte counts with leading spaces
ok - a captured process-event result wakes a healthy watcher proactively, with no manual drain
ok - an unacknowledged process-event result re-drains until handling is acknowledged
ok - complete process-event queue keys map to distinct seen markers
WAKE_ACK_REQUIRED: after handling completes run bin/fm-wake-drain.sh --ack-through 1 --recovery-generation 54075.1786607689.G6W1fh
ok - queue revalidation, proactive output, and marker commit serialize with drain
/Users/muhammadzahmed/.no-mistakes/worktrees/c3eda0adcf59/01KZWYTG20J314KFDG0FZVD9SM/bin/fm-push-transition-lib.sh: line 96: echo: write error: Broken pipe
tests/wake-helpers.sh: line 387: 56250 Killed: 9               PATH="$dir/fakebin:$PATH" FM_HOME="$dir" FM_PROCEVENT_CLAIM_ROOT="$dir/claims" FM_CREW_STATE_BIN="$dir/fakebin/fm-crew-state.sh" FM_POLL=0.2 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out"
tests/wake-helpers.sh: line 387: 59061 Killed: 9               PATH="$dir/fakebin:$PATH" FM_HOME="$dir" FM_PROCEVENT_CLAIM_ROOT="$dir/claims" FM_CREW_STATE_BIN="$dir/fakebin/fm-crew-state.sh" FM_POLL=0.2 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out"
ok - surfacing failures replay until post-handling acknowledgement
ok - marker failure exits through the shared wake owner, releases its lock, and replays later
ok - a heartbeat with no captain-relevant change is absorbed and backs off the cadence
ok - heartbeat backstop fail-safe surfaces a captain-relevant status the per-wake path missed
ok - the liveness beacon stays fresh while the watcher absorbs benign wakes (fm-guard never false-alarms)
ok - with .afk present the watcher reverts to one-shot so the daemon owns triage (no double-triage)
ok - AFK changed paused panes hand off plain stale identities for daemon-owned pause triage

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 1 issue found → auto-fixed (2) ✅
  • ⚠️ bin/fm-wake-lib.sh:489 - fm_child_stop_bounded's poll loop (and its post-loop escalation check) treats an unreaped zombie child as still alive, because kill -0 "$pid" succeeds for a zombie (its process-table entry persists until the parent calls wait(), which this function only does at the very end, after both the poll loop and the escalation decision). A child that exits immediately on TERM but hasn't been reaped yet will make the loop spin the full limit ticks and then incorrectly set escalated=1 and send a needless KILL. The codebase already has a dedicated workaround for exactly this pattern - tests/wake-helpers.sh:402's is_live_non_zombie (checks ps -o stat= for Z) - used precisely because a plain kill -0 poll-for-child-exit loop is known to misreport zombies as live. This new production helper (shared by bin/fm-supervise-daemon.sh's stop_watcher_child and bin/fm-watch-arm.sh's handle_arm_signal) reintroduces the naive check instead of reusing that convention.

🔧 Fix: {"summary": "Fix zombie/kill-0 race in fm_child_stop_bounded liveness check"}
1 warning still open:

  • ⚠️ bin/fm-supervise-daemon.sh:1368 - stop_watcher_child's fm_child_stop_bounded ... || log "warn: ..." makes the wrapper's own return code equal to log()'s return code whenever an escalation happens, not fm_child_stop_bounded's. log() is [ -n "${LOG:-}" ] && printf ... >> "$LOG" (line 1343), which returns 0 whenever $LOG is set (the normal daemon runtime) - so stop_watcher_child silently reports success (rc=0) even after a KILL escalation, breaking the '0 when TERM alone sufficed, 1 when escalation was needed' contract that fm_child_stop_bounded's own header documents for exactly this kind of caller ('a caller with somewhere to log can say which happened'). It only happens to return 1 today when $LOG is unset, which is coincidental, not the contract working. Currently harmless - fm_super_main (line 1498) never inspects stop_watcher_child's return value, and the warn log line itself still fires correctly since that happens inside the || regardless - but it is a latent trap for any future caller that does check the return value. The new daemon test in tests/fm-watcher-signal-safety.test.sh (test_daemon_child_stop_is_bounded_against_a_term_ignoring_child, added in this fix round) only asserts rc==0 for the non-escalating/ordinary-child branch; it never asserts what rc is for the TERM-ignoring/escalating branch, so this masking would not be caught by the new regression coverage. Mechanical fix: capture fm_child_stop_bounded's status explicitly (e.g. local rc=0; fm_child_stop_bounded "$pid" "..." || { rc=1; log "warn: ..."; }; return "$rc") instead of letting log()'s own status leak through.

🔧 Fix: {"summary": "Fix stop_watcher_child return code masked by log()"}
✅ Re-checked - no issues remain.

✅ **Test** - passed

✅ No issues found.

  • bash tests/fm-watcher-signal-safety.test.sh (target commit) — all 5 cases pass, including the real-process deadlock reproduction and TERM-from-every-position sweep
  • bash tests/fm-watcher-signal-safety.test.sh re-run against base commit 6d074eca5c7dc5b148db1db9df595a5115d2c78b's bin/fm-watch.sh and bin/fm-wake-lib.sh in an isolated /tmp scratch checkout — fails with not ok - a watcher signalled inside a contended critical section never exited (self-deadlock), reproducing the reported wild symptom (red side of red/green)
  • bash tests/fm-liveness-source.test.sh — confirms assert_reaped_on_term fires as a named failure over a real TERM-ignoring process
  • bash tests/fm-watch-triage.test.sh — full suite passes, including both restored-quarantine cases (test_nonterminal_stale_paused_absorbed_then_resurfaced, test_exited_declared_pause_is_bounded_but_live_gate_surfaces)
  • bash tests/fm-wake-queue.test.sh — passes, including the BASHPID-gated subshell reclaim case on this machine's Bash 3.2
  • bash tests/fm-watch-arm.test.sh — passes (covers the arm-signal handler's shared fm_child_stop_bounded path)
  • bash tests/fm-daemon.test.sh — passes (covers daemon shutdown surfaces)
  • Manual grep confirming no remaining references to the removed FM_SIGNAL_DEADLOCK_SKIP quarantine anywhere in the tree
  • ps aux checks and worktree git status after each run confirming no orphaned processes or transient artifacts were left behind
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

Local verification

All results below are measured against the pushed head 12484d739da833dc1a18c117a176a0eb353e3c8c, on macOS 26.4.1 (arm64), /bin/bash 3.2.57, ShellCheck 0.11.0 (pinned).

Commands and results

bash bin/fm-lint.sh                          # clean
bash bin/fm-doc-audience-check.sh            # ok surfaces=67 local_links=243
bash tests/fm-watcher-signal-safety.test.sh  #  5 ok, 0 not ok
bash tests/fm-watch-triage.test.sh           # 49 ok, 0 not ok
bash tests/fm-liveness-source.test.sh        # 15 ok, 0 not ok
bash tests/fm-wake-queue.test.sh             # 19 ok, 0 not ok
bash tests/fm-watcher-lock.test.sh           # 31 ok, 0 not ok
bash tests/fm-daemon.test.sh                 # 99 ok, 0 not ok
bash tests/fm-watch-arm.test.sh              # 14 ok, 0 not ok

232 assertions, 0 failures. fm-watch-triage was confirmed green on two independent runs. fm-watcher-signal-safety, fm-daemon and fm-watch-arm were re-run against the pushed head after the pipeline's two fix commits landed, since those commits changed the shared bounded-stop helper.

Deterministic red-then-green

Each case was run with bin/ at the merge base 6d074ec and tests/ from this branch, in an isolated clone, then against this branch. Determinism comes from an external process holding the recovery-marker lock, which pins the watcher inside the contended section rather than racing a signal against a sub-millisecond section.

case base 6d074ec this branch
a signal inside a deferred section is held, then fires through the caller's own handler not ok - ...unwound the process instead of being held ok
a watcher signalled inside a contended critical section exits and abandons no queue lock not ok - ...never exited (self-deadlock) ok
TERM stops the watcher from every position in the poll cycle ok (control) ok
watcher shutdown stays bounded against a lock another live process never releases not ok - shutdown blocked indefinitely on a marker lock held by another live process ok
daemon shutdown bounds a TERM-ignoring watcher child not ok - daemon shutdown left a TERM-ignoring child alive ok

The third row is the row that judges the design: it passes on both sides. That is the guard against the mirror-image defect - the deferral did not buy correctness by making the watcher harder to stop.

TERM-stops proof, specifically

test_term_stops_the_watcher_from_every_loop_position starts a fresh watcher per round on its own state (a stopped watcher leaves a downtime marker that makes the next one resurface and exit early, which would quietly hollow out the sweep), waits for a completed poll so the offset is measured from inside the loop with the exit path installed, then delivers TERM at ten offsets that do not divide evenly into the 0.2s poll. Every round must exit within the bound and leave no .watch.lock behind, so a re-arm can always take over.

Restored-quarantine proof

FM_SIGNAL_DEADLOCK_SKIP and both skip sites are gone. tests/fm-watch-triage.test.sh runs 49 ok, 0 not ok with the previously-skipped assertions restored, and both formerly-quarantined cases now call assert_reaped_on_term, which turns the old skip condition into a named failure rather than a silent step-aside.

What was not run, and why

  • CI. Excluded at run start. This fork has zero Actions runs and checks never register (bin/fm-ci-probe.sh zeeshaanahmad/firstmate → none; PR fix(bin): show registered decision key and fold every stated key #3 shows "no CI checks configured"), so the ci step has no reachable end condition. Local verification above stands in for it.
  • The full bin/fm-test-run.sh suite. The pipeline's own test step covers it; the sweep above targets every suite that touches the changed files.
  • live-harness-optin guards. Nothing here is harness-dependent - the verdict comes from signal disposition and lock ownership, not from anything a vendor emits - so no live-harness guard or per-harness verification record is owed.

Pre-existing issues found, not caused by this change

  • bin/fm-ci-probe.sh resolves the wrong repository. With no argument it uses gh repo view, which returns the fork's parent (kunchenguid/firstmate) rather than origin, so it answers present on a fork where checks never register - the opposite of the answer it exists to give, and it would let the ci step poll for a check that can never arrive. bash bin/fm-ci-probe.sh → present; bash bin/fm-ci-probe.sh zeeshaanahmad/firstmate → none; gh repo view --json nameWithOwner → kunchenguid/firstmate while origin is zeeshaanahmad/firstmate. Reported separately; not fixed here.
  • tests/fm-wake-queue.test.sh's subshell half cannot hold on stock macOS Bash 3.2. Lock ownership is keyed on ${BASHPID:-$$}; BASHPID is Bash 4.0+, so on 3.2 a subshell reports its parent's pid. Confirmed failing identically at origin/main. Now gated on the capability and says so, rather than failing for the wrong reason.
  • tests/fm-watch-triage.test.sh flakes under heavy load, on both sides. Its absorb gates are wait_live "$pid" 30 - a fixed 3-second wall-clock slice, which is what the harness's own wait_watcher_beat comment warns against. Measured alternating in isolated checkouts at load ~20: this branch failed 2/2 and origin/main failed 2/2, each at a different case. At load ~5.7 the suite is 49/49 green here. Converting ~26 gates changes what those cases assert and belongs in its own task.

A hardening that was written, proved wrong, and removed

While chasing an unrelated symptom I saw fm_lock_try_acquire recurse into .wake-queue.lock.steal.steal.steal... until the filename exceeded the OS limit, and added a recursion depth bound. Its own red/green proof showed the bound was harmful: an unbounded walk of a fully abandoned steal chain legitimately reclaims the lock, and capping it turns slow, noisy recovery into a permanent refusal that fm_lock_acquire_wait would then spin on forever - a deadlock this PR would have introduced. Reverted. The observation is real and needs a fix that cleans up an abandoned chain rather than refusing it; that is its own task, deliberately not smuggled in here.

Merge reasoning (firstmate)

  • Closes the root cause of the watcher self-deadlock (upstream Watcher wedges on every arm; supervision stays down; only a session restart recovers kunchenguid/firstmate#2251/afk daemon: SIGTERM shutdown can hang and the watcher child never retries after a stale-lock refusal kunchenguid/firstmate#2270 family): signal delivery during a lock critical section now defers through fm-wake-lib.sh's single owner, and shutdown is bounded in every script the mechanical trap audit found sharing those locks (watcher cleanup, afk daemon, watch-arm handler, pr-check-migrate).
  • Proof quality: red-then-green against the merge base (self-deadlock and unbounded-shutdown red at base, green on branch) with a TERM-stops-from-every-loop-position control green on BOTH sides; 232 assertions across 7 supervision suites, 0 failures, re-run at exactly the merged head 12484d7; the stopgap watch-triage quarantine is removed and replaced with a hard assertion.
  • The pipeline independently caught two real defects in the new helper (kill -0 treating an unreaped zombie as alive; a masked escalation return code) - both auto-fixed and pinned by assertions, which is the review layer doing its job.
  • Final diff verified via the raw files API: 12 files matching the described scope; ci excluded at run start per fork facts.
  • Known follow-ups deliberately excluded and already queued as their own tasks: fixed-wall-clock gate flakiness (pre-existing on main) and a safe lock-recursion bound (first design disproven by its own test).

…the watcher

Firstmate's locks are not reentrant, and the long-running supervision
processes trap HUP/INT/TERM into an exit whose cleanup re-enters those same
locks. A signal delivered mid-section unwound the shell out of a critical
section and straight back into it, producing a watcher that held .watch.lock
and .wake-queue.lock with a frozen beacon until the session was restarted.

Make the section uninterruptible rather than the unwind survivable: record the
signal, finish the short section, release, then re-raise it against the
caller's own disposition. Deferral is opt-in per section, so a long-held
ownership lock cannot turn into a TERM-immune process, and a nested
acquisition inside a deferred section becomes bounded once a signal is pending
so shutdown never waits on another process's lifetime.

Back it with a liveness bound: watcher cleanup and the away-mode daemon's child
reap now terminate loudly instead of blocking forever.
…nd footgun

Restoring the caller's signal disposition by clearing first left a window, hit
on every section exit, where HUP/INT/TERM carried their default action; a
signal landing there killed the process with no EXIT trap and its locks still
held. Overwrite each trap instead, and clear only signals the caller had no
handler for.

Stop seeding the acquire bound from the environment: an exported value would
bound every lock wait in the process, and the callers that ignore the return
value would then proceed without the lock they believe they hold.

Record the signal-safety contract and its regression pointer in the watcher
continuity owner.
An audit of every trap-registered handler in bin/ - rather than a reading pass -
turned up two more instances of the same shape the watcher had:

- fm-pr-check-migrate.sh traps into a cleanup that re-enters the recovery-marker
  lock while holding the watch lock, so it gets the same bound and keeps its
  existing stale-lock-evidence outcome.
- fm-watch-arm.sh's signal handler stopped its child with an unbounded wait, so
  the handler that exists to stop the arm could itself hang on a child that was
  slow to go.

The bounded child stop now has one owner in fm-wake-lib.sh instead of a copy per
caller. Everything else the audit listed only releases locks in cleanup, which
never blocks, and is genuinely unaffected.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant