fix(watch): keep supervision alive across delivered wakes and parked lanes - #3
Merged
Merged
Conversation
Captured 2026-07-31 from a live home's state/.watch-cycle-exits.log. It accumulates over days and cannot be regenerated, so it is committed as raw evidence before any analysis or fix is written against it.
… lanes a check-in cadence wip checkpoint: implementation only, regression tests still to come. The arm layer never started the next watcher cycle after delivering an actionable wake (successor=none on 499 of 499 captured cycles), so supervision ended on every wake and resumed only if some adapter above it happened to re-arm. Arming is now the first act of consuming a wake. Parked lanes each carried a standing per-cycle trigger, so a fresh watcher died within seconds of every start on whichever lane was next overdue. Due lanes now batch into one home-wide check-in rate-limited by .last-parked-checkin mtime.
… parked check-in cadence Six cases across the two suites that already own this behavior: - a delivered wake arms its successor before the wake is handled - an attached arm delivers the wake its successor cycle closed on - no successor is armed for a home with nothing left to supervise - the preserved 499-record capture still shows the defect a fresh cycle no longer reproduces - parked lanes batch into one check-in per cadence window, and the cadence still comes due - a parked lane whose wait has cleared is surfaced at once inside a closed cadence
…adence Corrects the ownership claim that continuity lives only in the per-harness adapters, and records the forked-session case where Claude's Stop hook is inert by design because state/.lock still names the live pre-fork harness.
fm_test_tmproot is called through a command substitution, so its array append landed in the subshell and every suite leaked its entire temp root. Registration now goes to a file that survives the subshell, and only the sourcing shell may run the cleanup. This became load-bearing once an arm can leave a detached successor: a leaked temp root kept a live watcher supervising a home no test owned any more, and the accumulated processes destabilised timing-sensitive cases in the same suite. Also keep the away-mode daemon on the unbatched parked wake form, since it classifies one window per printed reason.
tests/fixtures/<name>/ is the repo's fixture convention: the changed-test selector resolves a fixture through the directory its consuming suite names, and a file sitting directly under tests/fixtures/ has no mapping at all.
…p root A named per-owner path is also what lets the command-substitution subshell append to the registry its parent will read, and a suite that never takes a temp root now leaves nothing behind.
…e watcher comments
This was referenced Jul 31, 2026
mattadams-dev
added a commit
that referenced
this pull request
Aug 5, 2026
…te every kill through a verified helper (#8) * test(fixtures): preserve the supervisor-duplication lab reproduction and its baseline capture Three probes decide the three unproven hypotheses from the 2026-08-04 endogenous supervisor duplication, captured against the pre-fix tree: - H2 (release-by-path unlinks a rotated successor's lock): KILLED. - H3 (pre-lock work widens the window): PROVEN, and stronger than stated - the pre-lock PR-check migration SIGTERMs the incumbent watcher, so a newcomer that has passed no singleton gate evicts the incumbent as its first act. - H1 (lock-path identity): PROVEN - two state addresses for one home each take their own lock and both survive, and both locks record the same fm-home. The lock records an identity it is never keyed by. Committed before the fix so the reproduction survives handoff and the post-fix regression tests remain falsifiable against it. * wip: singleton-at-birth for both supervision layers, plus the safe-kill owner * feat(policy): route every termination through a verified safe-kill helper pkill, killall, and grep/pgrep pipelines into kill select their target by matching text. In this fleet an agent's brief travels on argv, so those patterns match the crewmates working on the problem and the shell running the search - a recorded census counted four supervise daemons where zero existed. They are denied outright. A bare kill by pid is denied separately and routed to bin/fm-safe-kill.sh, which takes its authority from the lock that names the target rather than from process inspection. kill -0, kill -l, and job specs stay allowed: a guard that refuses every termination is the same failure one mirror over. * test(safe-kill): cover the termination owner in both directions * test+docs: regression cases for the proven cause and the singleton-at-birth contract * docs: record the singleton-at-birth contract, termination policy, and evidence * fix(singleton): keep a reused-pid lock reclaimable and stop wedging pre-upgrade locks Three corrections found by running the existing suites against the new gate: - The reclaim path dropped fm_lock_try_create's steal-owner argument, so the steal token it was holding blocked its own re-create. It removed the lock and then failed to replace it, leaving the home with no lock at all. - Treating a missing or differing fm-home as undecidable wedged every lock written before that field existed - away mode could never start again. A supervision lock lives inside the state directory it governs, so reaching it already means sharing that fleet; identity alone decides the holder, and cross-home authority stays with bin/fm-safe-kill.sh. - fm_harness_pid_alive forked basename on a path that now runs before a stop is delivered; every subprocess there is time the target keeps running after the decision to stop it. * feat(watcher): let proven health reset the wedge alarm, never elapsed time The escalation ratchet had no counterpart: proven health could not count it back down. A test-heavy lane goes static in the ordinary course of working, so it reached demand-deep-inspection just by doing its job and stayed there. Four deep inspections of one healthy lane inside twenty minutes, all four confirming health. Evidence is rendered output changing, process-tree CPU advancing, or live descendants appearing - each seeing through a blind spot of the others, and any one of them enough. Frozen on all three is what a wedge looks like. Elapsed time is never a signal: a time-based amnesty would pardon the slow wedge the alarm exists to catch. Unknown resets nothing and accelerates nothing. Resets log the count they cleared. * fix: scope the health reset to liveness alarms and stop wedging pre-upgrade daemon stops - The busy-no-completed-turn alarm asks a different question than the stale alarm: a pane producing output while completing no turn is the shape it exists to catch, so evidence-of-movement is the wrong counterpart there. Its counterpart is a completed turn, which already clears the timer. - fm-safe-kill treated a lock that records NO home as a foreign home. A lock predating that field carries none, which would leave every pre-upgrade home unable to end its own daemon. A differing home is still a proven refusal. * refactor(afk): route the away-mode daemon stop through the termination owner The stop took its authority from the daemon lock and then signalled inline. It now goes through bin/fm-safe-kill.sh, which re-derives that authority and refuses a session, an ancestor, or a pid the lock does not name. The test seam moves with it: intercepting the shell builtin no longer models the stop path, so the fixture substitutes the helper instead. * docs: name the arming authorities and the one-authority-per-mode rule * docs: one sentence per line * test(fixtures): carry the cancelled round's unresolved findings forward The review gate raised two ask-user findings and three auto-fix findings, and the round was cancelled before any of them was ruled on. The captain's order supersedes the round, not the findings, so they are recorded beside the lab reproduction rather than left to die with the run. The most serious is mine: Part 4's scoping comment says the health reset is confined to the stale alarms, but the code applies it to the busy alarm too - an absent rendered hash only blanks that one component while CPU and children are still sampled, and a busy pane's CPU advances by definition. So the busy alarm can never escalate. That is the mirror-image failure from the unwatched direction: not a guard that refuses everything, an alarm that can never sound. The test that should have caught it is vacuous on exactly that property, because its fake backend makes the evidence unreadable for an unrelated reason. Also records the sequencing the next round must follow: wait for the synced base, then classify upstream's rewrite of the six shared files as subsumed, conflicting, or complementary before writing any more code. * test(fixtures): classify this branch against upstream before rebasing The round's deliverable, written before the rebase and before any further code, per the captain's order. Two of the round's leads do not survive verification, and the same check kills both: the harness-independent successor floor and the forked-session inertness sentence are BOTH already present at the merge-base. They are the fork's own #3 work, not upstream's - the upstream diff adds neither, and upstream leaves the session-lock gate untouched. Part 1's finding is therefore not subsumed and nothing is dropped for it. The overlap set is twelve files, not the six named; a rebase does not care which list a file was on. Per-file: fm-wake-lib.sh conflicts on a "role" field both sides introduced with disjoint vocabularies and opposite write timing - resolved by renaming this lane's to supervisor-role, because its before-publish timing is load-bearing for singleton-at-birth and cannot move to a setter. This lane's fm-session-lock-lib.sh line is subsumed by an upstream rewrite and is dropped rather than re-applied to code that no longer exists. Everything else is complementary. The Part 4 reconciliation is the consequential item: upstream's positive recovery reset clears a different alarm, but it already fires on verified health and never on elapsed time. So Part 4's principle is an established property of this seam, not a parallel invention - and a beacon FRESHNESS requirement is not a time-based amnesty, because requiring recent evidence is the opposite of pardoning elapsed silence. * perf(session-lock): drop the pre-stop basename fork from the harness matcher Re-measured against upstream's rewrite rather than re-applied to it. The classification deferred this deliberately: the fork-cost change this branch carried was written for a function upstream replaced, so it was dropped in the rebase and the question re-asked of the new shape. The answer is that the new shape still forks, in fm_harness_process_matches. That matters because the predicate is a universal refusal inside fm-safe-kill.sh, evaluated after the decision to stop a process but before the signal is delivered, so every subprocess there is time the target keeps running. The window is observable, not theoretical: tests/fm-pr-check-security.test.sh catches a legacy check executing inside it, and that test failed on the rebased tree until this landed. ${comm##*/} is exactly basename for the paths ps reports here, so upstream's matching semantics are unchanged; its own fm-session-lock-ancestry suite stays green. * test(fixtures): record the one-reset-owner amendment and the design it forces The classification called the two reset mechanisms complementary. That asserted the separation rather than proving it, and left both live on one seam with nothing enforcing that they stay out of each other's state. The answer: they are genuinely different seams and must not be composed into one function. Neither signal can answer the other's question - a healthy watcher says nothing about whether one crewmate is stuck, and a computing crewmate says nothing about whether the home is supervised. A single spanning function would be a dispatcher with two disjoint branches, which hides the boundary instead of making it checkable. So one-owner is satisfied at the rule, not the function: one documented principle - alarm state clears only on positive evidence about its own subject, never on elapsed time - which upstream's reset already obeys, stated once and applied twice. The boundary becomes a fixture rather than prose, and the mutation proof grows from two mutants to four, adding the two crossings that only exist because the amendment was made. * fix(watcher): make the alarm-reset rule one owner with a proven boundary Closes both ask-user defects, the three auto-fix findings, and rebuilds Part 4 under the one-reset-owner amendment. The seam's reset rule is now stated once - alarm state clears only on positive evidence about that alarm's own subject, never on elapsed time - with two applications that have disjoint subjects and disjoint state. They stay two functions deliberately: one reset spanning both would dispatch into branches sharing no evidence and no state, hiding the boundary instead of making it checkable. The boundary is a fixture, not a comment - each reset is driven with the other's state present and must leave it untouched. Which alarm a wedge_timer_check call belongs to is now stated explicitly and fails closed. Inferring it from a withheld rendered hash was wrong invisibly: that blanks one component while CPU and descendants are still sampled, and a busy pane's CPU advances by definition, so the busy alarm could never sound. The reclaim path returns undecidable instead of announcing a peer it never verified. fm-safe-kill gives an already-gone target its own outcome, so a one-shot watcher exiting inside the check window no longer reads as an unauthorized stop. Away-mode start now requires positive evidence that nothing holds the lock before wiping a live daemon's escalation buffer. The health tree walk is portable and reports unknown where it used to fabricate a constant - twice: a GNU-only option made every tree report 1 on macOS, and a missing root reported 1 and 0 CPU for a process that does not exist. * test(health-evidence): observe the pane each case claims to observe The drivers pointed at the test shell, a live process doing work, so the frozen-pane assertions held only while the harness burned less than one clock tick between two calls. They passed standalone and failed under a full-suite sweep - a verdict that depended on ambient load rather than on its subject, which is the third way this fixture has managed to be green without measuring what it claimed. Each case now observes the process it names: a sleep for a frozen pane, a spin loop for a busy one. Verified by six consecutive runs against four competing CPU loaders with zero failures, and all five alarm-reset mutants still kill exactly their own case. * fix(policy): keep the mandated kill helper reachable in unmodelled grammar The raw fallback that scans grammar the parser cannot model tested for a bare kill verb. bin/fm-safe-kill.sh - the one command the policy directs every termination through - ends in those same bytes, so any command mentioning it inside a loop or conditional was denied. A recovery kill is written in exactly that grammar, which made the escape hatch unreachable precisely when it was needed, and silently: the refusal names a rule, not a missing capability. Drop the helper's own token before the scan, with a trailing boundary so fm-safe-killall cannot launder a killall through the exemption. The parsed path already allowed the helper; this only aligns the fallback with it. L13-L17 cover the recovery direction, K21-K23 the laundering direction. * docs(verification): record the helper-name exemption and its honest proof status Adds the refuses-too-much mutation row for the unmodelled-grammar fallback and a subsection covering how the defect surfaced - twice, against ordinary work, presenting as a rule rather than as a fault. K22 is recorded as redundantly guarded rather than dressed up as a single-mutation proof: the unstripped pattern scan and the exemption boundary each deny it alone, so only the combined mutation kills it. Baseline counts corrected to the measured 13 and 186. * test(fixtures): stop the h1 probe reporting success as exit 127 The watchers are started through a command substitution, so they are not this shell's children and the trailing bare `wait` returned 127. That leaked as the probe's exit status, so a documented verification entry point exited nonzero on a fully successful run while h2 and h3 both ended on an explicit verdict. Give h1 the same shape: derive the verdict from the two measurements it already takes and exit on it, so a caller reading the exit code reads the probe rather than the last cleanup command. * no-mistakes(review): exempt signal-0 probes and stop non-watchers reading as watchers * no-mistakes(review): exempt signal-0 from the watcher-pid guard and name the real lock holder * no-mistakes(document): document safe-kill knobs, helper exemption, and wedge-alarm reset * no-mistakes(document): point delivery-record coexistence at its real backlog owner
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Fix two defects in firstmate's supervision watcher, diagnosed by the captain from a preserved 499-record capture of a live home's watcher cycle-exit ledger (committed at tests/fixtures/watch-cycle-exits/cycle-exits.log because it accumulates over days and cannot be regenerated).
FRAMING THE CAPTAIN SET, WHICH MATTERS FOR REVIEW: the watchers were never failing. All 499 records show exit_code=0 and signal=none - the watcher is one-shot BY DESIGN and completes in order to deliver a wake. Earlier attempts treated this as a supervision failure and repaired it by hand; that framing was wrong. This is a redesign, not a fourth patch.
DEFECT 1: successor=none on 499 of 499 records. Nothing armed the next watcher cycle after a wake was delivered, so supervision ended on every delivered wake and only resumed if some adapter above the arm layer happened to re-arm. The captain specified the fix shape verbatim: 'arming the successor is the first act of consuming a wake, before handling what it woke you for.' That ordering IS the design - handling first and arming afterwards leaves a window as wide as the handling turn, which is exactly what was swallowing supervision. So in bin/fm-watch-arm.sh the successor is armed and verified before print_watch_output, deliberately. The successor is intentionally NOT the arm's child (double-forked) because it must outlive the arm that delivers the wake; that is a deliberate departure from the file's existing 'the child stays our child' comment, and it is documented in the header. It is still verified through the same live-process-plus-fresh-beacon honesty gate as an owned child, so it is not the fire-and-forget shell-& pattern the header warns about. Gated on supervision actually being needed and away mode being off, so an idle home is never left with a detached watcher.
Root cause of the secondary path, recorded but deliberately NOT fixed here: after a session fork, state/.lock still names the pre-fork harness pid, which stays alive, so the forked session's Claude Stop auto-arm hook can never claim the home and is permanently inert. Evidence from the live home: .lock=17907, a live claude process. The arm-layer fix makes continuity independent of that path; the lock behaviour itself belongs to the existing fm-upstream-issue-dossier backlog item, whose first line is already 'watcher re-arm after fork'.
A third exit shape the captain observed and told me to record rather than chase is resolved as a side effect: an ATTACHED arm could only report a genuine wake as 'watcher: FAILED - cycle ended without an actionable reason', because it could not see its watcher's output. The successor now writes its one reason to state/.watch-successor-output. before releasing the singleton, and the attached arm claims and delivers it.
DEFECT 2: 112 of 499 cycles (22%) exited within 5 seconds, all on actionable-stale, all driven by lanes deliberately parked awaiting a captain decision with idle ages in the thousands of seconds. Each parked lane carried its own independently-phased standing trigger, so a fresh watcher died within seconds of every start on whichever lane was next overdue. The captain specified the fix as his own pins-for-conditions rule: 'parked-awaiting-captain is not actionable staleness - give it a check-in cadence, not a standing trigger', and explicitly said NOT to solve it by suppressing the lane, because a declared wait can stop holding: the distinction is CADENCE, not silence. So due lanes now batch into one check-in wake, rate-limited home-wide by a state/.last-parked-checkin mtime - deliberately copying the shape the slow per-task checks already use (.last-check mtime) so the cadence survives watcher restarts, which the captain pointed at as the model. A lane whose wait has actually cleared still surfaces at once, independent of the cadence, and that is covered by its own test. Away mode keeps the unbatched one-shot form because the daemon classifies one window per printed reason.
This is guard-class code, so the captain applied the fleet mutation bar: each protection is proven by the mutation that breaks exactly its own test. That matrix is recorded in docs/verification/supervision.md - six mutations, each run twice (once to see which case it kills, once with that case removed to prove nothing else breaks). One row is honestly reported as shared by two cases guarding the same pause release.
Also in the branch, because it turned out to be load-bearing rather than incidental: tests/lib.sh's fm_test_tmproot is called through a command substitution, so its array append landed in the subshell and EVERY suite leaked its entire temp root. That was harmless until an arm could leave a detached successor - then a leaked temp root kept a live watcher supervising a home no test owned any more, and the accumulated processes destabilised a timing-sensitive case (test_watch_restart_attaches_to_healthy_peer) in the same suite. Registration now goes to a named per-owner registry file that survives the subshell, and only the sourcing shell may run the cleanup. After the fix both watcher suites run with zero stray watcher processes and zero leftover temp roots.
Known pre-existing failure, NOT caused by this branch and deliberately not fixed here: tests/fm-session-start.test.sh's Herdr husk-recovery case fails identically on the unmodified base commit f7d0d0a ('the later fleet read did not confirm the relaunched Herdr endpoint'). Verified by checking out the base commit and re-running. It is unrelated to supervision successor arming.
What Changed
bin/fm-watch-arm.shnow arms and verifies a successor watcher as the first act of consuming a wake, before the wake is printed, so supervision no longer ends on every delivered cycle. The successor is deliberately detached (double-forked) so it outlives the arm that delivers the wake, is still confirmed through the same live-process-plus-fresh-beacon gate as an owned child, and is armed only when supervision is actually needed and away mode is off. Each successor writes its one reason to a pid-keyedstate/.watch-successor-output.<pid>file, which an attached arm claims and delivers - an attached close carrying a genuine wake is no longer reported ascycle ended without an actionable reason. The outcome is recorded per cycle in thestate/.watch-cycle-exits.logledger.bin/fm-watch.shgives parked lanes (declared external waits and captain holds) a check-in cadence instead of a standing per-cycle trigger: every lane due in one poll is batched into a single wake, rate-limited home-wide by thestate/.last-parked-checkinmtime (FM_PARKED_CHECKIN_SECS, defaulting toFM_PAUSE_RESURFACE_SECS), and only the lanes carried in an emitted check-in get their per-lane cadence stamped. A lane whose pause has actually cleared still surfaces immediately, and away mode keeps the unbatched one-shot form the daemon's one-window-per-reason triage expects.tests/lib.shregisters temp roots through a named per-owner registry file that survives the command substitutionfm_test_tmprootis called through, so suites stop leaking temp roots (which, once an arm could leave a detached successor, kept live watchers supervising abandoned homes); the 499-record live cycle-exit capture is preserved attests/fixtures/watch-cycle-exits/cycle-exits.logbecause it accumulates over days and cannot be regenerated; new cases intests/fm-watcher-lock.test.shandtests/fm-watch-triage.test.shcover successor arming, attached-cycle delivery, the away-mode gate and the parked cadence; a pre-existing race intest_watch_restart_attaches_to_healthy_peer(the TERM-resistant peer was signalled before its handler was installed) is fixed; anddocs/verification/supervision.mdrecords the capture analysis plus an eight-mutation guard-class matrix, each mutation run at least twice, with the sharedclear_pause_staterow and a surviving mutant at the redundant pause-release site reported honestly.Note on diff size: this branch is cut from a point behind the gate's
main, so the delta also carries eight already-merged upstream commits (kunchenguid#1303, kunchenguid#1327, kunchenguid#1328, kunchenguid#1339, kunchenguid#1349, kunchenguid#1350, kunchenguid#1356, kunchenguid#1358). The supervision change itself is the delta fromf7d0d0a.Risk Assessment
✅ Low: The round-1 blocking defect is fixed with a test that can no longer pass over it, the two arm-layer race gaps are closed with a pid handshake I verified from source (exec preserves the pid, and nothing re-execs before the lock claim), the registry is now unpredictable and swept by every lib-sourcing suite, and only a trivial prune-reachability tidiness item remains.
Testing
I ran the two suites that own this contract repeatedly - fm-watcher-lock four times and fm-watch-triage three times, all clean - then proved the intent at the operator surface by driving the real arm and watcher against an identical staged home under base-commit bin and branch bin, capturing arm stdout, the cycle-exit ledger record, and whether any watcher still supervises the home afterwards. The base reproduces all three captured shapes exactly and the branch fixes each. I also found the away-mode half of the successor gate had no test, so I added a focused case and mutation-proved it kills exactly that protection, taking the suite to 35 cases. The change is shell and CLI only with no rendered surface, so evidence is CLI transcripts rather than screenshots. Every run ended with zero leftover test temp roots and zero stray watcher processes, confirming the tests/lib.sh registry fix. The only working-tree change left is the added test case; bin/fm-watch-arm.sh was restored byte-identical after the mutation check.
Evidence: DEFECT 1 - successor arming, base vs branch
Evidence: Third exit shape - attached arm delivers its successor cycle wake
Evidence: DEFECT 2 - parked check-in cadence, base vs branch
Evidence: Away-mode successor gate
Evidence: Mutation proof for the added away-mode case
Evidence: Preserved 499-record capture stats
Evidence: Evidence index
Evidence: Suite log - fm-watcher-lock 35/35
Evidence: Suite log - fm-watch-triage 42/42
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
⏭️ **Rebase** - skipped
tests/fm-secondmate-harness.test.sh- merge conflict rebasing onto origin/mainbin/fm-watch.sh:357- The parked-lane accumulator loses its record separator, so the batched check-in only ever carries one lane.parked_due=$(printf '%s%s\t%s\n' "$parked_due" "$win" "$reason")runs through a command substitution, which strips the trailing newline; the next append then concatenates directly onto the previous line. With 3 due lanes the value isw1<TAB>reason1w2<TAB>reason2w3<TAB>reason3(verified by running the exact loop in bash).flush_parked_checkin(bin/fm-watch.sh:378) therefore performs exactly onereaditeration: only the first lane is enqueued viafm_wake_append stale "$win" "$reason", only its.paused-resurfaced-<key>marker is stamped, thecombined="$combined; $reason"join never executes, andwakeprints a tab-glued mash of every lane's reason attributed to the first window. This is the exact multi-lane scenario the fix targets (the capture's 112 sub-5s exits were independently-phased parked lanes), and it also silently violates the function's own documented contract that lanes not included in an emitted check-in stay unstamped and due - lanes 2..N are never stamped, so they remain permanently due. tests/fm-watch-triage.test.sh:709 does not catch it becausegrep -F "test:fm-park$i"matches the glued line andgrep -c '^stale:'is 1 either way. Fix: preserve the newline, e.g.parked_due="${parked_due}${win}"$'\t'"${reason}"$'\n', and tighten the test to assert one wake-queue record per lane and a;-joined reason.bin/fm-watch-arm.sh:385- When two arms are attached to the same watcher - a topology this file explicitly supports ("A live cycle already present means re-arm attaches", header line 42) - both reach the close path inattach_and_waitsimultaneously. Both passwatch_output_has_wake "$out", thenmv -f "$out" "$claimed"succeeds for one and fails for the other. The loser falls back toclaimed=$out, which no longer exists, soprint_watch_outputemits nothing and the function still doesreturn 0. The arm exits 0 with completely empty stdout - precisely the "clean empty completion that an adapter could mistake for a no-op" the function comment (line 361) says it must never produce, and which the non-wake path below deliberately turns into a typed nonzero failure. Fix: onmvfailure, do not treat the wake as claimed; fall through towait_for_healthy_successorso the arm either attaches to the successor the winner just armed or fails loudly.bin/fm-watch-arm.sh:325-arm_successorbinds its pending output file to whichever pid wins the singleton (HEALTHY_PID), not to the process it actually forked. When another watcher wins - a concurrentarm_successorfrom a second attached arm, or an adapter-launched successor (docs/architecture.md:63 notes Pi and OpenCode launch their own singleton successor from child-close handlers) - the arm still renames its own tmp to<winner-pid>. The winner's stdout goes elsewhere, so the published file ends up holding the loser watcher'swatcher: already running pid Nline instead of the winner's reason, while the winner's real reason (whose fd points at the unlinked tmp on the[ -e "$target" ]branch) never lands on disk. A later arm attaching to that pid then failswatch_output_has_wakeand can reportwatcher: FAILED - cycle ended without an actionable reasonfor a genuine wake - the third exit shape this change set out to resolve. The reason still reachesstate/.wake-queue, so nothing is lost durably; only the arm-layer delivery degrades. Minimal fix: capture the pid of the process this arm actually started and only publish the output file when the confirmed healthy pid matches it, recordingattached:<pid>otherwise. Cleaner shared boundary: have fm-watch.sh write its own reason tostate/.watch-successor-output.$WATCHER_PID, so the binding is authoritative regardless of who forked it.bin/fm-watch-arm.sh:305- Recording the residual bound of the arm-layer fix, which the intent explicitly authorizes as containment for the deferredstate/.lockfork defect. A successor is armed only from an arm process, so with a permanently inert adapter (the observed forked-Claude-session case, where the Stop auto-arm hook can never claim the home) the chain is exactly one cycle deep: the detached successor covers the handling turn, wakes, writes its reason to.watch-successor-output.<pid>andstate/.wake-queue, exits, and no further successor is armed. The blind window is bounded by the successor's own cycle rather than eliminated, so the intent's "makes continuity independent of that path" is stronger than what the arm layer alone delivers. No action requested - the durable fix is correctly deferred to the fm-upstream-issue-dossier item, and the unclaimed output file is reaped byprune_successor_outputsat FM_WATCH_SUCCESSOR_OUT_TTL.tests/lib.sh:73- The cleanup registry path is fully predictable (${TMPDIR:-/tmp}/fm-test-cleanup.<pid>) and its contents driverm -rfon any line starting with/(tests/lib.sh:84). On a shared host with a sticky/tmp, another user can pre-create that path as a symlink; the sticky bit prevents therm -fat line 74 from removing it, the>>infm_test_tmprootthen follows the symlink and writes as the test user, andfm_test_cleanuprecursively deletes every absolute path it reads back. The comment's two stated requirements (subshell-visible path, nothing created for a suite that takes no temp root) are both satisfiable safely:FM_TEST_CLEANUP_REGISTRY=$(mktemp "${TMPDIR:-/tmp}/fm-test-cleanup.XXXXXX")plusexportgives an unpredictable, atomically-created path that subshells still see, andfm_test_cleanupalready removes it so it leaves no litter.tests/fm-kimi-harness.test.sh:22-trap fm_test_cleanup EXITis now installed at source time, so any suite that installs its own EXIT trap afterwards replaces it. tests/fm-kimi-harness.test.sh:22 and tests/fm-afk-pi-herdr-return-e2e.test.sh:62 both do this without callingfm_test_cleanup. Their temp roots are still removed (eachrm -rf "$TMP_ROOT"directly), so the branch's zero-leftover-temp-roots claim holds, but the registry file itself is never unlinked, leaving one stray${TMPDIR:-/tmp}/fm-test-cleanup.<pid>per run. Addingfm_test_cleanupto both cleanup functions - which the library comment at tests/lib.sh:57 already prescribes - closes it.🔧 Fix: fix parked check-in batching and successor output binding
1 info still open:
bin/fm-watch-arm.sh:309-prune_successor_outputsis only ever called fromarm_successor(the sole call site, line 309), and it sits after thesuccessor_wanted || { printf 'none'; return 0; }early return on line 308. Once a home stops needing supervision - the exact state where leftover files accumulate and nothing new claims them - the sweeper becomes unreachable, so any unclaimed.watch-successor-output.<pid>(plus orphan.pending.*/.forkpid.*from an arm killed mid-arm_successor) stays instate/indefinitely rather than aging out at FM_WATCH_SUCCESSOR_OUT_TTL. Not a correctness or supervision risk - the files are small, AGENTS.md:114 marks them watcher internals, and every live path still cleans up its own - but it defeats the sweeper's stated purpose in a codebase that otherwise size-caps every state artifact. Moving theprune_successor_outputscall above thesuccessor_wantedgate makes it run on every arm cycle, including the ones that decline to arm.🔧 **Test** - 1 issue found → auto-fixed ✅
tests/fm-watcher-lock.test.sh:447- tests/fm-watcher-lock.test.sh:447test_watch_restart_attaches_to_healthy_peerwas flaky for a cause unrelated to the leaked-temp-root explanation recorded in docs/verification/supervision.md. The case stages a TERM-resistant peer withnode -e 'process.on("SIGTERM", ...)'and immediately launchesfm-watch-arm.sh --restart, whose first act is to TERM the recorded lock pid. Nothing waited for node to finish starting, so when node lost that race the TERM landed before the handler existed, the peer died, and the arm correctly started a fresh watcher - reported asrestart did not attach to the verified healthy peer. Reproduced 3 times in 18 runs on an idle 32-core host (load 0.15), with a diagnostic confirmingpeer DEAD (TERM landed before its handler). The case is unchanged since a323c2b, so it predates this branch. Fixed by having the peer publish a readiness marker from inside the handler-installed process and having the case wait for it before anything can signal the peer; 15/15 isolated runs and two clean full-suite runs since.bash tests/fm-watcher-lock.test.sh- full suite, two consecutive clean runs (34 cases each), includingtest_delivered_wake_arms_its_successor_before_it_is_handled,test_attached_arm_delivers_its_successor_cycle_wake,test_successor_is_not_armed_without_supervision_need,test_preserved_capture_still_shows_the_defect_a_fresh_cycle_no_longer_hasbash tests/fm-watch-triage.test.sh- full suite clean, includingtest_parked_lanes_batch_into_one_checkin_on_a_shared_cadence,test_a_cleared_park_is_noticed_promptly_inside_a_closed_cadence,test_afk_present_reverts_watcher_to_one_shotManual E2E before/after:git archive f7d0d0a bin | tar -x -C <evidence>/base-tree, then ranbin/fm-watch-arm.shandbin/fm-watch.shfrom BOTH the base tree and the branch against identically staged homes (/tmp/no-mistakes-evidence/01KYVWHDGTDH63YJJC091K85AP/supervision-e2e-demo.sh), capturing arm stdout, thestate/.watch-cycle-exits.logrecord, andps -o pid=,ppid=,args=proof of the surviving detached successorManual E2E parked-lane cadence: three overdue parked lanes run through three watcher cycles per variant with per-lane throttles cleared between cycles, then.last-parked-checkinaged past its window, then one lane's pause verb cleared inside a closed cadenceLeak check after every run:ls -d /tmp/fm-watcher-lock-tests.* /tmp/fm-watch-triage* /tmp/fm-test-cleanup.* /tmp/fm-supervision-e2e.*andps -eo pid,args | grep c5e834eacdf2.*fm-watch- all emptyFlake reproduction and fix verification fortest_watch_restart_attaches_to_healthy_peer: 18 runs before the fix (3 failures, with a peer-liveness diagnostic), 15 isolated runs after the fix (0 failures)🔧 Fix: wait for TERM-resistant peer handler before restarting arm
✅ Re-checked - no issues remain.
tests/fm-watcher-lock.test.sh - 4 consecutive full runs, 34/34 cases each, all cleantests/fm-watch-triage.test.sh - 3 consecutive full runs, 42/42 cases each, all cleanManual E2E defect-1 contrast via demo-successor-arming.sh against base-commit bin and branch binManual E2E attached-delivery contrast via demo-attached-delivery.shManual E2E defect-2 contrast via demo-parked-cadence.sh over four successive watcher cyclesManual verification of the away-mode successor gate with state/.afk presentAdded test_successor_is_not_armed_in_away_mode to tests/fm-watcher-lock.test.sh; suite now 35/35 okMutation proof: removed only the away-mode line from successor_wanted() in bin/fm-watch-arm.sh, reran the suite, then restored the fileRe-derived preserved capture stats from tests/fixtures/watch-cycle-exits/cycle-exits.logLeak hygiene after every run: zero fm-* test temp roots in /tmp and no test-owned watcher processes🔧 **Document** - 1 issue found → auto-fixed (3) ✅
docs/verification/supervision.md:222- The guard-class mutation matrix has no row for the away-mode half of the successor gate. The working tree adds test_successor_is_not_armed_in_away_mode to tests/fm-watcher-lock.test.sh (uncommitted), which guards a distinct protection - .afk suppressing the successor so a detached watcher never competes with the supervise-daemon - and the intent states the fleet mutation bar requires each protection be proven by the mutation that breaks exactly its own test. That new case also shifts the '33 cases clean' figure recorded for the three lock-suite rows, which was measured before it existed. Closing this needs the mutation matrix re-run (mutating bin/fm-watch-arm.sh and running both suites), which is outside a documentation phase, so I did not invent a row or adjust the recorded counts.🔧 Fix: no doc changes; mutation matrix re-measurement unfinished
1 error still open:
docs/verification/supervision.md:225- The guard-class mutation matrix was NOT updated this round: the re-measurement run was still on its first mutation when I was required to finalize, so I have suite baselines but no per-mutation kill results. I made no edit rather than publish a partially re-measured table, which the captain explicitly ruled worse than an obviously old one.WHAT IS ESTABLISHED BY MEASUREMENT: both suites were re-run clean against the current tree (HEAD 58dae2b). tests/fm-watcher-lock.test.sh emits 35 'ok -' cases, exit 0. tests/fm-watch-triage.test.sh emits 42 'ok -' cases, exit 0. Since the table's 'N cases clean' figures are baseline-minus-the-removed-case, this already proves the three lock-suite rows reading '33 cases clean' (lines 227, 228, 229, 232 - four rows total) are stale: 33 = a 34-case baseline minus one, and the baseline is now 35 because test_successor_is_not_armed_in_away_mode was added. Those rows must read 34 once confirmed by an actual mutation run. The triage figures (41 on lines 230/231, and 40 on line 233) are arithmetically consistent with the current 42-case baseline, but consistency is not measurement and they were not re-run.
WHAT IS BUILT AND READY: /tmp/no-mistakes-evidence/mutate.py implements the full matrix with the two-run method. All eight mutations are defined and every anchor was verified to match exactly once in its target file. The new away-mode row mutates successor_wanted in bin/fm-watch-arm.sh:297, replacing the '[ -e "$STATE/.afk" ] && return 1' guard with ':' while leaving fm_supervision_needed intact, so a successor is armed even in away mode. The shared-pause-release row targets clear_pause_state in bin/fm-watch.sh:398-403 (not the similar rm at line 477), because that is the release primitive both test_a_cleared_park_is_noticed_promptly_inside_a_closed_cadence and test_secondmate_unpause_clears_pause_tracking depend on. The harness removes each killed case's invocation line and re-runs, looping until the suite goes clean, so a mutation shared by more than one case is recorded honestly instead of being forced into a one-to-one claim. One run was observed working end to end before I stopped it.
TO CLOSE: run 'python3 /tmp/no-mistakes-evidence/mutate.py' (about 35 minutes; roughly 18 suite runs at ~2 minutes each) and transcribe matrix.json into the table, including the new away-mode row. Caveat: the harness lives in ephemeral /tmp and may not survive; if it is gone it must be rebuilt before the numbers can be re-measured.
SAFETY STATE: the interrupted run left mutation 1 applied to bin/fm-watch-arm.sh. I inspected the diff, confirmed it was exactly that mutation and nothing else, and reverted it. The worktree is verified clean at 58dae2b with zero modified files, so no behavioral change escaped this phase.
🔧 Fix: re-measure guard-class mutation matrix and add away-mode row
1 info still open:
bin/fm-watch.sh:477- Surviving mutant found while re-measuring the matrix, reported rather than fixed because this is a documentation phase and closing it means adding a test.While determining which release site the 'Stop releasing a lane whose pause verb has gone' row actually describes, I measured both candidates. The
elsebranch insidesurface_nonterminal_stale(bin/fm-watch.sh:477) - the site that literally matches that label, releasing.paused-<key>,.paused-rechecked-<key>and.paused-resurfaced-<key>whenstatus_is_paused_or_captain_heldis false - can be replaced with:and tests/fm-watch-triage.test.sh still passes in full: 42ok -lines, exit 0, zero cases killed. Measured in a private copy; log at /tmp/no-mistakes-evidence/01KYVWHDGTDH63YJJC091K85AP/matrix/keep-cleared-pause-verb.run1.log.The row was therefore published against the site that IS guarded,
clear_pause_state(bin/fm-watch.sh:403), which four cases depend on. That is honest and measured, but it means the record's 'one mutation per protection' framing now has a known hole: a lane that surfaces as non-terminal stale after its declared pause verb disappears has its pause markers cleared by an untested branch.Suggested follow-up, outside this change: a triage case driving a window through
surface_nonterminal_stalewith a status whose pause verb has been replaced by a non-paused one, asserting the three.paused-*markers are gone. Then the matrix can carry a genuine row for that protection instead of relying on the broaderclear_pause_statemutation.🔧 Fix: record surviving mutant at redundant pause-release site
✅ Re-checked - no issues remain.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.