Skip to content

fix: reclaim orphaned watcher arms on the next park - #6335

Merged
kunchenguid merged 8 commits into
mainfrom
fm/fm-restart-orphan-watcher-arm-r1
Oct 2, 2026
Merged

kunchenguid merged 8 commits into
mainfrom
fm/fm-restart-orphan-watcher-arm-r1

Conversation

@kunchenguid

@kunchenguid kunchenguid commented Oct 1, 2026 •

Copy link
Copy Markdown
Owner

Intent

"agree", to: do item 7 first, then items 1-3 together as one PR, leave 4-6 for later.
Item 7, as authorized: recheck that the orphaned pre-restart watcher arms persist in axi-a1, eddies-wallet-e1, mini-default, sshhip-h7; if they do, fix the root cause so a second-mate restart does not leave the old arm as an orphan owner with the new arm attaching and polling (a leak fix, its own PR; do not kill the live orphans by hand as the fix, though cleaning them up after the fix lands is fine through the guarded path).

Context (data/fm-poll-churn-post-update-measure-s1/report.md, finding 4): after the 2026-10-01 fleet update to 549e07f restarted every second mate, ps at 10:04Z showed two fm-watch-arm.sh processes in each of axi-a1, mini-default, sshhip-h7 and eddies-wallet-e1. The owner arm is an orphan (ppid 1) started about 04:28Z by the pre-restart park host and is still the parent of the live fm-watch.sh. The second arm is a child of the new park host running attach_and_wait (fm-watch-arm.sh about lines 378-400), calling fm_watcher_healthy (fm-wake-lib.sh about line 187: four $(cat ...), fm_pid_identity via ps, fm_path_age) then sleep "$ATTACH_POLL" (0.5 s) at about 1.67 Hz, forking 7.4-7.9 children/s each - about 100-170 process creations/s host-wide, 8-13% of the new idle floor. fmdev-f1 has a single arm under its park host, the pre-update topology; before the update every home had exactly one arm, a child of its park host. fm-watch-arm.sh did not change in 549e07f. The report did not test whether the orphan ends at the watcher's next natural turnover or whether the attached arm still delivers wakes correctly.

Decisions and constraints from the earlier validation of this PR, which stand: the take-over restores the acknowledged recovery episode only when the taken-over arm's cycle-ledger row for that exact arm and watcher shows signal=TERM, waited for within a short bound; a missing row or any other signal deliberately leaves the downtime for the fresh cycle's ordinary recovery wake (one extra recovery wake is the accepted safe direction, never a hidden failure). The --restart first-cycle branch used by the OpenCode and omp hosts predates this PR (#5941) and stays out of scope as a documented follow-up, as does the supervision-host-off Stop-hook handling successor. No watcher, signal-trap, or polling-cadence change; no shared sleep abstraction.

This update: the PR was closed by accident and reopened for a merge decision. Current origin/main, which now contains #6338 (the Pi 1.0.0 Calm test fix that cleared an unrelated red check), was merged into this branch with an ordinary merge commit - no rebase, no amend, no force-push - so the pipeline publishes the update and fresh CI runs. Landing timing is coordinated with the main firstmate.

What Changed

  • Persist detached successor arm identities and take over their watcher cycles on the next non---restart park; retain interrupted takeovers for retry and stop successors that cannot be recorded.
  • Restore acknowledged recovery state only after bounded confirmation of TERM for the exact arm/watcher pair and no intervening wake append; otherwise preserve recovery and delivered wakes.
  • Add regression coverage for takeover ownership, interruption, persistence failure, queued work, and watcher self-exit; document unchanged --restart and host-off Stop-hook paths.

Risk Assessment

✅ Low: The takeover and persistence changes are bounded, preserve the authorized recovery fallback, and introduce no substantiated material defects.

Testing

Focused baseline checks and seven live lifecycle scenarios passed. The pre-fix negative control reproduced acknowledgement erasure. Product transcripts, topology, persisted state, and ledgers were captured; all disposable resources were removed without touching the live fleet.

  • Live validation: ✅ go - 7 of 7 scenarios driven live against the product
Scenario Result Live Evidence
Restart the primary after a main-only close; the next park replaces the orphan with a host-owned cycle and delivers the next decision ✅ pass live live-topology.log; live-after-restart.txt
Interrupt a takeover while its marker lock is held; a later park retries using the retained successor identity ✅ pass live live-interruption.log
Make the successor-record destination a writable directory; the close remains unchanged and the unrecorded successor is cleaned up ✅ pass live live-record-failure.log; live-record-failure.txt; live-record-rearmed.txt
Request takeover using a process that does not own the watcher; the arm attaches without stopping that watcher ✅ pass live live-watch-scenarios.log
Take over a cycle with work still queued; the fresh cycle surfaces that work ✅ pass live live-watch-scenarios.log
Let the old watcher fail after the handover snapshot but before TERM; genuine downtime produces a recovery wake ✅ pass live live-watch-scenarios.log; self-exit-negative-control.log
Prevent the owner from recording its TERM row; takeover proceeds but conservatively surfaces recovery ✅ pass live live-missing-row.log
Evidence: Restart replaces the orphan owner and preserves acknowledgement

Source: Restart replaces the orphan owner and preserves acknowledgement

Before primary restart:
  PID  PPID ARGS
91946     1 bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch-arm.sh
91986 91946 bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh

left record: 91946	Thu Oct  1 14:09:57 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch-arm.sh

recovery: acked:handling:56763.1790888997.CrEpp7

After primary restart, before park:
  PID  PPID ARGS
91946     1 bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch-arm.sh
91986 91946 bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh

After restarted primary parks:
  PID  PPID ARGS
66732 66716 bash bin/fm-supervision-host.sh park
67145 66732 bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch-arm.sh --take-over 91946
68273 67145 bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh

host record: host	66732	Thu Oct  1 14:12:25 2026     bash bin/fm-supervision-host.sh park
arm	67145	Thu Oct  1 14:12:25 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch-arm.sh --take-over 91946

recovery: acked:handling:56763.1790888997.CrEpp7

host ledger: 1790888907	start	gen=host-3491-1790888907	primary=claude
1790888910	to-main	the supervision session could not take this wake: the dialog mirror could not be read; this wake is yours
1790888960	start	gen=host-56093-1790888960	primary=claude
1790888997	pass-through	attended	main-only	signal: ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/.l/state/demo.status
1790889145	start	gen=host-66732-1790889145	primary=claude
1790889145	take-over	arm=91946

cycle ledger: arm_pid=3840	watcher_pid=3883	origin=started	started_at=1790888907	ended_at=1790888908	exit_code=0	signal=none	reason=actionable-signal	beacon_age=1	lock_before=pid:3883|identity:Thu Oct  1 14:08:27 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=started:6472
arm_pid=6439	watcher_pid=6472	origin=started	started_at=1790888908	ended_at=1790888910	exit_code=143	signal=TERM	reason=arm-interrupted	beacon_age=1	lock_before=pid:6472|identity:Thu Oct  1 14:08:28 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=none
arm_pid=56634	watcher_pid=56763	origin=started	started_at=1790888960	ended_at=1790888997	exit_code=0	signal=none	reason=actionable-signal	beacon_age=1	lock_before=pid:56763|identity:Thu Oct  1 14:09:20 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=started:91986
arm_pid=91946	watcher_pid=91986	origin=started	started_at=1790888997	ended_at=1790889145	exit_code=143	signal=TERM	reason=signal-exit	beacon_age=0	lock_before=pid:91986|identity:Thu Oct  1 14:09:57 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=none
arm_pid=67145	watcher_pid=91986	origin=attached	started_at=1790889145	ended_at=1790889145	exit_code=unknown	signal=unknown	reason=taken-over	beacon_age=0	lock_before=pid:91986|identity:Thu Oct  1 14:09:57 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=started:68273

Observed: previous orphan arm exited; fresh watcher belongs to new host arm; acknowledgement retained, no recovery close.
Evidence: Real Claude reports the next decision close after restart

Source: Real Claude reports the next decision close after restart


⏺ Running the corrected command (no trailing argument), leaving it running.

  Ran 1 shell command

⏺ The corrected park command is running in the background. I'll leave it running and take no other actions, reporting
  its line only if it closes.

✻ Worked for 7s · done 2:12 PM

⏺ Background command "Run supervision host in park mode" completed (exit code 0)

⏺ The park command closed. Let me read its output line to report it.

  Read 1 file

⏺ The park command closed cleanly (exit code 0). Reporting its line as instructed - not acting on it and not stopping
  any successor:

  watcher: started pid=68273 (beacon fresh)
  signal: ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/.l/state/demo.status

  [exited with code 0]

  The supervision host started a watcher (pid=68273, beacon fresh), then exited at code 0 after a signal event on
  demo.status. Per your instruction I'm reporting this line only - not acting on the signal and not stopping any
  successor.

✻ Churned for 8s · done 2:13 PM

────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
❯ restart it
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
  Opus 4.8 ctx window   ▉               63.1k / 1.0M
  ⏵⏵ auto mode on (shift+tab to cycle) · ← for agents
Evidence: Interrupted takeover retains its record and succeeds on retry

Source: Interrupted takeover retains its record and succeeds on retry

Held recovery-marker lock during takeover; interrupted lab host 81907 with TERM; left arm 50979 watcher 51140
Host cleanup retained byte-identical successor record and old arm/watcher remain alive.
1790888907	start	gen=host-3491-1790888907	primary=claude
1790888910	to-main	the supervision session could not take this wake: the dialog mirror could not be read; this wake is yours
1790888960	start	gen=host-56093-1790888960	primary=claude
1790888997	pass-through	attended	main-only	signal: ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/.l/state/demo.status
1790889145	start	gen=host-66732-1790889145	primary=claude
1790889145	take-over	arm=91946
1790889182	pass-through	attended	main-only	signal: ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/.l/state/demo.status
1790889282	start	gen=host-81907-1790889282	primary=claude
1790889282	take-over	arm=50979

Next park retries takeover: old arm gone, watcher parent is host-owned arm, acknowledged episode unchanged.
  PID  PPID ARGS
41930 41922 bash bin/fm-supervision-host.sh park
42344 41930 bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch-arm.sh --take-over 50979
43507 42344 bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh

1790888907	start	gen=host-3491-1790888907	primary=claude
1790888910	to-main	the supervision session could not take this wake: the dialog mirror could not be read; this wake is yours
1790888960	start	gen=host-56093-1790888960	primary=claude
1790888997	pass-through	attended	main-only	signal: ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/.l/state/demo.status
1790889145	start	gen=host-66732-1790889145	primary=claude
1790889145	take-over	arm=91946
1790889182	pass-through	attended	main-only	signal: ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/.l/state/demo.status
1790889282	start	gen=host-81907-1790889282	primary=claude
1790889282	take-over	arm=50979
1790889337	start	gen=host-41930-1790889337	primary=claude
1790889337	take-over	arm=50979

arm_pid=3840	watcher_pid=3883	origin=started	started_at=1790888907	ended_at=1790888908	exit_code=0	signal=none	reason=actionable-signal	beacon_age=1	lock_before=pid:3883|identity:Thu Oct  1 14:08:27 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=started:6472
arm_pid=6439	watcher_pid=6472	origin=started	started_at=1790888908	ended_at=1790888910	exit_code=143	signal=TERM	reason=arm-interrupted	beacon_age=1	lock_before=pid:6472|identity:Thu Oct  1 14:08:28 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=none
arm_pid=56634	watcher_pid=56763	origin=started	started_at=1790888960	ended_at=1790888997	exit_code=0	signal=none	reason=actionable-signal	beacon_age=1	lock_before=pid:56763|identity:Thu Oct  1 14:09:20 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=started:91986
arm_pid=91946	watcher_pid=91986	origin=started	started_at=1790888997	ended_at=1790889145	exit_code=143	signal=TERM	reason=signal-exit	beacon_age=0	lock_before=pid:91986|identity:Thu Oct  1 14:09:57 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=none
arm_pid=67145	watcher_pid=91986	origin=attached	started_at=1790889145	ended_at=1790889145	exit_code=unknown	signal=unknown	reason=taken-over	beacon_age=0	lock_before=pid:91986|identity:Thu Oct  1 14:09:57 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=started:68273
arm_pid=67145	watcher_pid=68273	origin=started	started_at=1790889145	ended_at=1790889181	exit_code=0	signal=none	reason=actionable-signal	beacon_age=1	lock_before=pid:68273|identity:Thu Oct  1 14:12:25 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=started:51140
arm_pid=82196	watcher_pid=51140	origin=attached	started_at=1790889282	ended_at=1790889298	exit_code=143	signal=TERM	reason=arm-interrupted	beacon_age=2	lock_before=pid:51140|identity:Thu Oct  1 14:13:02 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:51140|identity:Thu Oct  1 14:13:02 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	successor=none
arm_pid=50979	watcher_pid=51140	origin=started	started_at=1790889182	ended_at=1790889337	exit_code=143	signal=TERM	reason=signal-exit	beacon_age=1	lock_before=pid:51140|identity:Thu Oct  1 14:13:02 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=none
arm_pid=42344	watcher_pid=51140	origin=attached	started_at=1790889337	ended_at=1790889338	exit_code=unknown	signal=unknown	reason=taken-over	beacon_age=2	lock_before=pid:51140|identity:Thu Oct  1 14:13:02 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=started:43507
Evidence: Directory destination retains cleanup ownership and allows fresh rearm

Source: Directory destination retains cleanup ownership and allows fresh rearm

Writable directory at left-record destination: close passes through unchanged, successor-unrecorded logged, directory empty, all recorded lab arms/watchers exited.
1790888907	start	gen=host-3491-1790888907	primary=claude
1790888910	to-main	the supervision session could not take this wake: the dialog mirror could not be read; this wake is yours
1790888960	start	gen=host-56093-1790888960	primary=claude
1790888997	pass-through	attended	main-only	signal: ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/.l/state/demo.status
1790889145	start	gen=host-66732-1790889145	primary=claude
1790889145	take-over	arm=91946
1790889182	pass-through	attended	main-only	signal: ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/.l/state/demo.status
1790889282	start	gen=host-81907-1790889282	primary=claude
1790889282	take-over	arm=50979
1790889337	start	gen=host-41930-1790889337	primary=claude
1790889337	take-over	arm=50979
1790889353	pass-through	attended	main-only	signal: ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/.l/state/demo.status
1790889354	pass-through	successor-unrecorded	signal: ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/.l/state/demo.status

arm_pid=3840	watcher_pid=3883	origin=started	started_at=1790888907	ended_at=1790888908	exit_code=0	signal=none	reason=actionable-signal	beacon_age=1	lock_before=pid:3883|identity:Thu Oct  1 14:08:27 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=started:6472
arm_pid=6439	watcher_pid=6472	origin=started	started_at=1790888908	ended_at=1790888910	exit_code=143	signal=TERM	reason=arm-interrupted	beacon_age=1	lock_before=pid:6472|identity:Thu Oct  1 14:08:28 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=none
arm_pid=56634	watcher_pid=56763	origin=started	started_at=1790888960	ended_at=1790888997	exit_code=0	signal=none	reason=actionable-signal	beacon_age=1	lock_before=pid:56763|identity:Thu Oct  1 14:09:20 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=started:91986
arm_pid=91946	watcher_pid=91986	origin=started	started_at=1790888997	ended_at=1790889145	exit_code=143	signal=TERM	reason=signal-exit	beacon_age=0	lock_before=pid:91986|identity:Thu Oct  1 14:09:57 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=none
arm_pid=67145	watcher_pid=91986	origin=attached	started_at=1790889145	ended_at=1790889145	exit_code=unknown	signal=unknown	reason=taken-over	beacon_age=0	lock_before=pid:91986|identity:Thu Oct  1 14:09:57 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=started:68273
arm_pid=67145	watcher_pid=68273	origin=started	started_at=1790889145	ended_at=1790889181	exit_code=0	signal=none	reason=actionable-signal	beacon_age=1	lock_before=pid:68273|identity:Thu Oct  1 14:12:25 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=started:51140
arm_pid=82196	watcher_pid=51140	origin=attached	started_at=1790889282	ended_at=1790889298	exit_code=143	signal=TERM	reason=arm-interrupted	beacon_age=2	lock_before=pid:51140|identity:Thu Oct  1 14:13:02 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:51140|identity:Thu Oct  1 14:13:02 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	successor=none
arm_pid=50979	watcher_pid=51140	origin=started	started_at=1790889182	ended_at=1790889337	exit_code=143	signal=TERM	reason=signal-exit	beacon_age=1	lock_before=pid:51140|identity:Thu Oct  1 14:13:02 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=none
arm_pid=42344	watcher_pid=51140	origin=attached	started_at=1790889337	ended_at=1790889338	exit_code=unknown	signal=unknown	reason=taken-over	beacon_age=2	lock_before=pid:51140|identity:Thu Oct  1 14:13:02 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=started:43507
arm_pid=42344	watcher_pid=43507	origin=started	started_at=1790889338	ended_at=1790889353	exit_code=0	signal=none	reason=actionable-signal	beacon_age=1	lock_before=pid:43507|identity:Thu Oct  1 14:15:38 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=started:65865
arm_pid=65813	watcher_pid=65865	origin=started	started_at=1790889353	ended_at=1790889354	exit_code=143	signal=TERM	reason=arm-interrupted	beacon_age=0	lock_before=pid:65865|identity:Thu Oct  1 14:15:53 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=none


After handling close, next park owns fresh cycle with no left record:
  PID  PPID ARGS
77811 77807 bash bin/fm-supervision-host.sh park
78052 77811 bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch-arm.sh
78109 78052 bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh

recovery: acked:handling:43507.1790889352.xg5vg3

host	77811	Thu Oct  1 14:17:52 2026     bash bin/fm-supervision-host.sh park
arm	78052	Thu Oct  1 14:17:53 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch-arm.sh
Evidence: Real watcher ownership guard, queued work, and self-exit recovery

Source: Real watcher ownership guard, queued work, and self-exit recovery

ok - watch-arm: --take-over attaches to a cycle the named arm does not own and leaves it running

REAL PRODUCT STATE take-over-not-owner after test_take_over_attaches_to_a_cycle_the_named_arm_does_not_own

--- arm.out ---
watcher: attached pid=3678 (beacon 0s)

--- watch.out ---

--- .watcher-down ---
pending:downtime:3678.1790889401.TZOz0R

--- .watch-cycle-exits.log ---
arm_pid=4263	watcher_pid=3678	origin=attached	started_at=1790889399	ended_at=1790889401	exit_code=143	signal=TERM	reason=arm-interrupted	beacon_age=2	lock_before=pid:3678|identity:Thu Oct  1 14:16:39 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=none
ok - watch-arm: --take-over owns a fresh cycle without a recovery wake and still surfaces queued work

REAL PRODUCT STATE take-over-not-owner after test_take_over_owns_a_fresh_cycle_and_keeps_queued_work_surfacing

--- arm.out ---
watcher: attached pid=3678 (beacon 0s)

--- watch.out ---

--- .watcher-down ---
pending:downtime:3678.1790889401.TZOz0R

--- .watch-cycle-exits.log ---
arm_pid=4263	watcher_pid=3678	origin=attached	started_at=1790889399	ended_at=1790889401	exit_code=143	signal=TERM	reason=arm-interrupted	beacon_age=2	lock_before=pid:3678|identity:Thu Oct  1 14:16:39 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=none

REAL PRODUCT STATE take-over-owner after test_take_over_owns_a_fresh_cycle_and_keeps_queued_work_surfacing

--- arm.out ---
check: rearm-resurface

--- owner.out ---
watcher: started pid=6207 (beacon fresh)
watcher: FAILED - watcher cycle exited 143 without an actionable reason

--- owner2.out ---
watcher: started pid=12320 (beacon fresh) recovery-generation=8418.1790889406.ZN3CBg
watcher: FAILED - watcher cycle exited 143 without an actionable reason

--- .watcher-down ---
announced:downtime:13033.1790889407.iHEEMm

--- .watch-cycle-exits.log ---
arm_pid=6187	watcher_pid=6207	origin=started	started_at=1790889401	ended_at=1790889402	exit_code=143	signal=TERM	reason=signal-exit	beacon_age=1	lock_before=pid:6207|identity:Thu Oct  1 14:16:41 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=none
arm_pid=7876	watcher_pid=6207	origin=attached	started_at=1790889402	ended_at=1790889402	exit_code=unknown	signal=unknown	reason=taken-over	beacon_age=1	lock_before=pid:6207|identity:Thu Oct  1 14:16:41 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=started:8418
arm_pid=7876	watcher_pid=8418	origin=started	started_at=1790889402	ended_at=1790889406	exit_code=143	signal=TERM	reason=arm-interrupted	beacon_age=1	lock_before=pid:8418|identity:Thu Oct  1 14:16:42 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=none
arm_pid=12252	watcher_pid=12320	origin=started	started_at=1790889406	ended_at=1790889407	exit_code=143	signal=TERM	reason=signal-exit	beacon_age=1	lock_before=pid:12320|identity:Thu Oct  1 14:16:46 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=none
arm_pid=13204	watcher_pid=12320	origin=attached	started_at=1790889407	ended_at=1790889407	exit_code=unknown	signal=unknown	reason=taken-over	beacon_age=1	lock_before=pid:12320|identity:Thu Oct  1 14:16:46 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=none
arm_pid=13204	watcher_pid=14108	origin=started	started_at=1790889407	ended_at=1790889408	exit_code=0	signal=none	reason=actionable-check	beacon_age=1	lock_before=pid:none|identity:none	lock_after=pid:none|identity:none	successor=none

--- .wake-queue ---
1790889407	2	signal	take-over	signal: fixture still queued for main
watcher: secondmate liveness check failed
ok - watch-arm: takeover preserves self-exit downtime and surfaces a recovery wake

REAL PRODUCT STATE take-over-not-owner after test_take_over_preserves_downtime_from_watcher_self_exit

--- arm.out ---
watcher: attached pid=3678 (beacon 0s)

--- watch.out ---

--- .watcher-down ---
pending:downtime:3678.1790889401.TZOz0R

--- .watch-cycle-exits.log ---
arm_pid=4263	watcher_pid=3678	origin=attached	started_at=1790889399	ended_at=1790889401	exit_code=143	signal=TERM	reason=arm-interrupted	beacon_age=2	lock_before=pid:3678|identity:Thu Oct  1 14:16:39 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=none

REAL PRODUCT STATE take-over-owner after test_take_over_preserves_downtime_from_watcher_self_exit

--- arm.out ---
check: rearm-resurface

--- owner.out ---
watcher: started pid=6207 (beacon fresh)
watcher: FAILED - watcher cycle exited 143 without an actionable reason

--- owner2.out ---
watcher: started pid=12320 (beacon fresh) recovery-generation=8418.1790889406.ZN3CBg
watcher: FAILED - watcher cycle exited 143 without an actionable reason

--- .watcher-down ---
announced:downtime:13033.1790889407.iHEEMm

--- .watch-cycle-exits.log ---
arm_pid=6187	watcher_pid=6207	origin=started	started_at=1790889401	ended_at=1790889402	exit_code=143	signal=TERM	reason=signal-exit	beacon_age=1	lock_before=pid:6207|identity:Thu Oct  1 14:16:41 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=none
arm_pid=7876	watcher_pid=6207	origin=attached	started_at=1790889402	ended_at=1790889402	exit_code=unknown	signal=unknown	reason=taken-over	beacon_age=1	lock_before=pid:6207|identity:Thu Oct  1 14:16:41 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=started:8418
arm_pid=7876	watcher_pid=8418	origin=started	started_at=1790889402	ended_at=1790889406	exit_code=143	signal=TERM	reason=arm-interrupted	beacon_age=1	lock_before=pid:8418|identity:Thu Oct  1 14:16:42 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=none
arm_pid=12252	watcher_pid=12320	origin=started	started_at=1790889406	ended_at=1790889407	exit_code=143	signal=TERM	reason=signal-exit	beacon_age=1	lock_before=pid:12320|identity:Thu Oct  1 14:16:46 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=none
arm_pid=13204	watcher_pid=12320	origin=attached	started_at=1790889407	ended_at=1790889407	exit_code=unknown	signal=unknown	reason=taken-over	beacon_age=1	lock_before=pid:12320|identity:Thu Oct  1 14:16:46 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=none
arm_pid=13204	watcher_pid=14108	origin=started	started_at=1790889407	ended_at=1790889408	exit_code=0	signal=none	reason=actionable-check	beacon_age=1	lock_before=pid:none|identity:none	lock_after=pid:none|identity:none	successor=none

--- .wake-queue ---
1790889407	2	signal	take-over	signal: fixture still queued for main

REAL PRODUCT STATE take-over-self-exit after test_take_over_preserves_downtime_from_watcher_self_exit

--- arm.out ---
watcher: started pid=18856 (beacon fresh)
check: rearm-resurface

--- owner.out ---
watcher: started pid=14941 (beacon fresh)
watcher: FAILED - watcher cycle exited 1 without an actionable reason

--- .watcher-down ---
announced:downtime:14941.1790889411.6xxRM9

--- .watch-cycle-exits.log ---
arm_pid=14913	watcher_pid=14941	origin=started	started_at=1790889408	ended_at=1790889411	exit_code=1	signal=none	reason=nonzero-exit	beacon_age=0	lock_before=pid:14941|identity:Thu Oct  1 14:16:48 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=none
arm_pid=16882	watcher_pid=14941	origin=attached	started_at=1790889409	ended_at=1790889411	exit_code=unknown	signal=unknown	reason=taken-over-unconfirmed-stop	beacon_age=0	lock_before=pid:14941|identity:Thu Oct  1 14:16:48 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=started:18856
arm_pid=16882	watcher_pid=18856	origin=started	started_at=1790889411	ended_at=1790889412	exit_code=0	signal=none	reason=actionable-check	beacon_age=0	lock_before=pid:18856|identity:Thu Oct  1 14:16:51 2026     bash ~/.no-mistakes/worktrees/52b07e9083e7/01M3WMCNX8QFSJCB7Q3K5VB23W/bin/fm-watch.sh	lock_after=pid:none|identity:none	successor=none

--- .wake-queue ---
Evidence: Missing owner ledger row leaves recovery visible

Source: Missing owner ledger row leaves recovery visible

Missing owner row deliberately induced with real ledger-lock contention. Old owner and watcher exited; fresh cycle reports recovery rather than restoring acknowledgement.
Before: acked:handling:56538.1790889545.Y89EBD
After: announced:downtime:55820.1790889546.bQIG6M
watcher: started pid=55820 (beacon fresh)
watcher: FAILED - watcher cycle exited 143 without an actionable reason
watcher: started pid=69430 (beacon fresh)
check: rearm-resurface
Evidence: Pre-fix negative control reproduces hidden downtime

Source: Pre-fix negative control reproduces hidden downtime

watcher: secondmate liveness check failed
not ok - takeover restored the old acknowledgement
Evidence: Disposable lab teardown

Source: Disposable lab teardown

Private fm-lab tmux server stopped. All recorded arms and watchers in disposable validation homes have exited. No live fleet sessions were touched.

Evidence and trade-offs

Recheck (read-only, 2026-10-01 17:33Z)

  • axi-a1, eddies-wallet-e1, mini-default, and sshhip-h7 each showed a single arm, a child of a park host started 11:59-12:10Z.
    Their host ledgers show the orphans ended at the park boundary (boundary after 27001s, 11:59-12:10Z), which stops the home's watcher; they had lived about 7.5 h since the 04:28-04:40Z restart-persist pass-throughs (pass-through attended main-only signal: .../parent-replies.status).
  • The control home fmdev-f1 showed the same topology live: an orphan arm (ppid 1) owning the watcher plus the park host's attached arm, re-created by every attended main-only pass-through in its host ledger.
    So the leak recurs on every such close and on every second-mate restart, rather than being a one-off of the 549e07f update.

Root cause

  • Since fix: keep remote reply listeners and watcher cycles running #5941 (2026-09-27), an attended main-only pass-through starts a successor watcher cycle, disowns it, and exits so the Claude Stop hook can exit 2 and wake main.
    The successor arm is reparented to init and stays in the exited hook's process group, which nothing stops.
  • The next park host's first arm found that healthy watcher and attached to it (attach_and_wait, polling every 0.5 s) instead of owning it, and nothing ever reclaimed ownership before the watcher's next close or the 27000 s park boundary.
  • A second-mate restart always hits this: the persist request sent before the restart is a main-only close in the mate, and a quiet mate may see no further close for hours.

Fix

  • The host records the successor it leaves for main (state/.supervision-host-left, pid and identity), and the next host's activation consumes that record.
    The pipeline's review rounds hardened that record: it persists until the recorded arm is gone or has a different identity, so a later host retries an interrupted take-over; the write is atomic and read back before the host exits; and when the write fails the host keeps the successor tracked and stops it at cleanup rather than leaving an unrecorded orphan.
  • While that arm still runs under its identity, the host's first cycle runs fm-watch-arm.sh --take-over <arm>.
    Only when the healthy watcher is that arm's own child, it stops the watcher by its locked identity, reports a reason the cycle delivered before the stop landed, and otherwise owns a fresh cycle.
  • The stopped watcher publishes downtime like any close; the taking arm undoes that publication over an acknowledged episode when no wake was appended in between (handover-restore, docs/watcher-continuity.md "Generation reuse"), so the handover wakes nobody and hides no queued work.
  • That undo applies only when the taken-over arm, which waits on its watcher, recorded that exact watcher ending by the take-over's TERM in the cycle ledger (state/.watch-cycle-exits.log, waited for within a short bound).
    A watcher that failed on its own just before the stop records signal=none, so its downtime stays and the fresh cycle surfaces it as an ordinary recovery wake (Greptile review finding; a missing record costs one extra recovery wake, never a hidden one).
  • The watcher, signal traps, and attached-arm cadence are unchanged; a cycle the named arm does not own, or one that outlives the stop, is attached to exactly as before.

Regression test and counterfactual

  • tests/fm-supervision-host.test.sh test_next_park_takes_over_the_cycle_a_pass_through_left_for_main drives the real Stop hook, host, arm, and watcher through a main-only pass-through, main's drain and acknowledgement, and a session restart.
    Before the fix it fails with exactly the fleet topology (one arm with ppid 1 owning the watcher, one attached arm under the new host); after the fix it passes, with one arm owned by the host, no wake to main, the episode still acknowledged, and the next close delivered.
  • With the downtime undo disabled, the same test fails with check: rearm-resurface waking main on every takeover.
  • tests/fm-watch-arm.test.sh covers the ownership guard, a fresh cycle without a recovery wake, and queued work still resurfacing; tests/fm-wake-queue.test.sh covers that the undo applies only to the stop's own publication.
  • test_take_over_preserves_downtime_from_watcher_self_exit makes the owned watcher fail on its own between the take-over's snapshot and its stop.
    Before the confirmed-stop check it fails (takeover restored the old acknowledgement); after it, it passes with check: rearm-resurface, while the ordinary take-over test and the host regression above still pass with no wake.

Herdr lab before and after

A disposable primary home ran a fake Claude second mate in an isolated Herdr lab pane, firing the real Stop hook in its own process group; the restart killed the session and its current hook group outright (the fleet's cycle log has no record for the old park at the real restart) and started a new session in the same pane.

  • Unfixed (HEAD): after the restart the orphan arm (ppid 1, in the first hook's dead process group) still owned the watcher, and the restarted session's park had an attached arm polling it.
    The next close was still delivered through the attached arm, which answers the report's open question about wake delivery.
  • Fixed: the old session's next park took the cycle over, so one arm owned the watcher before the restart.
    After the restart, the restarted session's park ended with one arm, its own child, owning the watcher, and the next close was delivered.

Trade-off: one recovery wake per second-mate restart

  • Before this change, the orphan carried the old watcher across a second-mate restart, which is the leak itself.
  • Now the watcher belongs to the session's own park, so a restart that kills the session's hook group kills that watcher too, and the restarted mate gets one check: rearm-resurface recovery wake before its next park owns a fresh cycle.
  • That is the ordinary recovery path for a watcher that died with its session, and it is how restarts behaved before fix: keep remote reply listeners and watcher cycles running #5941; it costs one short main turn per restart (restarts happen at fleet updates), against an attached poller forking about 25-40 processes per second per home for up to 7.5 h.
  • Continuity through main's handling turn after a main-only close is unchanged: the left successor still covers that turn, and only the next park replaces it.

Follow-up, not in this PR

  • Homes that opt out of the supervision host (config/supervision-host-off) still use the Claude Stop hook's own detached handling successor (start_handling_successor in bin/fm-claude-stop-autoarm.sh), and the next Stop's foreground arm attaches to it the same way.
    That variant is unfixed here and should get the same take-over in a separate change.
  • Hosts whose owner relaunches them with park --restart (the OpenCode plugin and omp) start their first cycle with fm-watch-arm.sh --restart, which stops a successor left by a main-only pass-through without the take-over's acknowledgement-preserving handover, so that stop can surface a recovery wake.
    That branch predates this PR (it has stopped the left successor's watcher since fix: keep remote reply listeners and watcher cycles running #5941; this PR only adds the take-over branch after it), and those hosts never attach to the successor, so they do not have the orphan-plus-attached-poller leak fixed here.
    The successor record those hosts write is dropped at the next activation once the stopped successor has exited.
    Giving that restart the same handover is a separate change.

Landing

This changes supervision machinery that every live home runs, so landing timing is coordinated with the main firstmate; the live orphans are cleaned up through the guarded path after it lands, not by hand as part of this fix.

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

✅ **Review** - passed

✅ No issues found.

✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 7 of 7 scenarios driven live against the product
Scenario Result Live Evidence
Restart the primary after a main-only close; the next park replaces the orphan with a host-owned cycle and delivers the next decision ✅ pass live live-topology.log; live-after-restart.txt
Interrupt a takeover while its marker lock is held; a later park retries using the retained successor identity ✅ pass live live-interruption.log
Make the successor-record destination a writable directory; the close remains unchanged and the unrecorded successor is cleaned up ✅ pass live live-record-failure.log; live-record-failure.txt; live-record-rearmed.txt
Request takeover using a process that does not own the watcher; the arm attaches without stopping that watcher ✅ pass live live-watch-scenarios.log
Take over a cycle with work still queued; the fresh cycle surfaces that work ✅ pass live live-watch-scenarios.log
Let the old watcher fail after the handover snapshot but before TERM; genuine downtime produces a recovery wake ✅ pass live live-watch-scenarios.log; self-exit-negative-control.log
Prevent the owner from recording its TERM row; takeover proceeds but conservatively surfaces recovery ✅ pass live live-missing-row.log
  • Ran selected executable regressions from tests/fm-watch-arm.test.sh, tests/fm-supervision-host.test.sh, and tests/fm-wake-queue.test.sh through temporary focused drivers.
  • Started real claude on the disposable home's private fm-lab tmux socket; drove bin/fm-lock.sh, bin/fm-supervision-host.sh park, and generation-bound wake acknowledgements.
  • Restarted the lab primary and inspected persisted ownership records, process parentage, recovery markers, cycle ledgers, and subsequent decision delivery.
  • Held the recovery-marker lock, interrupted the lab host with TERM, released the lock, and drove another park.
  • Introduced a writable directory at the successor-record destination; verified cleanup and the next fresh park.
  • Drove real watcher-arm interfaces without mock tmux or crew-state implementations for non-owner attachment, queued work, watcher self-exit, and missing owner-ledger evidence.
  • Ran the self-exit regression against 99601828^; it failed with takeover restored the old acknowledgement, while the current implementation passed.
  • Stopped the private tmux server, verified recorded lab processes exited, removed disposable homes and temporary drivers, and confirmed a clean worktree.
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

An attended main-only pass-through leaves a successor watcher cycle
running through main's handling turn. The session's next park attached
to that cycle instead of owning it, so the successor's arm, orphaned by
its host's exit, kept owning the watcher while the new park's arm polled
it twice a second until the next close or the park boundary, hours later
in a quiet second mate. A second-mate restart hit this every time, since
its persist request is a main-only close.

The host now records the successor it leaves for main, and the next
host's first cycle runs bin/fm-watch-arm.sh --take-over on it: when that
arm still owns the healthy watcher, the new arm stops it, reports a
reason the cycle delivered first, and otherwise owns a fresh cycle. The
stop's own downtime publication is undone over an acknowledged episode
when no wake was appended in between, so the handover wakes nobody.
@greptile-apps

greptile-apps Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

[High risk] Adds watcher cycle takeover and handover recovery logic.

The PR appears safe to merge; no outstanding blocking finding or new actionable issue was identified.

Reviews (4) · Last reviewed commit: "no-mistakes(document): Clarify watcher t..."

Comment thread bin/fm-watch-arm.sh Outdated
Comment thread bin/fm-supervision-host.sh
… exact taken-over arm/watcher ledger row with signal=TERM, awaited within a short bound. Otherwise takeover proceeds without erasing downtime. Added a self-exit regression confirmed failing before the fix and passing afterward; ordinary takeover tests and the full watcher-arm suite pass. Updated Generation reuse documentation. Syntax, diff checks, and ShellCheck pass with existing SC1091/SC2034 warnings excluded. ci-2 remains unchanged
Comment thread bin/fm-watch-arm.sh
@kunchenguid kunchenguid changed the title fix: prevent orphan watcher arms after session restarts fix: reclaim orphaned watcher arms on the next park Oct 1, 2026
@kunchenguid
kunchenguid merged commit 8690c41 into main Oct 2, 2026
22 checks passed
@kunchenguid
kunchenguid deleted the fm/fm-restart-orphan-watcher-arm-r1 branch October 2, 2026 00:08
d-ploutarchos added a commit to OK-LG/firstmate that referenced this pull request Oct 2, 2026
* test: preserve Pi calm transcript captures with Pi 1.0 (kunchenguid#6338)

* test(calm): pin Pi's regular TUI mode where pane assertions read scrollback

Pi 1.0.0 defaults its TUI to a fullscreen alternate-screen mode whose
scrollable transcript is application-owned, so rows that leave the viewport
never enter terminal scrollback and tmux capture-pane -S can no longer see
them. The Pi Calm e2e launches now pass --tui-mode regular wherever the flag
exists so the transcript assertions keep reading real scrollback on both the
Pi 1.0.0 line and earlier Pi lines, which have no such flag and render
regular-only anyway.

* no-mistakes(document): Correct Pi TUI documentation and scrollback rationale

* fix(bin): preserve hold reasons and reject invalid completion inventories (kunchenguid#6331)

* fix(bin): encode captain-hold reasons and reject self-inventory in complete

hold now stores a reason with parentheses, line breaks, or percent signs
through a reversible percent encoding that every reader decodes, instead of
refusing it. hold --origin records the origin on the held task, and complete
refuses the origin as its own inventory entry and an entry held for a
different origin; holds with no recorded origin are accepted and flagged.

* fix(review): Decode marked hold reasons consistently across readers

* fix(review): Remove unnecessary lifecycle test dispatch

* fix(review): Correct hold origin identity and inventory recovery

* fix(review): Record origins before placing backend holds

* fix(document): Clarify captain-hold validation and reason reader documentation

* fix(ci): Fixed both findings: failed backend holds restore the previous origin, and invalid base64/UTF-8 reasons remain verbatim. Added regressions and documented valid-literal ambiguity. Both failures were reproduced before fixes. Verification: 54 lifecycle tests and 9 wrapper tests passed; 7 Beads-specific cases skipped because tasks-axi is markdown-only. Focused lint and diff checks passed. No pipeline or publication actions performed

* fix: reclaim orphaned watcher arms on the next park (kunchenguid#6335)

* fix(bin): take over the watcher cycle a main-only pass-through leaves

An attended main-only pass-through leaves a successor watcher cycle
running through main's handling turn. The session's next park attached
to that cycle instead of owning it, so the successor's arm, orphaned by
its host's exit, kept owning the watcher while the new park's arm polled
it twice a second until the next close or the park boundary, hours later
in a quiet second mate. A second-mate restart hit this every time, since
its persist request is a main-only close.

The host now records the successor it leaves for main, and the next
host's first cycle runs bin/fm-watch-arm.sh --take-over on it: when that
arm still owns the healthy watcher, the new arm stops it, reports a
reason the cycle delivered first, and otherwise owns a fresh cycle. The
stop's own downtime publication is undone over an acknowledged episode
when no wake was appended in between, so the handover wakes nobody.

* no-mistakes(review): Keep left-arm record until the orphaned arm is gone

* no-mistakes(review): Relinquish successor arm only after durably recording it

* no-mistakes(review): Relinquish successor only after its record reads back

* no-mistakes(document): Clarify successor takeover guarantees and authoritative documentation

* no-mistakes(ci): Fixed ci-1: acknowledgement restore now requires the exact taken-over arm/watcher ledger row with signal=TERM, awaited within a short bound. Otherwise takeover proceeds without erasing downtime. Added a self-exit regression confirmed failing before the fix and passing afterward; ordinary takeover tests and the full watcher-arm suite pass. Updated Generation reuse documentation. Syntax, diff checks, and ShellCheck pass with existing SC1091/SC2034 warnings excluded. ci-2 remains unchanged

* no-mistakes(document): Clarify watcher take-over recovery and restart limits

* fix(bin): restore downtime on supervision-host hand-back when the successor already closed (kunchenguid#6355)

* fix(bin): restore supervision host hand-back continuity

* no-mistakes(review): Scope host hand-back failure fallback to lost pending:handling

* no-mistakes(review): Remove stray scratch test copy tests/.tmp-rest.test.sh

* no-mistakes(test): Initialise successor globals so early hand-back survives set -u

* no-mistakes(document): Document host hand-back downtime failure and Claude lost-handback notice

* no-mistakes(ci): I fixed the Greptile finding. The rule that must hold: when the supervision host hands back an actionable wake, its rewake is refused, and no watcher is healthy, the hand-back still has to reach main as a delivered notice. That must be true whether the recovery marker is `pending:handling` or `announced:handling`. Only one place applies this check: the lost hand-back fallback in `bin/fm-claude-stop-autoarm.sh`. **Fix:** that check now accepts both `pending:handling:*` and `announced:handling:*` tokens (a one-line change). Nothing else in the fallback changed: - Refusals on any other marker, such as an already acknowledged one, still exit 0 silently and open no failure episode. - The notice is still sent once per episode, and repeats are recorded as `failed-suppressed`. **Tests:** - `tests/fm-claude-stop-autoarm.test.sh`: the lost hand-back test now runs as a shared helper with two variants, one writing a `pending:handling` marker and a new one writing `announced:handling` (`test_host_lost_announced_handback_notifies_once_per_episode`). - `tests/fm-supervision-host.test.sh`: the end-to-end test where downtime restoration fails is now a shared helper too, with a new `announced` variant (`test_claude_stop_hook_notifies_when_closed_announced_successor_downtime_restore_fails`). It moves the handling episode to `announced` before the host hands back. Without the fix the hook would exit 0 here; the test requires exit 2, `outcome=failed` and a delivered failure notice. **Verification (all under nice -n 10):** - The full `tests/fm-claude-stop-autoarm.test.sh` suite passed (rc=0), including both lost hand-back variants and the benign-refusal test. - In `tests/fm-supervision-host.test.sh` I ran only the four hand-back test functions, all passing (rc=0). The suite can't run single functions, so I used a temporary copy with a trimmed test list and deleted it afterwards; `git status` shows only the 3 intended files changed. - shellcheck is clean on all three changed files. I did not run `bin/fm-lint.sh`. - I did not run the new tests against the unfixed code; the claim that they fail without the fix comes from reading the old check

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: Mickaël Rémond <mremond@process-one.net>
Co-authored-by: Tiago <tiagop@hey.com>
roderik added a commit to roderik/firstmate-tweaked that referenced this pull request Oct 3, 2026
* fix(bin): converge every open owner onto a known terminal contribution (#6112)

* fix(bin): converge every open owner onto a known terminal contribution

settle_final only cleared a stale error on retry, so an owner whose saved
row still said open kept projecting a merged or closed pull request as
open after another owner's row had already recorded the terminal
observation. Copy the known terminal observation to every owner whose
saved row is not itself terminal, keeping that owner's own pending and
notified state, and clear its error.

* no-mistakes(review): Carry terminal checked_at when converging existing owner rows

* no-mistakes(ci): I fixed Greptile finding ci-2 as you asked, with a change to tests/fm-contributions.test.sh only. The rule it enforces: when a retry converges an owner onto a URL that is already merged or closed, that owner gets the terminal owner's whole observation, not just its state. The same weak check appeared twice in test_interrupted_multi_owner_poll_settles_every_owner, so I fixed both: - **Open owner (line 784):** the check now also requires `.observation == $terminal[0].records[0].observation`. The existing checks for error, checked_at, pending and notified are unchanged. - **Errored owner (just below):** it only checked state and error before. It now reads the terminal owner's file and makes the same full-observation comparison. Adding the comparison alone would not have caught anything. The test fixtures gave both owners identical observations apart from `state`, so copying only the state would still have passed. In both cases I also set the terminal owner's observation head to HEAD_B, so the two observations now really differ. Verification: - The focused test passes against the current bin/fm-contributions.sh. - I temporarily changed `settle_final` so it copied only the state. The test then failed, reporting the owner still on the old head (HEAD_A). I restored the file afterwards, and `git status` shows only the test file modified. - The full tests/fm-contributions.test.sh suite exits 0. No product code changed. The other CI finding (ci-1, "Behavior portable serial 9") was left alone because you chose to ignore it

* feat: enable supervision host by default for Claude primaries (#6124)

* feat: run the supervision host by default on a Claude primary

An absent config/supervision-host on a Claude primary now reads as on with
the default engine, and a file holding `off` opts any home out. Cursor,
OpenCode, omp, Grok, and Codex stay file-gated, with `off` read as disabled
there too. Every reader asks fm_supervision_host_enabled instead of testing
the file, and non-bash readers query it through the lib's `enabled` entry.
A primary's `off` is not inherited by secondmates: each home keeps its own
supervision posture.

* test: pin the watcher-path posture in fixtures that assume no supervision host

Fixtures that drive the watcher arm or assert a non-host drain now write
an explicit off file, and fixtures that copy the Stop auto-arm or the
supervision instructions carry the engine lib they now source. The two
drain suites also stop reading the code root's config.

* fix: name the opt-out when an off home passes an attended wake to main

A host parked when the home writes off now logs that the home does not run
the supervision host, rather than claiming it has no engine.

* no-mistakes(document): Clarify Claude supervision defaults and historical evidence

* no-mistakes(ci): Fixed process leaks in the two added host tests. Each case now stops its recorded watcher and host/arm processes; fake hook sessions exit through session.stop. The full host suite passed before the final cleanup refinement, and both affected cases, bash syntax, ShellCheck, and diff checks passed afterward. CI runtime still needs confirmation

* fix(bin): create the state dir on a fresh primary before the session-start scope check (#6125)

* fix(bin): create the state dir on a fresh primary before the session-start scope check

fm_primary_scope_matches required an already-existing state directory, so
bin/fm-sessionstart-run.sh stood down on a fresh clone before anything could
create it. Split out fm_primary_root_matches so the run wrapper can confirm
primary-home identity first, create the gitignored state dir when it is
missing, and only then run the unchanged scope check.

* no-mistakes(document): Document session-start state dir creation on fresh clones

* no-mistakes(ci): I fixed the Greptile P1 the way you asked. When a fresh primary can't create `state/`, the run wrapper no longer stands down silently. **Invariant:** when an otherwise eligible fresh primary cannot create `state/`, startup must never fail silently. This path has only one site: the mkdir in `bin/fm-sessionstart-run.sh`. Other hooks and the nudge wrapper never create `state/`, so they have no equivalent failure. **What changed:** - **Run wrapper** (`bin/fm-sessionstart-run.sh`): it captures mkdir's error and prints one line to stderr before standing down as before (exit 0, or 3 for the Pi prerequisite). The line looks like `fm-sessionstart-run: startup could not create the state directory <path>: <reason>`. - **Test** (`tests/fm-sessionstart-nudge.test.sh`): the new case `test_run_reports_a_state_dir_it_cannot_create` uses a fresh primary with no `state/` and a read-only (0500) root. It checks four things: exit 0, no digest on stdout, no state dir created, and exactly one stderr line ending in "Permission denied". It fails without the fix and passes with it. - **Docs** (`docs/sessionstart-nudge.md`): I added one sentence describing the stderr line and one describing what the new test proves. **Verification:** I ran `tests/fm-sessionstart-nudge.test.sh`, and every test passes. `bin/fm-lint.sh` on the changed scripts (pinned ShellCheck 0.11.0) and `tests/fm-documentation-audiences.test.sh` also pass. As you asked, the wrapper still stands down with the ineligible-checkout status afterwards. It does not report this as a failed eligible startup, which is what the bot suggested

* fix(bin): measure pending-reply grace from turn completion, not delivery (#6126)

* fix(bin): measure pending-reply grace from turn completion, not delivery

Fixes #6057

The pending-reply guard demanded a repost ("REPOST REQUIRED: previous
marked request had no correlated parent report") while the second
mate's correlated reply was already on its way.
fm_pending_reply_send_recovery measured its grace window from delivery
instead of from the request turn's completion, so any turn longer than
the grace fired the demand the moment the turn ended, before the reply
could have landed. The missed-report escalation had the same gap: it
fired the instant the recovery turn's completion was observed, with no
grace at all.

Both now measure grace from the relevant turn's completion (request
turn for the recovery repost, recovery turn for the escalation), and
both take one fresh, uncached read of the parent status file
immediately before firing, accepting a correlated line regardless of
its verb. Transport-failure escalations stay immediate, and the
one-repost limit is unchanged.

* no-mistakes(review): Document grace window as measured from turn completion

* no-mistakes(ci): Both Greptile findings were real and caused by this PR, so I fixed them. The full `tests/fm-pending-reply.test.sh` suite passes. **ci-1 (a reply could be overwritten by a repost).** The rule that must hold: a recovery send is recorded only if the record is still unresolved, checked under the same per-correlation lock that resolution uses. The escalation path already did this (`_fm_pending_reply_maybe_escalate_locked` reads fresh and publishes under one lock). The recovery path did not: `fm_pending_reply_send_recovery` did its fresh read through `fm_pending_reply_try_resolve`, which let go of the lock before the send was recorded. A reply landing in that gap could be overwritten, and the repost would go out anyway. Now `send_recovery` takes the lock once and, while holding it, re-checks that the phase is still `awaiting_report`, runs the fresh uncached read, and records the send (sender pid and identity, attempt time, phase `recovery_sending`). It releases the lock before actually sending, so the lock is not held during the send. It uses the same lock helpers the other lock wrappers use. Grace timing, the one-repost limit and the escalation path are unchanged. **ci-2 (the test would pass even without the fix).** In `test_recovery_fresh_status_read_resolves_before_firing`, the reply is still appended to the status file, but the stored file signature is then set to the file's new signature. That stands in for a same-size rewrite that the signature cache cannot see. The test first checks that a normal cached read misses the reply, then that the fresh read before sending catches it. I also added the same check for the fresh read before escalation, which the review said was uncovered. The test now sets its own send hook, so it no longer depends on one left over from an earlier test (that leftover had made failures exit silently). **Checks:** - I removed the fresh-read bypass at each site in turn and reran the suite. With it gone from recovery, the test fails with "recovery must not fire once a correlated reply has landed". With it gone from escalation, it fails with "the fresh pre-escalation read should have resolved the record, got escalated". With both in place, all tests pass. - Shellcheck with `-x` timed out locally. Without `-x` and ignoring SC1091, the only warnings are SC2034 on the existing `maybe_escalate` lock wrapper, which is not part of this change. The new code adds no warnings. Changes are in `bin/fm-pending-reply-lib.sh` and `tests/fm-pending-reply.test.sh`. Nothing is committed yet; a plain commit message such as "fix(bin): record the pending-reply recovery send under the fresh-read lock" fits the instruction

* no-mistakes(ci): ci-1 was real and caused by this PR. The same bug was also in the escalation path, so both are fixed. The full tests/fm-pending-reply.test.sh suite passes. The rule that must hold: a recovery repost or an escalation goes out only if the record's phase, read after the fresh-read resolve, is still what it was before. The resolver writes phase=resolved first and only then writes the other resolution fields. If one of those later writes fails, it returns an error even though the record is already resolved. Places this rule applies, both fixed: - Recovery (fm_pending_reply_send_recovery): the fresh-read resolve now runs first, and the phase is re-read right after it, whatever it returned. The send is recorded and made only if the phase is still exactly awaiting_report. This replaces the earlier phase check rather than adding a second one. - Escalation (_fm_pending_reply_maybe_escalate_locked): same bug. After a failed resolve it went on to publish the blocked line and set phase=escalated. One added line after the resolve call returns 1 without publishing if the phase has changed. Test: added test_partial_resolve_write_blocks_firing. It forces a failure on the resolved_epoch write after a correlated reply has landed. It checks that the recovery send hook is never called, that no escalation line is published, and that the phase stays resolved. The forced failure runs in a subshell so it can't affect later tests. Checks: - With the recovery fix reverted, the new test fails with "recovery must not fire after a partial resolve". - With the escalation fix reverted, it fails with "partial resolve should block escalation, got escalated". - With both fixes in, every test passes. - Shellcheck was run with SC1091 excluded and without -x, not through the repo's lint script. The only new message is one SC2329 info on the test's override function; other test overrides in the same file already get that same info, unsuppressed. Changed files: bin/fm-pending-reply-lib.sh and tests/fm-pending-reply.test.sh. Nothing is committed. Suggested plain commit message: "fix(bin): recheck pending-reply phase after the fresh read before sending

* fix(bin): stop provider-table lookup from writing broken-pipe errors to stderr (#6001)

* fix: provider-table lookup never writes a broken-pipe error to stderr

Fixes #5956

fm_quota_single_provider_for_harness returned from its while read loop
as soon as it found a match, closing the pipe while
fm_quota_single_provider_table's printf could still be writing.
Where SIGPIPE is ignored, as on GitHub Actions runners, bash then
prints "printf: write error: Broken pipe" on the resolver's stderr,
which intermittently broke the one-diagnostic-line assertions in
tests/fm-dispatch-resolve.test.sh.

Read the whole table before answering, the way
fm_control_harness_supported already does, so the writer always
finishes. Return values and output are unchanged.

Reproduced by running tests/fm-dispatch-resolve.test.sh with SIGPIPE
ignored on a single pinned core under CPU contention: 30 of 30 runs
failed before the fix, 0 of 30 after. Note: reproducing requires
setting the trap inside the tested shell because nice(1) resets an
inherited SIGPIPE ignore to SIG_DFL. tests/fm-quota-choose.test.sh
passes and bin/fm-lint.sh is clean.

* no-mistakes(ci): Fixed both Greptile findings the user chose to address. ci-1 (bin/fm-quota-axi-lib.sh:154). Invariant: looking up a harness must always end with status 0 and print the provider, even when the caller runs under `set -e`. The loop body `[ -z "$found" ] && [ "$harness" = "$1" ] && found=$provider` now ends in `|| :`. Every iteration succeeds and the whole table is still read. Only `fm_quota_single_provider_for_harness` loops over the table this way, so this is the one place the fix was needed. One caveat: on bash 5.3 the old code did not actually exit under `set -e`, because the `while` loop is not the function's last command, so the new `set -e` test would have passed before this fix too. The change makes the loop's success explicit, as the user asked. ci-2 (regression coverage). I added three cases to the existing `tests/fm-quota-choose.test.sh`, all calling the public lookup function after sourcing the library: 1. With SIGPIPE ignored (`trap "" PIPE`), it looks up every harness 200 times and checks that nothing reaches stderr. 2. A deterministic version of the race: the table function is wrapped so it writes the first row, pauses 0.2 s, then writes the rest. With SIGPIPE ignored, it checks that looking up `claude` prints `claude` and writes nothing to stderr. The stress loop alone reproduced the bug in only about 1 of 5 local runs, which is why this case exists. 3. A direct call under `set -e` prints `claude`. Verification: - `bash tests/fm-quota-choose.test.sh`: all pass. - Same test against the pre-PR library (fa48367, via `FM_ROOT_OVERRIDE`): fails with `printf: write error: Broken pipe`. The deterministic case failed in one run and the stress loop caught it in another. - `shellcheck` on both files: clean. - `tests/fm-dispatch-resolve.test.sh`: passes

* fix: restore portable CI behavior across Pi rendering and remote provisioning (#6162)

* fix: survive Pi 0.99 rendering and Git 2.55 local-clone races

Pi 0.99 puts arguments on the stock tool header and leaves hidden custom messages in the export conversation column. Match that header, and keep Calm's boundary on the visible column. Clone a remote home with --no-local so a prune during Git's loose-object copy cannot fail the seed.

* no-mistakes(review): Stop SIGPIPE write errors; cover older Pi export and project clones

* no-mistakes(document): Clarify Calm export visibility and tool rendering

* no-mistakes(ci): Fixed the dispatch diagnostic to list every provider-less use/default profile in one line and added a multi-profile behavior test. Shortened supervision fixtures using the existing engine-grace and park-clock knobs; removed stray scratch files. Dispatch tests, syntax checks, and three targeted supervision cases passed. CI’s prior supervision duration was 751s; the single permitted local full-suite run timed out at 1200s, so an after-duration is not established. The cancelled serial check had no failure verdict. The outer executor should record the measured before/after duration in the PR body when available

* no-mistakes(review): Gate Pi 0.99 call headers by version; drop hidden-row assertion

* no-mistakes(review): Test stock call headers under Pi 0.87 and 0.99 stubs

* no-mistakes(test): Fix older-Pi queued-row test and verify park-boundary behavior

* no-mistakes(document): Clarify Pi Calm export and queued-turn documentation

* no-mistakes(ci): Fixed the stock macOS Bash 3.2 parse failure in tests/fm-calm-pi-extension.test.sh; its parse check passes. The watcher CI failure is in unchanged code: the isolated five-minute/66-minute case passes locally, but the CI log omits the drain error needed to establish its cause. No speculative watcher fix was made. The full local watcher suite timed out after 500 seconds

* fix: confirm Lavish board replies before worker handoff (#6169)

* Prevent premature Lavish board handoffs

* Prove Lavish arm lacks reply acknowledgement

* Confirm Lavish replies before arming worker boards

* no-mistakes(review): Post Lavish reply only after locked arm eligibility checks

* no-mistakes(review): Fail Lavish reply closed on unknown version

* no-mistakes(document): Correct Lavish reply documentation and remove stale guidance

* no-mistakes(document): Clarify Lavish reply routing and remove duplicate version guidance

* fix: inherit supervision host opt-out across secondmates (#6154)

* feat: inherit the supervision-host opt-out from the primary

Move the supervision host's off opt-out out of config/supervision-host into
its own presence flag, config/supervision-host-off, and add that flag to the
primary-authoritative inherited config set. A primary that opts out now opts
every secondmate home out at spawn and convergence, and clearing it converges
them back. config/supervision-host stays the home-local engine choice.

Shape: config/supervision-host mixed two things, a fleet posture (off) and a
per-home engine and model. Only the posture should follow the primary, so it
becomes a separate presence flag that rides the existing inherited-config
mechanism (FM_INHERITABLE_CONFIG in bin/fm-config-inherit-lib.sh) with no new
machinery, while the engine line stays local. The parse stays in its one
owner, fm_supervision_host_enabled. There is no migration or compatibility
handling for a home that still holds off in config/supervision-host.

Primary off, mate on: inherited material is primary-authoritative by design,
so a mate cannot keep the host while the primary is opted out, and a mate's
own opt-out is removed at the next convergence while the primary has none.
Running the host on a mate is the primary's choice for the fleet; no override
mechanism is added.

Live validation (disposable bin/fm-live-lab.sh lab, Claude primary with a
real seeded secondmate, --supervision-host off):
- up: every readiness check ok, including "host: none running, as expected"
  and a live mate session; the spawned mate home held the inherited
  config/supervision-host-off and the gate read primary OFF, mate OFF.
- primary removed its opt-out, then bin/fm-config-push.sh reported
  "supervision-host-off: pushed - mirrored primary absence" and a config
  reread sent; the gate read primary ON, mate ON, and the live mate handled
  the reread.
- primary opted out again and pushed: "supervision-host-off: pushed", mate
  gate OFF.
- down stopped every lab process and left no lab process running.

Out of scope, follow-up: default-on for the other harnesses, away-daemon
retirement, rollout.

* no-mistakes(document): Document inherited supervision-host opt-out ownership

* no-mistakes(ci): Fixed ci-4: with `--supervision-host off --mate`, lab readiness now requires the inherited flag in the mate home and a disabled mate supervision-host gate. The focused behavior test, shellcheck, and diff checks pass. Left ci-1–ci-3 untouched as directed

* no-mistakes(test): Fix mate readiness HOST_OFF initialization in lab up

* no-mistakes(ci): Fixed Lint 2 by making the new test’s fixtures source resolvable to ShellCheck; its off/on readiness test and ShellCheck now pass locally. Behavior portable serial 5 failed in the unchanged remote-reply test at generation 7. That test passes locally, and no PR-caused defect was identified, so no remote-reply code was changed

* fix: reduce supervision exit latency and stabilize host tests (#6179)

* fix(tests): cut the fixed sleeps in supervision-host cycles

The serial CI lane keeps brushing its 30-minute cap because
fm-supervision-host.test.sh spends ~903s of the job, and per the
run-36635306527 case profile the top nine cases are all multi-cycle
ones (3-10 park/close/turn cycles each): every close waits out the
host's sleep $POLL in await_close plus a watcher sleep $FM_POLL scan
cycle, and every engine turn waits out the fixed sleep 1 descendant
snapshot. That is ~3s of pure sleep per cycle before any real work.

The host poll now accepts positive decimal seconds through a new
seconds_or validator (FM_SUPERVISION_HOST_POLL), and the engine turn's
snapshot loop takes FM_SUPERVISION_ENGINE_SNAPSHOT_SECONDS, also a
positive decimal defaulting to one second - the smallest seam at each
wait's single owner. The suite drives them at 0.2 alongside the
existing FM_POLL=0.5 and FM_ARM_ATTACH_POLL=0.2 knobs, so the real
poll loops still run. The park-boundary case moves onto the injected
test clock instead of a real 3s wait, per-case cleanup polls the host
pid rather than sleeping a full second, and the proof-by-absence
windows (flood re-escalation, successor re-announce, watcher
persistence, recovery staying off main) shrink from 2-3s to 1s, which
still spans two watcher polls at the test cadence.

Every assertion, process lifecycle, and reaping path is unchanged;
production defaults stay at one second. Isolated case timings on a
contended host, base vs branch: attended-latch 54.3->34.6s,
undelivered-dialog 67.7->59.1s, away-latch 46.5->30.5s, held-cadence
47.9->21.6s, unreadable-mirror 39.2->38.5s, park-limit 18.2->12.3s,
registration-fallback 14.1->10.0s, first-cycle-status 12.6->8.4s,
latch-scope 16.7->16.3s. Full suite: 65/65 pass. fm-lint and
shellcheck clean.

* no-mistakes(review): Wait for scan lock release before duplicate check

* no-mistakes(document): Correct supervision snapshot cadence documentation

* fix(tests): keep production poll cadence, probe exits at 0.1s

The fractional poll cadences multiplied the cost of each loop body:
full process-table scans in the engine turn and process refreshes in
await_close ran five times more often, which swamped the thin CI runner
and nearly doubled every multi-cycle case (serial 5 was cancelled at its
30-minute limit on run 36635306527's successor). Restore the production
cadence and notice arm/engine exits with a cheap kill -0 probe at a
tenth of a second between the one-second bodies instead: strictly less
dead time than baseline with no added CPU.

Also hold each injected-clock park bound well past its case's
wall-clock checks so a host that ignored the test clock fails instead
of silently passing at a real-time boundary, and restore the shortened
proof windows (watcher liveness, recovery-off-main absence, first-cycle
stream) to their baseline depth.

* no-mistakes(document): Clarify supervision engine snapshot documentation

* ci: rebalance portable test groups and enforce a packing budget (#6192)

* fix: rebalance portable CI from current duration measurements

* no-mistakes(test): Test serial packing boundary and verify endpoint timeout cleanup

* no-mistakes(document): Clarify timeout guidance and remove duplicated packing estimates

* fix(bin): run no repository hook when core.hooksPath is empty (#6216)

* fix(bin): run no repository hook when core.hooksPath is empty

The per-task hook wrapper refused every commit in a repository whose own
config sets core.hooksPath to the empty string, because git rev-parse
--git-path hooks fails on it. Plain git reads that setting as no hooks, so
the wrapper now runs none; every other lookup failure still refuses and
shows git's error.

Fixes #6171

* no-mistakes(review): Refuse commits when core.hooksPath is a valueless key

* no-mistakes(document): Document empty core.hooksPath handling in commit attribution docs

* no-mistakes(ci): When the wrapper refuses a commit, Git's hook-lookup error now shows up once instead of twice. That required changing one line in the wrapper, and the tests were extended so both bad-config cases would catch the duplicate. Invariant: when the wrapper refuses, Git's lookup error must appear exactly once. In the failure path, the only Git call besides the deliberate second lookup is the `git config --get --type=path core.hooksPath` check in `runtime_chain_body` (`bin/fm-git-strip-ai-trailers.sh:168`). That check prints the same error, so it was the one place to fix. I added `2>/dev/null` to it. Its exit status still decides the outcome: an empty value still runs no hook, and anything else goes on to the second lookup, which prints Git's error once, and the commit is refused. Tests (`tests/fm-git-strip-ai-trailers.test.sh`): - The unresolvable-path test (`~fm-no-such-user-6171/hooks`) now requires `failed to expand user dir` to appear exactly once in the refused commit's output. - The valueless-key test now requires `missing value for 'core.hookspath'` to appear exactly once. - Pre-existing bug in the unresolvable-path test: its `git add` ran after the bad config was set, so it failed silently (exit 128) and the "refused commit" had nothing staged. The test now stages the file before writing the config, the same way the valueless test does, so a real commit gets refused. - The empty-string test is unchanged and still passes, so an empty `core.hooksPath` still runs no hook. Verification: - With the wrapper change reverted, both new checks fail with `expected '1', got '2'`. With the change in place, the whole suite passes. - `bash -n` passes. shellcheck shows only an info-level SC1091 note about sourcing `lib.sh`, which was already there before this change. - `git status` lists only the two intended files

* fix(bin): let a stale record on a reassigned slot retire records-only (#6213)

* fix(bin): let a stale record on a reassigned slot retire records-only

When a pool slot's owner claim names another task, the stale record's
teardown touches nothing under the slot, so the exclusive-slot record scan
no longer refuses it. Full teardowns of a slot this task still claims, or
one with no claim, keep the refusal.

Fixes #6184

* no-mistakes(document): Note claim-over-record precedence for reassigned teardown slots

* fix(bin): keep the steering doorbell short under deep homes (#6240)

* fix(bin): keep the steering doorbell short under deep homes

The doorbell printed the task inbox's absolute path twice, so under a deep
home it grew to about 290 characters and a Herdr submit reported it never
reached the pane on every re-ring. It now names the inbox once by its short
<task>.inbox name and points at the full path the worker's brief already
gives, so its length no longer depends on the home's depth.

Fixes #6120

* no-mistakes(review): Export FM_TASK_INBOX at launch and name it in doorbell

* no-mistakes(ci): ci-1 (Behavior portable serial 9) was caused by this PR, and I fixed it in the test. tests/fm-claude-trust.test.sh failed with "the launch command did not carry a brief doorbell". Its claude_launch_doorbell helper stripped exactly two leading `export ...;` statements before reading the final prompt argument. This PR adds a third one (`export FM_TASK_INBOX=...`) to every launch, so the helper was reading the wrong command. The invariant: a test that parses the launch command must skip every leading export statement, however many there are. I checked every test that parses the launch this way. The only other ones are the two helpers in tests/fm-spawn-dispatch-profile.test.sh, and they already loop over all exports. The kimi and dispatch-profile exact-string checks were updated earlier in this PR. The fix makes claude_launch_doorbell use the same loop (`while [[ "$command" == export\ *\;* ]]; do command=${command#*; }; done`) and then take the last argument. The ordinary path still works: the claude spawn test and the secondmate-clone spawn test both resolve the brief record through the same helper. Verified locally: `bash tests/fm-claude-trust.test.sh` exits 0 with no failing cases. ci-2 (Behavior tests (Herdr)) was not caused by this change, and I made no code change for it. In tests/fm-backend-herdr-presentation-e2e.test.sh, the concurrent secondmate recovery failed with "herdr presentation recovery could not acquire its session lock; refusing a concurrent resume". Two reasons it is not this PR: - The same failure, in the same test and case, happened on run 36655209015 for the unrelated branch fm/fm-contributions-old-gh-compat about 14 hours earlier. - This PR's change cannot lengthen how long the lock is held. The launch is written to a file and sent to the pane as `. launch.N.sh`, so the extra export changes neither the pane submit nor the lock hold time. The cause is a race that was already there: spawn_herdr_presentation_order_lock_acquire gives up after 5 seconds, and a concurrent real-Herdr recovery can hold the lock longer. Fixing that means changing the product's lock timeout, which is outside this PR. It should be tracked separately, and a rerun of the Herdr job is expected to pass. The only file changed is tests/fm-claude-trust.test.sh

* feat(bin): add opt-in config/wait-no-turns so a waiting worker spends no turns (#4859)

* fix(dod): drive no-mistakes with one foreground call, not a background poll

The brief told workers to background the drive call and poll `axi status`
because one call "routinely outlives what your harness lets a single
command run". That advice contradicts the tool it drives: `no-mistakes
axi run --help` documents `--wait` with an 8m default, existing precisely
"so an agent harness with a 10-minute tool cap gets a structured return
instead of an unbounded hang".

Following the old text, a worker could never idle - a backgrounded call
returns in milliseconds, so it does not wait at all - and each attempt
leaked a live timer that later fired as a paid wake. Tell workers to make
one foreground call, let it block, and repeat it when it returns on
elapsed wait rather than on a gate or outcome.

Also drops the generalisation that told workers on any unestablished
harness to assume a command cap and use the same shape, which exported
the defect to harnesses with no such cap.

* fix(bin): let a waiting worker spend no turns until it is answered

A worker waiting on a decision, a pipeline gate, CI, or a heavy-test slot
kept taking model turns: the brief told it to list its inbox at any natural
checkpoint, and six automatic senders nudged secondmates whatever their open
decisions.

- The ship and scout briefs gain one Waiting section: end the turn after
  needs-decision or blocked, and hold an external wait inside ONE blocking
  command bounded by the harness's own command ceiling. The checkpoint clause
  is deleted. Forbidding the wrong shapes is not enough on its own, so the
  section also names the blocking foreground `until` loop as the wait a Claude
  Code worker may use, because that harness can refuse a sleep-then-check
  command while pointing at backgrounding, which is the one shape a waiting
  worker must not take.
- fm-send --automatic defers (exit 4, nothing written or rung) while the
  target has an open decision or blocker of its own; every automatic sender
  passes it and keeps its retry state, and the pending-reply recovery waits
  the same way.
- The two senders that report the result classified it by matching the text of
  the send's captured output against `deferred:*`. fm-send runs bin/fm-guard.sh
  as a supervision warning, and that guard prints its worktree-tangle banner
  whenever the primary checkout is on a feature branch, which is exactly what a
  CI pull-request checkout is. The banner lands ahead of the `deferred:` line,
  so the match fell through and a waiting mate was reported as a failed send,
  with the banner as the reason. Both senders now classify on fm-send's exit
  status, which is the contract the deferral is actually stated in, and select
  the `deferred:` line out of the output rather than assuming it came first.

The third root cause, a no-mistakes definition of done that backgrounded the
drive call and polled axi status, is fixed by this branch's parent commit
"drive no-mistakes with one foreground call, not a background poll"; this
commit takes that text as is and adds the regression test.

Upstream's spawn abort path no longer calls the lease-return helper at all,
so the fork's missing-helper guard and its pin-feature test line are moot
here and are not ported.

The command ceilings each harness enforces, and the probes behind the named
Claude Code wait, are recorded in docs/verification/runtime-backends.md.

* no-mistakes(review): Exempt captain holds, quiet deferred reconcile, clarify worker pauses

* no-mistakes(document): Document deferred automatic nudges, rereads, and reply recovery

* no-mistakes(document): Ring unlanded fire-and-forget steers exactly once more

* no-mistakes(ci): The failing check, "PR must be raised via no-mistakes", reads the pipeline's attestation record, which says document=skipped. No file in the repository can change that record, so I did not touch the check or the PR body. As you said, the no-mistakes rerun after this run finishes will re-execute the document step and record document=completed. The one change is the documentation sentence you ordered. It adds a line to docs/remote-secondmates.md, right after the line saying the remote host runs no re-ring ladder of its own: "A fire-and-forget record, such as a reconcile ask, gets its single retry ring only on the local plane: the remote steer leg owes no re-ring, so a swallowed remote doorbell for one waits for the next ring into that inbox, and a remote-side retry is known follow-up scope." No behavior changed. Checks: tests/fm-documentation-audiences.test.sh passes (4/4) and bin/fm-lint.sh is clean. The change is left uncommitted in the working tree for the pipeline to pick up

* no-mistakes(review): Hold automatic wakes until a mate's own decision closes

* no-mistakes(document): Document watcher delivery of deferred remote re-read nudges

* no-mistakes(review): Merge duplicate elapsed-wait reattach instructions in DOD

* no-mistakes(test): Resolve merged default decision in remote-reply recovery fixture

* no-mistakes(test): Source classify lib so config-push retry-deferred honors open decisions

* no-mistakes(ci): Fixed a flaky test that also fails on main. Neither this PR's bin/fm-brief.sh nor its bin/fm-dod-lib.sh change is involved: bin/fm-dispatch-resolve.sh sources neither file. Another branch (fm-attended-cutover-smoothing-s1, run 36343879084) failed the same shard 8 check the same way, on a different case ("a rule-criterion match prints one diagnostic line, got 2"). Root cause: `fm_quota_single_provider_for_harness` in bin/fm-quota-axi-lib.sh returned from its `while read` loop as soon as it found a match. That closed the pipe while `fm_quota_single_provider_table`'s `printf` was sometimes still writing. GitHub Actions runners ignore SIGPIPE, so bash printed `fm-quota-axi-lib.sh: line 138: printf: write error: Broken pipe` to the resolver's stderr. That is the extra line. I reproduced it locally by running the test with SIGPIPE ignored: 2 of 20 runs failed, one with the resolver's diagnostic line plus two broken-pipe lines. Invariant: looking up a harness in the provider table must never make the table writer fail. The only reader of that table is this function, and all of the resolver's lookups (line 208 without stderr redirected, line 222 with it) go through it. So the fix is in that one place: read the whole table, then print the match. The same file now shows it reads the full table first, like `fm_control_harness_supported` does. Return values and output are unchanged. Verification: with SIGPIPE ignored, tests/fm-dispatch-resolve.test.sh failed 0 of 30 runs after the fix (2 of 20 before). tests/fm-dispatch-resolve.test.sh, tests/fm-brief.test.sh, tests/fm-send-inbox.test.sh, tests/fm-quota-choose.test.sh and tests/fm-quota-array-dispatch-live-e2e.test.sh all pass, and shellcheck is clean. tests/fm-procevent-quota.test.sh fails locally with or without the change ("process-event state root is not a private directory"), so that failure comes from the local environment, not from this fix. No new test was added: the existing one-diagnostic-line assertions already catch this whenever SIGPIPE is ignored, as it is in CI

* Revert "no-mistakes(ci): Fixed a flaky test that also fails on main. Neither this PR's bin/fm-brief.sh nor its bin/fm-dod-lib.sh change is involved: bin/fm-dispatch-resolve.sh sources neither file. Another branch (fm-attended-cutover-smoothing-s1, run 36343879084) failed the same shard 8 check the same way, on a different case ("a rule-criterion match prints one diagnostic line, got 2"). Root cause: `fm_quota_single_provider_for_harness` in bin/fm-quota-axi-lib.sh returned from its `while read` loop as soon as it found a match. That closed the pipe while `fm_quota_single_provider_table`'s `printf` was sometimes still writing. GitHub Actions runners ignore SIGPIPE, so bash printed `fm-quota-axi-lib.sh: line 138: printf: write error: Broken pipe` to the resolver's stderr. That is the extra line. I reproduced it locally by running the test with SIGPIPE ignored: 2 of 20 runs failed, one with the resolver's diagnostic line plus two broken-pipe lines. Invariant: looking up a harness in the provider table must never make the table writer fail. The only reader of that table is this function, and all of the resolver's lookups (line 208 without stderr redirected, line 222 with it) go through it. So the fix is in that one place: read the whole table, then print the match. The same file now shows it reads the full table first, like `fm_control_harness_supported` does. Return values and output are unchanged. Verification: with SIGPIPE ignored, tests/fm-dispatch-resolve.test.sh failed 0 of 30 runs after the fix (2 of 20 before). tests/fm-dispatch-resolve.test.sh, tests/fm-brief.test.sh, tests/fm-send-inbox.test.sh, tests/fm-quota-choose.test.sh and tests/fm-quota-array-dispatch-live-e2e.test.sh all pass, and shellcheck is clean. tests/fm-procevent-quota.test.sh fails locally with or without the change ("process-event state root is not a private directory"), so that failure comes from the local environment, not from this fix. No new test was added: the existing one-diagnostic-line assertions already catch this whenever SIGPIPE is ignored, as it is in CI"

This reverts commit c7199284297ea278c8da7d0a698cc5823ac37cec.

* no-mistakes(review): Retry deferred local instruction nudges via the watcher

* no-mistakes(review): Document watcher retry for deferred local instruction nudges

* no-mistakes(ci): I fixed both review findings you selected (ci-1 and ci-3). I did not touch the deferral check in bin/fm-send.sh. ci-1 (bin/fm-config-push.sh, retry_deferred_rereads) - Rule that must hold: a deferred reread stays flagged until it is actually delivered. - Before the fix, the flag was removed before any of the steps that can skip a mate: the remote lock-path lookup, validate_secondmate_home, the local lock-path lookup, and the lock acquire. A skip at any of those dropped the flag, so the watcher lost track of the reread. - Now the flag is removed in one place only, when the send succeeds (rc 0). A skipped home, a busy lock, a deferred send (rc 4) or a failed send all leave it in place. The re-mark calls on a busy lock and on rc 4 were no longer needed, so I removed them. I updated the comment above the function to match. - Side effect: a send that keeps failing now stays flagged, so the watcher retries it on every poll and logs each failure. That follows your "don't clear until delivered" rule, but it replaces the old behaviour of leaving a failed send to the next config push or session start. - New test in tests/fm-secondmate-sync.test.sh: T8j "a deferred flag survives a skipped invalid home and is retried once it validates". It takes the home's marker away to make validation fail, checks that nothing is sent and the flag stays, then puts the marker back and checks that the nudge is delivered and both the flag and the retry marker are cleared. It fails on the old code and passes now. ci-3 (bin/fm-secondmate-restart.sh) - Rule that must hold: no automatic send wakes a mate that is waiting on its own open decision. - The two automatic sends in this script are the fallback reread nudge (fall_back_to_nudge) and the persist request. Both now pass --automatic. If a persist request is deferred, its correlation is discarded and the mate goes to the fallback nudge, which is also deferred, so the mate is reported as unreached. - New test in tests/fm-secondmate-restart.test.sh: T3b. It gives a mate an open needs-decision and runs a restart. It checks that both sends report as deferred, the mate's doorbell is never rung, its inbox gets no message, nothing is stopped, and the mate is reported as unreached with exit status 3. It fails on the old code and passes now. - The test marks the watcher as alive first. Without that, the watcher-down warning is printed first and becomes the reported reason instead of the deferral message. Verification - tests/fm-secondmate-sync.test.sh passes. - tests/fm-secondmate-restart.test.sh passes. - tests/fm-secondmate-harness.test.sh (the other test that exercises --retry-deferred) passes. - The fm-send-inbox test that covers automatic deferral passes. I only looked at the last lines of that run, not the whole file. - `shellcheck -x` on the four changed files is clean

* Pin autoarm supervision model in secondmate restart T3b

The fresh watcher beat the test writes proves a live watcher only under the
autoarm model; on CI hosts with no detected harness the persistent model
demands a lock-holding watcher, so the watcher-down banner became the
reported reason and the deferral assertion failed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Keep deferred secondmate nudges retryable under the inheritance lock.

A bootstrap instruction nudge could write its deferral flag outside the lock the watcher retry holds, so a concurrent retry could delete a flag that had just been set. A restart fallback that is deferred now records the same marker and flag, so the watcher delivers it once the decision closes.

* no-mistakes(document): Document watcher retry of deferred restart re-read nudges

* Send secondmate reread and restart nudges immediately again.

Deferring those nudges let a later config push drop an incomplete transfer once the decision closed. They now send as they do on main.

* Make the no-turn wait opt-in behind config/wait-no-turns.

Homes that do not create the file keep the previous briefs, drive text, and sends.

* no-mistakes(document): Document wait-no-turns inbox wording change in configuration

* no-mistakes(review): Keep checkpoint inbox check; forbid only polling while waiting

* no-mistakes(ci): Fixed ci-2 (Greptile: a concurrent retry marker gets lost). The rule that was broken: the watcher may remove only the `.retry-ring` mark for the record it just processed. A newer mark written in the meantime is owed its own retry. `fm_task_inbox_clear_retry` is the one shared function that removes the mark, and I fixed it there. In `bin/fm-task-inbox-lib.sh` it now takes the record path. It compares the mark's content with that record's name and removes the mark only when they match. When the mark names a different record it returns success and leaves the mark alone. It still fails only when the processed record's own mark can't be removed. Both callers in `bin/fm-watch.sh` now pass `"$rec"`: the dead or missing pane path and the path after a retry ring. So the fix holds at both removal sites. Tests, in `tests/fm-task-inbox.test.sh`: - I added an optional `FM_RING_MARKS_RETRY` hook to the fake tmux. It writes a newer record's mark while the doorbell is being typed, which reproduces the race deterministically. - I added `test_watcher_retry_keeps_a_newer_mark`. The owed retry rings once, the newer mark survives, and a later check rings the newer record once and then clears its mark. The test fails without the fix ("the spent retry removed a newer record's mark written during its ring") and passes with it. - I updated the direct `clear_retry` call in the existing unit test to pass the record. Results: `tests/fm-task-inbox.test.sh` passes in full and `tests/fm-send-inbox.test.sh` passes 15/15. Shellcheck reports only SC1091 "not following sourced file" notices. As instructed, I didn't change the brief inbox wording

* no-mistakes(document): Fix stale wait-no-turns inbox wording in inbox lib comment

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

* fix(bin): close Gerrit-landed backlog items with the change URL as a note (#6140)

* fix(bin): record Gerrit change URLs as close notes

Teardown's backlog_done_args hands every ship's recorded pr= URL to
fm_backlog_done as --pr, and tasks-axi refuses any --pr that is not a
canonical GitHub or Forgejo pull request. A Gerrit change URL therefore
left the item In flight after cleanup, and the pending backlog-close
record replayed into the same refusal at every session start.

fm_backlog_done now rewrites a --pr whose value fm_pr_url_parse reads
as a Gerrit change into --note "Gerrit change <url>". The mapping sits
at the tasks-axi call rather than in the pending-close record, so
records already written with --pr replay to a close unchanged. The
captain-held retain path records the URL in its deliverable line and
skips the update --pr it cannot make.

* no-mistakes(review): Note retained Gerrit change URL when captain answers early

* no-mistakes(document): Document Gerrit change URL handling in captain-hold retention

* fix: reduce remote-job and supervision polling churn (#6255)

* perf(remote): separate active job sampling from dispatcher cadence

* no-mistakes(document): Link remote wait timing to its authoritative contract

* no-mistakes(ci): Fixed ci-1 with two narrowly scoped SC2030 annotations documenting intentional subshell-local legacy and active cadence overrides in tests/fm-remote-job.test.sh. Runtime behavior is unchanged. Reproduced the lint failure before the fix; afterward ShellCheck 0.11.0 with source following, Bash syntax validation, the complete remote-job behavior suite, and git diff --check all passed

* perf(supervision): reduce park, delta and dispatcher polling

* no-mistakes(document): Clarify poll latency contracts and authoritative documentation pointers

* fix(bin): load backend sibling libraries when sourced under zsh (#6221)

* fix(bin): load backend sibling libraries under zsh

fm_backend_source kept each backend's sibling list in one space-separated
string and iterated it unquoted. zsh does not word-split an unquoted
expansion, so the readability check saw the whole list as one path and
refused every backend with more than one sibling. Hold the list in the
function's positional parameters instead, which needs no word splitting
in Bash 3.2, Bash 5, or zsh.

The existing zsh case in tests/fm-backend.test.sh covers it wherever zsh
is installed.

* test: run the Calm mod suite on stock Bash 3.2

The suite injected shell values into its generated Node scripts with the
${value@Q} transformation, which needs Bash 4.4. Stock macOS Bash 3.2
reports a bad substitution, so every case failed before it asserted
anything. Build each JavaScript string literal with JSON.stringify
through a small helper instead, which works on any Bash and is a valid
literal for any value.

* no-mistakes(review): fix(bin): rename zsh-special path local in fm_backend_source

* test: narrow the zsh backend claim to name matching

Under zsh the adapters locate their siblings through BASH_SOURCE, so a
successful fm_backend_source is not a full load. Assert only what the
contract states, and pass js_string values after -- so node never reads
a leading-dash value as its own option.

---------

Co-authored-by: Nova Agent B <novaagentb@gmail.com>

* fix(bin): exclude a remote mate's own parent channel from self-home status scans (#5263)

* fix(bin): exclude a remote mate's own parent channel from self-home scans

A remote secondmate home's outbound parent channel lives at state/parent-replies.status inside its own state dir, so the watcher's signal scan enumerated it as a task status file and the open-decisions fold classified it as a phantom task named parent-replies: every parent-channel append spun a spurious signal wake and a phantom open decision in the mate's own home.
fm-parent-channel-lib.sh gains fm_parent_channel_outbound_status, which resolves the channel into the mate's own state dir for the remote route only, and fm-classify-lib.sh's status_scan_parent_channel_exclude wraps it for the fleet-wide scans.
The watcher's scan_signals and heartbeat fail-safe backstop, the whole-file and incremental open-decisions folds, the presentation snapshot, and the unread-surface scan now skip exactly that resolved path.
The exclusion is home-shape-aware: a parent-replies.status in a main home or a local mate is an ordinary task log and keeps waking and folding, and every other status file is untouched.

* no-mistakes(review): exclude a remote mate's parent channel from the daemon heartbeat scan

* no-mistakes(document): Document remote mate parent-channel scan exclusion

* ci: retrigger portable serial 4

* no-mistakes(ci): CI check 'Behavior portable serial 7' failed in tests/fm-contributions.test.sh ('reservation poll failed'). CI stderr showed bin/fm-contributions.sh:345 arithmetic 'DEADLINE - 6\n90077104: syntax error in expression': the fixture's fake date returned a torn two-line clock value. Root cause: the fake forge wrapper in wrap_forge advances the shared controllable clock via a non-atomic read-modify-write ('$(cat $FORGE/clock) + 6' with truncate-in-place '> $FORGE/clock') while concurrent background gh calls run and the fake date reads the same file; an interleaved truncate+write publishes a half-written value (CI's torn '6\n90077104', tail of 1790077104) or an emptied-read value ('6'), which either breaks the poll's arithmetic (nonzero exit -> 'reservation poll failed') or defeats the 15-second reservation defer. This is a pre-existing test-fixture race, not caused by the PR's diff (base..target touches no contributions code; the same commit passed this shard in run 35711207830 earlier the same day). Fixed the flaky fixture at its root: clock_bump() now writes each new value to a per-process mktemp file in the same directory and publishes it with mv (atomic rename), so concurrent forge callers and the fake date always read one complete old-or-new clock; fault patterns and deltas are unchanged. Verified: minimal 3-way concurrency repro shows the old wrapper corrupting (12/32/38 outcomes incl. empty-read) while the rename-based wrapper never corrupts (20/20 clean); the full tests/fm-contributions.test.sh passes twice (all 38 assertions ok, incl. the reservation, budget-exhaustion, genuine-failure, shared-once, and latency tests); 10 isolated reservation runs pass; shellcheck rc=0; worktree contains only this one-file change

* no-mistakes(document): drop stale file-set copy in daemon catch-all comment

* fix(bin): document accepted contribution verdict actors (#6307)

* fix(bin): name the accepted verdict actors in fm-contributions help and refusal

* fix(ci): Updated tests/fm-contributions.test.sh to assert exactly captain, fleet, maintainer, and nobody in command-emitted help and refusal output. Three focused regressions passed; all three extra-actor mutations were rejected. ShellCheck, syntax, and diff checks passed. Production code remains unchanged

* fix(bin): recognize clone roots across path spelling differences (#6306)

* fix(bin): recognise a clone root git names with different path spelling

fm-fleet-sync compared git's --show-toplevel with pwd -P as strings, so a clone
root that git recorded with different casing (case-insensitive volume) was
skipped as not a clone root and never refreshed. Compare filesystem identity
instead, which also covers symlink spelling.

* fix(document): Remove stale clone-root comparison comment

* test: preserve Pi calm transcript captures with Pi 1.0 (#6338)

* test(calm): pin Pi's regular TUI mode where pane assertions read scrollback

Pi 1.0.0 defaults its TUI to a fullscreen alternate-screen mode whose
scrollable transcript is application-owned, so rows that leave the viewport
never enter terminal scrollback and tmux capture-pane -S can no longer see
them. The Pi Calm e2e launches now pass --tui-mode regular wherever the flag
exists so the transcript assertions keep reading real scrollback on both the
Pi 1.0.0 line and earlier Pi lines, which have no such flag and render
regular-only anyway.

* no-mistakes(document): Correct Pi TUI documentation and scrollback rationale

* fix(bin): preserve hold reasons and reject invalid completion inventories (#6331)

* fix(bin): encode captain-hold reasons and reject self-inventory in complete

hold now stores a reason with parentheses, line breaks, or percent signs
through a reversible percent encoding that every reader decodes, instead of
refusing it. hold --origin records the origin on the held task, and complete
refuses the origin as its own inventory entry and an entry held for a
different origin; holds with no recorded origin are accepted and flagged.

* fix(review): Decode marked hold reasons consistently across readers

* fix(review): Remove unnecessary lifecycle test dispatch

* fix(review): Correct hold origin identity and inventory recovery

* fix(review): Record origins before placing backend holds

* fix(document): Clarify captain-hold validation and reason reader documentation

* fix(ci): Fixed both findings: failed backend holds restore the previous origin, and invalid base64/UTF-8 reasons remain verbatim. Added regressions and documented valid-literal ambiguity. Both failures were reproduced before fixes. Verification: 54 lifecycle tests and 9 wrapper tests passed; 7 Beads-specific cases skipped because tasks-axi is markdown-only. Focused lint and diff checks passed. No pipeline or publication actions performed

* fix: reclaim orphaned watcher arms on the next park (#6335)

* fix(bin): take over the watcher cycle a main-only pass-through leaves

An attended main-only pass-through leaves a successor watcher cycle
running through main's handling turn. The session's next park attached
to that cycle instead of owning it, so the successor's arm, orphaned by
its host's exit, kept owning the watcher while the new park's arm polled
it twice a second until the next close or the park boundary, hours later
in a quiet second mate. A second-mate restart hit this every time, since
its persist request is a main-only close.

The host now records the successor it leaves for main, and the next
host's first cycle runs bin/fm-watch-arm.sh --take-over on it: when that
arm still owns the healthy watcher, the new arm stops it, reports a
reason the cycle delivered first, and otherwise owns a fresh cycle. The
stop's own downtime publication is undone over an acknowledged episode
when no wake was appended in between, so the handover wakes nobody.

* no-mistakes(review): Keep left-arm record until the orphaned arm is gone

* no-mistakes(review): Relinquish successor arm only after durably recording it

* no-mistakes(review): Relinquish successor only after its record reads back

* no-mistakes(document): Clarify successor takeover guarantees and authoritative documentation

* no-mistakes(ci): Fixed ci-1: acknowledgement restore now requires the exact taken-over arm/watcher ledger row with signal=TERM, awaited within a short bound. Otherwise takeover proceeds without erasing downtime. Added a self-exit regression confirmed failing before the fix and passing afterward; ordinary takeover tests and the full watcher-arm suite pass. Updated Generation reuse documentation. Syntax, diff checks, and ShellCheck pass with existing SC1091/SC2034 warnings excluded. ci-2 remains unchanged

* no-mistakes(document): Clarify watcher take-over recovery and restart limits

* fix(bin): restore downtime on supervision-host hand-back when the successor already closed (#6355)

* fix(bin): restore supervision host hand-back continuity

* no-mistakes(review): Scope host hand-back failure fallback to lost pending:handling

* no-mistakes(review): Remove stray scratch test copy tests/.tmp-rest.test.sh

* no-mistakes(test): Initialise successor globals so early hand-back survives set -u

* no-mistakes(document): Document host hand-back downtime failure and Claude lost-handback notice

* no-mistakes(ci): I fixed the Greptile finding. The rule that must hold: when the supervision host hands back an actionable wake, its rewake is refused, and no watcher is healthy, the hand-back still has to reach main as a delivered notice. That must be true whether the recovery marker is `pending:handling` or `announced:handling`. Only one place applies this check: the lost hand-back fallback in `bin/fm-claude-stop-autoarm.sh`. **Fix:** that check now accepts both `pending:handling:*` and `announced:handling:*` tokens (a one-line change). Nothing else in the fallback changed: - Refusals on any other marker, such as an already acknowledged one, still exit 0 silently and open no failure episode. - The notice is still sent once per episode, and repeats are recorded as `failed-suppressed`. **Tests:** - `tests/fm-claude-stop-autoarm.test.sh`: the lost hand-back test now runs as a shared helper with two variants, one writing a `pending:handling` marker and a new one writing `announced:handling` (`test_host_lost_announced_handback_notifies_once_per_episode`). - `tests/fm-supervision-host.test.sh`: the end-to-end test where downtime restoration fails is now a shared helper too, with a new `announced` variant (`test_claude_stop_hook_notifies_when_closed_announced_successor_downtime_restore_fails`). It moves the handling episode to `announced` before the host hands back. Without the fix the hook would exit 0 here; the test requires exit 2, `outcome=failed` and a delivered failure notice. **Verification (all under nice -n 10):** - The full `tests/fm-claude-stop-autoarm.test.sh` suite passed (rc=0), including both lost hand-back variants and the benign-refusal test. - In `tests/fm-supervision-host.test.sh` I ran only the four hand-back test functions, all passing (rc=0). The suite can't run single functions, so I used a temporary copy with a trimmed test list and deleted it afterwards; `git status` shows only the 3 intended files changed. - shellcheck is clean on all three changed files. I did not run `bin/fm-lint.sh`. - I did not run the new tests against the unfixed code; the claim that they fail without the fix comes from reading the old check

* fix: reduce remote-job polling process churn (#6363)

* perf: cut remote-job idle process creation in the three hot loops

Post-update host measurement still attributes most idle churn to three
per-sample loops: result-consumer state reads and date calls, the delta
reader's capture/hash pass on every poll, and the lane preemption scan's
per-field pipelines. This drops each to its minimum without touching the
contracts around them.

* fm_remote_job_read_state gains an optional result-variable form backed
  by fm_remote_job_read_line, a builtin-only bounded record read (regular
  non-symlink file, byte bound, one newline-terminated line, tolerated
  unterminated tail, no carriage returns). fm_remote_job_wait samples
  state and the SECONDS clock with no per-sample children; one date call
  converts the epoch deadline once.
* fm-remote-delta-read stats the log each poll and re-runs the bounded
  capture and hashing only when size, mtime, ctime, inode, or device
  change. The snapshot's own stat writes the comparison key, so a log
  that moves between the gate and the capture is never read as stable.
* worker_preempting_waiter_exists reads state, home, and the staged argv
  head with builtins only. The now-unused worker_job_command goes away.

The bounded reads use -d '' -n, which behaves identically on the macOS
stock bash 3.2 and current bash; -N does not exist on 3.2. Tests cover
the malformed-record corpus, delta identity gating, fork-free lane
scanning through counting PATH shims, and same-home versus cross-home
preemption. No signal traps or sleep contracts change.

* no-mistakes(review): Restore subsecond delta keys and byte-bounded builtin record reads

* no-mistakes(document): Clarify delta snapshot caching and coarse-timestamp fallback

* no-mistakes(lint): Scope UTF-8 regression locales to individual function calls

* no-mistakes(ci): Fixed both lint failures by applying the documented production-library analysis boundary at the two affected test imports. Runtime behavior is unchanged; the library remains independently linted. Canonical full-analysis lint passed for the library and both suites, as did bash syntax checks and git diff --check

* fix(bin): recognize titled Claude top rules and preserve grey slash commands (#5963)

* fix(composer): read a titled Claude top rule as the composer's edge

A named Claude Code session draws its title into the composer's top rule.
The strict separator predicate rejected that row, so the closing rule read
as a lower unmatched separator and an idle, empty composer classified
unknown on every cursorless backend, refusing fm-send, exit, and relaunch.

Spare a bare agent-glyph row sandwiched between a width-proven titled rule
and the screen's only unmatched separator directly below it. The strict
separator predicate, dead-shell rule, and blank-row posture are unchanged.

Fixes #5601
Fixes #5558

* no-mistakes(test): Keep Claude's grey slash command in Herdr payload proof

* no-mistakes(test): Make missing-herdr version check hermetic to installed herdr

* fix: prevent Pi trust prompts in seeded secondmate homes (#6387)

* fix: pre-approve Pi trust for seeded secondmate homes

Unattended first launches of Firstmate-seeded Pi secondmate homes stalled on
"Trust project folder?" until Enter. Probe --approve like --tui-mode and pass
it only for --secondmate when help advertises it (.fm-secondmate-home signal),
leaving ordinary workers and older Pi unchanged.

* no-mistakes(document): Consolidate Pi seeded-home trust documentation ownership

* no-mistakes(ci): Fixed Lint 1’s unused polling counter. Diagnosed Behavior portable serial 4 as a pre-existing delta-reader test clock race; replaced timing-dependent rewrite and deletion with deterministic executable-boundary synchronization. ShellCheck, Bash syntax checks, and git diff --check passed. Delta-reader tests passed three consecutive runs; all three live Pi trust cases passed. Production behavior unchanged

* fix(bin): retry ShellCheck roots that hit the memory ceiling without --external-sources (#6443)

* fix(lint): retry memory-bound roots without external sources

* no-mistakes(review): Make fallback tests portable and correct source-following telemetry

* no-mistakes(review): Remove committed parity fixtures and use disposable test roots

* no-mistakes(document): Document ShellCheck memory fallback and telemetry

* no-mistakes(review): Cover bounded and unbounded fallback RSS behavior

* no-mistakes(document): Correct stale lint fallback documentation

* no-mistakes(document): Correct stale lint test documentation

* no-mistakes(ci): The memory fallback (the retry without --external-sources) now gets only the time left in its root's original deadline, so it can no longer outlast the CI job. Invariant: one root's first attempt plus its fallback must fit inside a single FM_LINT_ROOT_SECONDS deadline, plus the cleanup grace. Only one site started a new deadline: the fallback call in fm_lint_run_root. The deadline is the only budget involved, because the memory limit already applies to each process separately. Changes in bin/fm-lint.sh: - fm_lint_exec_root now takes a <seconds> argument instead of always reading FM_LINT_INTERNAL_ROOT_SECS. - The first attempt passes the full deadline. - The fallback passes floor((start + deadline - now) / 1000) seconds. - When bounds are enforced and less than 1 second is left, no retry starts. fm_exec_timed rejects 0 seconds, so the retry cannot run with no time. The root keeps reason=memory, and the shard output says "no time left in its Ns deadline to retry without it". - Unbounded local runs have no deadline and behave as before. - The header comment now describes the shared deadline. Changes in tests/fm-lint.test.sh: a new test, test_memory_fallback_spends_only_the_remaining_root_deadline, runs only on hosts that can enforce bounds. It uses a 6 s deadline and 1 s grace. - Case 1…
yehezkieled added a commit to yehezkieled/firstmate that referenced this pull request Oct 5, 2026
* fix: reclaim orphaned watcher arms on the next park (kunchenguid#6335)

* fix(bin): take over the watcher cycle a main-only pass-through leaves

An attended main-only pass-through leaves a successor watcher cycle
running through main's handling turn. The session's next park attached
to that cycle instead of owning it, so the successor's arm, orphaned by
its host's exit, kept owning the watcher while the new park's arm polled
it twice a second until the next close or the park boundary, hours later
in a quiet second mate. A second-mate restart hit this every time, since
its persist request is a main-only close.

The host now records the successor it leaves for main, and the next
host's first cycle runs bin/fm-watch-arm.sh --take-over on it: when that
arm still owns the healthy watcher, the new arm stops it, reports a
reason the cycle delivered first, and otherwise owns a fresh cycle. The
stop's own downtime publication is undone over an acknowledged episode
when no wake was appended in between, so the handover wakes nobody.

* no-mistakes(review): Keep left-arm record until the orphaned arm is gone

* no-mistakes(review): Relinquish successor arm only after durably recording it

* no-mistakes(review): Relinquish successor only after its record reads back

* no-mistakes(document): Clarify successor takeover guarantees and authoritative documentation

* no-mistakes(ci): Fixed ci-1: acknowledgement restore now requires the exact taken-over arm/watcher ledger row with signal=TERM, awaited within a short bound. Otherwise takeover proceeds without erasing downtime. Added a self-exit regression confirmed failing before the fix and passing afterward; ordinary takeover tests and the full watcher-arm suite pass. Updated Generation reuse documentation. Syntax, diff checks, and ShellCheck pass with existing SC1091/SC2034 warnings excluded. ci-2 remains unchanged

* no-mistakes(document): Clarify watcher take-over recovery and restart limits

* fix(bin): restore downtime on supervision-host hand-back when the successor already closed (kunchenguid#6355)

* fix(bin): restore supervision host hand-back continuity

* no-mistakes(review): Scope host hand-back failure fallback to lost pending:handling

* no-mistakes(review): Remove stray scratch test copy tests/.tmp-rest.test.sh

* no-mistakes(test): Initialise successor globals so early hand-back survives set -u

* no-mistakes(document): Document host hand-back downtime failure and Claude lost-handback notice

* no-mistakes(ci): I fixed the Greptile finding. The rule that must hold: when the supervision host hands back an actionable wake, its rewake is refused, and no watcher is healthy, the hand-back still has to reach main as a delivered notice. That must be true whether the recovery marker is `pending:handling` or `announced:handling`. Only one place applies this check: the lost hand-back fallback in `bin/fm-claude-stop-autoarm.sh`. **Fix:** that check now accepts both `pending:handling:*` and `announced:handling:*` tokens (a one-line change). Nothing else in the fallback changed: - Refusals on any other marker, such as an already acknowledged one, still exit 0 silently and open no failure episode. - The notice is still sent once per episode, and repeats are recorded as `failed-suppressed`. **Tests:** - `tests/fm-claude-stop-autoarm.test.sh`: the lost hand-back test now runs as a shared helper with two variants, one writing a `pending:handling` marker and a new one writing `announced:handling` (`test_host_lost_announced_handback_notifies_once_per_episode`). - `tests/fm-supervision-host.test.sh`: the end-to-end test where downtime restoration fails is now a shared helper too, with a new `announced` variant (`test_claude_stop_hook_notifies_when_closed_announced_successor_downtime_restore_fails`). It moves the handling episode to `announced` before the host hands back. Without the fix the hook would exit 0 here; the test requires exit 2, `outcome=failed` and a delivered failure notice. **Verification (all under nice -n 10):** - The full `tests/fm-claude-stop-autoarm.test.sh` suite passed (rc=0), including both lost hand-back variants and the benign-refusal test. - In `tests/fm-supervision-host.test.sh` I ran only the four hand-back test functions, all passing (rc=0). The suite can't run single functions, so I used a temporary copy with a trimmed test list and deleted it afterwards; `git status` shows only the 3 intended files changed. - shellcheck is clean on all three changed files. I did not run `bin/fm-lint.sh`. - I did not run the new tests against the unfixed code; the claim that they fail without the fix comes from reading the old check

* fix: reduce remote-job polling process churn (kunchenguid#6363)

* perf: cut remote-job idle process creation in the three hot loops

Post-update host measurement still attributes most idle churn to three
per-sample loops: result-consumer state reads and date calls, the delta
reader's capture/hash pass on every poll, and the lane preemption scan's
per-field pipelines. This drops each to its minimum without touching the
contracts around them.

* fm_remote_job_read_state gains an optional result-variable form backed
  by fm_remote_job_read_line, a builtin-only bounded record read (regular
  non-symlink file, byte bound, one newline-terminated line, tolerated
  unterminated tail, no carriage returns). fm_remote_job_wait samples
  state and the SECONDS clock with no per-sample children; one date call
  converts the epoch deadline once.
* fm-remote-delta-read stats the log each poll and re-runs the bounded
  capture and hashing only when size, mtime, ctime, inode, or device
  change. The snapshot's own stat writes the comparison key, so a log
  that moves between the gate and the capture is never read as stable.
* worker_preempting_waiter_exists reads state, home, and the staged argv
  head with builtins only. The now-unused worker_job_command goes away.

The bounded reads use -d '' -n, which behaves identically on the macOS
stock bash 3.2 and current bash; -N does not exist on 3.2. Tests cover
the malformed-record corpus, delta identity gating, fork-free lane
scanning through counting PATH shims, and same-home versus cross-home
preemption. No signal traps or sleep contracts change.

* no-mistakes(review): Restore subsecond delta keys and byte-bounded builtin record reads

* no-mistakes(document): Clarify delta snapshot caching and coarse-timestamp fallback

* no-mistakes(lint): Scope UTF-8 regression locales to individual function calls

* no-mistakes(ci): Fixed both lint failures by applying the documented production-library analysis boundary at the two affected test imports. Runtime behavior is unchanged; the library remains independently linted. Canonical full-analysis lint passed for the library and both suites, as did bash syntax checks and git diff --check

* fix(bin): recognize titled Claude top rules and preserve grey slash commands (kunchenguid#5963)

* fix(composer): read a titled Claude top rule as the composer's edge

A named Claude Code session draws its title into the composer's top rule.
The strict separator predicate rejected that row, so the closing rule read
as a lower unmatched separator and an idle, empty composer classified
unknown on every cursorless backend, refusing fm-send, exit, and relaunch.

Spare a bare agent-glyph row sandwiched between a width-proven titled rule
and the screen's only unmatched separator directly below it. The strict
separator predicate, dead-shell rule, and blank-row posture are unchanged.

Fixes kunchenguid#5601
Fixes kunchenguid#5558

* no-mistakes(test): Keep Claude's grey slash command in Herdr payload proof

* no-mistakes(test): Make missing-herdr version check hermetic to installed herdr

* fix: prevent Pi trust prompts in seeded secondmate homes (kunchenguid#6387)

* fix: pre-approve Pi trust for seeded secondmate homes

Unattended first launches of Firstmate-seeded Pi secondmate homes stalled on
"Trust project folder?" until Enter. Probe --approve like --tui-mode and pass
it only for --secondmate when help advertises it (.fm-secondmate-home signal),
leaving ordinary workers and older Pi unchanged.

* no-mistakes(document): Consolidate Pi seeded-home trust documentation ownership

* no-mistakes(ci): Fixed Lint 1’s unused polling counter. Diagnosed Behavior portable serial 4 as a pre-existing delta-reader test clock race; replaced timing-dependent rewrite and deletion with deterministic executable-boundary synchronization. ShellCheck, Bash syntax checks, and git diff --check passed. Delta-reader tests passed three consecutive runs; all three live Pi trust cases passed. Production behavior unchanged

* fix(bin): retry ShellCheck roots that hit the memory ceiling without --external-sources (kunchenguid#6443)

* fix(lint): retry memory-bound roots without external sources

* no-mistakes(review): Make fallback tests portable and correct source-following telemetry

* no-mistakes(review): Remove committed parity fixtures and use disposable test roots

* no-mistakes(document): Document ShellCheck memory fallback and telemetry

* no-mistakes(review): Cover bounded and unbounded fallback RSS behavior

* no-mistakes(document): Correct stale lint fallback documentation

* no-mistakes(document): Correct stale lint test documentation

* no-mistakes(ci): The memory fallback (the retry without --external-sources) now gets only the time left in its root's original deadline, so it can no longer outlast the CI job. Invariant: one root's first attempt plus its fallback must fit inside a single FM_LINT_ROOT_SECONDS deadline, plus the cleanup grace. Only one site started a new deadline: the fallback call in fm_lint_run_root. The deadline is the only budget involved, because the memory limit already applies to each process separately. Changes in bin/fm-lint.sh: - fm_lint_exec_root now takes a <seconds> argument instead of always reading FM_LINT_INTERNAL_ROOT_SECS. - The first attempt passes the full deadline. - The fallback passes floor((start + deadline - now) / 1000) seconds. - When bounds are enforced and less than 1 second is left, no retry starts. fm_exec_timed rejects 0 seconds, so the retry cannot run with no time. The root keeps reason=memory, and the shard output says "no time left in its Ns deadline to retry without it". - Unbounded local runs have no deadline and behave as before. - The header comment now describes the shared deadline. Changes in tests/fm-lint.test.sh: a new test, test_memory_fallback_spends_only_the_remaining_root_deadline, runs only on hosts that can enforce bounds. It uses a 6 s deadline and 1 s grace. - Case 1: the first attempt runs 3 s and then fails with memory status 251. The test asserts one fallback ran, reported reason=timeout, and the root's recorded duration is under 7000 ms. - Case 2: the first attempt runs 5.2 s. The test asserts no fallback starts, the skip is explained, and the sidecar records memory with source-following 1. Verification: - Full `nice -n 10 bash tests/fm-lint.test.sh` passed, including the new test, in about 5 minutes. - Case 1 run against the HEAD script: the root took 9168 ms, so the under-7000 ms check fails before the fix. - `bin/fm-lint.sh bin/fm-lint.sh tests/fm-lint.test.sh` reported no findings. - The CI workflow is unchanged, so the Test step still runs only tests/fm-lint.test.sh with nice -n 10 and the 12 GiB ShellCheck limit

* fix(pi): restore watcher continuity across successor gaps and make extension log opt-in (kunchenguid#5489)

* Fix Pi watcher successor-gap confirmations and add extension log

Accept an already-acknowledged handling confirmation as a no-op when the
generation matches, confirm the restoration's own recovery token with a
superseded (not rejected) outcome on generation mismatch, retire an arm on
confirm failure only when the failed token names that exact pid, and record
restore attempts, readiness timeouts, and confirm results in the bounded
state/.watch-extension.log. Regression tests: already-acked no-op plus
mismatch/dead-pid/lock-mismatch rejections and the manual-restart churn
contract in fm-watch-arm.test.sh, and a mid-restore marker advance with no
rejection appendix in fm-pi-watch-extension.test.sh.

* Treat a dead arm child as an empty slot so repair and retry recover

startArm and scheduleRetry answered unchanged while holding a ChildProcess
whose OS process was already gone but whose close had not fired, so neither
the repair tool nor the retry timer started anything until that close fired.
Gate slot occupancy on a liveness check (exit/signal codes plus pid probe)
and start a fresh arm instead, with a regression test driving the repair
tool against a dead-but-unclosed child.

* no-mistakes(document): Document new Pi extension log knob

* no-mistakes(review): Fix confirm-failure retire token match, add distinct-pid test

* no-mistakes(document): Clarify retire guard needs pid and generation

* Make the Pi extension diagnostic log opt-in and default-off

Only a positive FM_WATCH_EXTENSION_LOG_KEEP_LINES enables
state/.watch-extension.log. Unset, empty, non-numeric, zero, and
negative values disable logging entirely, so the default run writes
nothing and never creates the file. The shared positiveInteger
fallback semantics stay untouched for the retry and timeout knobs.
docs/configuration.md owns the knob contract and
docs/watcher-continuity.md points at it. Tests: the superseded-delivery
case runs opted in, and a new case proves unset, zero, and non-numeric
values create no log file while delivery still succeeds.

* no-mistakes(document): Qualify extension-log coverage bullet as opt-in

* no-mistakes(ci): The two reported checks (CI run 36372002913, Require no-mistakes run 36372069208) show conclusion action_required with 0 jobs and no logs because this is a fork PR (RibatTRW/firstmate) and GitHub is holding the workflow runs awaiting maintainer approval; that gate is external to the code and needs a maintainer to approve the runs. While verifying the change locally I found a real defect the approved CI run would hit: the PR's new test test_handling_delivered_rejects_a_superseded_generation failed deterministically. Invariant violated: reopen-announced (a non-successor/manual arm start) mints a fresh recovery generation only when the durable wake queue holds unrecovered work; an announced episode with an empty queue must be left untouched so idle arm starts never churn generations (the [ -s queue ] guard from kunchenguid#4819, relied on by bin/fm-watch.sh:2413 and covered by the append-reopens and announcement-bound sibling tests). The test called reopen with an empty queue and expected churn, so the fix establishes the queued-work precondition (append one wake, re-announce, re-read the generation) before asserting the reopen mints and the old confirmation mismatches. Test-only change, 14 lines in tests/fm-watch-arm.test.sh. Verified: fm-watch-arm.test.sh 25/25 ok on two consecutive runs, fm-pi-watch-extension.test.sh 55/55 ok, and shellcheck reports only one pre-existing warning outside the edited region

* no-mistakes(document): Restore blank line in watcher-continuity docs

* Route superseded Pi deliveries like confirmed ones and cover the retire guard

A superseded handling confirmation now falls through to the normal delivery
path, so an accepting supervision branch owns the wake instead of main.
Scope the watcher-continuity token-pinned confirmation and narrowed retire
rule to Pi, since omp and OpenCode still confirm against the current
successor. Add tests that fail when the retire guard, the scheduled-retry
gate, or the deferred-close gate is reverted, relabel the churned-generation
characterization test, and use a reaped pid for the dead-pid rejection.

* fix(pi): hide duplicate assistant finals from hidden processing retries (kunchenguid#5863)

* fix(pi): silence unacknowledged processing retry replies

Suppress autonomous processing prose before persistence and during streaming while retaining tool calls, signed reasoning, and retryable outcomes. Restore ordinary output after acknowledgement or a user message.

Fixes kunchenguid#4954

* no-mistakes(review): Silence only processing retries, keep first presentation visible

* fix(pi): preserve differing processing retry replies

* test(pi): accept Pi 1.0.1's renamed HTML export renderer lookup (kunchenguid#6530)

Pi 1.0.1's createToolHtmlRenderer reads getToolRenderers and ignores
getToolDefinition. Calm /export still includes stock grep HTML; the
fixture has to pass the lookup key the installed Pi actually reads.

* fix(bin): tolerate transient quota read failures (kunchenguid#6490)

* fix(procevent-quota): tolerate consecutive slow quota-axi reads

The quota allowance poll treated any quota_json failure as terminal, so one
slow quota-axi --json (measured max ~29s under a 48s derived bound) shut the
watch down until someone re-armed it, and the detail always said
"missing/incompatible". Tolerate three consecutive failed or timed-out reads
before going terminal, reset the streak on any good read, and report a
timeout distinctly from a missing or incompatible tool. Each timed poll runs
exactly one bounded --version and one bounded --json: validate the captured
version text through fm_quota_axi_version_compatible rather than launching a
second probe, and describe a mixed failure streak by count plus last cause.

* no-mistakes(document): Document quota polling failure tolerance

* no-mistakes(ci): Fixed ci-2 and ci-3. Permanent quota read failures (rc 2 missing, rc 3 incompatible) now report on the first poll, while transient rc 1/4 failures retain the existing three-failure retry behavior. The missing-binary test now uses an isolated PATH without quota-axi and asserts both permanent failures stop at condition_polls: 1. Verification passed: tests/fm-procevent-quota.test.sh, canonical fast lint for both changed files, bash syntax checks, and git diff --check

* no-mistakes(review): Classify untimed quota version failures as transient

* no-mistakes(document): Clarify quota polling failure budget

* fix(bin): clarify scratch guidance and dirty teardown refusals (kunchenguid#6505)

* fix(teardown): clarify scratch guidance and dirty worktree refusals

Keep ship proof material outside the task worktree and distinguish untracked-only leftovers from tracked edits without changing cleanup guards.

Fixes kunchenguid#6319

* fix(ci): Fixed ci-3 only. Both promotion outputs now replace the scout restriction and require external scratch storage and a clean worktree before done. Verification: 27 delivery tests passed, five mutations caught, restored test passed, pinned ShellCheck and syntax/whitespace checks passed. ci-1, ci-2, and ci-4 remain untouched

* fix(bin): escalate inbox instructions blocked by busy workers (kunchenguid#6518)

* Escalate inbox instructions stuck behind a busy worker

Count consecutive busy-deferred due doorbells durably and escalate at the configured bound without typing into the worker pane.

Fixes kunchenguid#6445

* fix(review): Fix inbox escalation deduplication and busy streak resets

* fix(review): Preserve busy inbox escalations through daemon supervision

* fix(document): Correct busy-inbox escalation documentation

* fix(ci): Fixed SC2034 in tests/fm-task-inbox.test.sh by including the loop counter in the failure diagnostic. Source-aware lint, all 34 inbox tests, and git diff --check pass. Behavior portable serial 6 reproduces identically on base 1f3e769 and target 78156b8 with Pi 1.0.1: an unrelated renderer API change breaks the unchanged Calm test. No Calm changes made; that failure is addressed separately by kunchenguid#6516. Logs retained in scratchpad-ci/

* fix(ci): Fixed ci-2, ci-3, and ci-4: successor failures surface, reset alerts deduplicate, and oversized busy limits fall back to two. Passed 47 inbox tests, 7 focused daemon checks, all 13 mutation checks, lint, documentation checks, and diff checks. Evidence: scratchpad-ci-selected/summary.json. ci-1 remains unchanged and unwaived. Fresh live Herdr proof remains with the outer driver

* fix: prevent unnecessary remote worker turnover (kunchenguid#6431)

* fix: prevent healthy remote job worker turnover

* no-mistakes(review): Serialize full LaunchAgent repair and verify launchd-tracked owners

* no-mistakes(review): Let launchd-tracked unpublished spawns start before reloading

* no-mistakes(document): Document remote worker heartbeat and serialized LaunchAgent recovery

* no-mistakes(ci): Full CI log showed the idle-worker regression exceeded its outdated command budget (82 versus 80) after independent heartbeat ownership checks were added. Raised the budget to 120 while retaining the separate busy-poll sleep limit. Remote-job and LaunchAgent executable tests passed, as did bash syntax validation and git diff --check. No production behavior changed

* no-mistakes(ci): Fixed missing readiness recovery under verified live ownership, preserving the serving PID and private file mode. Added executable regressions for deletion during a blocked sweep and stale readiness diagnostics without LaunchAgent reload. Deletion regression failed before the fix. Both remote-job test suites, bash syntax validation, and git diff --check passed

* no-mistakes(ci): Full CI log identified a flaky ownership-loss test racing an already-authorized heartbeat refresh. Replaced backdating and a fixed sleep with bounded observation of readiness expiry through the public probe. Production behavior unchanged. Remote-job and LaunchAgent executable suites passed; bash syntax validation and git diff --check passed

* fix: wait for launchd bootout cleanup

* no-mistakes(document): Document remote worker recovery and read-only turnover verification

* no-mistakes(review): Publish worker identity before lock owner records

* no-mistakes(document): Document worker identity publication safety invariant

* no-mistakes(ci): Fixed ci-1: replacement workers discard predecessor readiness before publishing identity and roll back identity if lock-owner recording fails. Added executable regressions reproducing both defects. LaunchAgent and orphan-reap suites, shell syntax checks, and git diff --check passed. No live service state was modified

* no-mistakes(ci): Restored lock-owner-before-identity publication and removed the identity rollback and reordering-only tests. Retained an executable regression proving predecessor readiness is rejected until replacement startup completes. LaunchAgent and orphan-reap suites, shell syntax checks, and git diff --check passed. Other changes remain intact; the retained-identity interrupted-repair edge remains out of scope. No live service state was modified

* fix: speed up remote-job sequence claim cleanup (kunchenguid#6575)

* fix(remote-job): reap expired seq claims with one directory walk

The hourly claim sweep forked uname+stat per .seq-claims entry and blocked
serving for ~85s at ~17k dirs. Delete expired empty claim dirs with a single
find -exec rmdir batch and cache the host uname for remaining mtime reads.

* no-mistakes(review): Restore original path mtime helper and drop uname cache

* no-mistakes(test): Restore claim retention eligibility; focused retention and serving tests pass

* no-mistakes(document): Document single-walk claim cleanup and regression entrypoints

* no-mistakes(ci): Fixed Lint 2’s reproduced SC1091 by adding the tests/lib.sh ShellCheck source directive to the retention test. Runtime behavior is unchanged. ShellCheck passed for both new claim tests; Bash syntax, the retention behavior test, and git diff --check passed

* fix(tests): keep fixture registries out of git worktree roots

A TMPDIR pointed at a repository root placed live .fm-test-* registries
beside tracked files, and a concurrent git add during the claim-walk CI
fix round committed three of them. Route registries and fixture roots
through a TMPDIR that refuses git worktree roots, remove the stray files,
and pin the escape with a behavioral cleanup test.

* no-mistakes(review): Preserve whole-second claim expiry in single-walk sweep

* no-mistakes(document): Correct temporary-directory resolution documentation

* no-mistakes(ci): Fixed ci-3 by changing only the stale, fresh, and read-only orphan fixture paths in tests/fm-test-fixture-cleanup.test.sh to use $FM_TEST_TMPDIR. All seven tests passed both normally and with TMPDIR set to the worktree root. Shell syntax and git diff --check passed

* no-mistakes(document): Align merged docs with upstream hand-back and policy owners

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: Tiago <tiagop@hey.com>
Co-authored-by: RibatTRW <aydinhrrs@gmail.com>
Co-authored-by: Ian Brown <742554+zestysoft@users.noreply.github.com>
Co-authored-by: Joseph Kim <jokim1@gmail.com>
Co-authored-by: Mickaël Rémond <mremond@process-one.net>
Co-authored-by: Hezki <hezki@users.noreply.github.com>
max-metaplanet pushed a commit to max-metaplanet/firstmate that referenced this pull request Oct 6, 2026
* fix(bin): take over the watcher cycle a main-only pass-through leaves

An attended main-only pass-through leaves a successor watcher cycle
running through main's handling turn. The session's next park attached
to that cycle instead of owning it, so the successor's arm, orphaned by
its host's exit, kept owning the watcher while the new park's arm polled
it twice a second until the next close or the park boundary, hours later
in a quiet second mate. A second-mate restart hit this every time, since
its persist request is a main-only close.

The host now records the successor it leaves for main, and the next
host's first cycle runs bin/fm-watch-arm.sh --take-over on it: when that
arm still owns the healthy watcher, the new arm stops it, reports a
reason the cycle delivered first, and otherwise owns a fresh cycle. The
stop's own downtime publication is undone over an acknowledged episode
when no wake was appended in between, so the handover wakes nobody.

* no-mistakes(review): Keep left-arm record until the orphaned arm is gone

* no-mistakes(review): Relinquish successor arm only after durably recording it

* no-mistakes(review): Relinquish successor only after its record reads back

* no-mistakes(document): Clarify successor takeover guarantees and authoritative documentation

* no-mistakes(ci): Fixed ci-1: acknowledgement restore now requires the exact taken-over arm/watcher ledger row with signal=TERM, awaited within a short bound. Otherwise takeover proceeds without erasing downtime. Added a self-exit regression confirmed failing before the fix and passing afterward; ordinary takeover tests and the full watcher-arm suite pass. Updated Generation reuse documentation. Syntax, diff checks, and ShellCheck pass with existing SC1091/SC2034 warnings excluded. ci-2 remains unchanged

* no-mistakes(document): Clarify watcher take-over recovery and restart limits
mlim1972 added a commit to mlim1972/firstmate that referenced this pull request Oct 8, 2026
* fix(bin): keep the steering doorbell short under deep homes (#6240)

* fix(bin): keep the steering doorbell short under deep homes

The doorbell printed the task inbox's absolute path twice, so under a deep
home it grew to about 290 characters and a Herdr submit reported it never
reached the pane on every re-ring. It now names the inbox once by its short
<task>.inbox name and points at the full path the worker's brief already
gives, so its length no longer depends on the home's depth.

Fixes #6120

* no-mistakes(review): Export FM_TASK_INBOX at launch and name it in doorbell

* no-mistakes(ci): ci-1 (Behavior portable serial 9) was caused by this PR, and I fixed it in the test. tests/fm-claude-trust.test.sh failed with "the launch command did not carry a brief doorbell". Its claude_launch_doorbell helper stripped exactly two leading `export ...;` statements before reading the final prompt argument. This PR adds a third one (`export FM_TASK_INBOX=...`) to every launch, so the helper was reading the wrong command. The invariant: a test that parses the launch command must skip every leading export statement, however many there are. I checked every test that parses the launch this way. The only other ones are the two helpers in tests/fm-spawn-dispatch-profile.test.sh, and they already loop over all exports. The kimi and dispatch-profile exact-string checks were updated earlier in this PR. The fix makes claude_launch_doorbell use the same loop (`while [[ "$command" == export\ *\;* ]]; do command=${command#*; }; done`) and then take the last argument. The ordinary path still works: the claude spawn test and the secondmate-clone spawn test both resolve the brief record through the same helper. Verified locally: `bash tests/fm-claude-trust.test.sh` exits 0 with no failing cases. ci-2 (Behavior tests (Herdr)) was not caused by this change, and I made no code change for it. In tests/fm-backend-herdr-presentation-e2e.test.sh, the concurrent secondmate recovery failed with "herdr presentation recovery could not acquire its session lock; refusing a concurrent resume". Two reasons it is not this PR: - The same failure, in the same test and case, happened on run 36655209015 for the unrelated branch fm/fm-contributions-old-gh-compat about 14 hours earlier. - This PR's change cannot lengthen how long the lock is held. The launch is written to a file and sent to the pane as `. launch.N.sh`, so the extra export changes neither the pane submit nor the lock hold time. The cause is a race that was already there: spawn_herdr_presentation_order_lock_acquire gives up after 5 seconds, and a concurrent real-Herdr recovery can hold the lock longer. Fixing that means changing the product's lock timeout, which is outside this PR. It should be tracked separately, and a rerun of the Herdr job is expected to pass. The only file changed is tests/fm-claude-trust.test.sh

* feat(bin): add opt-in config/wait-no-turns so a waiting worker spends no turns (#4859)

* fix(dod): drive no-mistakes with one foreground call, not a background poll

The brief told workers to background the drive call and poll `axi status`
because one call "routinely outlives what your harness lets a single
command run". That advice contradicts the tool it drives: `no-mistakes
axi run --help` documents `--wait` with an 8m default, existing precisely
"so an agent harness with a 10-minute tool cap gets a structured return
instead of an unbounded hang".

Following the old text, a worker could never idle - a backgrounded call
returns in milliseconds, so it does not wait at all - and each attempt
leaked a live timer that later fired as a paid wake. Tell workers to make
one foreground call, let it block, and repeat it when it returns on
elapsed wait rather than on a gate or outcome.

Also drops the generalisation that told workers on any unestablished
harness to assume a command cap and use the same shape, which exported
the defect to harnesses with no such cap.

* fix(bin): let a waiting worker spend no turns until it is answered

A worker waiting on a decision, a pipeline gate, CI, or a heavy-test slot
kept taking model turns: the brief told it to list its inbox at any natural
checkpoint, and six automatic senders nudged secondmates whatever their open
decisions.

- The ship and scout briefs gain one Waiting section: end the turn after
  needs-decision or blocked, and hold an external wait inside ONE blocking
  command bounded by the harness's own command ceiling. The checkpoint clause
  is deleted. Forbidding the wrong shapes is not enough on its own, so the
  section also names the blocking foreground `until` loop as the wait a Claude
  Code worker may use, because that harness can refuse a sleep-then-check
  command while pointing at backgrounding, which is the one shape a waiting
  worker must not take.
- fm-send --automatic defers (exit 4, nothing written or rung) while the
  target has an open decision or blocker of its own; every automatic sender
  passes it and keeps its retry state, and the pending-reply recovery waits
  the same way.
- The two senders that report the result classified it by matching the text of
  the send's captured output against `deferred:*`. fm-send runs bin/fm-guard.sh
  as a supervision warning, and that guard prints its worktree-tangle banner
  whenever the primary checkout is on a feature branch, which is exactly what a
  CI pull-request checkout is. The banner lands ahead of the `deferred:` line,
  so the match fell through and a waiting mate was reported as a failed send,
  with the banner as the reason. Both senders now classify on fm-send's exit
  status, which is the contract the deferral is actually stated in, and select
  the `deferred:` line out of the output rather than assuming it came first.

The third root cause, a no-mistakes definition of done that backgrounded the
drive call and polled axi status, is fixed by this branch's parent commit
"drive no-mistakes with one foreground call, not a background poll"; this
commit takes that text as is and adds the regression test.

Upstream's spawn abort path no longer calls the lease-return helper at all,
so the fork's missing-helper guard and its pin-feature test line are moot
here and are not ported.

The command ceilings each harness enforces, and the probes behind the named
Claude Code wait, are recorded in docs/verification/runtime-backends.md.

* no-mistakes(review): Exempt captain holds, quiet deferred reconcile, clarify worker pauses

* no-mistakes(document): Document deferred automatic nudges, rereads, and reply recovery

* no-mistakes(document): Ring unlanded fire-and-forget steers exactly once more

* no-mistakes(ci): The failing check, "PR must be raised via no-mistakes", reads the pipeline's attestation record, which says document=skipped. No file in the repository can change that record, so I did not touch the check or the PR body. As you said, the no-mistakes rerun after this run finishes will re-execute the document step and record document=completed. The one change is the documentation sentence you ordered. It adds a line to docs/remote-secondmates.md, right after the line saying the remote host runs no re-ring ladder of its own: "A fire-and-forget record, such as a reconcile ask, gets its single retry ring only on the local plane: the remote steer leg owes no re-ring, so a swallowed remote doorbell for one waits for the next ring into that inbox, and a remote-side retry is known follow-up scope." No behavior changed. Checks: tests/fm-documentation-audiences.test.sh passes (4/4) and bin/fm-lint.sh is clean. The change is left uncommitted in the working tree for the pipeline to pick up

* no-mistakes(review): Hold automatic wakes until a mate's own decision closes

* no-mistakes(document): Document watcher delivery of deferred remote re-read nudges

* no-mistakes(review): Merge duplicate elapsed-wait reattach instructions in DOD

* no-mistakes(test): Resolve merged default decision in remote-reply recovery fixture

* no-mistakes(test): Source classify lib so config-push retry-deferred honors open decisions

* no-mistakes(ci): Fixed a flaky test that also fails on main. Neither this PR's bin/fm-brief.sh nor its bin/fm-dod-lib.sh change is involved: bin/fm-dispatch-resolve.sh sources neither file. Another branch (fm-attended-cutover-smoothing-s1, run 36343879084) failed the same shard 8 check the same way, on a different case ("a rule-criterion match prints one diagnostic line, got 2"). Root cause: `fm_quota_single_provider_for_harness` in bin/fm-quota-axi-lib.sh returned from its `while read` loop as soon as it found a match. That closed the pipe while `fm_quota_single_provider_table`'s `printf` was sometimes still writing. GitHub Actions runners ignore SIGPIPE, so bash printed `fm-quota-axi-lib.sh: line 138: printf: write error: Broken pipe` to the resolver's stderr. That is the extra line. I reproduced it locally by running the test with SIGPIPE ignored: 2 of 20 runs failed, one with the resolver's diagnostic line plus two broken-pipe lines. Invariant: looking up a harness in the provider table must never make the table writer fail. The only reader of that table is this function, and all of the resolver's lookups (line 208 without stderr redirected, line 222 with it) go through it. So the fix is in that one place: read the whole table, then print the match. The same file now shows it reads the full table first, like `fm_control_harness_supported` does. Return values and output are unchanged. Verification: with SIGPIPE ignored, tests/fm-dispatch-resolve.test.sh failed 0 of 30 runs after the fix (2 of 20 before). tests/fm-dispatch-resolve.test.sh, tests/fm-brief.test.sh, tests/fm-send-inbox.test.sh, tests/fm-quota-choose.test.sh and tests/fm-quota-array-dispatch-live-e2e.test.sh all pass, and shellcheck is clean. tests/fm-procevent-quota.test.sh fails locally with or without the change ("process-event state root is not a private directory"), so that failure comes from the local environment, not from this fix. No new test was added: the existing one-diagnostic-line assertions already catch this whenever SIGPIPE is ignored, as it is in CI

* Revert "no-mistakes(ci): Fixed a flaky test that also fails on main. Neither this PR's bin/fm-brief.sh nor its bin/fm-dod-lib.sh change is involved: bin/fm-dispatch-resolve.sh sources neither file. Another branch (fm-attended-cutover-smoothing-s1, run 36343879084) failed the same shard 8 check the same way, on a different case ("a rule-criterion match prints one diagnostic line, got 2"). Root cause: `fm_quota_single_provider_for_harness` in bin/fm-quota-axi-lib.sh returned from its `while read` loop as soon as it found a match. That closed the pipe while `fm_quota_single_provider_table`'s `printf` was sometimes still writing. GitHub Actions runners ignore SIGPIPE, so bash printed `fm-quota-axi-lib.sh: line 138: printf: write error: Broken pipe` to the resolver's stderr. That is the extra line. I reproduced it locally by running the test with SIGPIPE ignored: 2 of 20 runs failed, one with the resolver's diagnostic line plus two broken-pipe lines. Invariant: looking up a harness in the provider table must never make the table writer fail. The only reader of that table is this function, and all of the resolver's lookups (line 208 without stderr redirected, line 222 with it) go through it. So the fix is in that one place: read the whole table, then print the match. The same file now shows it reads the full table first, like `fm_control_harness_supported` does. Return values and output are unchanged. Verification: with SIGPIPE ignored, tests/fm-dispatch-resolve.test.sh failed 0 of 30 runs after the fix (2 of 20 before). tests/fm-dispatch-resolve.test.sh, tests/fm-brief.test.sh, tests/fm-send-inbox.test.sh, tests/fm-quota-choose.test.sh and tests/fm-quota-array-dispatch-live-e2e.test.sh all pass, and shellcheck is clean. tests/fm-procevent-quota.test.sh fails locally with or without the change ("process-event state root is not a private directory"), so that failure comes from the local environment, not from this fix. No new test was added: the existing one-diagnostic-line assertions already catch this whenever SIGPIPE is ignored, as it is in CI"

This reverts commit c7199284297ea278c8da7d0a698cc5823ac37cec.

* no-mistakes(review): Retry deferred local instruction nudges via the watcher

* no-mistakes(review): Document watcher retry for deferred local instruction nudges

* no-mistakes(ci): I fixed both review findings you selected (ci-1 and ci-3). I did not touch the deferral check in bin/fm-send.sh. ci-1 (bin/fm-config-push.sh, retry_deferred_rereads) - Rule that must hold: a deferred reread stays flagged until it is actually delivered. - Before the fix, the flag was removed before any of the steps that can skip a mate: the remote lock-path lookup, validate_secondmate_home, the local lock-path lookup, and the lock acquire. A skip at any of those dropped the flag, so the watcher lost track of the reread. - Now the flag is removed in one place only, when the send succeeds (rc 0). A skipped home, a busy lock, a deferred send (rc 4) or a failed send all leave it in place. The re-mark calls on a busy lock and on rc 4 were no longer needed, so I removed them. I updated the comment above the function to match. - Side effect: a send that keeps failing now stays flagged, so the watcher retries it on every poll and logs each failure. That follows your "don't clear until delivered" rule, but it replaces the old behaviour of leaving a failed send to the next config push or session start. - New test in tests/fm-secondmate-sync.test.sh: T8j "a deferred flag survives a skipped invalid home and is retried once it validates". It takes the home's marker away to make validation fail, checks that nothing is sent and the flag stays, then puts the marker back and checks that the nudge is delivered and both the flag and the retry marker are cleared. It fails on the old code and passes now. ci-3 (bin/fm-secondmate-restart.sh) - Rule that must hold: no automatic send wakes a mate that is waiting on its own open decision. - The two automatic sends in this script are the fallback reread nudge (fall_back_to_nudge) and the persist request. Both now pass --automatic. If a persist request is deferred, its correlation is discarded and the mate goes to the fallback nudge, which is also deferred, so the mate is reported as unreached. - New test in tests/fm-secondmate-restart.test.sh: T3b. It gives a mate an open needs-decision and runs a restart. It checks that both sends report as deferred, the mate's doorbell is never rung, its inbox gets no message, nothing is stopped, and the mate is reported as unreached with exit status 3. It fails on the old code and passes now. - The test marks the watcher as alive first. Without that, the watcher-down warning is printed first and becomes the reported reason instead of the deferral message. Verification - tests/fm-secondmate-sync.test.sh passes. - tests/fm-secondmate-restart.test.sh passes. - tests/fm-secondmate-harness.test.sh (the other test that exercises --retry-deferred) passes. - The fm-send-inbox test that covers automatic deferral passes. I only looked at the last lines of that run, not the whole file. - `shellcheck -x` on the four changed files is clean

* Pin autoarm supervision model in secondmate restart T3b

The fresh watcher beat the test writes proves a live watcher only under the
autoarm model; on CI hosts with no detected harness the persistent model
demands a lock-holding watcher, so the watcher-down banner became the
reported reason and the deferral assertion failed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Keep deferred secondmate nudges retryable under the inheritance lock.

A bootstrap instruction nudge could write its deferral flag outside the lock the watcher retry holds, so a concurrent retry could delete a flag that had just been set. A restart fallback that is deferred now records the same marker and flag, so the watcher delivers it once the decision closes.

* no-mistakes(document): Document watcher retry of deferred restart re-read nudges

* Send secondmate reread and restart nudges immediately again.

Deferring those nudges let a later config push drop an incomplete transfer once the decision closed. They now send as they do on main.

* Make the no-turn wait opt-in behind config/wait-no-turns.

Homes that do not create the file keep the previous briefs, drive text, and sends.

* no-mistakes(document): Document wait-no-turns inbox wording change in configuration

* no-mistakes(review): Keep checkpoint inbox check; forbid only polling while waiting

* no-mistakes(ci): Fixed ci-2 (Greptile: a concurrent retry marker gets lost). The rule that was broken: the watcher may remove only the `.retry-ring` mark for the record it just processed. A newer mark written in the meantime is owed its own retry. `fm_task_inbox_clear_retry` is the one shared function that removes the mark, and I fixed it there. In `bin/fm-task-inbox-lib.sh` it now takes the record path. It compares the mark's content with that record's name and removes the mark only when they match. When the mark names a different record it returns success and leaves the mark alone. It still fails only when the processed record's own mark can't be removed. Both callers in `bin/fm-watch.sh` now pass `"$rec"`: the dead or missing pane path and the path after a retry ring. So the fix holds at both removal sites. Tests, in `tests/fm-task-inbox.test.sh`: - I added an optional `FM_RING_MARKS_RETRY` hook to the fake tmux. It writes a newer record's mark while the doorbell is being typed, which reproduces the race deterministically. - I added `test_watcher_retry_keeps_a_newer_mark`. The owed retry rings once, the newer mark survives, and a later check rings the newer record once and then clears its mark. The test fails without the fix ("the spent retry removed a newer record's mark written during its ring") and passes with it. - I updated the direct `clear_retry` call in the existing unit test to pass the record. Results: `tests/fm-task-inbox.test.sh` passes in full and `tests/fm-send-inbox.test.sh` passes 15/15. Shellcheck reports only SC1091 "not following sourced file" notices. As instructed, I didn't change the brief inbox wording

* no-mistakes(document): Fix stale wait-no-turns inbox wording in inbox lib comment

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

* fix(bin): close Gerrit-landed backlog items with the change URL as a note (#6140)

* fix(bin): record Gerrit change URLs as close notes

Teardown's backlog_done_args hands every ship's recorded pr= URL to
fm_backlog_done as --pr, and tasks-axi refuses any --pr that is not a
canonical GitHub or Forgejo pull request. A Gerrit change URL therefore
left the item In flight after cleanup, and the pending backlog-close
record replayed into the same refusal at every session start.

fm_backlog_done now rewrites a --pr whose value fm_pr_url_parse reads
as a Gerrit change into --note "Gerrit change <url>". The mapping sits
at the tasks-axi call rather than in the pending-close record, so
records already written with --pr replay to a close unchanged. The
captain-held retain path records the URL in its deliverable line and
skips the update --pr it cannot make.

* no-mistakes(review): Note retained Gerrit change URL when captain answers early

* no-mistakes(document): Document Gerrit change URL handling in captain-hold retention

* fix: reduce remote-job and supervision polling churn (#6255)

* perf(remote): separate active job sampling from dispatcher cadence

* no-mistakes(document): Link remote wait timing to its authoritative contract

* no-mistakes(ci): Fixed ci-1 with two narrowly scoped SC2030 annotations documenting intentional subshell-local legacy and active cadence overrides in tests/fm-remote-job.test.sh. Runtime behavior is unchanged. Reproduced the lint failure before the fix; afterward ShellCheck 0.11.0 with source following, Bash syntax validation, the complete remote-job behavior suite, and git diff --check all passed

* perf(supervision): reduce park, delta and dispatcher polling

* no-mistakes(document): Clarify poll latency contracts and authoritative documentation pointers

* fix(bin): load backend sibling libraries when sourced under zsh (#6221)

* fix(bin): load backend sibling libraries under zsh

fm_backend_source kept each backend's sibling list in one space-separated
string and iterated it unquoted. zsh does not word-split an unquoted
expansion, so the readability check saw the whole list as one path and
refused every backend with more than one sibling. Hold the list in the
function's positional parameters instead, which needs no word splitting
in Bash 3.2, Bash 5, or zsh.

The existing zsh case in tests/fm-backend.test.sh covers it wherever zsh
is installed.

* test: run the Calm mod suite on stock Bash 3.2

The suite injected shell values into its generated Node scripts with the
${value@Q} transformation, which needs Bash 4.4. Stock macOS Bash 3.2
reports a bad substitution, so every case failed before it asserted
anything. Build each JavaScript string literal with JSON.stringify
through a small helper instead, which works on any Bash and is a valid
literal for any value.

* no-mistakes(review): fix(bin): rename zsh-special path local in fm_backend_source

* test: narrow the zsh backend claim to name matching

Under zsh the adapters locate their siblings through BASH_SOURCE, so a
successful fm_backend_source is not a full load. Assert only what the
contract states, and pass js_string values after -- so node never reads
a leading-dash value as its own option.

---------

Co-authored-by: Nova Agent B <novaagentb@gmail.com>

* fix(bin): exclude a remote mate's own parent channel from self-home status scans (#5263)

* fix(bin): exclude a remote mate's own parent channel from self-home scans

A remote secondmate home's outbound parent channel lives at state/parent-replies.status inside its own state dir, so the watcher's signal scan enumerated it as a task status file and the open-decisions fold classified it as a phantom task named parent-replies: every parent-channel append spun a spurious signal wake and a phantom open decision in the mate's own home.
fm-parent-channel-lib.sh gains fm_parent_channel_outbound_status, which resolves the channel into the mate's own state dir for the remote route only, and fm-classify-lib.sh's status_scan_parent_channel_exclude wraps it for the fleet-wide scans.
The watcher's scan_signals and heartbeat fail-safe backstop, the whole-file and incremental open-decisions folds, the presentation snapshot, and the unread-surface scan now skip exactly that resolved path.
The exclusion is home-shape-aware: a parent-replies.status in a main home or a local mate is an ordinary task log and keeps waking and folding, and every other status file is untouched.

* no-mistakes(review): exclude a remote mate's parent channel from the daemon heartbeat scan

* no-mistakes(document): Document remote mate parent-channel scan exclusion

* ci: retrigger portable serial 4

* no-mistakes(ci): CI check 'Behavior portable serial 7' failed in tests/fm-contributions.test.sh ('reservation poll failed'). CI stderr showed bin/fm-contributions.sh:345 arithmetic 'DEADLINE - 6\n90077104: syntax error in expression': the fixture's fake date returned a torn two-line clock value. Root cause: the fake forge wrapper in wrap_forge advances the shared controllable clock via a non-atomic read-modify-write ('$(cat $FORGE/clock) + 6' with truncate-in-place '> $FORGE/clock') while concurrent background gh calls run and the fake date reads the same file; an interleaved truncate+write publishes a half-written value (CI's torn '6\n90077104', tail of 1790077104) or an emptied-read value ('6'), which either breaks the poll's arithmetic (nonzero exit -> 'reservation poll failed') or defeats the 15-second reservation defer. This is a pre-existing test-fixture race, not caused by the PR's diff (base..target touches no contributions code; the same commit passed this shard in run 35711207830 earlier the same day). Fixed the flaky fixture at its root: clock_bump() now writes each new value to a per-process mktemp file in the same directory and publishes it with mv (atomic rename), so concurrent forge callers and the fake date always read one complete old-or-new clock; fault patterns and deltas are unchanged. Verified: minimal 3-way concurrency repro shows the old wrapper corrupting (12/32/38 outcomes incl. empty-read) while the rename-based wrapper never corrupts (20/20 clean); the full tests/fm-contributions.test.sh passes twice (all 38 assertions ok, incl. the reservation, budget-exhaustion, genuine-failure, shared-once, and latency tests); 10 isolated reservation runs pass; shellcheck rc=0; worktree contains only this one-file change

* no-mistakes(document): drop stale file-set copy in daemon catch-all comment

* fix(bin): document accepted contribution verdict actors (#6307)

* fix(bin): name the accepted verdict actors in fm-contributions help and refusal

* fix(ci): Updated tests/fm-contributions.test.sh to assert exactly captain, fleet, maintainer, and nobody in command-emitted help and refusal output. Three focused regressions passed; all three extra-actor mutations were rejected. ShellCheck, syntax, and diff checks passed. Production code remains unchanged

* fix(bin): recognize clone roots across path spelling differences (#6306)

* fix(bin): recognise a clone root git names with different path spelling

fm-fleet-sync compared git's --show-toplevel with pwd -P as strings, so a clone
root that git recorded with different casing (case-insensitive volume) was
skipped as not a clone root and never refreshed. Compare filesystem identity
instead, which also covers symlink spelling.

* fix(document): Remove stale clone-root comparison comment

* test: preserve Pi calm transcript captures with Pi 1.0 (#6338)

* test(calm): pin Pi's regular TUI mode where pane assertions read scrollback

Pi 1.0.0 defaults its TUI to a fullscreen alternate-screen mode whose
scrollable transcript is application-owned, so rows that leave the viewport
never enter terminal scrollback and tmux capture-pane -S can no longer see
them. The Pi Calm e2e launches now pass --tui-mode regular wherever the flag
exists so the transcript assertions keep reading real scrollback on both the
Pi 1.0.0 line and earlier Pi lines, which have no such flag and render
regular-only anyway.

* no-mistakes(document): Correct Pi TUI documentation and scrollback rationale

* fix(bin): preserve hold reasons and reject invalid completion inventories (#6331)

* fix(bin): encode captain-hold reasons and reject self-inventory in complete

hold now stores a reason with parentheses, line breaks, or percent signs
through a reversible percent encoding that every reader decodes, instead of
refusing it. hold --origin records the origin on the held task, and complete
refuses the origin as its own inventory entry and an entry held for a
different origin; holds with no recorded origin are accepted and flagged.

* fix(review): Decode marked hold reasons consistently across readers

* fix(review): Remove unnecessary lifecycle test dispatch

* fix(review): Correct hold origin identity and inventory recovery

* fix(review): Record origins before placing backend holds

* fix(document): Clarify captain-hold validation and reason reader documentation

* fix(ci): Fixed both findings: failed backend holds restore the previous origin, and invalid base64/UTF-8 reasons remain verbatim. Added regressions and documented valid-literal ambiguity. Both failures were reproduced before fixes. Verification: 54 lifecycle tests and 9 wrapper tests passed; 7 Beads-specific cases skipped because tasks-axi is markdown-only. Focused lint and diff checks passed. No pipeline or publication actions performed

* fix: reclaim orphaned watcher arms on the next park (#6335)

* fix(bin): take over the watcher cycle a main-only pass-through leaves

An attended main-only pass-through leaves a successor watcher cycle
running through main's handling turn. The session's next park attached
to that cycle instead of owning it, so the successor's arm, orphaned by
its host's exit, kept owning the watcher while the new park's arm polled
it twice a second until the next close or the park boundary, hours later
in a quiet second mate. A second-mate restart hit this every time, since
its persist request is a main-only close.

The host now records the successor it leaves for main, and the next
host's first cycle runs bin/fm-watch-arm.sh --take-over on it: when that
arm still owns the healthy watcher, the new arm stops it, reports a
reason the cycle delivered first, and otherwise owns a fresh cycle. The
stop's own downtime publication is undone over an acknowledged episode
when no wake was appended in between, so the handover wakes nobody.

* no-mistakes(review): Keep left-arm record until the orphaned arm is gone

* no-mistakes(review): Relinquish successor arm only after durably recording it

* no-mistakes(review): Relinquish successor only after its record reads back

* no-mistakes(document): Clarify successor takeover guarantees and authoritative documentation

* no-mistakes(ci): Fixed ci-1: acknowledgement restore now requires the exact taken-over arm/watcher ledger row with signal=TERM, awaited within a short bound. Otherwise takeover proceeds without erasing downtime. Added a self-exit regression confirmed failing before the fix and passing afterward; ordinary takeover tests and the full watcher-arm suite pass. Updated Generation reuse documentation. Syntax, diff checks, and ShellCheck pass with existing SC1091/SC2034 warnings excluded. ci-2 remains unchanged

* no-mistakes(document): Clarify watcher take-over recovery and restart limits

* fix(bin): restore downtime on supervision-host hand-back when the successor already closed (#6355)

* fix(bin): restore supervision host hand-back continuity

* no-mistakes(review): Scope host hand-back failure fallback to lost pending:handling

* no-mistakes(review): Remove stray scratch test copy tests/.tmp-rest.test.sh

* no-mistakes(test): Initialise successor globals so early hand-back survives set -u

* no-mistakes(document): Document host hand-back downtime failure and Claude lost-handback notice

* no-mistakes(ci): I fixed the Greptile finding. The rule that must hold: when the supervision host hands back an actionable wake, its rewake is refused, and no watcher is healthy, the hand-back still has to reach main as a delivered notice. That must be true whether the recovery marker is `pending:handling` or `announced:handling`. Only one place applies this check: the lost hand-back fallback in `bin/fm-claude-stop-autoarm.sh`. **Fix:** that check now accepts both `pending:handling:*` and `announced:handling:*` tokens (a one-line change). Nothing else in the fallback changed: - Refusals on any other marker, such as an already acknowledged one, still exit 0 silently and open no failure episode. - The notice is still sent once per episode, and repeats are recorded as `failed-suppressed`. **Tests:** - `tests/fm-claude-stop-autoarm.test.sh`: the lost hand-back test now runs as a shared helper with two variants, one writing a `pending:handling` marker and a new one writing `announced:handling` (`test_host_lost_announced_handback_notifies_once_per_episode`). - `tests/fm-supervision-host.test.sh`: the end-to-end test where downtime restoration fails is now a shared helper too, with a new `announced` variant (`test_claude_stop_hook_notifies_when_closed_announced_successor_downtime_restore_fails`). It moves the handling episode to `announced` before the host hands back. Without the fix the hook would exit 0 here; the test requires exit 2, `outcome=failed` and a delivered failure notice. **Verification (all under nice -n 10):** - The full `tests/fm-claude-stop-autoarm.test.sh` suite passed (rc=0), including both lost hand-back variants and the benign-refusal test. - In `tests/fm-supervision-host.test.sh` I ran only the four hand-back test functions, all passing (rc=0). The suite can't run single functions, so I used a temporary copy with a trimmed test list and deleted it afterwards; `git status` shows only the 3 intended files changed. - shellcheck is clean on all three changed files. I did not run `bin/fm-lint.sh`. - I did not run the new tests against the unfixed code; the claim that they fail without the fix comes from reading the old check

* fix: reduce remote-job polling process churn (#6363)

* perf: cut remote-job idle process creation in the three hot loops

Post-update host measurement still attributes most idle churn to three
per-sample loops: result-consumer state reads and date calls, the delta
reader's capture/hash pass on every poll, and the lane preemption scan's
per-field pipelines. This drops each to its minimum without touching the
contracts around them.

* fm_remote_job_read_state gains an optional result-variable form backed
  by fm_remote_job_read_line, a builtin-only bounded record read (regular
  non-symlink file, byte bound, one newline-terminated line, tolerated
  unterminated tail, no carriage returns). fm_remote_job_wait samples
  state and the SECONDS clock with no per-sample children; one date call
  converts the epoch deadline once.
* fm-remote-delta-read stats the log each poll and re-runs the bounded
  capture and hashing only when size, mtime, ctime, inode, or device
  change. The snapshot's own stat writes the comparison key, so a log
  that moves between the gate and the capture is never read as stable.
* worker_preempting_waiter_exists reads state, home, and the staged argv
  head with builtins only. The now-unused worker_job_command goes away.

The bounded reads use -d '' -n, which behaves identically on the macOS
stock bash 3.2 and current bash; -N does not exist on 3.2. Tests cover
the malformed-record corpus, delta identity gating, fork-free lane
scanning through counting PATH shims, and same-home versus cross-home
preemption. No signal traps or sleep contracts change.

* no-mistakes(review): Restore subsecond delta keys and byte-bounded builtin record reads

* no-mistakes(document): Clarify delta snapshot caching and coarse-timestamp fallback

* no-mistakes(lint): Scope UTF-8 regression locales to individual function calls

* no-mistakes(ci): Fixed both lint failures by applying the documented production-library analysis boundary at the two affected test imports. Runtime behavior is unchanged; the library remains independently linted. Canonical full-analysis lint passed for the library and both suites, as did bash syntax checks and git diff --check

* fix(bin): recognize titled Claude top rules and preserve grey slash commands (#5963)

* fix(composer): read a titled Claude top rule as the composer's edge

A named Claude Code session draws its title into the composer's top rule.
The strict separator predicate rejected that row, so the closing rule read
as a lower unmatched separator and an idle, empty composer classified
unknown on every cursorless backend, refusing fm-send, exit, and relaunch.

Spare a bare agent-glyph row sandwiched between a width-proven titled rule
and the screen's only unmatched separator directly below it. The strict
separator predicate, dead-shell rule, and blank-row posture are unchanged.

Fixes #5601
Fixes #5558

* no-mistakes(test): Keep Claude's grey slash command in Herdr payload proof

* no-mistakes(test): Make missing-herdr version check hermetic to installed herdr

* fix: prevent Pi trust prompts in seeded secondmate homes (#6387)

* fix: pre-approve Pi trust for seeded secondmate homes

Unattended first launches of Firstmate-seeded Pi secondmate homes stalled on
"Trust project folder?" until Enter. Probe --approve like --tui-mode and pass
it only for --secondmate when help advertises it (.fm-secondmate-home signal),
leaving ordinary workers and older Pi unchanged.

* no-mistakes(document): Consolidate Pi seeded-home trust documentation ownership

* no-mistakes(ci): Fixed Lint 1’s unused polling counter. Diagnosed Behavior portable serial 4 as a pre-existing delta-reader test clock race; replaced timing-dependent rewrite and deletion with deterministic executable-boundary synchronization. ShellCheck, Bash syntax checks, and git diff --check passed. Delta-reader tests passed three consecutive runs; all three live Pi trust cases passed. Production behavior unchanged

* fix(bin): retry ShellCheck roots that hit the memory ceiling without --external-sources (#6443)

* fix(lint): retry memory-bound roots without external sources

* no-mistakes(review): Make fallback tests portable and correct source-following telemetry

* no-mistakes(review): Remove committed parity fixtures and use disposable test roots

* no-mistakes(document): Document ShellCheck memory fallback and telemetry

* no-mistakes(review): Cover bounded and unbounded fallback RSS behavior

* no-mistakes(document): Correct stale lint fallback documentation

* no-mistakes(document): Correct stale lint test documentation

* no-mistakes(ci): The memory fallback (the retry without --external-sources) now gets only the time left in its root's original deadline, so it can no longer outlast the CI job. Invariant: one root's first attempt plus its fallback must fit inside a single FM_LINT_ROOT_SECONDS deadline, plus the cleanup grace. Only one site started a new deadline: the fallback call in fm_lint_run_root. The deadline is the only budget involved, because the memory limit already applies to each process separately. Changes in bin/fm-lint.sh: - fm_lint_exec_root now takes a <seconds> argument instead of always reading FM_LINT_INTERNAL_ROOT_SECS. - The first attempt passes the full deadline. - The fallback passes floor((start + deadline - now) / 1000) seconds. - When bounds are enforced and less than 1 second is left, no retry starts. fm_exec_timed rejects 0 seconds, so the retry cannot run with no time. The root keeps reason=memory, and the shard output says "no time left in its Ns deadline to retry without it". - Unbounded local runs have no deadline and behave as before. - The header comment now describes the shared deadline. Changes in tests/fm-lint.test.sh: a new test, test_memory_fallback_spends_only_the_remaining_root_deadline, runs only on hosts that can enforce bounds. It uses a 6 s deadline and 1 s grace. - Case 1: the first attempt runs 3 s and then fails with memory status 251. The test asserts one fallback ran, reported reason=timeout, and the root's recorded duration is under 7000 ms. - Case 2: the first attempt runs 5.2 s. The test asserts no fallback starts, the skip is explained, and the sidecar records memory with source-following 1. Verification: - Full `nice -n 10 bash tests/fm-lint.test.sh` passed, including the new test, in about 5 minutes. - Case 1 run against the HEAD script: the root took 9168 ms, so the under-7000 ms check fails before the fix. - `bin/fm-lint.sh bin/fm-lint.sh tests/fm-lint.test.sh` reported no findings. - The CI workflow is unchanged, so the Test step still runs only tests/fm-lint.test.sh with nice -n 10 and the 12 GiB ShellCheck limit

* fix(pi): restore watcher continuity across successor gaps and make extension log opt-in (#5489)

* Fix Pi watcher successor-gap confirmations and add extension log

Accept an already-acknowledged handling confirmation as a no-op when the
generation matches, confirm the restoration's own recovery token with a
superseded (not rejected) outcome on generation mismatch, retire an arm on
confirm failure only when the failed token names that exact pid, and record
restore attempts, readiness timeouts, and confirm results in the bounded
state/.watch-extension.log. Regression tests: already-acked no-op plus
mismatch/dead-pid/lock-mismatch rejections and the manual-restart churn
contract in fm-watch-arm.test.sh, and a mid-restore marker advance with no
rejection appendix in fm-pi-watch-extension.test.sh.

* Treat a dead arm child as an empty slot so repair and retry recover

startArm and scheduleRetry answered unchanged while holding a ChildProcess
whose OS process was already gone but whose close had not fired, so neither
the repair tool nor the retry timer started anything until that close fired.
Gate slot occupancy on a liveness check (exit/signal codes plus pid probe)
and start a fresh arm instead, with a regression test driving the repair
tool against a dead-but-unclosed child.

* no-mistakes(document): Document new Pi extension log knob

* no-mistakes(review): Fix confirm-failure retire token match, add distinct-pid test

* no-mistakes(document): Clarify retire guard needs pid and generation

* Make the Pi extension diagnostic log opt-in and default-off

Only a positive FM_WATCH_EXTENSION_LOG_KEEP_LINES enables
state/.watch-extension.log. Unset, empty, non-numeric, zero, and
negative values disable logging entirely, so the default run writes
nothing and never creates the file. The shared positiveInteger
fallback semantics stay untouched for the retry and timeout knobs.
docs/configuration.md owns the knob contract and
docs/watcher-continuity.md points at it. Tests: the superseded-delivery
case runs opted in, and a new case proves unset, zero, and non-numeric
values create no log file while delivery still succeeds.

* no-mistakes(document): Qualify extension-log coverage bullet as opt-in

* no-mistakes(ci): The two reported checks (CI run 36372002913, Require no-mistakes run 36372069208) show conclusion action_required with 0 jobs and no logs because this is a fork PR (RibatTRW/firstmate) and GitHub is holding the workflow runs awaiting maintainer approval; that gate is external to the code and needs a maintainer to approve the runs. While verifying the change locally I found a real defect the approved CI run would hit: the PR's new test test_handling_delivered_rejects_a_superseded_generation failed deterministically. Invariant violated: reopen-announced (a non-successor/manual arm start) mints a fresh recovery generation only when the durable wake queue holds unrecovered work; an announced episode with an empty queue must be left untouched so idle arm starts never churn generations (the [ -s queue ] guard from #4819, relied on by bin/fm-watch.sh:2413 and covered by the append-reopens and announcement-bound sibling tests). The test called reopen with an empty queue and expected churn, so the fix establishes the queued-work precondition (append one wake, re-announce, re-read the generation) before asserting the reopen mints and the old confirmation mismatches. Test-only change, 14 lines in tests/fm-watch-arm.test.sh. Verified: fm-watch-arm.test.sh 25/25 ok on two consecutive runs, fm-pi-watch-extension.test.sh 55/55 ok, and shellcheck reports only one pre-existing warning outside the edited region

* no-mistakes(document): Restore blank line in watcher-continuity docs

* Route superseded Pi deliveries like confirmed ones and cover the retire guard

A superseded handling confirmation now falls through to the normal delivery
path, so an accepting supervision branch owns the wake instead of main.
Scope the watcher-continuity token-pinned confirmation and narrowed retire
rule to Pi, since omp and OpenCode still confirm against the current
successor. Add tests that fail when the retire guard, the scheduled-retry
gate, or the deferred-close gate is reverted, relabel the churned-generation
characterization test, and use a reaped pid for the dead-pid rejection.

* fix(pi): hide duplicate assistant finals from hidden processing retries (#5863)

* fix(pi): silence unacknowledged processing retry replies

Suppress autonomous processing prose before persistence and during streaming while retaining tool calls, signed reasoning, and retryable outcomes. Restore ordinary output after acknowledgement or a user message.

Fixes #4954

* no-mistakes(review): Silence only processing retries, keep first presentation visible

* fix(pi): preserve differing processing retry replies

* test(pi): accept Pi 1.0.1's renamed HTML export renderer lookup (#6530)

Pi 1.0.1's createToolHtmlRenderer reads getToolRenderers and ignores
getToolDefinition. Calm /export still includes stock grep HTML; the
fixture has to pass the lookup key the installed Pi actually reads.

* fix(bin): tolerate transient quota read failures (#6490)

* fix(procevent-quota): tolerate consecutive slow quota-axi reads

The quota allowance poll treated any quota_json failure as terminal, so one
slow quota-axi --json (measured max ~29s under a 48s derived bound) shut the
watch down until someone re-armed it, and the detail always said
"missing/incompatible". Tolerate three consecutive failed or timed-out reads
before going terminal, reset the streak on any good read, and report a
timeout distinctly from a missing or incompatible tool. Each timed poll runs
exactly one bounded --version and one bounded --json: validate the captured
version text through fm_quota_axi_version_compatible rather than launching a
second probe, and describe a mixed failure streak by count plus last cause.

* no-mistakes(document): Document quota polling failure tolerance

* no-mistakes(ci): Fixed ci-2 and ci-3. Permanent quota read failures (rc 2 missing, rc 3 incompatible) now report on the first poll, while transient rc 1/4 failures retain the existing three-failure retry behavior. The missing-binary test now uses an isolated PATH without quota-axi and asserts both permanent failures stop at condition_polls: 1. Verification passed: tests/fm-procevent-quota.test.sh, canonical fast lint for both changed files, bash syntax checks, and git diff --check

* no-mistakes(review): Classify untimed quota version failures as transient

* no-mistakes(document): Clarify quota polling failure budget

* fix(bin): clarify scratch guidance and dirty teardown refusals (#6505)

* fix(teardown): clarify scratch guidance and dirty worktree refusals

Keep ship proof material outside the task worktree and distinguish untracked-only leftovers from tracked edits without changing cleanup guards.

Fixes #6319

* fix(ci): Fixed ci-3 only. Both promotion outputs now replace the scout restriction and require external scratch storage and a clean worktree before done. Verification: 27 delivery tests passed, five mutations caught, restored test passed, pinned ShellCheck and syntax/whitespace checks passed. ci-1, ci-2, and ci-4 remain untouched

* fix(bin): escalate inbox instructions blocked by busy workers (#6518)

* Escalate inbox instructions stuck behind a busy worker

Count consecutive busy-deferred due doorbells durably and escalate at the configured bound without typing into the worker pane.

Fixes #6445

* fix(review): Fix inbox escalation deduplication and busy streak resets

* fix(review): Preserve busy inbox escalations through daemon supervision

* fix(document): Correct busy-inbox escalation documentation

* fix(ci): Fixed SC2034 in tests/fm-task-inbox.test.sh by including the loop counter in the failure diagnostic. Source-aware lint, all 34 inbox tests, and git diff --check pass. Behavior portable serial 6 reproduces identically on base 1f3e7696 and target 78156b86 with Pi 1.0.1: an unrelated renderer API change breaks the unchanged Calm test. No Calm changes made; that failure is addressed separately by https://github.com/kunchenguid/firstmate/pull/6516. Logs retained in scratchpad-ci/

* fix(ci): Fixed ci-2, ci-3, and ci-4: successor failures surface, reset alerts deduplicate, and oversized busy limits fall back to two. Passed 47 inbox tests, 7 focused daemon checks, all 13 mutation checks, lint, documentation checks, and diff checks. Evidence: scratchpad-ci-selected/summary.json. ci-1 remains unchanged and unwaived. Fresh live Herdr proof remains with the outer driver

* fix: prevent unnecessary remote worker turnover (#6431)

* fix: prevent healthy remote job worker turnover

* no-mistakes(review): Serialize full LaunchAgent repair and verify launchd-tracked owners

* no-mistakes(review): Let launchd-tracked unpublished spawns start before reloading

* no-mistakes(document): Document remote worker heartbeat and serialized LaunchAgent recovery

* no-mistakes(ci): Full CI log showed the idle-worker regression exceeded its outdated command budget (82 versus 80) after independent heartbeat ownership checks were added. Raised the budget to 120 while retaining the separate busy-poll sleep limit. Remote-job and LaunchAgent executable tests passed, as did bash syntax validation and git diff --check. No production behavior changed

* no-mistakes(ci): Fixed missing readiness recovery under verified live ownership, preserving the serving PID and private file mode. Added executable regressions for deletion during a blocked sweep and stale readiness diagnostics without LaunchAgent reload. Deletion regression failed before the fix. Both remote-job test suites, bash syntax validation, and git diff --check passed

* no-mistakes(ci): Full CI log identified a flaky ownership-loss test racing an already-authorized heartbeat refresh. Replaced backdating and a fixed sleep with bounded observation of readiness expiry through the public probe. Production behavior unchanged. Remote-job and LaunchAgent executable suites passed; bash syntax validation and git diff --check passed

* fix: wait for launchd bootout cleanup

* no-mistakes(document): Document remote worker recovery and read-only turnover verification

* no-mistakes(review): Publish worker identity before lock owner records

* no-mistakes(document): Document worker identity publication safety invariant

* no-mistakes(ci): Fixed ci-1: replacement workers discard predecessor readiness before publishing identity and roll back identity if lock-owner recording fails. Added executable regressions reproducing both defects. LaunchAgent and orphan-reap suites, shell syntax checks, and git diff --check passed. No live service state was modified

* no-mistakes(ci): Restored lock-owner-before-identity publication and removed the identity rollback and reordering-only tests. Retained an executable regression proving predecessor readiness is rejected until replacement startup completes. LaunchAgent and orphan-reap suites, shell syntax checks, and git diff --check passed. Other changes remain intact; the retained-identity interrupted-repair edge remains out of scope. No live service state was modified

* fix: speed up remote-job sequence claim cleanup (#6575)

* fix(remote-job): reap expired seq claims with one directory walk

The hourly claim sweep forked uname+stat per .seq-claims entry and blocked
serving for ~85s at ~17k dirs. Delete expired empty claim dirs with a single
find -exec rmdir batch and cache the host uname for remaining mtime reads.

* no-mistakes(review): Restore original path mtime helper and drop uname cache

* no-mistakes(test): Restore claim retention eligibility; focused retention and serving tests pass

* no-mistakes(document): Document single-walk claim cleanup and regression entrypoints

* no-mistakes(ci): Fixed Lint 2’s reproduced SC1091 by adding the tests/lib.sh ShellCheck source directive to the retention test. Runtime behavior is unchanged. ShellCheck passed for both new claim tests; Bash syntax, the retention behavior test, and git diff --check passed

* fix(tests): keep fixture registries out of git worktree roots

A TMPDIR pointed at a repository root placed live .fm-test-* registries
beside tracked files, and a concurrent git add during the claim-walk CI
fix round committed three of them. Route registries and fixture roots
through a TMPDIR that refuses git worktree roots, remove the stray files,
and pin the escape with a behavioral cleanup test.

* no-mistakes(review): Preserve whole-second claim expiry in single-walk sweep

* no-mistakes(document): Correct temporary-directory resolution documentation

* no-mistakes(ci): Fixed ci-3 by changing only the stale, fresh, and read-only orphan fixture paths in tests/fm-test-fixture-cleanup.test.sh to use $FM_TEST_TMPDIR. All seven tests passed both normally and with TMPDIR set to the worktree root. Shell syntax and git diff --check passed

* feat(bin): record no-mistakes pipeline spend per task at teardown, opt-in (#5354)

* feat(bin): record each task's no-mistakes pipeline spend at cleanup

no-mistakes keeps every pipeline agent invocation's token usage only in its
local agent_invocations records, and cold pipeline agents (review, test,
document) leave no session log, so a task's review-loop cost never reached
Firstmate's records and could not be attributed after cleanup.

bin/fm-pipeline-spend.sh attributes a task's runs by the repository
no-mistakes resolves for the task copy, the task branch, and the branch's
creation (a relaunch mints a new spawn_gen while the branch keeps
validating), then sums every invocation, failed and cancelled included.
Token fields use no-mistakes' per-round deltas so resumed review rounds are
not counted twice, and unrecorded values stay unknown rather than zero.
show prints the record read-only; record appends it once per task
incarnation to data/pipeline-spend.jsonl. Teardown records it for every
ship task it cleans up, before deleting the branch and the task record.

fm_nm_state_db becomes the one owner of where no-mistakes' state database
lives, shared with the capped run-inventory reader.

* no-mistakes(review): Drop spend show command and timeout override

* style(bin): rewrap fm-pipeline-spend header comment

* Make pipeline spend recording opt-in

* feat(bin): record each task's no-mistakes pipeline spend at cleanup

no-mistakes keeps every pipeline agent invocation's token usage only in its
local agent_invocations records, and cold pipeline agents (review, test,
document) leave no session log, so a task's review-loop cost never reached
Firstmate's records and could not be attributed after cleanup.

bin/fm-pipeline-spend.sh attributes a task's runs by the repository
no-mistakes resolves for the task copy, the task branch, and the branch's
creation (a relaunch mints a new spawn_gen while the branch keeps
validating), then sums every invocation, failed and cancelled included.
Token fields use no-mistakes' per-round deltas so resumed review rounds are
not counted twice, and unrecorded values stay unknown rather than zero.
show prints the record read-only; record appends it once per task
incarnation to data/pipeline-spend.jsonl. Teardown records it for every
ship task it cleans up, before deleting the branch and the task record.

fm_nm_state_db becomes the one owner of where no-mistakes' state database
lives, shared with the capped run-inventory reader.

* no-mistakes(review): Drop spend show command and timeout override

* style(bin): rewrap fm-pipeline-spend header comment

* Make pipeline spend recording opt-in

* no-mistakes(document): Note pipeline-spend opt-in gate in teardown comments

* no-mistakes(ci): I fixed all three Greptile findings following your decision. Both test suites pass: tests/fm-teardown.test.sh (105 ok, exit 0) and tests/fm-pipeline-spend.test.sh (8 ok). I didn't run either new test against the old code, so the claim that they fail before the fix is from reading the code, not a run. Nothing was committed or pushed. **ci-1 / ci-2 (bin/fm-teardown.sh):** The rule is that every owned ship task in a home that has opted in gets a durable spend record before its task record is deleted. The `[ -d "$WT" ]` check skipped the recorder when the worktree folder was already gone, so no record was written. I removed that check and kept the other conditions: the task must be a ship task, teardown must own its worktree, and `config/pipeline-spend` must exist. `bin/fm-pipeline-spend.sh` already writes an unavailable-source line when the worktree is missing. The new test `test_teardown_records_unavailable_spend_for_a_gone_worktree` in tests/fm-teardown.test.sh sets up an owned ship task whose worktree is missing, with recording turned on. It checks that teardown succeeds, writes a line with `source=="unavailable"`, `total==null` and a reason saying the task copy is gone, and still removes the task record. Under the old check no ledger would have been written. **ci-3 (bin/fm-nm-run-lib.sh):** When `NM_HOME` and `HOME` were both unset, `fm_nm_state_db` fell back to `${HOME:-}/.no-mistakes`, which resolves to `/.no-mistakes`. It now uses `root=~/.no-mistakes`. Bash expands `~` to the account's home directory even when `HOME` is unset, which matches the previous Python `Path.home()` lookup. The new test `test_state_db_without_nm_home_or_home_uses_the_account_home` in tests/fm-pipeline-spend.test.sh runs the function with both variables unset and compares the result to the home directory from the system's password database. The old code would have returned `/.no-mistakes/state.sqlite`. shellcheck reports nothing new on the changed files. I added one `disable=SC2016` comment in the new test, where the single-quoted `$1` is meant to expand in the child shell

* feat(bin): start tasks from a named base branch (#6442)

* feat(bin): start tasks from a named base branch

Spawns always reset a task's pooled copy to origin's default branch,
so work that belongs on a feature, integration, or release branch
started from the wrong code and opened its PR against the default.

fm-brief.sh --base-branch records the base in the brief, fm-spawn.sh
resets the copy to origin/<base> and records base_branch= in task
meta, and the worker targets its PR at that branch. Review diffs,
the cleanup content check, and scout promotion read the recorded
base. local-only and Gerrit deliveries refuse a named base.

* no-mistakes(review): Anchor brief base-branch parse on scaffold Setup line

* no-mistakes(review): Require explicit --base-branch on spawn, matching brief lines

* no-mistakes(document): Note base-branch reset in fm-spawn freshness header

* no-mistakes(ci): I fixed all four Greptile findings the way you specified. The affected suites pass locally: fm-brief, fm-dod-lib, fm-spawn-pool-base-freshen, fm-review-diff, fm-teardown and fm-task-delivery. shellcheck on the changed files reports only info-level notes, no warnings or errors. I didn't revert the code to watch each new test fail before the fix. - **ci-1 (brief text blocked or redirected spawns):** In bin/fm-dod-lib.sh, `fm_brief_base_branches` now only accepts a `Base branch: <value>` line that comes directly after the base-variant Setup sentence fm-brief.sh writes. Any other `Base branch:` line is ignored. `--base-branch` is still the only authority, and the existing checks in fm-spawn.sh are unchanged. With the flag, the spawn is refused unless a matching pair exists and no pair names a different branch. Without the flag, it is refused only if such a pair exists. The header comment in fm-spawn.sh is updated to match. - The `brief_with_base` test helper now writes the real two-line pair. - The decoy refusal case now uses a full decoy pair in the captain's intent. - New test `test_prose_base_branch_line_is_ignored`: a brief with no base whose intent quotes `Base branch: release/1.2` spawns normally with no `base_branch=` recorded. A spawn with `--base-branch feature/hub` whose intent names release/1.2 starts from `origin/feature/hub` and records it. - **ci-2 (scout base not checked against the forge):** fm-spawn.sh now looks up the project's registered forge with `bin/fm-project-mode.sh --forge`, like the ship path does, before `fm_base_branch_valid` runs. Previously it passed `none`. New test `test_scout_base_branch_refused_on_gerrit_forge`: a scout with a base on a forge=gerrit project is refused at spawn and no meta file is written. - **ci-3 (base not shell-quoted in worker commands):** `fm_dod_block` now quotes the base with `printf %q` in the `--base` and `--base-branch` arguments, and the branch name in the prose stays readable. The test in fm-brief.test.sh uses the base `release/$HOTFIX` and checks that both direct-PR and no-mistakes briefs render `release/\$HOTFIX` in the commands. - **ci-4 (ambiguous PR sentence):** The direct-PR sentence now reads "open a PR with `gh-axi` that is ready for review, not a draft, against the base branch `X` (`--base X`), not the repository default." A brief with no base renders exactly as before. The existing assertion in fm-brief.test.sh is updated

* no-mistakes(ci): Both real failures on PR 6442 (run 37064756768) come from CI tooling and a Pi version bump, not from the base-branch feature. Fixes for both already exist on upstream main, so I applied those two commits to the working tree with `git cherry-pick --no-commit`. Nothing is committed or pushed, and no feature file changed. - **Lint 2:** ShellCheck ran out of memory on `bin/fm-teardown.sh` (rss about 8 GB, exit 251, "shellcheck: out of memory"). The rule that broke is that the lint gate must finish on any valid shell root without exceeding the runner's memory ceiling. Upstream commit d719ef3 (#6443) fixes this in `bin/fm-lint.sh`: a root that hits the memory ceiling is retried without `--external-sources`. It also updates `tests/fm-lint.test.sh` and `docs/fm-test-portable-shards.md`. - **Behavior portable serial 6:** `tests/fm-calm-pi-extension.test.sh` failed with "grep disappeared from /export calm.html HTML while calm mode was on". Pi 1.0.1 renamed its HTML export renderer lookup, and upstream commit fede619 (#6530) updates the test to accept the new name. This change is test-only. Both commits applied cleanly. The four touched files are `bin/fm-lint.sh`, `docs/fm-test-portable-shards.md`, `tests/fm-lint.test.sh` and `tests/fm-calm-pi-extension.test.sh`. None of them is in the feature diff, so the feature contract is unchanged. **Local checks:** - `tests/fm-lint.test.sh` passes. - `bin/fm-lint.sh` on `bin/fm-teardown.sh`, `bin/fm-spawn.sh` and `bin/fm-dod-lib.sh` exits 0 with full ShellCheck analysis. - `tests/fm-calm-pi-extension.test.sh` exits 0 with 7 ok and 0 not ok. One case skipped because the Pi package isn't installed locally, so the Pi 1.0.1 export path that failed in CI couldn't be reproduced here. CI needs to confirm it. - I didn't reproduce the 8 GB OOM locally. The PR has to be updated through the pipeline with a normal push: no force push, no replacement PR, no merge

* fix(bin): name the Firstmate skill file as a fallback in the worker role (#6647)

* fix(bin): point project workers at the Firstmate skill file

The Skill tool cannot resolve a Firstmate skill from another project's worktree, so the launch role names the readable skill file instead.

* no-mistakes(review): fix(bin): name Firstmate skill file as fallback only

* feat(bin): retire a contribution whose forge object is permanently gone (#6655)

* feat(bin): retire a contribution whose forge object is permanently gone

Add fm-contributions.sh retire <task> <url> <captain|fleet> <reason>,
which records actor, reason and time on the saved record and removes
that task/url pair from known, poll rotation and coverage even while a
backlog link remains. Repeating a retire keeps the first provenance;
an unrecorded pair, unknown actor, empty reason or unacknowledged
pending signal is refused.

* no-mistakes(review): Keep retirement per task when settling final owners

* no-mistakes(ci): Made the three Greptile fixes the captain chose. ci-1, retirement needs the captain's word: `retire` now accepts only the actor `captain`. Any other actor is refused, including `fleet`. I removed `fleet` from the usage line, the script header (`bin/fm-contributions.sh`), the check in `bin/fm-contributions.sh` and the retired-record validation in `bin/fm-contributions.jq`, which now requires `.actor == "captain"`. The script header now says a retirement always records the captain's word, because the script cannot verify who runs it and the authority to retire is the captain's. The Bearings skill line in `.agents/skills/bearings/SKILL.md` now says `retire` records the captain's word. In the tests, a `fleet` retire is now a refusal case and must leave the record unretired. Every other retire call in the tests uses `captain`. The PR description has not been changed yet: the ruling asks for the same captain's-word statement there, and that belongs to the PR phase. ci-2, blank reasons: the reason check is now `[ -n "${5//[[:space:]]/}" ]`, so a reason made only of spaces or tabs is refused. I added a refusal test that passes a space-tab-space reason. The header also lists a blank reason among the refusals. ci-3, the gone-object test: the test forge wrapper has a new `not-found` fault that returns HTTP 404 for every `api repos/o/r/...` read. The happy-path test `test_retire_ends_observation_of_a_gone_contribution` now uses it in place of the 502 `down` fault, so it models a permanently gone repository rather than an outage. Verification: `bash tests/fm-contributions.test.sh` passed with exit 0. The last lines of its output include the three retire tests passing and no failures. `shellcheck` flagged only SC2034 (`FM_WAKE_QUEUE` unused) and SC1091 (an unfollowed source of `bin/fm-path-lib.sh`), neither on a line this round changed

* test: extend remote-reply whole-log recapture waits (#6639)

* test: give the remote-reply whole-log recapt…
yehezkieled added a commit to yehezkieled/firstmate that referenced this pull request Oct 9, 2026
* fix: reclaim orphaned watcher arms on the next park (#6335)

* fix(bin): take over the watcher cycle a main-only pass-through leaves

An attended main-only pass-through leaves a successor watcher cycle
running through main's handling turn. The session's next park attached
to that cycle instead of owning it, so the successor's arm, orphaned by
its host's exit, kept owning the watcher while the new park's arm polled
it twice a second until the next close or the park boundary, hours later
in a quiet second mate. A second-mate restart hit this every time, since
its persist request is a main-only close.

The host now records the successor it leaves for main, and the next
host's first cycle runs bin/fm-watch-arm.sh --take-over on it: when that
arm still owns the healthy watcher, the new arm stops it, reports a
reason the cycle delivered first, and otherwise owns a fresh cycle. The
stop's own downtime publication is undone over an acknowledged episode
when no wake was appended in between, so the handover wakes nobody.

* no-mistakes(review): Keep left-arm record until the orphaned arm is gone

* no-mistakes(review): Relinquish successor arm only after durably recording it

* no-mistakes(review): Relinquish successor only after its record reads back

* no-mistakes(document): Clarify successor takeover guarantees and authoritative documentation

* no-mistakes(ci): Fixed ci-1: acknowledgement restore now requires the exact taken-over arm/watcher ledger row with signal=TERM, awaited within a short bound. Otherwise takeover proceeds without erasing downtime. Added a self-exit regression confirmed failing before the fix and passing afterward; ordinary takeover tests and the full watcher-arm suite pass. Updated Generation reuse documentation. Syntax, diff checks, and ShellCheck pass with existing SC1091/SC2034 warnings excluded. ci-2 remains unchanged

* no-mistakes(document): Clarify watcher take-over recovery and restart limits

* fix(bin): restore downtime on supervision-host hand-back when the successor already closed (#6355)

* fix(bin): restore supervision host hand-back continuity

* no-mistakes(review): Scope host hand-back failure fallback to lost pending:handling

* no-mistakes(review): Remove stray scratch test copy tests/.tmp-rest.test.sh

* no-mistakes(test): Initialise successor globals so early hand-back survives set -u

* no-mistakes(document): Document host hand-back downtime failure and Claude lost-handback notice

* no-mistakes(ci): I fixed the Greptile finding. The rule that must hold: when the supervision host hands back an actionable wake, its rewake is refused, and no watcher is healthy, the hand-back still has to reach main as a delivered notice. That must be true whether the recovery marker is `pending:handling` or `announced:handling`. Only one place applies this check: the lost hand-back fallback in `bin/fm-claude-stop-autoarm.sh`. **Fix:** that check now accepts both `pending:handling:*` and `announced:handling:*` tokens (a one-line change). Nothing else in the fallback changed: - Refusals on any other marker, such as an already acknowledged one, still exit 0 silently and open no failure episode. - The notice is still sent once per episode, and repeats are recorded as `failed-suppressed`. **Tests:** - `tests/fm-claude-stop-autoarm.test.sh`: the lost hand-back test now runs as a shared helper with two variants, one writing a `pending:handling` marker and a new one writing `announced:handling` (`test_host_lost_announced_handback_notifies_once_per_episode`). - `tests/fm-supervision-host.test.sh`: the end-to-end test where downtime restoration fails is now a shared helper too, with a new `announced` variant (`test_claude_stop_hook_notifies_when_closed_announced_successor_downtime_restore_fails`). It moves the handling episode to `announced` before the host hands back. Without the fix the hook would exit 0 here; the test requires exit 2, `outcome=failed` and a delivered failure notice. **Verification (all under nice -n 10):** - The full `tests/fm-claude-stop-autoarm.test.sh` suite passed (rc=0), including both lost hand-back variants and the benign-refusal test. - In `tests/fm-supervision-host.test.sh` I ran only the four hand-back test functions, all passing (rc=0). The suite can't run single functions, so I used a temporary copy with a trimmed test list and deleted it afterwards; `git status` shows only the 3 intended files changed. - shellcheck is clean on all three changed files. I did not run `bin/fm-lint.sh`. - I did not run the new tests against the unfixed code; the claim that they fail without the fix comes from reading the old check

* fix: reduce remote-job polling process churn (#6363)

* perf: cut remote-job idle process creation in the three hot loops

Post-update host measurement still attributes most idle churn to three
per-sample loops: result-consumer state reads and date calls, the delta
reader's capture/hash pass on every poll, and the lane preemption scan's
per-field pipelines. This drops each to its minimum without touching the
contracts around them.

* fm_remote_job_read_state gains an optional result-variable form backed
  by fm_remote_job_read_line, a builtin-only bounded record read (regular
  non-symlink file, byte bound, one newline-terminated line, tolerated
  unterminated tail, no carriage returns). fm_remote_job_wait samples
  state and the SECONDS clock with no per-sample children; one date call
  converts the epoch deadline once.
* fm-remote-delta-read stats the log each poll and re-runs the bounded
  capture and hashing only when size, mtime, ctime, inode, or device
  change. The snapshot's own stat writes the comparison key, so a log
  that moves between the gate and the capture is never read as stable.
* worker_preempting_waiter_exists reads state, home, and the staged argv
  head with builtins only. The now-unused worker_job_command goes away.

The bounded reads use -d '' -n, which behaves identically on the macOS
stock bash 3.2 and current bash; -N does not exist on 3.2. Tests cover
the malformed-record corpus, delta identity gating, fork-free lane
scanning through counting PATH shims, and same-home versus cross-home
preemption. No signal traps or sleep contracts change.

* no-mistakes(review): Restore subsecond delta keys and byte-bounded builtin record reads

* no-mistakes(document): Clarify delta snapshot caching and coarse-timestamp fallback

* no-mistakes(lint): Scope UTF-8 regression locales to individual function calls

* no-mistakes(ci): Fixed both lint failures by applying the documented production-library analysis boundary at the two affected test imports. Runtime behavior is unchanged; the library remains independently linted. Canonical full-analysis lint passed for the library and both suites, as did bash syntax checks and git diff --check

* fix(bin): recognize titled Claude top rules and preserve grey slash commands (#5963)

* fix(composer): read a titled Claude top rule as the composer's edge

A named Claude Code session draws its title into the composer's top rule.
The strict separator predicate rejected that row, so the closing rule read
as a lower unmatched separator and an idle, empty composer classified
unknown on every cursorless backend, refusing fm-send, exit, and relaunch.

Spare a bare agent-glyph row sandwiched between a width-proven titled rule
and the screen's only unmatched separator directly below it. The strict
separator predicate, dead-shell rule, and blank-row posture are unchanged.

Fixes #5601
Fixes #5558

* no-mistakes(test): Keep Claude's grey slash command in Herdr payload proof

* no-mistakes(test): Make missing-herdr version check hermetic to installed herdr

* fix: prevent Pi trust prompts in seeded secondmate homes (#6387)

* fix: pre-approve Pi trust for seeded secondmate homes

Unattended first launches of Firstmate-seeded Pi secondmate homes stalled on
"Trust project folder?" until Enter. Probe --approve like --tui-mode and pass
it only for --secondmate when help advertises it (.fm-secondmate-home signal),
leaving ordinary workers and older Pi unchanged.

* no-mistakes(document): Consolidate Pi seeded-home trust documentation ownership

* no-mistakes(ci): Fixed Lint 1’s unused polling counter. Diagnosed Behavior portable serial 4 as a pre-existing delta-reader test clock race; replaced timing-dependent rewrite and deletion with deterministic executable-boundary synchronization. ShellCheck, Bash syntax checks, and git diff --check passed. Delta-reader tests passed three consecutive runs; all three live Pi trust cases passed. Production behavior unchanged

* fix(bin): retry ShellCheck roots that hit the memory ceiling without --external-sources (#6443)

* fix(lint): retry memory-bound roots without external sources

* no-mistakes(review): Make fallback tests portable and correct source-following telemetry

* no-mistakes(review): Remove committed parity fixtures and use disposable test roots

* no-mistakes(document): Document ShellCheck memory fallback and telemetry

* no-mistakes(review): Cover bounded and unbounded fallback RSS behavior

* no-mistakes(document): Correct stale lint fallback documentation

* no-mistakes(document): Correct stale lint test documentation

* no-mistakes(ci): The memory fallback (the retry without --external-sources) now gets only the time left in its root's original deadline, so it can no longer outlast the CI job. Invariant: one root's first attempt plus its fallback must fit inside a single FM_LINT_ROOT_SECONDS deadline, plus the cleanup grace. Only one site started a new deadline: the fallback call in fm_lint_run_root. The deadline is the only budget involved, because the memory limit already applies to each process separately. Changes in bin/fm-lint.sh: - fm_lint_exec_root now takes a <seconds> argument instead of always reading FM_LINT_INTERNAL_ROOT_SECS. - The first attempt passes the full deadline. - The fallback passes floor((start + deadline - now) / 1000) seconds. - When bounds are enforced and less than 1 second is left, no retry starts. fm_exec_timed rejects 0 seconds, so the retry cannot run with no time. The root keeps reason=memory, and the shard output says "no time left in its Ns deadline to retry without it". - Unbounded local runs have no deadline and behave as before. - The header comment now describes the shared deadline. Changes in tests/fm-lint.test.sh: a new test, test_memory_fallback_spends_only_the_remaining_root_deadline, runs only on hosts that can enforce bounds. It uses a 6 s deadline and 1 s grace. - Case 1: the first attempt runs 3 s and then fails with memory status 251. The test asserts one fallback ran, reported reason=timeout, and the root's recorded duration is under 7000 ms. - Case 2: the first attempt runs 5.2 s. The test asserts no fallback starts, the skip is explained, and the sidecar records memory with source-following 1. Verification: - Full `nice -n 10 bash tests/fm-lint.test.sh` passed, including the new test, in about 5 minutes. - Case 1 run against the HEAD script: the root took 9168 ms, so the under-7000 ms check fails before the fix. - `bin/fm-lint.sh bin/fm-lint.sh tests/fm-lint.test.sh` reported no findings. - The CI workflow is unchanged, so the Test step still runs only tests/fm-lint.test.sh with nice -n 10 and the 12 GiB ShellCheck limit

* fix(pi): restore watcher continuity across successor gaps and make extension log opt-in (#5489)

* Fix Pi watcher successor-gap confirmations and add extension log

Accept an already-acknowledged handling confirmation as a no-op when the
generation matches, confirm the restoration's own recovery token with a
superseded (not rejected) outcome on generation mismatch, retire an arm on
confirm failure only when the failed token names that exact pid, and record
restore attempts, readiness timeouts, and confirm results in the bounded
state/.watch-extension.log. Regression tests: already-acked no-op plus
mismatch/dead-pid/lock-mismatch rejections and the manual-restart churn
contract in fm-watch-arm.test.sh, and a mid-restore marker advance with no
rejection appendix in fm-pi-watch-extension.test.sh.

* Treat a dead arm child as an empty slot so repair and retry recover

startArm and scheduleRetry answered unchanged while holding a ChildProcess
whose OS process was already gone but whose close had not fired, so neither
the repair tool nor the retry timer started anything until that close fired.
Gate slot occupancy on a liveness check (exit/signal codes plus pid probe)
and start a fresh arm instead, with a regression test driving the repair
tool against a dead-but-unclosed child.

* no-mistakes(document): Document new Pi extension log knob

* no-mistakes(review): Fix confirm-failure retire token match, add distinct-pid test

* no-mistakes(document): Clarify retire guard needs pid and generation

* Make the Pi extension diagnostic log opt-in and default-off

Only a positive FM_WATCH_EXTENSION_LOG_KEEP_LINES enables
state/.watch-extension.log. Unset, empty, non-numeric, zero, and
negative values disable logging entirely, so the default run writes
nothing and never creates the file. The shared positiveInteger
fallback semantics stay untouched for the retry and timeout knobs.
docs/configuration.md owns the knob contract and
docs/watcher-continuity.md points at it. Tests: the superseded-delivery
case runs opted in, and a new case proves unset, zero, and non-numeric
values create no log file while delivery still succeeds.

* no-mistakes(document): Qualify extension-log coverage bullet as opt-in

* no-mistakes(ci): The two reported checks (CI run 36372002913, Require no-mistakes run 36372069208) show conclusion action_required with 0 jobs and no logs because this is a fork PR (RibatTRW/firstmate) and GitHub is holding the workflow runs awaiting maintainer approval; that gate is external to the code and needs a maintainer to approve the runs. While verifying the change locally I found a real defect the approved CI run would hit: the PR's new test test_handling_delivered_rejects_a_superseded_generation failed deterministically. Invariant violated: reopen-announced (a non-successor/manual arm start) mints a fresh recovery generation only when the durable wake queue holds unrecovered work; an announced episode with an empty queue must be left untouched so idle arm starts never churn generations (the [ -s queue ] guard from #4819, relied on by bin/fm-watch.sh:2413 and covered by the append-reopens and announcement-bound sibling tests). The test called reopen with an empty queue and expected churn, so the fix establishes the queued-work precondition (append one wake, re-announce, re-read the generation) before asserting the reopen mints and the old confirmation mismatches. Test-only change, 14 lines in tests/fm-watch-arm.test.sh. Verified: fm-watch-arm.test.sh 25/25 ok on two consecutive runs, fm-pi-watch-extension.test.sh 55/55 ok, and shellcheck reports only one pre-existing warning outside the edited region

* no-mistakes(document): Restore blank line in watcher-continuity docs

* Route superseded Pi deliveries like confirmed ones and cover the retire guard

A superseded handling confirmation now falls through to the normal delivery
path, so an accepting supervision branch owns the wake instead of main.
Scope the watcher-continuity token-pinned confirmation and narrowed retire
rule to Pi, since omp and OpenCode still confirm against the current
successor. Add tests that fail when the retire guard, the scheduled-retry
gate, or the deferred-close gate is reverted, relabel the churned-generation
characterization test, and use a reaped pid for the dead-pid rejection.

* fix(pi): hide duplicate assistant finals from hidden processing retries (#5863)

* fix(pi): silence unacknowledged processing retry replies

Suppress autonomous processing prose before persistence and during streaming while retaining tool calls, signed reasoning, and retryable outcomes. Restore ordinary output after acknowledgement or a user message.

Fixes #4954

* no-mistakes(review): Silence only processing retries, keep first presentation visible

* fix(pi): preserve differing processing retry replies

* test(pi): accept Pi 1.0.1's renamed HTML export renderer lookup (#6530)

Pi 1.0.1's createToolHtmlRenderer reads getToolRenderers and ignores
getToolDefinition. Calm /export still includes stock grep HTML; the
fixture has to pass the lookup key the installed Pi actually reads.

* fix(bin): tolerate transient quota read failures (#6490)

* fix(procevent-quota): tolerate consecutive slow quota-axi reads

The quota allowance poll treated any quota_json failure as terminal, so one
slow quota-axi --json (measured max ~29s under a 48s derived bound) shut the
watch down until someone re-armed it, and the detail always said
"missing/incompatible". Tolerate three consecutive failed or timed-out reads
before going terminal, reset the streak on any good read, and report a
timeout distinctly from a missing or incompatible tool. Each timed poll runs
exactly one bounded --version and one bounded --json: validate the captured
version text through fm_quota_axi_version_compatible rather than launching a
second probe, and describe a mixed failure streak by count plus last cause.

* no-mistakes(document): Document quota polling failure tolerance

* no-mistakes(ci): Fixed ci-2 and ci-3. Permanent quota read failures (rc 2 missing, rc 3 incompatible) now report on the first poll, while transient rc 1/4 failures retain the existing three-failure retry behavior. The missing-binary test now uses an isolated PATH without quota-axi and asserts both permanent failures stop at condition_polls: 1. Verification passed: tests/fm-procevent-quota.test.sh, canonical fast lint for both changed files, bash syntax checks, and git diff --check

* no-mistakes(review): Classify untimed quota version failures as transient

* no-mistakes(document): Clarify quota polling failure budget

* fix(bin): clarify scratch guidance and dirty teardown refusals (#6505)

* fix(teardown): clarify scratch guidance and dirty worktree refusals

Keep ship proof material outside the task worktree and distinguish untracked-only leftovers from tracked edits without changing cleanup guards.

Fixes #6319

* fix(ci): Fixed ci-3 only. Both promotion outputs now replace the scout restriction and require external scratch storage and a clean worktree before done. Verification: 27 delivery tests passed, five mutations caught, restored test passed, pinned ShellCheck and syntax/whitespace checks passed. ci-1, ci-2, and ci-4 remain untouched

* fix(bin): escalate inbox instructions blocked by busy workers (#6518)

* Escalate inbox instructions stuck behind a busy worker

Count consecutive busy-deferred due doorbells durably and escalate at the configured bound without typing into the worker pane.

Fixes #6445

* fix(review): Fix inbox escalation deduplication and busy streak resets

* fix(review): Preserve busy inbox escalations through daemon supervision

* fix(document): Correct busy-inbox escalation documentation

* fix(ci): Fixed SC2034 in tests/fm-task-inbox.test.sh by including the loop counter in the failure diagnostic. Source-aware lint, all 34 inbox tests, and git diff --check pass. Behavior portable serial 6 reproduces identically on base 1f3e7696 and target 78156b86 with Pi 1.0.1: an unrelated renderer API change breaks the unchanged Calm test. No Calm changes made; that failure is addressed separately by https://github.com/kunchenguid/firstmate/pull/6516. Logs retained in scratchpad-ci/

* fix(ci): Fixed ci-2, ci-3, and ci-4: successor failures surface, reset alerts deduplicate, and oversized busy limits fall back to two. Passed 47 inbox tests, 7 focused daemon checks, all 13 mutation checks, lint, documentation checks, and diff checks. Evidence: scratchpad-ci-selected/summary.json. ci-1 remains unchanged and unwaived. Fresh live Herdr proof remains with the outer driver

* fix: prevent unnecessary remote worker turnover (#6431)

* fix: prevent healthy remote job worker turnover

* no-mistakes(review): Serialize full LaunchAgent repair and verify launchd-tracked owners

* no-mistakes(review): Let launchd-tracked unpublished spawns start before reloading

* no-mistakes(document): Document remote worker heartbeat and serialized LaunchAgent recovery

* no-mistakes(ci): Full CI log showed the idle-worker regression exceeded its outdated command budget (82 versus 80) after independent heartbeat ownership checks were added. Raised the budget to 120 while retaining the separate busy-poll sleep limit. Remote-job and LaunchAgent executable tests passed, as did bash syntax validation and git diff --check. No production behavior changed

* no-mistakes(ci): Fixed missing readiness recovery under verified live ownership, preserving the serving PID and private file mode. Added executable regressions for deletion during a blocked sweep and stale readiness diagnostics without LaunchAgent reload. Deletion regression failed before the fix. Both remote-job test suites, bash syntax validation, and git diff --check passed

* no-mistakes(ci): Full CI log identified a flaky ownership-loss test racing an already-authorized heartbeat refresh. Replaced backdating and a fixed sleep with bounded observation of readiness expiry through the public probe. Production behavior unchanged. Remote-job and LaunchAgent executable suites passed; bash syntax validation and git diff --check passed

* fix: wait for launchd bootout cleanup

* no-mistakes(document): Document remote worker recovery and read-only turnover verification

* no-mistakes(review): Publish worker identity before lock owner records

* no-mistakes(document): Document worker identity publication safety invariant

* no-mistakes(ci): Fixed ci-1: replacement workers discard predecessor readiness before publishing identity and roll back identity if lock-owner recording fails. Added executable regressions reproducing both defects. LaunchAgent and orphan-reap suites, shell syntax checks, and git diff --check passed. No live service state was modified

* no-mistakes(ci): Restored lock-owner-before-identity publication and removed the identity rollback and reordering-only tests. Retained an executable regression proving predecessor readiness is rejected until replacement startup completes. LaunchAgent and orphan-reap suites, shell syntax checks, and git diff --check passed. Other changes remain intact; the retained-identity interrupted-repair edge remains out of scope. No live service state was modified

* fix: speed up remote-job sequence claim cleanup (#6575)

* fix(remote-job): reap expired seq claims with one directory walk

The hourly claim sweep forked uname+stat per .seq-claims entry and blocked
serving for ~85s at ~17k dirs. Delete expired empty claim dirs with a single
find -exec rmdir batch and cache the host uname for remaining mtime reads.

* no-mistakes(review): Restore original path mtime helper and drop uname cache

* no-mistakes(test): Restore claim retention eligibility; focused retention and serving tests pass

* no-mistakes(document): Document single-walk claim cleanup and regression entrypoints

* no-mistakes(ci): Fixed Lint 2’s reproduced SC1091 by adding the tests/lib.sh ShellCheck source directive to the retention test. Runtime behavior is unchanged. ShellCheck passed for both new claim tests; Bash syntax, the retention behavior test, and git diff --check passed

* fix(tests): keep fixture registries out of git worktree roots

A TMPDIR pointed at a repository root placed live .fm-test-* registries
beside tracked files, and a concurrent git add during the claim-walk CI
fix round committed three of them. Route registries and fixture roots
through a TMPDIR that refuses git worktree roots, remove the stray files,
and pin the escape with a behavioral cleanup test.

* no-mistakes(review): Preserve whole-second claim expiry in single-walk sweep

* no-mistakes(document): Correct temporary-directory resolution documentation

* no-mistakes(ci): Fixed ci-3 by changing only the stale, fresh, and read-only orphan fixture paths in tests/fm-test-fixture-cleanup.test.sh to use $FM_TEST_TMPDIR. All seven tests passed both normally and with TMPDIR set to the worktree root. Shell syntax and git diff --check passed

* feat(bin): record no-mistakes pipeline spend per task at teardown, opt-in (#5354)

* feat(bin): record each task's no-mistakes pipeline spend at cleanup

no-mistakes keeps every pipeline agent invocation's token usage only in its
local agent_invocations records, and cold pipeline agents (review, test,
document) leave no session log, so a task's review-loop cost never reached
Firstmate's records and could not be attributed after cleanup.

bin/fm-pipeline-spend.sh attributes a task's runs by the repository
no-mistakes resolves for the task copy, the task branch, and the branch's
creation (a relaunch mints a new spawn_gen while the branch keeps
validating), then sums every invocation, failed and cancelled included.
Token fields use no-mistakes' per-round deltas so resumed review rounds are
not counted twice, and unrecorded values stay unknown rather than zero.
show prints the record read-only; record appends it once per task
incarnation to data/pipeline-spend.jsonl. Teardown records it for every
ship task it cleans up, before deleting the branch and the task record.

fm_nm_state_db becomes the one owner of where no-mistakes' state database
lives, shared with the capped run-inventory reader.

* no-mistakes(review): Drop spend show command and timeout override

* style(bin): rewrap fm-pipeline-spend header comment

* Make pipeline spend recording opt-in

* feat(bin): record each task's no-mistakes pipeline spend at cleanup

no-mistakes keeps every pipeline agent invocation's token usage only in its
local agent_invocations records, and cold pipeline agents (review, test,
document) leave no session log, so a task's review-loop cost never reached
Firstmate's records and could not be attributed after cleanup.

bin/fm-pipeline-spend.sh attributes a task's runs by the repository
no-mistakes resolves for the task copy, the task branch, and the branch's
creation (a relaunch mints a new spawn_gen while the branch keeps
validating), then sums every invocation, failed and cancelled included.
Token fields use no-mistakes' per-round deltas so resumed review rounds are
not counted twice, and unrecorded values stay unknown rather than zero.
show prints the record read-only; record appends it once per task
incarnation to data/pipeline-spend.jsonl. Teardown records it for every
ship task it cleans up, before deleting the branch and the task record.

fm_nm_state_db becomes the one owner of where no-mistakes' state database
lives, shared with the capped run-inventory reader.

* no-mistakes(review): Drop spend show command and timeout override

* style(bin): rewrap fm-pipeline-spend header comment

* Make pipeline spend recording opt-in

* no-mistakes(document): Note pipeline-spend opt-in gate in teardown comments

* no-mistakes(ci): I fixed all three Greptile findings following your decision. Both test suites pass: tests/fm-teardown.test.sh (105 ok, exit 0) and tests/fm-pipeline-spend.test.sh (8 ok). I didn't run either new test against the old code, so the claim that they fail before the fix is from reading the code, not a run. Nothing was committed or pushed. **ci-1 / ci-2 (bin/fm-teardown.sh):** The rule is that every owned ship task in a home that has opted in gets a durable spend record before its task record is deleted. The `[ -d "$WT" ]` check skipped the recorder when the worktree folder was already gone, so no record was written. I removed that check and kept the other conditions: the task must be a ship task, teardown must own its worktree, and `config/pipeline-spend` must exist. `bin/fm-pipeline-spend.sh` already writes an unavailable-source line when the worktree is missing. The new test `test_teardown_records_unavailable_spend_for_a_gone_worktree` in tests/fm-teardown.test.sh sets up an owned ship task whose worktree is missing, with recording turned on. It checks that teardown succeeds, writes a line with `source=="unavailable"`, `total==null` and a reason saying the task copy is gone, and still removes the task record. Under the old check no ledger would have been written. **ci-3 (bin/fm-nm-run-lib.sh):** When `NM_HOME` and `HOME` were both unset, `fm_nm_state_db` fell back to `${HOME:-}/.no-mistakes`, which resolves to `/.no-mistakes`. It now uses `root=~/.no-mistakes`. Bash expands `~` to the account's home directory even when `HOME` is unset, which matches the previous Python `Path.home()` lookup. The new test `test_state_db_without_nm_home_or_home_uses_the_account_home` in tests/fm-pipeline-spend.test.sh runs the function with both variables unset and compares the result to the home directory from the system's password database. The old code would have returned `/.no-mistakes/state.sqlite`. shellcheck reports nothing new on the changed files. I added one `disable=SC2016` comment in the new test, where the single-quoted `$1` is meant to expand in the child shell

* feat(bin): start tasks from a named base branch (#6442)

* feat(bin): start tasks from a named base branch

Spawns always reset a task's pooled copy to origin's default branch,
so work that belongs on a feature, integration, or release branch
started from the wrong code and opened its PR against the default.

fm-brief.sh --base-branch records the base in the brief, fm-spawn.sh
resets the copy to origin/<base> and records base_branch= in task
meta, and the worker targets its PR at that branch. Review diffs,
the cleanup content check, and scout promotion read the recorded
base. local-only and Gerrit deliveries refuse a named base.

* no-mistakes(review): Anchor brief base-branch parse on scaffold Setup line

* no-mistakes(review): Require explicit --base-branch on spawn, matching brief lines

* no-mistakes(document): Note base-branch reset in fm-spawn freshness header

* no-mistakes(ci): I fixed all four Greptile findings the way you specified. The affected suites pass locally: fm-brief, fm-dod-lib, fm-spawn-pool-base-freshen, fm-review-diff, fm-teardown and fm-task-delivery. shellcheck on the changed files reports only info-level notes, no warnings or errors. I didn't revert the code to watch each new test fail before the fix. - **ci-1 (brief text blocked or redirected spawns):** In bin/fm-dod-lib.sh, `fm_brief_base_branches` now only accepts a `Base branch: <value>` line that comes directly after the base-variant Setup sentence fm-brief.sh writes. Any other `Base branch:` line is ignored. `--base-branch` is still the only authority, and the existing checks in fm-spawn.sh are unchanged. With the flag, the spawn is refused unless a matching pair exists and no pair names a different branch. Without the flag, it is refused only if such a pair exists. The header comment in fm-spawn.sh is updated to match. - The `brief_with_base` test helper now writes the real two-line pair. - The decoy refusal case now uses a full decoy pair in the captain's intent. - New test `test_prose_base_branch_line_is_ignored`: a brief with no base whose intent quotes `Base branch: release/1.2` spawns normally with no `base_branch=` recorded. A spawn with `--base-branch feature/hub` whose intent names release/1.2 starts from `origin/feature/hub` and records it. - **ci-2 (scout base not checked against the forge):** fm-spawn.sh now looks up the project's registered forge with `bin/fm-project-mode.sh --forge`, like the ship path does, before `fm_base_branch_valid` runs. Previously it passed `none`. New test `test_scout_base_branch_refused_on_gerrit_forge`: a scout with a base on a forge=gerrit project is refused at spawn and no meta file is written. - **ci-3 (base not shell-quoted in worker commands):** `fm_dod_block` now quotes the base with `printf %q` in the `--base` and `--base-branch` arguments, and the branch name in the prose stays readable. The test in fm-brief.test.sh uses the base `release/$HOTFIX` and checks that both direct-PR and no-mistakes briefs render `release/\$HOTFIX` in the commands. - **ci-4 (ambiguous PR sentence):** The direct-PR sentence now reads "open a PR with `gh-axi` that is ready for review, not a draft, against the base branch `X` (`--base X`), not the repository default." A brief with no base renders exactly as before. The existing assertion in fm-brief.test.sh is updated

* no-mistakes(ci): Both real failures on PR 6442 (run 37064756768) come from CI tooling and a Pi version bump, not from the base-branch feature. Fixes for both already exist on upstream main, so I applied those two commits to the working tree with `git cherry-pick --no-commit`. Nothing is committed or pushed, and no feature file changed. - **Lint 2:** ShellCheck ran out of memory on `bin/fm-teardown.sh` (rss about 8 GB, exit 251, "shellcheck: out of memory"). The rule that broke is that the lint gate must finish on any valid shell root without exceeding the runner's memory ceiling. Upstream commit d719ef3 (#6443) fixes this in `bin/fm-lint.sh`: a root that hits the memory ceiling is retried without `--external-sources`. It also updates `tests/fm-lint.test.sh` and `docs/fm-test-portable-shards.md`. - **Behavior portable serial 6:** `tests/fm-calm-pi-extension.test.sh` failed with "grep disappeared from /export calm.html HTML while calm mode was on". Pi 1.0.1 renamed its HTML export renderer lookup, and upstream commit fede619 (#6530) updates the test to accept the new name. This change is test-only. Both commits applied cleanly. The four touched files are `bin/fm-lint.sh`, `docs/fm-test-portable-shards.md`, `tests/fm-lint.test.sh` and `tests/fm-calm-pi-extension.test.sh`. None of them is in the feature diff, so the feature contract is unchanged. **Local checks:** - `tests/fm-lint.test.sh` passes. - `bin/fm-lint.sh` on `bin/fm-teardown.sh`, `bin/fm-spawn.sh` and `bin/fm-dod-lib.sh` exits 0 with full ShellCheck analysis. - `tests/fm-calm-pi-extension.test.sh` exits 0 with 7 ok and 0 not ok. One case skipped because the Pi package isn't installed locally, so the Pi 1.0.1 export path that failed in CI couldn't be reproduced here. CI needs to confirm it. - I didn't reproduce the 8 GB OOM locally. The PR has to be updated through the pipeline with a normal push: no force push, no replacement PR, no merge

* fix(bin): name the Firstmate skill file as a fallback in the worker role (#6647)

* fix(bin): point project workers at the Firstmate skill file

The Skill tool cannot resolve a Firstmate skill from another project's worktree, so the launch role names the readable skill file instead.

* no-mistakes(review): fix(bin): name Firstmate skill file as fallback only

* feat(bin): retire a contribution whose forge object is permanently gone (#6655)

* feat(bin): retire a contribution whose forge object is permanently gone

Add fm-contributions.sh retire <task> <url> <captain|fleet> <reason>,
which records actor, reason and time on the saved record and removes
that task/url pair from known, poll rotation and coverage even while a
backlog link remains. Repeating a retire keeps the first provenance;
an unrecorded pair, unknown actor, empty reason or unacknowledged
pending signal is refused.

* no-mistakes(review): Keep retirement per task when settling final owners

* no-mistakes(ci): Made the three Greptile fixes the captain chose. ci-1, retirement needs the captain's word: `retire` now accepts only the actor `captain`. Any other actor is refused, including `fleet`. I removed `fleet` from the usage line, the script header (`bin/fm-contributions.sh`), the check in `bin/fm-contributions.sh` and the retired-record validation in `bin/fm-contributions.jq`, which now requires `.actor == "captain"`. The script header now says a retirement always records the captain's word, because the script cannot verify who runs it and the authority to retire is the captain's. The Bearings skill line in `.agents/skills/bearings/SKILL.md` now says `retire` records the captain's word. In the tests, a `fleet` retire is now a refusal case and must leave the record unretired. Every other retire call in the tests uses `captain`. The PR description has not been changed yet: the ruling asks for the same captain's-word statement there, and that belongs to the PR phase. ci-2, blank reasons: the reason check is now `[ -n "${5//[[:space:]]/}" ]`, so a reason made only of spaces or tabs is refused. I added a refusal test that passes a space-tab-space reason. The header also lists a blank reason among the refusals. ci-3, the gone-object test: the test forge wrapper has a new `not-found` fault that returns HTTP 404 for every `api repos/o/r/...` read. The happy-path test `test_retire_ends_observation_of_a_gone_contribution` now uses it in place of the 502 `down` fault, so it models a permanently gone repository rather than an outage. Verification: `bash tests/fm-contributions.test.sh` passed with exit 0. The last lines of its output include the three retire tests passing and no failures. `shellcheck` flagged only SC2034 (`FM_WAKE_QUEUE` unused) and SC1091 (an unfollowed source of `bin/fm-path-lib.sh`), neither on a line this round changed

* test: extend remote-reply whole-log recapture waits (#6639)

* test: give the remote-reply whole-log recapture a longer wait

* no-mistakes(review): Extend both recapture waits and simplify retry handling

* fix(bin): reopen a pending-reply escalation after its resolve (#6654)

* fix(bin): reopen a pending-reply escalation after its resolve

A retry of an open decision still appends nothing. A later same-kind escalation after a resolved line for that key appends again, so the next loss is visible.

* no-mistakes(ci): Both Greptile findings are fixed in the working tree; the four directly affected test suites pass, and a wider run of related suites had not finished when I returned this result. Invariant: a line already in the status file is recorded, and only a caller with its own evidence of a new episode may append it again. It must hold for every caller of status_event_recorded: the continuity break and the other call in bin/fm-procevent-remote-reply.sh, fm_parent_channel_append_once, and the pending-reply escalation. Changes: - bin/fm-classify-lib.sh: status_event_recorded is back to the idempotent retry check. It returns at the first matching line, and a later resolved line no longer makes that line look new. This fixes ci-1 for every caller and restores the early exit for ci-2, with no separate scan optimization. - bin/fm-pending-reply-lib.sh: _fm_pending_reply_maybe_escalate_locked now owns the reopen. It appends a new blocked line when the record already has an escalated_epoch (it escalated before and was reset) and the keyed decision is no longer open. Otherwise it uses status_event_recorded, so a retry or a resend that leaves the decision open appends nothing. - Comments on status_event_recorded and the pending-reply header describe this scope. Tests: - tests/fm-pending-reply.test.sh: the regression for escalate, operator resolve, second escalate is unchanged and passes. - tests/fm-remote-reply.test.sh: new regression drives the real reader. After an operator resolve of remote-reply-continuity-ios, a repeated handle and ingest of the same break appends nothing and leaves the decision closed. I placed it beside the existing continuity-break test, not in the pending-reply file, because only that file has the reader harness. - tests/fm-classify-corr-token.test.sh and tests/fm-classify-decision-key.test.sh: the cases that expected a blocked line to append again after a resolve now assert it stays recorded. Verification: - fm-classify-decision-key, fm-classify-corr-token, fm-pending-reply and fm-remote-reply test suites all exit 0 with the fix. - With the two bin files reverted to the commit under review, the new continuity regression and the updated parent-publisher case both fail, so they reproduce ci-1. - shellcheck -x on the changed files reports nothing. - Not finished: a background run of every other suite that mentions the parent channel or pending replies had produced no output, pass or fail, when I returned. One known gap: the escalation writes escalated_epoch before phase. If the process dies between those two writes and an operator resolves the key before the next tick, that tick appends a blocked line for the same episode. I left it, because closing it needs new state. Issue 4755 and other escalation behavior are untouched. Nothing is committed

* fix: refuse a confirming Enter on the Claude background-task exit picker (#6666)

* fix: refuse a confirming Enter on the Claude exit picker

The background-task picker still classifies as pending, so a retried Enter confirms "Exit and stop tasks".
Stop after the Enter that opened it, and raise the existing stale wake with the dialog name.

A non-paused secondmate still skips that wake, so a secondmate parked on the picker still looks idle.
Model-downgrade confirmation, MCP approval, and any other Claude exit confirmation are not covered, because there is no recorded screen for them.

* no-mistakes(review): fix: anchor exit picker match and wake second mates

* fix: refuse a typed submit while the Claude exit picker is open

A pane that already shows the picker must not receive the message or a confirming Enter.
The match requires the recorded heading and footer lines, and it is skipped when no dialog name is being recorded.

* no-mistakes(review): restore watcher to main behaviour, drop dialog wake

* no-mistakes(test): test: align Herdr picker fixtures with the preflight read

* no-mistakes(document): document exit refusal on a recognised dialog

* fix: remove the dialog file when exit runs in a subshell

do_exit is invoked from a command substitution, so the parent cleanup never saw the path it is supposed to delete.

* no-mistakes(review): fix: set the dialog file path after the control lock

* no-mistakes(document): document why the dialog file path follows the lock

* no-mistakes(ci): The exit cleanup in bin/fm-control.sh now deletes the dialog file first and releases the control lock second. The changes are in the worktree and are not committed. Invariant: only the process that holds the control lock for a task may write or delete that task's dialog file ($STATE/<id>.composer-dialog, the file that carries the picker name from the composer read to the Enter checks). The old cleanup released the lock and then deleted the file, so the delete ran outside the lock. Places where this invariant must hold: control_cleanup is the only place that deletes the file, and the single assignment after CONTROL_LOCK_HELD=1 is the only place that names it. do_exit only truncates the file, and it runs while the parent holds the lock. A process that loses the lock never sets the path, so it deletes nothing (earlier fix, unchanged). No sibling site needed a change. Fix: in control_cleanup, the rm of the dialog file moved above the fm_lock_release call, with a two-line comment that gives the reason. No dialog recognition or supervision code changed. Test: added test_exit_removes_the_dialog_file_before_releasing_the_lock in tests/fm-control-relaunch.test.sh. It runs one `fm-control exit` with a recording rm on PATH. The lock release removes paths at or under the control lock with rm, so the recording rm writes whether the dialog file exists at that moment. The test asserts the last record is "absent". It starts no second command. Verification, all run locally: - Before the fix, the new test failed with "the dialog file must be gone when the control lock is released, got: present". - After the fix, tests/fm-control-relaunch.test.sh exits 0 with 75 ok lines and no "not ok" line, including the new test and the existing test that the file is gone after exit and after relaunch. - tests/fm-control.test.sh exits 0 with 44 ok lines. - shellcheck -x on both changed files exits 0 with no output. One thing I noticed and did not change: tests/fm-control-relaunch.test.sh prints 615 "rm: cannot remove ... Permission denied" lines during its temp-directory teardown. The count is the same with my change reverted, so this change does not cause it, and the suite still exits 0. I could not read the Greptile check log (the check run was not found), so I worked from the finding text and the user instructions

* fix: restore timeout watchdog compatibility with macOS Bash 3.2 (#6028)

* fix: support stock macOS Bash in timeout watchdog

* no-mistakes(ci): Fixed ci-1 in tests/fm-timeout-lib.test.sh: both startup-owner failure paths now kill the bounded command group before killing the watchdog. Fault injection reproduced both leaks before the fix and confirmed cleanup afterward. Bash 3.2 suite, syntax check, ShellCheck, and diff checks passed; GNU timeout coverage skipped because the binary is unavailable. Production code unchanged

* fix(project-management): use subshell form for Initialize command (#6699)

Wrap the cd command in a subshell to comply with the cd-guard policy
that blocks persistent top-level directory changes in the primary
firstmate checkout. The subshell form (cd projects/<name> && ...) is
accepted by the policy as documented in issue #6502.

Fixes #6502

* feat(bin): add armable daily startup growth check (#6725)

* Add daily startup growth check

* no-mistakes(review): watch printed startup memory, retain growth baselines, drop knobs

* no-mistakes(review): delegate budget verdict, report before publish, pin shim home

* no-mistakes(review): baseline first sightings silently, drop mtime, spare secondmates

* no-mistakes(review): cap the wake line, validate budget verdict fields

* no-mistakes(review): guard record schema, check appends, tighten assertions

* no-mistakes(review): drop arbitrary tracked library, make re-arm idempotent

* no-mistakes(review): correct tracked-set wording, restore shim guards, register artifacts

* no-mistakes(review): exit on signal instead of publishing partial record

* no-mistakes(document): correct startup-growth record removal cost in state registry

* no-mistakes(test): sweep orphaned empty startup-growth temp records on due evaluation

* no-mistakes(document): note watcher need for armed startup growth check

* no-mistakes(ci): Diagnosed "Behavior portable serial 5": the shard failed on tests/fm-contributions.test.sh in test_arm_plumbs_a_configured_budget_into_the_check_shim (inherited mode) with "generated check did not attempt a read" and a missing forge/calls file. Root cause: bin/fm-contributions.sh:379-382 computes DEADLINE=$(date +%s)+BUDGET and then gates each read on DEADLINE-$(date +%s) >= OBSERVATION_RESERVE (= min(BUDGET,15)). At the one-second budget this test uses, that gate demands zero elapsed whole seconds, so a wall-clock second boundary crossing between the two date calls makes the poll loop break before any forge read. The test calls wrap_forge (which installs a fake date reading $FORGE/clock only when that file exists) but never seeded forge/clock, so it ran against the real clock. The repo already documents this exact hazard at tests/fm-contributions.test.sh:675-677 and freezes the clock in every other budget-constrained test: test_budget_exhaustion_keeps_prior_record (both modes), test_genuine_failure_near_deadline_is_unavailable, test_reservation_defers_later_url_when_fifteen_seconds_do_not_remain, test_slow_read_deadline_kill_is_budget_refusal, test_budget_is_cut_down_to_the_watcher_check_bound. test_arm_plumbs was the only omission; its single loop body covers both the configured and inherited modes. The remaining wrap_forge users run the default 20-second budget (5 seconds of slack above the reserve) and are not reachable by this race. Fix (smallest, matching the sibling idiom): added the clock freeze `/bin/date +%s > "$home/forge/clock"` plus a two-line comment inside that test's loop, before the poll. Three added lines in tests/fm-contributions.test.sh; no production code and no other file touched. Verification: reproduced the exact CI failure message by running a scratch copy whose fake date advances one second per read (worst case of the unfrozen clock) - missing forge/calls, "not ok - generated check did not attempt a read". With the fix, the single test passes 10/10 standalone and the full tests/fm-contributions.test.sh passes all 46 assertions with exit 0. tests/fm-startup-growth-check.test.sh still exits 0. bash -n and shellcheck -S warning are clean on the edited file; scratch repro files were removed, leaving only the intended three-line change. Note for the author: this flake is not caused by this branch - tests/fm-contributions.test.sh is untouched by the PR and the race is pre-existing, surfacing only on a loaded runner (the shard's neighbouring assertions show 80+ second gaps). It was fixed because it is genuine nondeterminism with an established in-repo remedy and the check cannot otherwise go green. No work was done on the unselected greptile findings (ci-2..ci-5) or the cancelled "Behavior portable serial 2" check (ci-6)

* no-mistakes(ci): Diagnosed "Behavior portable serial 6": the omitted log section (fetched with gh api) shows the shard failed on tests/fm-calm-pi-extension.test.sh in test_hidden_block_geometry_e2e with "Pi Calm hidden-block geometry E2E did not complete the /reload viewport transition" (exit=1, duration_ms=23437; the family line pure-contract-unit count=3 duration_ms=51193 failed=1 matches devin-harness 3203 + calm-pi 23437 + task-delivery 24553). Root cause: tests/fm-calm-pi-extension.test.sh:2831 wait_for_geometry_transition detected the /reload by SAMPLING a single intermediate frame - it polled tmux capture-pane for Pi's transient "Reloading keybindings, extensions, skills, prompts, themes, and context files..." box and only accepted the rebuilt transcript afterwards (elif gated on saw_transient). Pi shows that box only while the reload runs and replaces it via dismissReloadBox when done, so on a loaded runner the box can live entirely between two polls; saw_transient then stays 0 forever and the 600-attempt wait times out although the reload fully succeeded. Reproduced locally under CPU oversubscription with CI's Pi 1.0.4: 1 failure in 8 runs, diagnostics proving the mechanism ("DIAG: transition timeout saw_transient=0 attempt=600") while the captured viewport already showed the completed reload status row and the intact collapsed transcript. Not a Pi-version regression: 1.0.4's handleReloadCommand and both banner strings are identical to the installed 0.85.1, and shard 6 of the previous pipeline run (job 112490351105) passed this same script on Pi 1.0.4. Invariant violated: a TUI E2E assertion must key off a durable state the UI retains, never off one intermediate frame polling may miss. Sibling enumeration: the helper was defined once and called once; grep for transient/saw_/two-stage waits across tests/, bin/ and .pi/ found no other one-shot-frame wait - every other wait in the file keys off durable text presence/absence or off continuously repeating animation frames (the working-ship motion checks) where sampling eventually connects; no doc or other test referenced the helper. Fix (smallest, one helper): renamed it wait_for_geometry_reload and made it wait for the durable status row Pi appends once the reload completed and the chat was rebuilt ("Reloaded keybindings, extensions, skills, prompts, themes, and context files", a persistent chat child via showStatus, present in both 0.85.1 and 1.0.4) together with CALM_GEOMETRY_FINAL. That marker is absent before the reload and never appears if the reload throws, so the check is strictly stronger than before (a failed reload now fails the test instead of passing on transient-then-final). One file changed: tests/fm-calm-pi-extension.test.sh, 2 hunks, no production code. Verification: deterministic before/after - with polls spaced 0.4s apart (the loaded-runner worst case) the pre-fix helper fails with exactly the CI message and the post-fix helper passes; 12/12 consecutive passes of the isolated test under load with Pi 1.0.4; full file 15/15 ok exit 0 under LANG=C.utf8 (matching the runner locale); bash -n, shellcheck -S warning and bin/fm-lint.sh clean; scratch repro files and the temp Pi 1.0.4 install removed, git status shows only the one intended modification. Notes: (1) the flake is not caused by this branch - that test file is untouched by the PR and the race is pre-existing; fixed because it is genuine nondeterminism and the check cannot otherwise go green. (2) Under far harsher load than CI applies a separate unrelated flake appeared in the same test: Pi's "Warning: tmux extended-keys is off..." row occasionally lands between the collapsed [skill] ahoy row and the final response, making assert_geometry_gap see 5 rows instead of 2. Different cause, never seen in CI, and its plausible remedies would change key delivery for the other TUI tests in the file, so it was left alone rather than growing this fix

---------

Co-authored-by: Quartermaster <quartermaster@users.noreply.github.com>

* fix(bin): reopen the remote-reply continuity decision on a later break (#6708)

* fix: reopen a remote-reply continuity break after repair

A later break for the same route and reason was swallowed after the
operator resolved the first one, because the status line matched for
the life of the log. The continuity ingest now appends again when the
cursor has moved or retirement has reset that episode, and an unchanged
re-read still appends nothing.

status_event_recorded is unchanged. Its other callers are the
pending-reply escalation, which already decides its own episode, the
parent-channel note append, and the remote document transfer note.

* no-mistakes(review): seed continuity episode for already recorded break line

* no-mistakes(document): document when a remote-reply continuity break reopens

* no-mistakes(ci): The adapter now stores the continuity episode record before it appends the `blocked` line, so a failed store appends nothing. The changes are uncommitted in the worktree, in `bin/fm-procevent-remote-reply.sh` and `tests/fm-remote-reply.test.sh`. **Invariant:** a continuity `blocked` line is on the parent status log only when the episode record for that break is already stored. Only the `continuity-broken` branch of `cmd_ingest` appends that line, so the fix is at that one place. **What changed in `cmd_ingest`:** - It decides first whether the break needs a line, then stores the episode record, then appends. - If storing the record fails, it stops with "cannot record continuity episode" and appends nothing. - If the line is already on the log and no episode record exists, it stores the record and does not append. **One addition you did not ask for:** if the append fails after the record is stored, the adapter puts the earlier record back (or deletes the new one when none existed). Without that, the retry would read as the unchanged repeat and append nothing, which is the issue 6701 failure again. **Deviation from your test instruction:** the test does not use `chmod` on the cursor directory. `write_continuity_episode` runs `chmod 700` on that directory before every write, so a read-only directory is made writable again. The test instead makes `mktemp` fail for the episode's temporary file in that directory, the same way the existing receipt-failure test does. **Tests added to `tests/fm-remote-reply.test.sh`:** - Store failure: the second break exits 1, appends no `blocked` line and opens no decision. One retry after storage recovers exits 3 and appends a single line. - Append failure: with the status log read-only, the third break exits 1 and appends nothing. One retry after the log is writable exits 3 and appends a single line. **Verification:** `bash tests/fm-remote-reply.test.sh` ends with "ALL TESTS PASSED" with the fix. Against the script at commit b750c828 the same test file fails at "a continuity break appended its line before its episode was stored". `shellcheck -S warning` reports only an unused loop variable at line 667 of the test file, which this change does not touch. I ran no other test files. The Greptile Review check log could not be retrieved, so I worked from the finding text alone

* fix: record a continuity break's reader position on its status line

A later break at another cursor is then a different line, so the existing
duplicate check appends it and reopens the decision. An unchanged re-read
builds the same line and appends nothing.

* fix: reopen a continuity break after an identical restore

A retirement that puts the same bytes back used to rebuild the recorded line, so the later break stayed closed. The retirement count on that line makes the later break distinct.

* no-mistakes(review): remove continuity match for full-prefix line without retirement count

* no-mistakes(document): clarify what a continuity break status line records

* fix: remove the reply cursor before recording retirement

A stop between those steps must leave the count unchanged, so an unchanged continuity break still builds the same line.

* fix(control): drop busy_gen from the task record when an incarnation is retired (#6733)

* fix(control): drop busy_gen when an incarnation is retired

A deliberate exit removed the busy sidecar and left busy_gen in the task record, so the two records disagreed about whether that incarnation was still observable.

* no-mistakes(review): drop GNU-only chmod and unreached sidecar-absent branch

* no-mistakes(review): correct lock comment to name the deadlock

* no-mistakes(ci): The test `test_exit_drops_meta_busy_gen_with_the_sidecar` in tests/fm-control.test.sh now compares the whole task record (the `state/<id>.meta` file), so the Greptile finding is fixed. Invariant: after `exit` retires an incarnation, the task record must equal the record from before `exit` with only the `busy_gen` line removed. This test is the only place in the change that asserts the record survives the rewrite, so it is the only site to fix. The other `busy_gen` tests assert that the line stays, and they do not go through the rewrite. What changed: before `exit`, the test writes the record without its `busy_gen` line to `expected.meta`. After `exit`, the test runs `diff` between that expected copy and the real record, and fails with the diff output if they differ. This one comparison replaces the two earlier checks (no `busy_gen` line left, and the `window` line present), because it covers both. I did not change bin/fm-control.sh or any other file. How I know it works: - I ran `bash tests/fm-control.test.sh`: exit code 0, 45 lines starting with `ok`, no other lines. - I temporarily changed the rewrite in bin/fm-control.sh to also drop the `harness` line. The test then failed with `not ok - exit should drop only busy_gen from the task record:` and the diff `< harness=codex`. The earlier `window`-only check would have passed that rewrite. I restored bin/fm-control.sh afterwards; `git status` shows only tests/fm-control.test.sh modified. - `bash -n` and `shellcheck` on the test file report no new warnings from the edit. The change is not committed; the working tree holds it

* test: cover PID collisions in harness ancestry detection (#6484)

* test(secondmate-harness): scope fake ps -codex label for pid 5252 to the liveness probe

Closes #6456

* no-mistakes(ci): Updated the collision test to log and assert that PID 5252 was queried before selecting 4242. Full fm-secondmate harness suite passes

---------

Co-authored-by: YifuGu <ironerumi@users.noreply.github.com>

* fix: share Pi Calm's working-ship widget slot (#1854)

* fix(calm): share the standalone Pi Calm working-ship widget slot

Firstmate Calm and the user-global standalone Pi Calm both install an
animated working-ship widget during agent runs. Each claimed its own Pi
widget key, so a session loading both (the main Firstmate home) rendered
two boats. Pi replaces widgets under one key, so claiming the shared
"calm-working-ship" slot keeps dual-install sessions to a single boat
while a Firstmate-only session is unchanged.

Pins the shared slot contract in the working-ship module test so the key
cannot silently diverge again.

* test(calm): pin the shared working-ship widget key in CI, document dual-install

The key-parity assertion inside the Pi fixture only runs where the
@earendil-works/pi-coding-agent package is installed, so CI never
exercised it. Add a source-level twin that needs nothing but the
tracked file, and note in docs/calm.md that the boat shares the
standalone Pi Calm working-row widget slot.

* no-mistakes(review): Add executable dual-install widget replacement coverage

* no-mistakes(review): Guard shared widget cleanup with disposal ownership

* no-mistakes(document): Document shared Calm working-ship slot behavior

* test(calm): read the standalone Calm slot from its own module

The dual-install check registered both boats itself under the shared
slot, so it could only prove that Pi replaces a widget under one key: it
would still pass if the standalone Pi Calm extension installed its boat
under a different key, which is the two-boat regression the check exists
to prevent.

Read the standalone extension's own working-ship module when it is
installed - FM_STANDALONE_CALM_SHIP, else ~/.pi/agent/extensions/calm -
and drive the check with the key that module exports, so a rename on
either side registers two widgets and fails naming both keys. A pinned
shared-slot contract still covers a machine without the extension, and
the run reports which side it used instead of passing silently over an
absent extension.

Verified: the touched Pi Calm suite passes and reads the installed
standalone extension; with a copy of it whose key is renamed to
calm-working-ship-v2 the suite fails naming the drift.

* no-mistakes(review): Gate stock-row restoration by shared-widget ownership

* no-mistakes(review): Removed redundant widget-key source assertions

* no-mistakes(document): Document shared Calm working-ship widget ownership

* feat(bin): add opt-in --herdr-resume-lock-wait to fm-spawn (#6649)

* fix(herdr): make exact-resume presentation-lock wait instead of a bounded timeout

The exact-resume path in bin/fm-spawn.sh used the same 50-attempt-then-
give-up lock acquire as the new-task-create path, but the two paths are
not equivalent on contention: a create has no prior state to strand and
can safely fall back to a flat layout, while a resume is recovering a
specific existing identity that a concurrent recovery may legitimately
be holding the lock for. Giving up there does not degrade gracefully,
it hard-fails the resume outright. The suite's own concurrent
cross-home recoveries test already asserts both concurrent recoveries
succeed with a genuine reclaim, and the file's header comment already
(inaccurately) claimed lock contention falls back to the ordinary flat
layout for both paths alike, so the intended contract was always that
recoveries serialize and both succeed, not that either one refuses
under a short bound.

Give spawn_herdr_presentation_order_lock_acquire a wait mode that uses
this file's own established fm_lock_acquire_wait idiom (already used
for its other fleet-shared locks) instead of the bounded loop, and use
it only at the exact-resume call site. The new-task-create call site
is unchanged and keeps its bounded-then-flat-fallback behavior, which
is already covered by its own passing test. Dead-owner PID-liveness
reclaim inside fm_lock_try_acquire still bounds the wait against a
holder that crashed mid-hold.

Adds a deterministic regression test that holds the shared session
lock from an unrelated process for well past the old bound, then
asserts the resume succeeds with a genuine reclaim and took close to
the full hold duration, so a fix that merely widens the bound rather
than genuinely waiting is still caught. The existing concurrent
cross-home recovery test exercises this under real timing but does not
reliably outlast a fixed bound on its own.

Corrects the header comment's claim that create and resume share one
bounded-then-flat-fallback behavior on lock contention; they no longer
do.

* no-mistakes(document): Document Herdr recovery waiting for presentation lock

* no-mistakes(document): Update stale hard-refusal claim in verification log

* no-mistakes(ci): Fixed the Greptile finding on tests/fm-backend-herdr-presentation-e2e.test.sh:1389 by bounding the resume lock-wait regression's spawn_task call. Added an optional 4th `deadline_seconds` arg to the `spawn_task` helper (defaults to empty, so all ~20 other existing call sites are unaffected and unwrapped by `timeout`). The lock-wait test now passes `LOCK_WAIT_HOLD_SECONDS + 60` (90s) as the deadline, and a dedicated check for exit code 124 emits a clear "hung for over Xs instead of waiting out a Ys lock hold" diagnostic before falling through to the existing pass/fail assertions, which are unchanged. No product code was touched. Verified with `bash -n`, `shellcheck -x` (no warnings), a standalone reproduction of the timeout/no-timeout/success paths, the project's `bin/fm-lint.sh --fast` on the file (clean), and the full `tests/fm-lint.test.sh` suite (all 46 assertions pass)

* no-mistakes(ci): Replaced the direct `timeout "$deadline_seconds"` call in `spawn_task()` (tests/fm-backend-herdr-presentation-e2e.test.sh) wi…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant