Skip to content

fix(bin): merge upstream main through fb75c1f9 and record the squashed #43 sync as merged - #104

Merged
MrGTV-love merged 80 commits into
mainfrom
fm/fm-upstream-sync-a2
Oct 11, 2026
Merged

MrGTV-love merged 80 commits into
mainfrom
fm/fm-upstream-sync-a2

Conversation

@MrGTV-love

@MrGTV-love MrGTV-love commented Oct 11, 2026 •

Copy link
Copy Markdown
Owner

Intent

firstmate, please compare the upstream Firstmate repo with fork, and commits awaiting merge on that fork. How do we get in sync

yes A2

Context for reading that ask. The comparison (data/fm-upstream-sync/report-20261010.md in the firstmate home ~/firstmate) found: fork origin/main d69b2c5 and upstream/main fb75c1f (kunchenguid/firstmate) share merge-base eb219c8; fork PR #43 (fd9b8b0) already brought in upstream through 2ce57d0 (upstream PRs kunchenguid#6064..kunchenguid#6784) as one squashed commit, so git does not know it, and a plain merge conflicts in 80 files / 220 hunks. The real merge history survives on origin/fm/fm-upstream-sync (e24c85a), whose tree is byte-identical to fd9b8b0 (tree 00cbc3bd); 2ce57d0 is its ancestor. Option A2, chosen: on a branch from origin/main, record #43 as a real merge with git merge -s ours --no-ff 2ce57d00 (no file change; the tree must stay identical to origin/main's), then git merge upstream/main to bring in the 7 new upstream commits (329ad4e kunchenguid#6814, 3c58ec8 kunchenguid#6787, 0ec1c5a kunchenguid#6823, b062eb9 kunchenguid#6818, 19fcbbd kunchenguid#6887, b75658b kunchenguid#6920, fb75c1f kunchenguid#6956), resolving the 5 conflicted files: tests/fm-supervision-host.test.sh (move upstream kunchenguid#6787's new test into the fork's split helpers/hook layout), tests/fm-control-herdr-smoke.test.sh (upstream fm_agent_standin vs fork's claude-named stand-in), bin/backends/herdr.sh (keep both the </dev/null server start and the COMPACT_ADVISER env names), docs/verification/runtime-backends.md, CONTRIBUTING.md (union of helper lists). The PR must merge as a merge commit, never squashed, so later syncs see the shared history.

What Changed

Risk Assessment

✅ Low: The change is an upstream sync that records #43 as a real no-change merge (tree 104878b5 equals origin/main's) and brings in seven upstream commits, with all five conflicts resolved as the intent specifies and no defect found in the fork-local adaptations.

Testing

I checked the A2 merge history again on HEAD c3df226 and wrote the results to history-sync-round2.txt. The ours-merge tree is the same as origin/main. All 8 upstream commits are ancestors of HEAD. A merge of upstream/main into HEAD is a clean no-op, and the same merge into the old origin/main still conflicts in 81 files. Since round 1, the only change is the test file, so the round-1 live Herdr probe still applies to the product code. I ran the fixed live Claude attended e2e against the real Claude CLI in its own isolated lab tmux socket, and it fully passed (ok, exit 0). Event 1 is now a main-only pass-through, which was the round-1 failure point. The turn end takes over the successor's cycle, and events 2 and 3 reach the idle primary. The late main-only close (event 4) is handed back with no engine turn, and no captain prompt is sent after setup. The lab cleaned itself up, and the worktree is clean. This change has no UI surface, so there are no screenshots.

  • Live validation: ✅ go - 10 of 10 scenarios driven live against the product
Scenario Result Live Evidence
Upstream #43 is recorded as a real merge with no file change (the 4b9be5e tree equals the d69b2c5 tree) ✅ pass live history-sync-round2.txt: both trees are 104878b5
All 7 new upstream commits and 2ce57d0 are ancestors of the sync branch ✅ pass live history-sync-round2.txt: all 8 commits are ancestors
The next upstream sync sees shared history: merging upstream/main into HEAD is a clean no-op ✅ pass live history-sync-round2.txt: merge-base = fb75c1f; merge-tree exit 0, result tree = HEAD tree
Adversarial: the same merge into the pre-A2 origin/main still conflicts, so the A2 history is what removes the conflicts ✅ pass live history-sync-round2.txt: exit 1, 81 conflicted files
Herdr probes stay observable-only and do not autostart a stopped lab server (kunchenguid#6956/b75658b6), and the lifecycle server starts with stdin from /dev/null ✅ pass live live-herdr-observable-probe.txt (round 1; bin/backends/herdr.sh has not changed since then)
Live Claude attended: event 1 passes through as main-only, leaves a live successor watcher, and the Stop hook rewake wakes the idle primary, which drains and acknowledges it (kunchenguid#6787 hand-back, round-1… ✅ pass live live-claude-attended-round2.log: 21:45:32 'pass-through attended main-only'; 21:45:33 'successor watcher pid 13466 alive'; 21:45:52 'rewake delivered ... outcome=rewake'; 'primary turn ran: bin/fm-wak…
Live Claude attended: the event 1 turn end takes over the successor's cycle, and event 2 is passed through and woken through the cycle that took over ✅ pass live live-claude-attended-round2.log: 21:45:56 'successor 13466: reason=taken-over ... now owning watcher 58589'; 21:46:08 'watcher 58589 closed: ... reason=actionable-signal'; 'its close was delivered: re…
Live Claude attended: the remote-reply listener delivers event 3 and stays owned throughout ✅ pass live live-claude-attended-round2.log: 21:46:54 'listener mirrored it (1 ingested) and it was delivered ... runner 4203 -> 4203, owned at every check'; 21:47:28 'listener runner 4203 still owned'
Live Claude attended: a routine close accepted for handling that turns main-only before its turn (event 4) is handed back with no engine turn, woken, drained and acknowledged, and its turn end re-arms ✅ pass live live-claude-attended-round2.log: 21:47:06 'pass-through attended main-only; no engine turn'; 'rewake delivered ... outcome=rewake'; 21:47:24 '--ack-through 10'; 21:47:28 'turn end re-armed ... watcher…
Live Claude attended e2e fully passes (no captain prompt is submitted after setup, the test prints ok, exit 0) ✅ pass live live-claude-attended-round2.log: 21:47:28 'captain prompts after setup: 0'; 'ok - attended live (2.1.296 (Claude Code)) ...'; 'exit=0'
Evidence: A2 history checks on HEAD c3df226

Source: A2 history checks on HEAD c3df2263

$ git log --first-parent --format="%h %p | %s" -7 HEAD
c3df2263 a095da77 | no-mistakes(test): test: align attended live e2e with lab setup and take-over
a095da77 d6613ba3 60775fa1 | merge: bring origin/main forward into the upstream sync branch
d6613ba3 20eaa665 8f9ac500 | merge: bring origin/main forward into the upstream sync branch
20eaa665 26fe5ffc | test: retry the interrupted winner read in the stale-lock steal race
26fe5ffc 41fbccd8 | fix: restore two upstream hunks the squashed sync (#43) dropped
41fbccd8 4b9be5e0 fb75c1f9 | merge: bring in upstream main through fb75c1f9 (upstream #6814, #6787, #6823, #6818, #6887, #6920, #6956)
4b9be5e0 d69b2c5a 2ce57d00 | merge: record upstream through 2ce57d00 as merged (content landed squashed in #43, fd9b8b02; real merge history on fm/fm-upstream-sync e24c85ae, identical tree 00cbc3bd)

## ours-merge 4b9be5e0 tree vs parent d69b2c5a
104878b5d2c15a39701cbe288be5a3579ed16bcc 104878b5d2c15a39701cbe288be5a3579ed16bcc

## Upstream commits are ancestors of HEAD
2ce57d00 ancestor
329ad4ed ancestor
3c58ec8e ancestor
0ec1c5a4 ancestor
b062eb94 ancestor
19fcbbde ancestor
b75658b6 ancestor
fb75c1f9 ancestor

## Next sync: merge-base(HEAD, fb75c1f9) = fb75c1f9
## merge-tree HEAD + upstream fb75c1f9
exit=0 result_tree=e605e1d34ff323c86ac8920d28dc08fa88ba83c0 head_tree=e605e1d34ff323c86ac8920d28dc08fa88ba83c0
## Adversarial: merge-tree old origin/main 60775fa1 + fb75c1f9 (pre-A2)
exit=1 conflicted_files=81

## Diff HEAD vs a095da77 (round-1 commit): only test file
 .../fm-supervision-host-attended-live-e2e.test.sh  | 38 ++++++++++++++--------
 1 file changed, 24 insertions(+), 14 deletions(-)
Evidence: Live Claude attended e2e transcript, round 2 (ok, exit 0)

Source: Live Claude attended e2e transcript, round 2 (ok, exit 0)

skip: control: set FM_SUPERVISION_HOST_ATTENDED_LIVE_CONTROL_REF to a pre-fix ref to run the negative control
# 21:45:28 positive step 1: primary idle (claude pid 73105, 2.1.296 (Claude Code), model haiku); tracked Stop hook registered; config/supervision-host present; host pid 99647 parked on watcher 1421; listener runner 4203; captain prompts so far: 1
# 21:45:29 positive transcript: ~/.claude/projects/-Users-charlesabrooker-tmp-fm-sh-attended-live-z9c1Zm-positive-fm/3697405c-2e59-4b3b-ac5d-841aef292335.jsonl
# 21:45:29 positive step 2: event 1 appended at 1791686729 (demo.status needs-decision)
# 21:45:32 positive step 2: host log: 1791686732	pass-through	attended	main-only
# 21:45:33 positive step 2: successor watcher pid 13466 alive; ledger: epoch=1 owner_pid=98137 outcome=arming updated_at=1791686726; marker: pending:downtime:1421.1791686730.2Rf3ra
# 21:45:52 positive step 2: rewake delivered at 1791686734 (Stop hook exited 2 with the banner); ledger: epoch=1 owner_pid=98137 outcome=rewake updated_at=1791686734 session_pid=73105 recovery_generation=1421.1791686730.2Rf3ra
# 21:45:52 positive step 2: primary turn ran: bin/fm-wake-drain.sh;bin/fm-wake-drain.sh --ack-through 2 --recovery-generation 1421.1791686730.2Rf3ra;
# 21:45:56 positive step 3: turn end re-armed: 1791686754	start	gen=host-54998-1791686754; take-over arm=13308; successor 13466: reason=taken-over successor=none; now owning watcher 58589
# 21:45:59 positive step 3: event 2 appended at 1791686759
# 21:46:08 positive step 3: watcher 58589 closed: arm_pid=56219 watcher_pid=58589 origin=started started_at=1791686755 ended_at=1791686760 exit_code=0 signal=none reason=actionable-signal
# 21:46:08 positive step 3: its close was delivered: rewake at 1791686763; host log: 1791686761	pass-through	attended	main-only
# 21:46:22 positive step 4: event 3 appended to the stand-in remote log at 1791686782
# 21:46:54 positive step 4: listener mirrored it (1 ingested) and it was delivered at 1791686789; listener runner 4203 -> 4203, owned at every check
# 21:46:59 positive step 5: event 4, a routine working line on task demo2, appended at 1791686819
# 21:47:05 positive step 5: the host accepted it and confirmed its successor's handling handoff (marker announced:handling:2900.1791686821.4YGAEz at 1791686825); a needs-decision on demo2 landed then, before the turn's start
# 21:47:06 positive step 5: host log: 1791686825	pass-through	attended	main-only; no engine turn
# 21:47:06 positive step 5: rewake delivered at 1791686826; ledger: epoch=4 owner_pid=94019 outcome=rewake updated_at=1791686826 session_pid=73105 recovery_generation=2900.1791686821.4YGAEz
# 21:47:24 positive step 5: primary turn ran: bin/fm-wake-drain.sh;bin/fm-wake-drain.sh --ack-through 10 --recovery-generation 2900.1791686821.4YGAEz;
# 21:47:28 positive step 5: turn end re-armed: 1791686847	start	gen=host-70465-1791686847; watcher 71390 live; listener runner 4203 still owned
# 21:47:28 positive: captain prompts after setup: 0 (mirror holds only the setup prompt)
ok - attended live (2.1.296 (Claude Code)): an idle primary is woken for four hand-offs, a take-over of the successor's cycle and a close that turned main-only at its turn included, with the listener owned throughout
exit=0
Evidence: Live attended e2e final result
ok - attended live (2.1.296 (Claude Code)): an idle primary is woken for four hand-offs, a take-over of the successor's cycle and a close that turned main-only at its turn included, with the listener owned throughout
exit=0
Evidence: Round 1 live Herdr observable-probe transcript

Source: Round 1 live Herdr observable-probe transcript

lab session: fm-lab-probe-39486-13473

== 1. Probe a target whose lab session has NO server (prepared only)
session before: []
visible_capture rc=1 elapsed=0s
composer_state: unknown
session after: []  (expect empty or running:false - no server autostarted)

== 2. Provision the lab server and probe a real pane
pane=w1:p2
visible_capture rc=0; marker seen: 3 line(s)
composer_state: unknown
session: [{"name":"fm-lab-probe-39486-13473","running":true}]

== 3. Stop the lab server (guarded) and probe the same pane again
stop rc=0
session after stop: [{"name":"fm-lab-probe-39486-13473","running":false}]
visible_capture rc=1 elapsed=0s
composer_state: unknown
session after probes: [{"name":"fm-lab-probe-39486-13473","running":false}]  (expect running:false - probe did not restart the server)

== 4. Contrast: lifecycle path (target_ready) still autostarts by design
caller env: COMPACT_ADVISER_DISABLE=1 FM_COMPACT_ADVISER_DISABLE=1 FM_COMPACT_ADVISER_HOOKS=probe 
target_ready rc=0
session: [{"name":"fm-lab-probe-39486-13473","running":true}]
server pid=58820; stdin (fd 0) of the server started by fm_backend_herdr_server_ensure:
  fd0 -> /dev/null
pid 58820: herdr server --session fm-lab-probe-39486-13473
  matching env: []
teardown rc=0
- Outcome: 🔧 3 issues found → auto-fixed ✅ across 2 runs (59m9s)

Sync notes

Merge as a merge commit. Never squash or rebase: later upstream syncs need the shared history.

Commits

  1. 4b9be5e - git merge -s ours --no-ff 2ce57d00. Records upstream through 2ce57d0 (upstream fix: treat quiet records as attended across supervision kunchenguid/firstmate#6064..docs(skills): add six missing skills to agent skill trigger index kunchenguid/firstmate#6784) as merged. No file change: its tree equals d69b2c5's tree. The content already landed squashed in feat: sync Firstmate upstream capabilities and reliability fixes #43 (fd9b8b0); the real merge history is on fm/fm-upstream-sync (e24c85a, tree 00cbc3bd).
  2. 41fbccd - git merge upstream/main through fb75c1f: upstream test: make agent process fixtures compatible with multicall sleep kunchenguid/firstmate#6814, fix(bin): keep the supervision host's successor watcher alive after the Stop hook's group is torn down kunchenguid/firstmate#6787, test: make tmux Claude readiness checks independent of permission footers kunchenguid/firstmate#6823, test: correct Herdr smoke assertions and cleanup kunchenguid/firstmate#6818, test: stabilize watcher lock race fixtures kunchenguid/firstmate#6887, fix(bin): keep herdr composer probes observable-only without autostarting server kunchenguid/firstmate#6920, fix(tests): clear inherited Git repository location in fixture helper kunchenguid/firstmate#6956. Conflicts resolved in 5 files:
  3. 26fe5ff - restores 2 upstream hunks that the squashed sync (feat: sync Firstmate upstream capabilities and reliability fixes #43) dropped:
    • bin/fm-control.sh: the picker comment line.
    • .agents/skills/afk/SKILL.md line 64.
  4. 20eaa66 - fork-local fix to the text of upstream test: stabilize watcher lock race fixtures kunchenguid/firstmate#6887's test test_lock_stale_steal_single_winner_under_concurrency (tests/fm-watcher-lock.test.sh). The FIFO open of the winner gate can be interrupted (EINTR) and bash does not retry it; the empty winner then deadlocked the parent and the winning contender. Now: bounded retry of the read; if no winner, kill the waiter and contenders, then fail. Test-only.
  5. d6613ba and a095da7 - ordinary merges of origin/main (8f9ac50, fork fix(bin): switch workers to declared model-matrix stand-ins when pooled OMP Codex quota is exhausted #91..fix(bin): raise Claude turn-end guard auto-arm wait to 12s from measured claim latency #102; then 60775fa, fork feat(bin): report worker and reviewer models older than the newest in their harness catalog #103) forward. No conflicts.
  6. c3df226 (pipeline test fix) - tests/fm-supervision-host-attended-live-e2e.test.sh, opt-in live test only: the lab marks its empty task status logs as read (an unmarked log is a fresh signal, so the host ran an engine turn that absorbed event 1), and step 3 now expects the turn end to take over the successor's cycle (fm-watch-arm.sh --take-over, which came in with feat: sync Firstmate upstream capabilities and reliability fixes #43) instead of attaching to it. This test failed the same way on origin/main 60775fa before the fix; with it, the live run against real Claude passes all 5 steps.
  7. ddb416c (pipeline doc fix) - docs/supervision-host.md: the table line for that test says "a take-over of the successor's cycle".

Review aid

See the "Review aid" section below: for each upstream-changed file in eb219c8..2ce57d0, whether origin/main differs and why.

Known pre-existing host failures (not caused by this PR)

  1. tests/fm-supervision-host.test.sh timing flake under high load: A/B merged 8/9 vs origin/main 8/9, same signature.
  2. tests/fm-watcher-lock.test.sh test_arm_self_eviction_is_loud_without_successor: 0/3 on branch, 0/3 on base d69b2c5, same message "self-evicted arm did not fail nonzero (status 124)".
  3. tests/fm-watcher-lock.test.sh test_arm_waits_for_peer_beacon_after_child_stands_down: 0/3 on branch, 0/3 on base d69b2c5, same message "attached arm did not fail after peer died (status 124): watcher: already running pid N".
  4. tests/fm-watcher-lock.test.sh test_wait_deadline_reaps_a_stopped_child: hit its 15 s outer bound under load average 80-140 in the pipeline test step, and fails the same way on base 60775fa.

Cases 2 and 3 are tracked as their own fix item; case 4 was first seen in the pipeline test step. Nothing is waived: Linux CI passed all 23 checks on ddb416c.

Review aid: upstream files changed between eb219c8 and 2ce57d0

Upstream changed 210 files in that range. For each file, this states whether origin/main (d69b2c5) differs from upstream at 2ce57d0, and why.

Method: (1) byte compare each file; (2) replay upstream's eb219c8..2ce57d0 change onto the fork file with a three-way merge, which shows upstream hunks the fork lacks; (3) read git show --remerge-diff for all six merges on fm/fm-upstream-sync (149 hand-resolved conflict regions plus edits made outside conflicts) and check every upstream line a resolution removed against the current fork tree.

Missing upstream hunks found: 2, both ported in this PR

  1. bin/fm-control.sh - the first line of the picker comment before the dialog check (# The submitting Enter can open the picker. The agent is still alive, and).
  2. .agents/skills/afk/SKILL.md - the supervision-host bullet condition: with config/supervision-host becomes that runs the supervision host.

No dropped code hunk was found. Every other difference is a deliberate fork change, listed below.

A. Identical to upstream (57 files)

.agents/skills/away-quiet-supervision/SKILL.md, .agents/skills/captain-hold-lifecycle/SKILL.md, .agents/skills/project-management/SKILL.md, .agents/skills/quiet/SKILL.md, .agents/skills/scout-completion/SKILL.md, .opencode/plugins/fm-primary-watch-arm.js, .pi/extensions/fm-calm.ts, .pi/extensions/lib/fm-calm-working-ship.ts, bin/fm-branch-report.sh, bin/fm-contributions.jq, bin/fm-exclude-tools-lib.sh, bin/fm-fleet-sync.sh, bin/fm-hold-reason-lib.sh, bin/fm-host-mirror.sh, bin/fm-pipeline-spend.sh, bin/fm-pr-check.sh, bin/fm-pr-merge.sh, bin/fm-primary-scope-lib.sh, bin/fm-procevent-quota.sh, bin/fm-project-capacity-lib.sh, bin/fm-promote.sh, bin/fm-quota-axi-lib.sh, bin/fm-remote-delta-read.sh, bin/fm-remote-job-worker.sh, bin/fm-review-diff.sh, bin/fm-sessionstart-run.sh, bin/fm-startup-growth-check.sh, bin/fm-supervision-host.sh, bin/fm-supervision-instructions.sh, bin/fm-turnend-guard-cursor.sh, bin/fm-watch-arm.sh, bin/fm-watch-checkpoint.sh, docs/secondmate-parent-channel.md, docs/supervision-protocols/claude.md, tests/fm-dod-lib.test.sh, tests/fm-fleet-sync.test.sh, tests/fm-host-mirror.test.sh, tests/fm-live-lab-up-mate.test.sh, tests/fm-live-lab.test.sh, tests/fm-parent-channel-scan-exclusion.test.sh, tests/fm-pi-seeded-home-trust-live-e2e.test.sh, tests/fm-pipeline-spend.test.sh, tests/fm-pr-merge.test.sh, tests/fm-procevent-quota.test.sh, tests/fm-quota-choose.test.sh, tests/fm-remote-job-claim-reap.test.sh, tests/fm-remote-job-claim-retention.test.sh, tests/fm-remote-job-launchagent.test.sh, tests/fm-remote-job-orphan-reap.test.sh, tests/fm-review-diff.test.sh, tests/fm-shared-captain-inheritance.test.sh, tests/fm-spawn-pool-base-freshen.test.sh, tests/fm-startup-growth-check.test.sh, tests/fm-startup-memory-budget.test.sh, tests/fm-supervision-instructions.test.sh, tests/fm-test-fixture-cleanup.test.sh, tests/fm-x-mode.test.sh

B. Differs; every upstream hunk is present; the fork changed other parts of the file (76 files)

File Why it differs
.agents/skills/bearings/SKILL.md fork after sync: #63
.agents/skills/process-event-sources/SKILL.md fork before sync: #9 #38; sync-branch fix commits: 70383df
.agents/skills/secondmate-provisioning/SKILL.md fork before sync: #13 #29 #44
.claude/mods/firstmate-calm/.claude-plugin/plugin.json fork before sync: #34
.claude/mods/firstmate-calm/hooks/register.ts fork before sync: #34
.omp/extensions/fm-primary-omp-watch.ts fork before sync: #48 #54; fork after sync: #73 #81 #22
AGENTS.md fork before sync: #4 #19 #13 #29 #45; fork after sync: #59
README.md fork before sync: #29 #41
bin/fm-afk-contract.sh fork before sync: #19
bin/fm-afk-launch.sh fork before sync: #29
bin/fm-afk-return.sh fork before sync: #19; fork after sync: #76
bin/fm-backend.sh fork before sync: #45; fork after sync: #67 #69 #22
bin/fm-backlog-transition-lib.sh Upstream content present; the fork places fm_backlog_pr_is_gerrit_change earlier in the file. fork before sync: #15 #26 #45 #51; fork after sync: #83
bin/fm-bootstrap.sh fork before sync: #11 #13 #29 #42 #44 #45 #51; fork after sync: #22
bin/fm-brief-heading-lib.sh fork before sync: #17 #24 #49
bin/fm-contributions.sh fork after sync: #61 #88
bin/fm-dispatch-resolve.sh fork before sync: #11 #16 #17 #24 #13 #42 #49
bin/fm-lease-lib.sh fork after sync: #76
bin/fm-nm-run-lib.sh Upstream content present; the fork places fm_nm_state_db elsewhere in the file. fork before sync: #47; fork after sync: #69
bin/fm-parent-channel-lib.sh fork before sync: #19
bin/fm-procevent-remote-reply.sh fork before sync: #19; sync-branch fix commits: 4afdea9
bin/fm-procevent.sh fork before sync: #9; sync-branch fix commits: 4afdea9; fork after sync: #78
bin/fm-remote-home-provision.sh fork before sync: #10; 1 hand-resolved conflict region(s) in the sync
bin/fm-remote-job-lib.sh sync-branch fix commits: 4afdea9
bin/fm-session-start.sh fork before sync: #52 #51; fork after sync: #67 #22
bin/fm-supervise-daemon.sh fork before sync: #19; sync-branch fix commits: 4afdea9; fork after sync: #73 #67
bin/fm-supervision-engine-lib.sh fork before sync: #29
bin/fm-task-inbox-lib.sh fork after sync: #22
bin/fm-timeout-lib.sh fork before sync: #4 #47; 1 hand-resolved conflict region(s) in the sync
bin/fm-wake-drain.sh fork before sync: #19 #52; fork after sync: #76 #79 #22
docs/herdr-backend.md fork before sync: #3 #12 #21 #25 #33 #27 #34 #37 #48; sync-branch fix commits: 70383df; fork after sync: #22
docs/pi-supervision-branch.md fork before sync: #26 #19
docs/sessionstart-nudge.md fork before sync: #52 #51; fork after sync: #78 #22
docs/supervision-host.md fork before sync: #54; sync-branch fix commits: df845c5; fork after sync: #66
docs/supervision-protocols/omp.md fork before sync: #54; fork after sync: #73
docs/supervision-protocols/supervision-host.md fork before sync: #54
docs/turnend-guard.md fork before sync: #4 #6 #9 #31 #48 #52; fork after sync: #62 #66
docs/verification/process-event-sources.md fork before sync: #12 #9 #38 #53; sync-branch fix commits: 4afdea9; fork after sync: #63 #78 #83
docs/verification/supervision.md fork before sync: #6 #47 #48; sync-branch fix commits: df845c5; fork after sync: #67
tests/fm-backend.test.sh fork after sync: #67
tests/fm-backlog-atomicity.test.sh fork before sync: #15 #47 #45 #51
tests/fm-backlog-read-bound.test.sh fork before sync: #47 #45; fork after sync: #78
tests/fm-bearings-board-render.test.sh fork before sync: #45
tests/fm-bootstrap.test.sh fork before sync: #11 #13 #47; fork after sync: #22
tests/fm-branch-supervision.test.sh fork before sync: #47 #45
tests/fm-brief.test.sh fork before sync: #17; fork after sync: #60 #66 #84
tests/fm-calm-claude-mod-live-e2e.test.sh fork before sync: #34
tests/fm-calm-claude-mod.test.sh fork before sync: #34; 1 hand-resolved conflict region(s) in the sync
tests/fm-classify-corr-token.test.sh fork before sync: #19
tests/fm-control.test.sh fork before sync: #4 #12 #46; sync-branch fix commits: 0fced65; fork after sync: #22
tests/fm-cursor-primary.test.sh fork before sync: #42 #47
tests/fm-dispatch-resolve.test.sh fork before sync: #11 #16 #17 #20 #24 #13
tests/fm-extension-binding.test.sh fork before sync: #42 #53
tests/fm-git-strip-ai-trailers.test.sh fork before sync: #23; fork after sync: #22
tests/fm-pi-branch-extension.test.sh fork before sync: #10 #19 #42
tests/fm-pi-branch-live-e2e.test.sh fork after sync: #73
tests/fm-pi-watch-extension.test.sh Upstream content present; the fork places the eight upstream test functions elsewhere in the file. fork after sync: #73
tests/fm-procevent.test.sh fork before sync: #9 #38
tests/fm-remote-reply.test.sh fork before sync: #12
tests/fm-secondmate-harness.test.sh fork before sync: #7 #13
tests/fm-secondmate-liveness.test.sh fork before sync: #29; fork after sync: #22
tests/fm-secondmate-sync.test.sh fork after sync: #22
tests/fm-send-inbox-doorbell-live-e2e.test.sh fork after sync: #22
tests/fm-send-inbox.test.sh fork after sync: #67 #22
tests/fm-session-lock-ancestry.test.sh fork before sync: #42 #47
tests/fm-session-start.test.sh fork before sync: #47 #52 #51; fork after sync: #67 #22
tests/fm-sessionstart-nudge.test.sh fork before sync: #47
tests/fm-task-delivery.test.sh fork after sync: #59
tests/fm-teardown-endpoint-safety.test.sh fork before sync: #47 #45 #55; fork after sync: #87
tests/fm-turnend-guard.test.sh fork before sync: #6 #31 #42 #47; fork after sync: #62
tests/fm-wake-drain-outcome-backstop.test.sh fork after sync: #79
tests/fm-wake-drain-unread-status.test.sh fork after sync: #79
tests/fm-watch-arm.test.sh fork before sync: #9
tests/fm-watch-checkpoint.test.sh fork after sync: #22
tests/fm-watch-triage.test.sh fork before sync: #9 #47; fork after sync: #77 #85
tests/lib.sh fork before sync: #8 #7 #34 #47 #45 #51; fork after sync: #76 #88 #84

C. Differs; the fork changed text that upstream also changed (77 files)

Deliberate fork change unless a row says otherwise. A row with no note had no upstream line removed by a conflict resolution.

File Why it differs
.agents/skills/afk/SKILL.md Missing upstream hunk, ported in this PR: the supervision-host bullet kept the old with config/supervision-host condition; upstream says that runs the supervision host. The rest of the bullet is the fork's shorter wording with a pointer to the supervision protocol. fork before sync: #19 #48 #54; 1 hand-resolved conflict region(s) in the sync; sync-branch fix commits: 70383df; fork after sync: #22
.agents/skills/agent-skill-trigger-index/SKILL.md fork before sync: #35 #51; 1 hand-resolved conflict region(s) in the sync; fork after sync: #59
.agents/skills/harness-adapters/references/harness/claude.md Upstream wording taken; fork keeps its own relative link path. fork before sync: #5 #12 #6 #31 #24 #42; 1 hand-resolved conflict region(s) in the sync; fork after sync: #65
.agents/skills/harness-adapters/references/harness/codex.md Upstream wording taken; fork keeps its own relative link path. fork before sync: #24; 1 hand-resolved conflict region(s) in the sync
.agents/skills/harness-adapters/references/harness/cursor.md Upstream wording taken; fork keeps its own relative link path. fork before sync: #24; 1 hand-resolved conflict region(s) in the sync
.agents/skills/harness-adapters/references/harness/grok.md Upstream wording taken; fork keeps its own relative link path. fork before sync: #24; 1 hand-resolved conflict region(s) in the sync
.agents/skills/harness-adapters/references/harness/omp.md Upstream wording taken; fork keeps its own relative link path, its "tracked supervision extensions" wording, and a pointer to watcher continuity. fork before sync: #28 #33 #31 #24 #13 #48 #54; 2 hand-resolved conflict region(s) in the sync; fork after sync: #81 #22
.agents/skills/harness-adapters/references/harness/opencode.md Upstream wording taken; fork keeps its own relative link path. fork before sync: #24 #48; 1 hand-resolved conflict region(s) in the sync
.agents/skills/harness-adapters/references/harness/pi.md Fork keeps its own relative link path; text otherwise upstream. fork before sync: #10
.agents/skills/operational-home-layout/SKILL.md fork before sync: #4 #8 #15 #19 #24 #34 #13 #42 #44 #49 #51; 1 hand-resolved conflict region(s) in the sync; sync-branch fix commits: 4afdea9; fork after sync: #73
.agents/skills/session-start-recovery/SKILL.md fork before sync: #29; fork after sync: #67 #22
.pi/extensions/fm-branch-supervision.ts Fork builds both supervision tool renderers from one helper (legacyStockRenderers, stockToolRenderers), so upstream's per-tool stockToolCallHeader(...) call lines appear once as stockToolCallHeader(definition.name, ...). fork before sync: #10; 3 hand-resolved conflict region(s) in the sync; sync-branch fix commits: 4d12456
.pi/extensions/fm-primary-pi-watch.ts fork after sync: #73
bin/backends/herdr.sh Fork selects the ghost-colour ceiling by agent identity (fm_backend_herdr_composer_ghost_luma, 0 for Claude), which covers upstream's Claude-only FM_COMPOSER_GHOST_LUMA_MAX=0. Also fork compact-adviser env names. fork before sync: #12 #25 #34 #48 #51; 6 hand-resolved conflict region(s) in the sync; fork after sync: #22
bin/fm-brief.sh fork before sync: #17 #19; 1 hand-resolved conflict region(s) in the sync; fork after sync: #60 #66 #84
bin/fm-captain-hold.sh fork after sync: #76 #78 #83
bin/fm-classify-lib.sh Fork split this library; the upstream-added comment lines live in bin/fm-status-record-lib.sh. fork before sync: #19 #51; 1 hand-resolved conflict region(s) in the sync; sync-branch fix commits: 4d12456; fork after sync: #73 #79
bin/fm-claude-stop-autoarm.sh fork before sync: #9 #52; 1 hand-resolved conflict region(s) in the sync; fork after sync: #66
bin/fm-claude-trust.sh fork before sync: #5; 1 hand-resolved conflict region(s) in the sync; fork after sync: #65
bin/fm-composer-lib.sh Fork has its own fix for session-titled Claude composer rules (fork #37, width-matched titled rule). Upstream's parallel fix (_fm_composer_bare_rule_sandwich) was left out during the sync. Behavior difference: a typed draft in a titled composer on a plain backend reads pending in the fork and unknown upstream. fork before sync: #12 #26 #33 #34 #37 #48; 2 hand-resolved conflict region(s) in the sync; sync-branch fix commits: 4afdea9 a9a9f98; fork after sync: #22
bin/fm-config-inherit-lib.sh Union: upstream list plus the fork's model-index.json and compact-adviser items. fork before sync: #8 #17 #34 #13 #29 #44; 4 hand-resolved conflict region(s) in the sync
bin/fm-control.sh Missing upstream hunk, ported in this PR: first line of the picker comment before the dialog check. Other differences are fork changes. fork before sync: #2 #4 #12 #21 #8 #7 #26 #13 #29 #46; 6 hand-resolved conflict region(s) in the sync; fork after sync: #76 #67 #87 #22
bin/fm-dod-lib.sh fork before sync: #26 #49; sync-branch fix commits: 4afdea9; fork after sync: #59
bin/fm-fleet-snapshot.sh fork before sync: #19 #45 #51; fork after sync: #61
bin/fm-git-strip-ai-trailers.sh fork before sync: #23; 1 hand-resolved conflict region(s) in the sync
bin/fm-lint.sh Fork runs each root through its lint cache (analysis_command); upstream's memory fallback is wired into that path, so upstream's direct "$FM_LINT_SHELLCHECK" "$@" lines and its changed-file-mode comments have no place. fork before sync: #19; 7 hand-resolved conflict region(s) in the sync; sync-branch fix commits: 70383df; fork after sync: #71
bin/fm-live-lab.sh sync-branch fix commits: 70383df; fork after sync: #76 #66
bin/fm-pending-reply-lib.sh Fork sources the wake library through _fm_pending_reply_source_wake; upstream's inline source and its comment were replaced. fork before sync: #9 #19; 1 hand-resolved conflict region(s) in the sync
bin/fm-procevent-lavish.sh fork before sync: #38; 3 hand-resolved conflict region(s) in the sync; fork after sync: #63 #78
bin/fm-send.sh fork before sync: #19; fork after sync: #22
bin/fm-spawn.sh Both sides kept at each conflict (fork FM_COMPACT_ADVISER_DISABLE and api_key with upstream capacity deferral and base_branch). The fork removed the "Raw commands must be POSIX sh compatible" comment sentence before the sync. fork before sync: #2 #5 #4 #21 #23 #8 #25 #28 #7 #26 #31 #27 #24 #34 #13 #29 #42 #46 #48 #45 #49 #51; 14 hand-resolved conflict region(s) in the sync; sync-branch fix commits: 4afdea9; fork after sync: #58 #62 #65 #68 #76 #87 #89 #22
bin/fm-tasks-axi.sh fork before sync: #45; 3 hand-resolved conflict region(s) in the sync; fork after sync: #76
bin/fm-teardown.sh Fork restructured the dirty-worktree refusal; report_worktree_dirt is still called. fork before sync: #4 #14 #15 #19 #27 #29 #45 #51; 6 hand-resolved conflict region(s) in the sync; fork after sync: #60 #70 #84 #89 #22
bin/fm-test-run.sh Fork re-measured duration hints, uses ten serial shards, and registers its split supervision suites. fork before sync: #8 #25 #6 #30 #33 #19 #27 #24 #37 #13 #29 #41 #42 #48 #45 #50 #49 #53 #51; 10 hand-resolved conflict region(s) in the sync; sync-branch fix commits: e24c85a df845c5; fork after sync: #62 #65 #74 #76 #73 #66 #69 #70 #88 #84 #89 #22
bin/fm-tmux-lib.sh fork before sync: #48; 1 hand-resolved conflict region(s) in the sync; fork after sync: #22
bin/fm-wake-lib.sh Fork moved fm_firstmate_root_home (with upstream's remote-parent rule) to bin/fm-secondmate-parent-lib.sh. fork before sync: #9 #19 #42 #52 #51; 1 hand-resolved conflict region(s) in the sync; fork after sync: #74 #76 #73 #69 #79 #87 #22
bin/fm-watch.sh fork before sync: #4 #9 #19 #29 #47 #48 #45; 3 hand-resolved conflict region(s) in the sync; fork after sync: #76 #77 #69 #70 #85 #22
docs/agent-control.md fork before sync: #4 #12 #21 #8 #7 #26 #13 #29 #46 #48; 1 hand-resolved conflict region(s) in the sync; fork after sync: #87 #22
docs/architecture.md Sentences merged from both sides (fork owner pointer bin/fm-status-event-lib.sh, upstream scratch-location contract). fork before sync: #4 #11 #26 #19 #31 #13 #29 #48 #56 #54 #51; 2 hand-resolved conflict region(s) in the sync; sync-branch fix commits: 70383df 4afdea9 a9a9f98; fork after sync: #61 #73 #77 #66 #70 #79 #85 #22
docs/calm-mode-feasibility.md Fork describes its stock-component delegation for the two supervision tools; upstream's self-renderer description does not apply to the fork code. fork before sync: #10 #34; 4 hand-resolved conflict region(s) in the sync; sync-branch fix commits: d42597a
docs/calm.md Sentence merged from both sides. fork before sync: #10 #34; 3 hand-resolved conflict region(s) in the sync
docs/captain-hold-lifecycle.md fork before sync: #15; sync-branch fix commits: df845c5; fork after sync: #63 #78 #70 #83
docs/configuration.md Rows merged from both sides (fork Claude launcher link and session-end relaunch note with upstream additions). fork before sync: #2 #4 #12 #23 #11 #16 #8 #17 #28 #7 #30 #9 #19 #31 #24 #34 #36 #38 #13 #29 #42 #48 #44 #45 #49 #56 #51; 5 hand-resolved conflict region(s) in the sync; sync-branch fix commits: 70383df d42597a 4afdea9; fork after sync: #61 #62 #65 #68 #74 #73 #77 #78 #70 #88 #85 #22
docs/fm-test-portable-shards.md Sentences merged from both sides; fork adds its own shard count, hint provenance, and cache coverage. fork before sync: #19 #29 #48 #45 #50 #53; 6 hand-resolved conflict region(s) in the sync; sync-branch fix commits: e24c85a df845c5 78ac31e; fork after sync: #60 #62 #71 #69 #88 #22
docs/remote-secondmates.md fork before sync: #10 #9 #13 #29; sync-branch fix commits: d42597a 4afdea9; fork after sync: #67 #70 #22
docs/scripts.md fork before sync: #6 #9 #26 #19 #24 #34 #13 #42 #45 #49 #53; 1 hand-resolved conflict region(s) in the sync; fork after sync: #60 #62 #76 #66 #70 #87 #84
docs/verification/runtime-backends.md Fork points to its own cross-home recovery section and records its own Herdr evidence. fork before sync: #5 #10 #12 #21 #23 #8 #25 #28 #26 #33 #31 #27 #34 #37 #13 #46 #48 #45 #54 #57; 2 hand-resolved conflict region(s) in the sync; sync-branch fix commits: df845c5 70383df f80c5b2; fork after sync: #62 #65 #81 #89 #22
docs/watcher-continuity.md Sentence merged from both sides. fork before sync: #6 #9 #48 #52 #54; 3 hand-resolved conflict region(s) in the sync; fork after sync: #76 #73 #66 #81 #22
tests/fm-afk-launch.test.sh fork before sync: #47; 1 hand-resolved conflict region(s) in the sync
tests/fm-afk-return.test.sh fork before sync: #19 #42; 1 hand-resolved conflict region(s) in the sync
tests/fm-backend-herdr-presentation-e2e.test.sh Upstream change taken with the fork's FIXTURE_ROOT in place of ROOT. fork before sync: #27 #34 #13 #47 #45; 3 hand-resolved conflict region(s) in the sync
tests/fm-backend-herdr.test.sh fork before sync: #12 #25 #34 #48; 2 hand-resolved conflict region(s) in the sync; sync-branch fix commits: c477635; fork after sync: #22
tests/fm-calm-pi-extension.test.sh Fork evaluates the conversation boundary in a browser (calm-boundary-result) and had replaced the Node-side check before the sync, so upstream's Node-side hidden-row parser has no target. fork before sync: #10; 6 hand-resolved conflict region(s) in the sync; fork after sync: #73
tests/fm-captain-hold-lifecycle.test.sh fork before sync: #47 #45; 1 hand-resolved conflict region(s) in the sync; sync-branch fix commits: df845c5; fork after sync: #63 #60 #78 #83
tests/fm-classify-decision-key.test.sh fork before sync: #45; 1 hand-resolved conflict region(s) in the sync
tests/fm-claude-stop-autoarm.test.sh fork before sync: #42 #47 #52; fork after sync: #66
tests/fm-claude-trust.test.sh fork before sync: #5 #23; 1 hand-resolved conflict region(s) in the sync
tests/fm-composer-lib.test.sh Follows the fork titled-composer design above (fixtures size the closing rule; plain-backend typed case expects pending). fork before sync: #26 #33 #34 #37 #48; 1 hand-resolved conflict region(s) in the sync; sync-branch fix commits: a9a9f98; fork after sync: #22
tests/fm-contributions.test.sh fork after sync: #88
tests/fm-control-relaunch.test.sh fork before sync: #2 #21 #23 #8 #7 #26 #13 #46 #47; 4 hand-resolved conflict region(s) in the sync; fork after sync: #58 #87 #22
tests/fm-daemon.test.sh fork before sync: #48; fork after sync: #67
tests/fm-kimi-harness.test.sh Fork replaced the exact launch-string assertions before the sync, so upstream's task_inbox_export edit to those assertions has no target. The FM_TASK_INBOX export itself is in bin/fm-spawn.sh. fork before sync: #23 #45 #51; 3 hand-resolved conflict region(s) in the sync
tests/fm-lint.test.sh Follows the fork lint cache design; the unbounded-fallback RSS assertion accepts a real reading because the fork reads RSS through /usr/bin/time. fork before sync: #19; fork after sync: #71
tests/fm-omp-harness.test.sh fork before sync: #28 #48 #54; 1 hand-resolved conflict region(s) in the sync; fork after sync: #62 #68 #73 #81 #89 #22
tests/fm-pending-reply.test.sh fork before sync: #9 #19
tests/fm-project-capacity.test.sh sync-branch fix commits: e24c85a; fork after sync: #87
tests/fm-remote-delta-read.test.sh fork after sync: #70
tests/fm-remote-job.test.sh fork before sync: #9; 1 hand-resolved conflict region(s) in the sync
tests/fm-spawn-compact-adviser-disable.test.sh fork before sync: #23 #34; fork after sync: #22
tests/fm-spawn-dispatch-profile.test.sh Fork replaced the exact launch-string helpers with claude_launch_arg before the sync, so upstream's task_inbox_export edit to those helpers has no target. The FM_TASK_INBOX export itself is in bin/fm-spawn.sh. fork before sync: #23 #8 #7 #13 #44; 3 hand-resolved conflict region(s) in the sync; fork after sync: #65
tests/fm-supervision-host.test.sh Fork split the Claude Stop-hook cases into tests/fm-supervision-host-hook.test.sh and shared helpers into tests/fm-supervision-host-helpers.sh; the upstream lines live there. sync-branch fix commits: df845c5
tests/fm-task-inbox.test.sh fork after sync: #22
tests/fm-teardown.test.sh fork before sync: #4 #15 #27 #34 #47 #45 #51; sync-branch fix commits: e24c85a df845c5; fork after sync: #60 #84
tests/fm-test-run.test.sh Follows the fork runner changes above. fork before sync: #19 #27 #24 #13 #41 #47 #45 #50 #49 #53; 1 hand-resolved conflict region(s) in the sync; sync-branch fix commits: e24c85a df845c5; fork after sync: #62 #76 #73 #66 #22
tests/fm-timeout-lib.test.sh fork before sync: #47; 1 hand-resolved conflict region(s) in the sync; fork after sync: #69
tests/fm-tmux-submit-busy.test.sh Fork extended the fake tmux -l branch. fork before sync: #48; 1 hand-resolved conflict region(s) in the sync; fork after sync: #22
tests/fm-wake-queue.test.sh fork before sync: #29 #47 #48 #52; fork after sync: #76 #73 #22

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

✅ **Review** - passed

✅ No issues found.

🔧 **Test** - 3 issues found → auto-fixed ✅
  • ⚠️ tests/fm-supervision-host-attended-live-e2e.test.sh:343 - The live Claude attended supervision-host e2e fails at step 2 ('event 1 was not a main-only pass-through') on both a095da7 (2 runs) and base origin/main 60775fa (1 run). An early actionable close arrives before event 1, the host takes it as handling, and that turn absorbs event 1, so there is no pass-through. The failure existed before this merge. It means the live proof of the #6787 attended hand-back does not complete on this machine. Decide whether to track a separate fix.
  • ⚠️ tests/fm-watcher-lock.test.sh:54 - Two unchanged cases in fm-watcher-lock fail the same way on base 60775fa under machine load average 80-140. test_wait_deadline_reaps_a_stopped_child hit its 15s outer bound, and test_arm_self_eviction_is_loud_without_successor timed out with status 124. The cases the merge changed pass when run past the first failure. These failures are not from this merge; remote CI owns that judgement.
  • 🚨 live validation verdict: no-go (9 of 13 scenarios were driven live against the product); failed: Live Claude attended e2e fully passes (an idle primary is woken for each main-only pass-through)
  • Live validation: ❌ no-go - 9 of 13 scenarios driven live against the product
Scenario Result Live Evidence
Record #43 as a real merge: 4b9be5e merges 2ce57d0 with -s ours and its tree equals origin/main's d69b2c5 (no file change) ✅ pass live history-sync.txt: both trees 104878b5, empty diff --stat
Bring in the 7 new upstream commits: all of 329ad4e..fb75c1f (and 2ce57d0) are ancestors of a095da7 ✅ pass live history-sync.txt ancestry block
Adversarial next-sync check: merging upstream/main fb75c1f into the old origin/main conflicts, but into a095da7 it is a clean no-op (merge-base = fb75c1f) ✅ pass live history-sync.txt: 81 conflicted files on 60775fa; exit=0 and result tree == HEAD tree on a095da7
No upstream content dropped: each upstream hunk is present, and the 5 conflicted files keep both sides (herdr.sh has </dev/null and the COMPACT_ADVISER names; CONTRIBUTING has the union helper list pl… ⏸️ untested no The prior payload established this only by git reverse-apply checks and manual file reading, not by driving the running product, so it did not record a live pass. The live Herdr lab covers only the he…
Herdr composer probe on a stopped or never-started lab session fails fast and does not start a server (#6920) ✅ pass live live-herdr-observable-probe.txt steps 1 and 3: rc=1 in 0s, composer_state unknown, session stays absent or running:false
Herdr probe on a running lab pane reads the screen ✅ pass live live-herdr-observable-probe.txt step 2: rc=0, marker seen
Merged herdr server start reads stdin from /dev/null and strips the COMPACT_ADVISER_DISABLE, FM_COMPACT_ADVISER_DISABLE and FM_COMPACT_ADVISER_HOOKS settings from the server env ✅ pass live live-herdr-observable-probe.txt step 4: fd0 -> /dev/null, matching env []
fm-control lifecycle against real herdr with the resolved stand-in fixture (conflict file tests/fm-control-herdr-smoke.test.sh) ✅ pass live suite-fm-control-herdr-smoke.log exit=0
Supervision host successor watcher survives the Stop hook's process group teardown (#6787, moved into the split hook suite) ⏸️ untested no The prior payload ran the suite with Claude's group teardown simulated and recorded live=false, so it did not establish a live result against the real running product.
Upstream test-fixture fixes: multicall-sleep stand-in (#6814), GIT_DIR isolation (#6956), tmux readiness without the permission footer (#6823), and backend-herdr fixtures ⏸️ untested no The prior payload ran only the fixture test suites and recorded live=false, so it did not establish a live result against the real running product.
Watcher-lock race fixtures changed by the merge (#6887 plus the fork's retry follow-up): stale-lock steal has a single winner, and a competing reaper cannot remove the successor ⏸️ untested no The prior payload ran the race-fixture cases in a test suite and recorded live=false, so it did not establish a live result against the real running product. Two unchanged cases in the same suite also…
Merge does not change attended hand-back: the live Claude idle primary run on a095da7 behaves the same as on base 60775fa ✅ pass live live-claude-supervision-host-attended.log, live-claude-attended-rerun2.log, live-claude-attended-BASE-60775fa1.log: all stop at the same step-2 'handled' then 'to-main' pattern
Live Claude attended e2e fully passes (an idle primary is woken for each main-only pass-through) ❌ fail live live-claude-attended-*.log: 'not ok - positive: event 1 was not a main-only pass-through' on both target and base
  • git log --first-parent, git rev-parse 4b9be5e0^{tree} d69b2c5a^{tree}, git merge-base --is-ancestor &lt;8 upstream commits&gt; a095da77
  • git merge-tree --write-tree 60775fa1 fb75c1f9 (81 conflicted files) vs git merge-tree --write-tree a095da77 fb75c1f9 (clean, tree unchanged)
  • Per-file git apply --reverse --check of each upstream commit (2ce57d00..fb75c1f9) against a095da77, plus a manual check of the non-applicable files
  • Live Herdr lab script: bin/fm-herdr-lab.sh name/prepare/provision/stop/teardown, driving fm_backend_herdr_visible_capture, fm_backend_herdr_composer_state, fm_backend_herdr_target_ready, and lsof/ps eww on the started server
  • bash tests/fm-supervision-host-hook.test.sh
  • bash tests/fm-backend-herdr.test.sh
  • bash tests/fm-test-fixtures.test.sh
  • bash tests/fm-tmux-agent-liveness.test.sh
  • bash tests/fm-control-herdr-smoke.test.sh (real herdr 0.9.1, lab session)
  • bash tests/fm-watcher-lock.test.sh (3 runs on target, 1 on base), plus a focused run that skips the old failing case, on target and on base
  • FM_SUPERVISION_HOST_ATTENDED_LIVE_E2E=1 bash tests/fm-supervision-host-attended-live-e2e.test.sh (2 runs on target, 1 on base 60775fa1, real Claude Code 2.1.296 with haiku)

🔧 Fix applied.
✅ Re-checked - no issues remain.

  • Live validation: ✅ go - 10 of 10 scenarios driven live against the product
Scenario Result Live Evidence
Upstream #43 is recorded as a real merge with no file change (the 4b9be5e tree equals the d69b2c5 tree) ✅ pass live history-sync-round2.txt: both trees are 104878b5
All 7 new upstream commits and 2ce57d0 are ancestors of the sync branch ✅ pass live history-sync-round2.txt: all 8 commits are ancestors
The next upstream sync sees shared history: merging upstream/main into HEAD is a clean no-op ✅ pass live history-sync-round2.txt: merge-base = fb75c1f; merge-tree exit 0, result tree = HEAD tree
Adversarial: the same merge into the pre-A2 origin/main still conflicts, so the A2 history is what removes the conflicts ✅ pass live history-sync-round2.txt: exit 1, 81 conflicted files
Herdr probes stay observable-only and do not autostart a stopped lab server (#6956/b75658b6), and the lifecycle server starts with stdin from /dev/null ✅ pass live live-herdr-observable-probe.txt (round 1; bin/backends/herdr.sh has not changed since then)
Live Claude attended: event 1 passes through as main-only, leaves a live successor watcher, and the Stop hook rewake wakes the idle primary, which drains and acknowledges it (#6787 hand-back, round-1… ✅ pass live live-claude-attended-round2.log: 21:45:32 'pass-through attended main-only'; 21:45:33 'successor watcher pid 13466 alive'; 21:45:52 'rewake delivered ... outcome=rewake'; 'primary turn ran: bin/fm-wak…
Live Claude attended: the event 1 turn end takes over the successor's cycle, and event 2 is passed through and woken through the cycle that took over ✅ pass live live-claude-attended-round2.log: 21:45:56 'successor 13466: reason=taken-over ... now owning watcher 58589'; 21:46:08 'watcher 58589 closed: ... reason=actionable-signal'; 'its close was delivered: re…
Live Claude attended: the remote-reply listener delivers event 3 and stays owned throughout ✅ pass live live-claude-attended-round2.log: 21:46:54 'listener mirrored it (1 ingested) and it was delivered ... runner 4203 -> 4203, owned at every check'; 21:47:28 'listener runner 4203 still owned'
Live Claude attended: a routine close accepted for handling that turns main-only before its turn (event 4) is handed back with no engine turn, woken, drained and acknowledged, and its turn end re-arms ✅ pass live live-claude-attended-round2.log: 21:47:06 'pass-through attended main-only; no engine turn'; 'rewake delivered ... outcome=rewake'; 21:47:24 '--ack-through 10'; 21:47:28 'turn end re-armed ... watcher…
Live Claude attended e2e fully passes (no captain prompt is submitted after setup, the test prints ok, exit 0) ✅ pass live live-claude-attended-round2.log: 21:47:28 'captain prompts after setup: 0'; 'ok - attended live (2.1.296 (Claude Code)) ...'; 'exit=0'
  • git log --first-parent --format='%h %p | %s' -7 HEAD
  • git rev-parse 4b9be5e0^{tree} vs git rev-parse d69b2c5a^{tree} (the ours-merge changes no files)
  • git merge-base --is-ancestor &lt;c&gt; HEAD for 2ce57d00 329ad4ed 3c58ec8e 0ec1c5a4 b062eb94 19fcbbde b75658b6 fb75c1f9
  • git merge-tree --write-tree HEAD fb75c1f9 (next sync is a clean no-op)
  • git merge-tree --write-tree --name-only 60775fa1 fb75c1f9 (adversarial pre-A2 control)
  • git diff --stat a095da77 HEAD (only the test file changed since round 1)
  • FM_SUPERVISION_HOST_ATTENDED_LIVE_E2E=1 bash tests/fm-supervision-host-attended-live-e2e.test.sh (live real Claude 2.1.296, haiku; ok, exit 0)
  • Teardown check: lab dir ~/tmp/fm-sh-attended-live.z9c1Zm removed, no lab processes left, worktree clean
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

kunchenguid and others added 30 commits September 28, 2026 17:19
…6064)

* fix(bin): read a live quiet record as a present captain at the host and watcher

A quiet record left without its daemon (a quiet start that never ran or was
interrupted) was read as away by the supervision host, so it parked a present
captain's main and held captain outcomes for a return that never comes, and
the watcher and daemon silenced captain-held rechecks on record presence.

The host's posture checks, the watcher's and daemon's captain-held silencing,
and the host's outcome path (branch report, drain BRANCH OUTCOMES, relocated
branch authority, the owners' away wake note, and the Codex checkpoint bound)
now ask the record owner's away-or-quiet reading, so only an away record is
away. A live away record keeps today's behavior.

* no-mistakes(document): Correct quiet-record documentation and supervision guidance

* no-mistakes(document): Clarify quiet-record posture and captain-held rechecks

* no-mistakes(document): Clarify quiet-record posture in documentation
…kunchenguid#6043)

* fix(bin): name an in-window engine latch in the return brief and drop the false handling GAP line

The away return brief said nothing had failed after the supervision host
latched on engine errors during the window, and printed a GAP: watcher
downtime line whenever a wake was merely being handled or queued at return.

The failures section now reads the host ledger and latch record and names
the latch time, the window's engine-error count, and whether the session
is still paused or recovered. An open recovery episode is reported as
information, and as a gap only when a queued episode outlived the return
grace or the marker cannot be read.

* no-mistakes(review): Fix latch trip time, drop marker-age grace, bound error count

* no-mistakes(review): Report paused latch without ledger trip row; bound errors

* no-mistakes(review): Never report a failed probe's latch row as trip time

* no-mistakes(review): Only a retained trip row marks a pre-window latch

* no-mistakes(document): Clarify return-brief latch and watcher-gap documentation

* no-mistakes(ci): Fixed Lint 1 by marking the shared cooldown constant as used by sourcing scripts. The repository lint command and diff check pass; the return test run was stopped by a 180-second timeout after its completed cases passed

* no-mistakes(ci): Fixed the return brief so the trip time and error count come from the same initial latch row, and ledger rows before the current session’s lock boundary cannot affect its latch report. Added real-script regressions for both findings. The return test suite, repository lint, and diff check pass

* no-mistakes(ci): Fixed the return brief’s restart cutoff so it retains in-window failures, prints one line per initial-trip row, and omits zero-error count wording. Added real-script restart regressions. The return test suite, ShellCheck, and diff check pass

* no-mistakes(ci): Fixed the return brief so a recorded trip followed by recovery stays recovered, while a later pause with a lost trip append gets a separate “trip time unavailable” line. Added a real-script regression that failed before the fix. The return test suite, ShellCheck, syntax checks, and diff check pass

* no-mistakes(ci): Fixed the false second latch during recovery. A real-script regression failed before the fix and passes now; the lost-second-trip test still passes. The return test suite, ShellCheck, syntax checks, and diff check pass
* fix(calm): name the Claude Code Calm plugin fm so supervision notes read "fm: "

Claude Code labels every mod transcript line with the plugin name, so the
notes rendered as "firstmate-calm: ⚓ ...". Rename the plugin to fm, update
the live guard to assert the fm: label, and document the one-time replay for
sessions resumed across the rename.

* no-mistakes(document): Clarify Calm plugin rename in documentation
…#6037)

* feat(bin): add fm-live-lab.sh, a one-command live supervision lab builder

* fix(bin): exact lab windows, per-lab task ids, self-safe teardown

* fix(bin): target lab windows by id, stop lab descendants, add readiness tests

* fix(bin): keep Claude's auto-updater off in live labs; list fm-live-lab.sh

* fix(bin): start the lab tmux server without user config

* no-mistakes(review): Scope lab teardown to its store, root, and task ids

* no-mistakes(review): Record selected user stores at up for check and down

* no-mistakes(document): Clarify live lab documentation and remove stale narratives

* no-mistakes(ci): Fixed the CI failure by checking for an existing lab root before looking up the harness executable. The affected behavioral test and shell syntax check pass; the refusal also works with Claude absent from PATH

* no-mistakes(ci): Fixed all four Greptile findings: teardown signals only recorded lab processes and their descendants; the worker gate is in its granted task directory and its path is exposed; readiness uses current crew state; and mate and worker IDs use 12 nonce hex digits. The CLI behavior tests pass, as do shell syntax, ShellCheck, and diff checks. The Claude no-host path is unchanged

* no-mistakes(ci): Fixed the CI test’s dependence on an installed Claude binary by supplying a test-local stub. The full fm-live-lab test, shell syntax check, and diff check pass

* no-mistakes(ci): Fixed all three selected findings in bin/fm-live-lab.sh: down waits for recorded processes and escalates before cleanup, PID roots are checked against recorded start times, and Claude primary trust is rechecked after mate/worker readiness. Added behavioral tests in tests/fm-live-lab.test.sh. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass

* no-mistakes(ci): Fixed the pre-primary settle wait, worker gate instructions, unused retry variable, and teardown PID revalidation in bin/fm-live-lab.sh. Added behavioral tests in tests/fm-live-lab.test.sh. Both requested commands pass: tests/fm-live-lab.test.sh and bin/fm-lint.sh

* no-mistakes(ci): Fixed teardown to track pre-kill lab processes by PID and start time, including children orphaned when a root exits. Up now rejects an empty pane PID before calling ps. Added regression tests and a Linux-safe worker fixture. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass

* no-mistakes(ci): Fixed teardown tracking for children spawned during shutdown and made the worker fixture verify its exact window with a Linux-available shell. Both requested checks pass. The lab test takes about 66 seconds locally, so the under-one-minute target remains unmet

* no-mistakes(ci): Fixed ci-2 and ci-4 in bin/fm-live-lab.sh and tests/fm-live-lab.test.sh. Teardown now tracks identity-checked members of captured lab process groups, including children orphaned during shutdown, without signaling the caller’s group or unrelated processes. Lint passed, and the lab test passed four times

* no-mistakes(ci): Fixed teardown so an observed-empty process group is permanently dropped, preventing a reused group ID from signalling unrelated work. Added a ps-shim regression test. The lab test, lint, and diff checks pass

* no-mistakes(ci): Fixed ci-1 in bin/fm-live-lab.sh and tests/fm-live-lab.test.sh. The TERM-born-child fixture now waits until its handler is installed before calling down. Down sends SIGKILL to identity-valid survivors on every pass from pass 20 onward and includes survivor process details if it must refuse cleanup. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass locally; Linux CI remains to be verified

* no-mistakes(ci): Fixed down’s teardown wait to require two empty identity-checked scans separated by 0.5 seconds, and removed the unused test loop variable without changing the TERM-born-child test. The lab test, lint, and diff check pass locally
…unchenguid#6103)

* fix(bin): keep slow watcher cycles and preempted reply polls from breaking supervision

- fm_pending_reply_tick selects the records it has work for in one awk pass,
  so settled records cost no lock or fork and the walk no longer grows with
  the never-pruned store.
- An attached arm keeps following a live, identity-matched holder whose beacon
  went stale until the lock changes or the shared stall bound
  (fm_watcher_stall_bound), then fails with a typed stalled-holder line so the
  retry replaces the holder.
- The remote-reply adapter reports the job worker's preemption (exit 76) as a
  closed window, so the listener keeps its claim and polls again instead of
  being relaunched every watcher cycle.

* no-mistakes(document): Clarify watcher grace and attached-arm documentation
…geable is UNKNOWN (kunchenguid#6110)

* fix(bin): retry a bounded number of times when GitHub mergeable is UNKNOWN

Fixes kunchenguid#6020

bin/fm-pr-merge.sh refused a GitHub merge whenever the pull request's
mergeable field was not literally MERGEABLE. GitHub reports UNKNOWN for
a short while after a push or a base-branch change while it recomputes
mergeability, so a green, conflict-free pull request was refused as if
it could not be merged.

github_verify_mergeable now returns a distinct status when mergeable is
the only failing condition and reads UNKNOWN. The caller retries up to
5 times, 3 seconds apart (overridable in tests), re-reading and
re-checking every live condition on each attempt. Once the bound is
spent it reports mergeability as still being computed rather than
unmergeable, with the same nonzero exit as before. Every other refusal
(closed, draft, conflicting, red or missing checks, away authority,
queue protection) is unchanged and never retried.

* no-mistakes(ci): I fixed both review findings the way you asked. The full suite (`bash tests/fm-pr-merge.test.sh`) ran to completion. Its last lines showed all `ok`, and any failure would have stopped the run early. I watched the output through `tail`, so I didn't see the new test's own `ok` line directly. **ci-2 (`bin/fm-pr-merge.sh`), retry delay not validated.** What must hold: the retry wait is always a short, valid `sleep` argument, so a bad `FM_PR_GITHUB_MERGEABLE_RETRY_DELAY` can never trip `set -e` or hold the task lock for a long time. The retry loop is the only place that reads this variable. The script now reads the value once before the loop and accepts only whole numbers from 0 to 10. Anything else (empty, `abc`, `-1`, `1.5`, `11`, a huge number, leading spaces) falls back to 3. I ran those values through the check by hand and each came out as expected. The loop now sleeps on that checked value. **ci-1 (`tests/fm-pr-merge.test.sh`), no test for a check changing between UNKNOWN reads.** What must hold: every retry re-checks all live conditions, not just mergeable. The fake `gh pr view` in the test can now take an optional second word on each line of the mergeable sequence, which sets the first check's result. The new test `test_github_mergeable_unknown_retry_rechecks_checks` feeds `UNKNOWN`, then `UNKNOWN FAILURE`. It asserts: - exit code 1 after exactly 2 reads, - the refusal names `check 'ci' is not green`, - the message does not say mergeability is still being computed, - `pr merge` was never called. If a later change made the retry look only at mergeable, the loop would read UNKNOWN 5 times, end with the "still being computed" message, and this test would fail. I didn't run it against a deliberately broken script to confirm that. `bash -n` passes. `shellcheck` reports only the existing info-level notes about files it can't follow. Only `bin/fm-pr-merge.sh` and `tests/fm-pr-merge.test.sh` changed
kunchenguid#6112)

* fix(bin): converge every open owner onto a known terminal contribution

settle_final only cleared a stale error on retry, so an owner whose saved
row still said open kept projecting a merged or closed pull request as
open after another owner's row had already recorded the terminal
observation. Copy the known terminal observation to every owner whose
saved row is not itself terminal, keeping that owner's own pending and
notified state, and clear its error.

* no-mistakes(review): Carry terminal checked_at when converging existing owner rows

* no-mistakes(ci): I fixed Greptile finding ci-2 as you asked, with a change to tests/fm-contributions.test.sh only. The rule it enforces: when a retry converges an owner onto a URL that is already merged or closed, that owner gets the terminal owner's whole observation, not just its state. The same weak check appeared twice in test_interrupted_multi_owner_poll_settles_every_owner, so I fixed both: - **Open owner (line 784):** the check now also requires `.observation == $terminal[0].records[0].observation`. The existing checks for error, checked_at, pending and notified are unchanged. - **Errored owner (just below):** it only checked state and error before. It now reads the terminal owner's file and makes the same full-observation comparison. Adding the comparison alone would not have caught anything. The test fixtures gave both owners identical observations apart from `state`, so copying only the state would still have passed. In both cases I also set the terminal owner's observation head to HEAD_B, so the two observations now really differ. Verification: - The focused test passes against the current bin/fm-contributions.sh. - I temporarily changed `settle_final` so it copied only the state. The test then failed, reporting the owner still on the old head (HEAD_A). I restored the file afterwards, and `git status` shows only the test file modified. - The full tests/fm-contributions.test.sh suite exits 0. No product code changed. The other CI finding (ci-1, "Behavior portable serial 9") was left alone because you chose to ignore it
…nguid#6124)

* feat: run the supervision host by default on a Claude primary

An absent config/supervision-host on a Claude primary now reads as on with
the default engine, and a file holding `off` opts any home out. Cursor,
OpenCode, omp, Grok, and Codex stay file-gated, with `off` read as disabled
there too. Every reader asks fm_supervision_host_enabled instead of testing
the file, and non-bash readers query it through the lib's `enabled` entry.
A primary's `off` is not inherited by secondmates: each home keeps its own
supervision posture.

* test: pin the watcher-path posture in fixtures that assume no supervision host

Fixtures that drive the watcher arm or assert a non-host drain now write
an explicit off file, and fixtures that copy the Stop auto-arm or the
supervision instructions carry the engine lib they now source. The two
drain suites also stop reading the code root's config.

* fix: name the opt-out when an off home passes an attended wake to main

A host parked when the home writes off now logs that the home does not run
the supervision host, rather than claiming it has no engine.

* no-mistakes(document): Clarify Claude supervision defaults and historical evidence

* no-mistakes(ci): Fixed process leaks in the two added host tests. Each case now stops its recorded watcher and host/arm processes; fake hook sessions exit through session.stop. The full host suite passed before the final cleanup refinement, and both affected cases, bash syntax, ShellCheck, and diff checks passed afterward. CI runtime still needs confirmation
…start scope check (kunchenguid#6125)

* fix(bin): create the state dir on a fresh primary before the session-start scope check

fm_primary_scope_matches required an already-existing state directory, so
bin/fm-sessionstart-run.sh stood down on a fresh clone before anything could
create it. Split out fm_primary_root_matches so the run wrapper can confirm
primary-home identity first, create the gitignored state dir when it is
missing, and only then run the unchanged scope check.

* no-mistakes(document): Document session-start state dir creation on fresh clones

* no-mistakes(ci): I fixed the Greptile P1 the way you asked. When a fresh primary can't create `state/`, the run wrapper no longer stands down silently. **Invariant:** when an otherwise eligible fresh primary cannot create `state/`, startup must never fail silently. This path has only one site: the mkdir in `bin/fm-sessionstart-run.sh`. Other hooks and the nudge wrapper never create `state/`, so they have no equivalent failure. **What changed:** - **Run wrapper** (`bin/fm-sessionstart-run.sh`): it captures mkdir's error and prints one line to stderr before standing down as before (exit 0, or 3 for the Pi prerequisite). The line looks like `fm-sessionstart-run: startup could not create the state directory <path>: <reason>`. - **Test** (`tests/fm-sessionstart-nudge.test.sh`): the new case `test_run_reports_a_state_dir_it_cannot_create` uses a fresh primary with no `state/` and a read-only (0500) root. It checks four things: exit 0, no digest on stdout, no state dir created, and exactly one stderr line ending in "Permission denied". It fails without the fix and passes with it. - **Docs** (`docs/sessionstart-nudge.md`): I added one sentence describing the stderr line and one describing what the new test proves. **Verification:** I ran `tests/fm-sessionstart-nudge.test.sh`, and every test passes. `bin/fm-lint.sh` on the changed scripts (pinned ShellCheck 0.11.0) and `tests/fm-documentation-audiences.test.sh` also pass. As you asked, the wrapper still stands down with the ineligible-checkout status afterwards. It does not report this as a failed eligible startup, which is what the bot suggested
…ery (kunchenguid#6126)

* fix(bin): measure pending-reply grace from turn completion, not delivery

Fixes kunchenguid#6057

The pending-reply guard demanded a repost ("REPOST REQUIRED: previous
marked request had no correlated parent report") while the second
mate's correlated reply was already on its way.
fm_pending_reply_send_recovery measured its grace window from delivery
instead of from the request turn's completion, so any turn longer than
the grace fired the demand the moment the turn ended, before the reply
could have landed. The missed-report escalation had the same gap: it
fired the instant the recovery turn's completion was observed, with no
grace at all.

Both now measure grace from the relevant turn's completion (request
turn for the recovery repost, recovery turn for the escalation), and
both take one fresh, uncached read of the parent status file
immediately before firing, accepting a correlated line regardless of
its verb. Transport-failure escalations stay immediate, and the
one-repost limit is unchanged.

* no-mistakes(review): Document grace window as measured from turn completion

* no-mistakes(ci): Both Greptile findings were real and caused by this PR, so I fixed them. The full `tests/fm-pending-reply.test.sh` suite passes. **ci-1 (a reply could be overwritten by a repost).** The rule that must hold: a recovery send is recorded only if the record is still unresolved, checked under the same per-correlation lock that resolution uses. The escalation path already did this (`_fm_pending_reply_maybe_escalate_locked` reads fresh and publishes under one lock). The recovery path did not: `fm_pending_reply_send_recovery` did its fresh read through `fm_pending_reply_try_resolve`, which let go of the lock before the send was recorded. A reply landing in that gap could be overwritten, and the repost would go out anyway. Now `send_recovery` takes the lock once and, while holding it, re-checks that the phase is still `awaiting_report`, runs the fresh uncached read, and records the send (sender pid and identity, attempt time, phase `recovery_sending`). It releases the lock before actually sending, so the lock is not held during the send. It uses the same lock helpers the other lock wrappers use. Grace timing, the one-repost limit and the escalation path are unchanged. **ci-2 (the test would pass even without the fix).** In `test_recovery_fresh_status_read_resolves_before_firing`, the reply is still appended to the status file, but the stored file signature is then set to the file's new signature. That stands in for a same-size rewrite that the signature cache cannot see. The test first checks that a normal cached read misses the reply, then that the fresh read before sending catches it. I also added the same check for the fresh read before escalation, which the review said was uncovered. The test now sets its own send hook, so it no longer depends on one left over from an earlier test (that leftover had made failures exit silently). **Checks:** - I removed the fresh-read bypass at each site in turn and reran the suite. With it gone from recovery, the test fails with "recovery must not fire once a correlated reply has landed". With it gone from escalation, it fails with "the fresh pre-escalation read should have resolved the record, got escalated". With both in place, all tests pass. - Shellcheck with `-x` timed out locally. Without `-x` and ignoring SC1091, the only warnings are SC2034 on the existing `maybe_escalate` lock wrapper, which is not part of this change. The new code adds no warnings. Changes are in `bin/fm-pending-reply-lib.sh` and `tests/fm-pending-reply.test.sh`. Nothing is committed yet; a plain commit message such as "fix(bin): record the pending-reply recovery send under the fresh-read lock" fits the instruction

* no-mistakes(ci): ci-1 was real and caused by this PR. The same bug was also in the escalation path, so both are fixed. The full tests/fm-pending-reply.test.sh suite passes. The rule that must hold: a recovery repost or an escalation goes out only if the record's phase, read after the fresh-read resolve, is still what it was before. The resolver writes phase=resolved first and only then writes the other resolution fields. If one of those later writes fails, it returns an error even though the record is already resolved. Places this rule applies, both fixed: - Recovery (fm_pending_reply_send_recovery): the fresh-read resolve now runs first, and the phase is re-read right after it, whatever it returned. The send is recorded and made only if the phase is still exactly awaiting_report. This replaces the earlier phase check rather than adding a second one. - Escalation (_fm_pending_reply_maybe_escalate_locked): same bug. After a failed resolve it went on to publish the blocked line and set phase=escalated. One added line after the resolve call returns 1 without publishing if the phase has changed. Test: added test_partial_resolve_write_blocks_firing. It forces a failure on the resolved_epoch write after a correlated reply has landed. It checks that the recovery send hook is never called, that no escalation line is published, and that the phase stays resolved. The forced failure runs in a subshell so it can't affect later tests. Checks: - With the recovery fix reverted, the new test fails with "recovery must not fire after a partial resolve". - With the escalation fix reverted, it fails with "partial resolve should block escalation, got escalated". - With both fixes in, every test passes. - Shellcheck was run with SC1091 excluded and without -x, not through the repo's lint script. The only new message is one SC2329 info on the test's override function; other test overrides in the same file already get that same info, unsuppressed. Changed files: bin/fm-pending-reply-lib.sh and tests/fm-pending-reply.test.sh. Nothing is committed. Suggested plain commit message: "fix(bin): recheck pending-reply phase after the fresh read before sending
…to stderr (kunchenguid#6001)

* fix: provider-table lookup never writes a broken-pipe error to stderr

Fixes kunchenguid#5956

fm_quota_single_provider_for_harness returned from its while read loop
as soon as it found a match, closing the pipe while
fm_quota_single_provider_table's printf could still be writing.
Where SIGPIPE is ignored, as on GitHub Actions runners, bash then
prints "printf: write error: Broken pipe" on the resolver's stderr,
which intermittently broke the one-diagnostic-line assertions in
tests/fm-dispatch-resolve.test.sh.

Read the whole table before answering, the way
fm_control_harness_supported already does, so the writer always
finishes. Return values and output are unchanged.

Reproduced by running tests/fm-dispatch-resolve.test.sh with SIGPIPE
ignored on a single pinned core under CPU contention: 30 of 30 runs
failed before the fix, 0 of 30 after. Note: reproducing requires
setting the trap inside the tested shell because nice(1) resets an
inherited SIGPIPE ignore to SIG_DFL. tests/fm-quota-choose.test.sh
passes and bin/fm-lint.sh is clean.

* no-mistakes(ci): Fixed both Greptile findings the user chose to address. ci-1 (bin/fm-quota-axi-lib.sh:154). Invariant: looking up a harness must always end with status 0 and print the provider, even when the caller runs under `set -e`. The loop body `[ -z "$found" ] && [ "$harness" = "$1" ] && found=$provider` now ends in `|| :`. Every iteration succeeds and the whole table is still read. Only `fm_quota_single_provider_for_harness` loops over the table this way, so this is the one place the fix was needed. One caveat: on bash 5.3 the old code did not actually exit under `set -e`, because the `while` loop is not the function's last command, so the new `set -e` test would have passed before this fix too. The change makes the loop's success explicit, as the user asked. ci-2 (regression coverage). I added three cases to the existing `tests/fm-quota-choose.test.sh`, all calling the public lookup function after sourcing the library: 1. With SIGPIPE ignored (`trap "" PIPE`), it looks up every harness 200 times and checks that nothing reaches stderr. 2. A deterministic version of the race: the table function is wrapped so it writes the first row, pauses 0.2 s, then writes the rest. With SIGPIPE ignored, it checks that looking up `claude` prints `claude` and writes nothing to stderr. The stress loop alone reproduced the bug in only about 1 of 5 local runs, which is why this case exists. 3. A direct call under `set -e` prints `claude`. Verification: - `bash tests/fm-quota-choose.test.sh`: all pass. - Same test against the pre-PR library (fa48367, via `FM_ROOT_OVERRIDE`): fails with `printf: write error: Broken pipe`. The deterministic case failed in one run and the stress loop caught it in another. - `shellcheck` on both files: clean. - `tests/fm-dispatch-resolve.test.sh`: passes
…isioning (kunchenguid#6162)

* fix: survive Pi 0.99 rendering and Git 2.55 local-clone races

Pi 0.99 puts arguments on the stock tool header and leaves hidden custom messages in the export conversation column. Match that header, and keep Calm's boundary on the visible column. Clone a remote home with --no-local so a prune during Git's loose-object copy cannot fail the seed.

* no-mistakes(review): Stop SIGPIPE write errors; cover older Pi export and project clones

* no-mistakes(document): Clarify Calm export visibility and tool rendering

* no-mistakes(ci): Fixed the dispatch diagnostic to list every provider-less use/default profile in one line and added a multi-profile behavior test. Shortened supervision fixtures using the existing engine-grace and park-clock knobs; removed stray scratch files. Dispatch tests, syntax checks, and three targeted supervision cases passed. CI’s prior supervision duration was 751s; the single permitted local full-suite run timed out at 1200s, so an after-duration is not established. The cancelled serial check had no failure verdict. The outer executor should record the measured before/after duration in the PR body when available

* no-mistakes(review): Gate Pi 0.99 call headers by version; drop hidden-row assertion

* no-mistakes(review): Test stock call headers under Pi 0.87 and 0.99 stubs

* no-mistakes(test): Fix older-Pi queued-row test and verify park-boundary behavior

* no-mistakes(document): Clarify Pi Calm export and queued-turn documentation

* no-mistakes(ci): Fixed the stock macOS Bash 3.2 parse failure in tests/fm-calm-pi-extension.test.sh; its parse check passes. The watcher CI failure is in unchanged code: the isolated five-minute/66-minute case passes locally, but the CI log omits the drain error needed to establish its cause. No speculative watcher fix was made. The full local watcher suite timed out after 500 seconds
…6169)

* Prevent premature Lavish board handoffs

* Prove Lavish arm lacks reply acknowledgement

* Confirm Lavish replies before arming worker boards

* no-mistakes(review): Post Lavish reply only after locked arm eligibility checks

* no-mistakes(review): Fail Lavish reply closed on unknown version

* no-mistakes(document): Correct Lavish reply documentation and remove stale guidance

* no-mistakes(document): Clarify Lavish reply routing and remove duplicate version guidance
…#6154)

* feat: inherit the supervision-host opt-out from the primary

Move the supervision host's off opt-out out of config/supervision-host into
its own presence flag, config/supervision-host-off, and add that flag to the
primary-authoritative inherited config set. A primary that opts out now opts
every secondmate home out at spawn and convergence, and clearing it converges
them back. config/supervision-host stays the home-local engine choice.

Shape: config/supervision-host mixed two things, a fleet posture (off) and a
per-home engine and model. Only the posture should follow the primary, so it
becomes a separate presence flag that rides the existing inherited-config
mechanism (FM_INHERITABLE_CONFIG in bin/fm-config-inherit-lib.sh) with no new
machinery, while the engine line stays local. The parse stays in its one
owner, fm_supervision_host_enabled. There is no migration or compatibility
handling for a home that still holds off in config/supervision-host.

Primary off, mate on: inherited material is primary-authoritative by design,
so a mate cannot keep the host while the primary is opted out, and a mate's
own opt-out is removed at the next convergence while the primary has none.
Running the host on a mate is the primary's choice for the fleet; no override
mechanism is added.

Live validation (disposable bin/fm-live-lab.sh lab, Claude primary with a
real seeded secondmate, --supervision-host off):
- up: every readiness check ok, including "host: none running, as expected"
  and a live mate session; the spawned mate home held the inherited
  config/supervision-host-off and the gate read primary OFF, mate OFF.
- primary removed its opt-out, then bin/fm-config-push.sh reported
  "supervision-host-off: pushed - mirrored primary absence" and a config
  reread sent; the gate read primary ON, mate ON, and the live mate handled
  the reread.
- primary opted out again and pushed: "supervision-host-off: pushed", mate
  gate OFF.
- down stopped every lab process and left no lab process running.

Out of scope, follow-up: default-on for the other harnesses, away-daemon
retirement, rollout.

* no-mistakes(document): Document inherited supervision-host opt-out ownership

* no-mistakes(ci): Fixed ci-4: with `--supervision-host off --mate`, lab readiness now requires the inherited flag in the mate home and a disabled mate supervision-host gate. The focused behavior test, shellcheck, and diff checks pass. Left ci-1–ci-3 untouched as directed

* no-mistakes(test): Fix mate readiness HOST_OFF initialization in lab up

* no-mistakes(ci): Fixed Lint 2 by making the new test’s fixtures source resolvable to ShellCheck; its off/on readiness test and ShellCheck now pass locally. Behavior portable serial 5 failed in the unchanged remote-reply test at generation 7. That test passes locally, and no PR-caused defect was identified, so no remote-reply code was changed
…nguid#6179)

* fix(tests): cut the fixed sleeps in supervision-host cycles

The serial CI lane keeps brushing its 30-minute cap because
fm-supervision-host.test.sh spends ~903s of the job, and per the
run-36635306527 case profile the top nine cases are all multi-cycle
ones (3-10 park/close/turn cycles each): every close waits out the
host's sleep $POLL in await_close plus a watcher sleep $FM_POLL scan
cycle, and every engine turn waits out the fixed sleep 1 descendant
snapshot. That is ~3s of pure sleep per cycle before any real work.

The host poll now accepts positive decimal seconds through a new
seconds_or validator (FM_SUPERVISION_HOST_POLL), and the engine turn's
snapshot loop takes FM_SUPERVISION_ENGINE_SNAPSHOT_SECONDS, also a
positive decimal defaulting to one second - the smallest seam at each
wait's single owner. The suite drives them at 0.2 alongside the
existing FM_POLL=0.5 and FM_ARM_ATTACH_POLL=0.2 knobs, so the real
poll loops still run. The park-boundary case moves onto the injected
test clock instead of a real 3s wait, per-case cleanup polls the host
pid rather than sleeping a full second, and the proof-by-absence
windows (flood re-escalation, successor re-announce, watcher
persistence, recovery staying off main) shrink from 2-3s to 1s, which
still spans two watcher polls at the test cadence.

Every assertion, process lifecycle, and reaping path is unchanged;
production defaults stay at one second. Isolated case timings on a
contended host, base vs branch: attended-latch 54.3->34.6s,
undelivered-dialog 67.7->59.1s, away-latch 46.5->30.5s, held-cadence
47.9->21.6s, unreadable-mirror 39.2->38.5s, park-limit 18.2->12.3s,
registration-fallback 14.1->10.0s, first-cycle-status 12.6->8.4s,
latch-scope 16.7->16.3s. Full suite: 65/65 pass. fm-lint and
shellcheck clean.

* no-mistakes(review): Wait for scan lock release before duplicate check

* no-mistakes(document): Correct supervision snapshot cadence documentation

* fix(tests): keep production poll cadence, probe exits at 0.1s

The fractional poll cadences multiplied the cost of each loop body:
full process-table scans in the engine turn and process refreshes in
await_close ran five times more often, which swamped the thin CI runner
and nearly doubled every multi-cycle case (serial 5 was cancelled at its
30-minute limit on run 36635306527's successor). Restore the production
cadence and notice arm/engine exits with a cheap kill -0 probe at a
tenth of a second between the one-second bodies instead: strictly less
dead time than baseline with no added CPU.

Also hold each injected-clock park bound well past its case's
wall-clock checks so a host that ignored the test clock fails instead
of silently passing at a real-time boundary, and restore the shortened
proof windows (watcher liveness, recovery-off-main absence, first-cycle
stream) to their baseline depth.

* no-mistakes(document): Clarify supervision engine snapshot documentation
…henguid#6192)

* fix: rebalance portable CI from current duration measurements

* no-mistakes(test): Test serial packing boundary and verify endpoint timeout cleanup

* no-mistakes(document): Clarify timeout guidance and remove duplicated packing estimates
…nguid#6216)

* fix(bin): run no repository hook when core.hooksPath is empty

The per-task hook wrapper refused every commit in a repository whose own
config sets core.hooksPath to the empty string, because git rev-parse
--git-path hooks fails on it. Plain git reads that setting as no hooks, so
the wrapper now runs none; every other lookup failure still refuses and
shows git's error.

Fixes kunchenguid#6171

* no-mistakes(review): Refuse commits when core.hooksPath is a valueless key

* no-mistakes(document): Document empty core.hooksPath handling in commit attribution docs

* no-mistakes(ci): When the wrapper refuses a commit, Git's hook-lookup error now shows up once instead of twice. That required changing one line in the wrapper, and the tests were extended so both bad-config cases would catch the duplicate. Invariant: when the wrapper refuses, Git's lookup error must appear exactly once. In the failure path, the only Git call besides the deliberate second lookup is the `git config --get --type=path core.hooksPath` check in `runtime_chain_body` (`bin/fm-git-strip-ai-trailers.sh:168`). That check prints the same error, so it was the one place to fix. I added `2>/dev/null` to it. Its exit status still decides the outcome: an empty value still runs no hook, and anything else goes on to the second lookup, which prints Git's error once, and the commit is refused. Tests (`tests/fm-git-strip-ai-trailers.test.sh`): - The unresolvable-path test (`~fm-no-such-user-6171/hooks`) now requires `failed to expand user dir` to appear exactly once in the refused commit's output. - The valueless-key test now requires `missing value for 'core.hookspath'` to appear exactly once. - Pre-existing bug in the unresolvable-path test: its `git add` ran after the bad config was set, so it failed silently (exit 128) and the "refused commit" had nothing staged. The test now stages the file before writing the config, the same way the valueless test does, so a real commit gets refused. - The empty-string test is unchanged and still passes, so an empty `core.hooksPath` still runs no hook. Verification: - With the wrapper change reverted, both new checks fail with `expected '1', got '2'`. With the change in place, the whole suite passes. - `bash -n` passes. shellcheck shows only an info-level SC1091 note about sourcing `lib.sh`, which was already there before this change. - `git status` lists only the two intended files
…kunchenguid#6213)

* fix(bin): let a stale record on a reassigned slot retire records-only

When a pool slot's owner claim names another task, the stale record's
teardown touches nothing under the slot, so the exclusive-slot record scan
no longer refuses it. Full teardowns of a slot this task still claims, or
one with no claim, keep the refusal.

Fixes kunchenguid#6184

* no-mistakes(document): Note claim-over-record precedence for reassigned teardown slots
…uid#6240)

* fix(bin): keep the steering doorbell short under deep homes

The doorbell printed the task inbox's absolute path twice, so under a deep
home it grew to about 290 characters and a Herdr submit reported it never
reached the pane on every re-ring. It now names the inbox once by its short
<task>.inbox name and points at the full path the worker's brief already
gives, so its length no longer depends on the home's depth.

Fixes kunchenguid#6120

* no-mistakes(review): Export FM_TASK_INBOX at launch and name it in doorbell

* no-mistakes(ci): ci-1 (Behavior portable serial 9) was caused by this PR, and I fixed it in the test. tests/fm-claude-trust.test.sh failed with "the launch command did not carry a brief doorbell". Its claude_launch_doorbell helper stripped exactly two leading `export ...;` statements before reading the final prompt argument. This PR adds a third one (`export FM_TASK_INBOX=...`) to every launch, so the helper was reading the wrong command. The invariant: a test that parses the launch command must skip every leading export statement, however many there are. I checked every test that parses the launch this way. The only other ones are the two helpers in tests/fm-spawn-dispatch-profile.test.sh, and they already loop over all exports. The kimi and dispatch-profile exact-string checks were updated earlier in this PR. The fix makes claude_launch_doorbell use the same loop (`while [[ "$command" == export\ *\;* ]]; do command=${command#*; }; done`) and then take the last argument. The ordinary path still works: the claude spawn test and the secondmate-clone spawn test both resolve the brief record through the same helper. Verified locally: `bash tests/fm-claude-trust.test.sh` exits 0 with no failing cases. ci-2 (Behavior tests (Herdr)) was not caused by this change, and I made no code change for it. In tests/fm-backend-herdr-presentation-e2e.test.sh, the concurrent secondmate recovery failed with "herdr presentation recovery could not acquire its session lock; refusing a concurrent resume". Two reasons it is not this PR: - The same failure, in the same test and case, happened on run 36655209015 for the unrelated branch fm/fm-contributions-old-gh-compat about 14 hours earlier. - This PR's change cannot lengthen how long the lock is held. The launch is written to a file and sent to the pane as `. launch.N.sh`, so the extra export changes neither the pane submit nor the lock hold time. The cause is a race that was already there: spawn_herdr_presentation_order_lock_acquire gives up after 5 seconds, and a concurrent real-Herdr recovery can hold the lock longer. Fixing that means changing the product's lock timeout, which is outside this PR. It should be tracked separately, and a rerun of the Herdr job is expected to pass. The only file changed is tests/fm-claude-trust.test.sh
… no turns (kunchenguid#4859)

* fix(dod): drive no-mistakes with one foreground call, not a background poll

The brief told workers to background the drive call and poll `axi status`
because one call "routinely outlives what your harness lets a single
command run". That advice contradicts the tool it drives: `no-mistakes
axi run --help` documents `--wait` with an 8m default, existing precisely
"so an agent harness with a 10-minute tool cap gets a structured return
instead of an unbounded hang".

Following the old text, a worker could never idle - a backgrounded call
returns in milliseconds, so it does not wait at all - and each attempt
leaked a live timer that later fired as a paid wake. Tell workers to make
one foreground call, let it block, and repeat it when it returns on
elapsed wait rather than on a gate or outcome.

Also drops the generalisation that told workers on any unestablished
harness to assume a command cap and use the same shape, which exported
the defect to harnesses with no such cap.

* fix(bin): let a waiting worker spend no turns until it is answered

A worker waiting on a decision, a pipeline gate, CI, or a heavy-test slot
kept taking model turns: the brief told it to list its inbox at any natural
checkpoint, and six automatic senders nudged secondmates whatever their open
decisions.

- The ship and scout briefs gain one Waiting section: end the turn after
  needs-decision or blocked, and hold an external wait inside ONE blocking
  command bounded by the harness's own command ceiling. The checkpoint clause
  is deleted. Forbidding the wrong shapes is not enough on its own, so the
  section also names the blocking foreground `until` loop as the wait a Claude
  Code worker may use, because that harness can refuse a sleep-then-check
  command while pointing at backgrounding, which is the one shape a waiting
  worker must not take.
- fm-send --automatic defers (exit 4, nothing written or rung) while the
  target has an open decision or blocker of its own; every automatic sender
  passes it and keeps its retry state, and the pending-reply recovery waits
  the same way.
- The two senders that report the result classified it by matching the text of
  the send's captured output against `deferred:*`. fm-send runs bin/fm-guard.sh
  as a supervision warning, and that guard prints its worktree-tangle banner
  whenever the primary checkout is on a feature branch, which is exactly what a
  CI pull-request checkout is. The banner lands ahead of the `deferred:` line,
  so the match fell through and a waiting mate was reported as a failed send,
  with the banner as the reason. Both senders now classify on fm-send's exit
  status, which is the contract the deferral is actually stated in, and select
  the `deferred:` line out of the output rather than assuming it came first.

The third root cause, a no-mistakes definition of done that backgrounded the
drive call and polled axi status, is fixed by this branch's parent commit
"drive no-mistakes with one foreground call, not a background poll"; this
commit takes that text as is and adds the regression test.

Upstream's spawn abort path no longer calls the lease-return helper at all,
so the fork's missing-helper guard and its pin-feature test line are moot
here and are not ported.

The command ceilings each harness enforces, and the probes behind the named
Claude Code wait, are recorded in docs/verification/runtime-backends.md.

* no-mistakes(review): Exempt captain holds, quiet deferred reconcile, clarify worker pauses

* no-mistakes(document): Document deferred automatic nudges, rereads, and reply recovery

* no-mistakes(document): Ring unlanded fire-and-forget steers exactly once more

* no-mistakes(ci): The failing check, "PR must be raised via no-mistakes", reads the pipeline's attestation record, which says document=skipped. No file in the repository can change that record, so I did not touch the check or the PR body. As you said, the no-mistakes rerun after this run finishes will re-execute the document step and record document=completed. The one change is the documentation sentence you ordered. It adds a line to docs/remote-secondmates.md, right after the line saying the remote host runs no re-ring ladder of its own: "A fire-and-forget record, such as a reconcile ask, gets its single retry ring only on the local plane: the remote steer leg owes no re-ring, so a swallowed remote doorbell for one waits for the next ring into that inbox, and a remote-side retry is known follow-up scope." No behavior changed. Checks: tests/fm-documentation-audiences.test.sh passes (4/4) and bin/fm-lint.sh is clean. The change is left uncommitted in the working tree for the pipeline to pick up

* no-mistakes(review): Hold automatic wakes until a mate's own decision closes

* no-mistakes(document): Document watcher delivery of deferred remote re-read nudges

* no-mistakes(review): Merge duplicate elapsed-wait reattach instructions in DOD

* no-mistakes(test): Resolve merged default decision in remote-reply recovery fixture

* no-mistakes(test): Source classify lib so config-push retry-deferred honors open decisions

* no-mistakes(ci): Fixed a flaky test that also fails on main. Neither this PR's bin/fm-brief.sh nor its bin/fm-dod-lib.sh change is involved: bin/fm-dispatch-resolve.sh sources neither file. Another branch (fm-attended-cutover-smoothing-s1, run 36343879084) failed the same shard 8 check the same way, on a different case ("a rule-criterion match prints one diagnostic line, got 2"). Root cause: `fm_quota_single_provider_for_harness` in bin/fm-quota-axi-lib.sh returned from its `while read` loop as soon as it found a match. That closed the pipe while `fm_quota_single_provider_table`'s `printf` was sometimes still writing. GitHub Actions runners ignore SIGPIPE, so bash printed `fm-quota-axi-lib.sh: line 138: printf: write error: Broken pipe` to the resolver's stderr. That is the extra line. I reproduced it locally by running the test with SIGPIPE ignored: 2 of 20 runs failed, one with the resolver's diagnostic line plus two broken-pipe lines. Invariant: looking up a harness in the provider table must never make the table writer fail. The only reader of that table is this function, and all of the resolver's lookups (line 208 without stderr redirected, line 222 with it) go through it. So the fix is in that one place: read the whole table, then print the match. The same file now shows it reads the full table first, like `fm_control_harness_supported` does. Return values and output are unchanged. Verification: with SIGPIPE ignored, tests/fm-dispatch-resolve.test.sh failed 0 of 30 runs after the fix (2 of 20 before). tests/fm-dispatch-resolve.test.sh, tests/fm-brief.test.sh, tests/fm-send-inbox.test.sh, tests/fm-quota-choose.test.sh and tests/fm-quota-array-dispatch-live-e2e.test.sh all pass, and shellcheck is clean. tests/fm-procevent-quota.test.sh fails locally with or without the change ("process-event state root is not a private directory"), so that failure comes from the local environment, not from this fix. No new test was added: the existing one-diagnostic-line assertions already catch this whenever SIGPIPE is ignored, as it is in CI

* Revert "no-mistakes(ci): Fixed a flaky test that also fails on main. Neither this PR's bin/fm-brief.sh nor its bin/fm-dod-lib.sh change is involved: bin/fm-dispatch-resolve.sh sources neither file. Another branch (fm-attended-cutover-smoothing-s1, run 36343879084) failed the same shard 8 check the same way, on a different case ("a rule-criterion match prints one diagnostic line, got 2"). Root cause: `fm_quota_single_provider_for_harness` in bin/fm-quota-axi-lib.sh returned from its `while read` loop as soon as it found a match. That closed the pipe while `fm_quota_single_provider_table`'s `printf` was sometimes still writing. GitHub Actions runners ignore SIGPIPE, so bash printed `fm-quota-axi-lib.sh: line 138: printf: write error: Broken pipe` to the resolver's stderr. That is the extra line. I reproduced it locally by running the test with SIGPIPE ignored: 2 of 20 runs failed, one with the resolver's diagnostic line plus two broken-pipe lines. Invariant: looking up a harness in the provider table must never make the table writer fail. The only reader of that table is this function, and all of the resolver's lookups (line 208 without stderr redirected, line 222 with it) go through it. So the fix is in that one place: read the whole table, then print the match. The same file now shows it reads the full table first, like `fm_control_harness_supported` does. Return values and output are unchanged. Verification: with SIGPIPE ignored, tests/fm-dispatch-resolve.test.sh failed 0 of 30 runs after the fix (2 of 20 before). tests/fm-dispatch-resolve.test.sh, tests/fm-brief.test.sh, tests/fm-send-inbox.test.sh, tests/fm-quota-choose.test.sh and tests/fm-quota-array-dispatch-live-e2e.test.sh all pass, and shellcheck is clean. tests/fm-procevent-quota.test.sh fails locally with or without the change ("process-event state root is not a private directory"), so that failure comes from the local environment, not from this fix. No new test was added: the existing one-diagnostic-line assertions already catch this whenever SIGPIPE is ignored, as it is in CI"

This reverts commit c719928.

* no-mistakes(review): Retry deferred local instruction nudges via the watcher

* no-mistakes(review): Document watcher retry for deferred local instruction nudges

* no-mistakes(ci): I fixed both review findings you selected (ci-1 and ci-3). I did not touch the deferral check in bin/fm-send.sh. ci-1 (bin/fm-config-push.sh, retry_deferred_rereads) - Rule that must hold: a deferred reread stays flagged until it is actually delivered. - Before the fix, the flag was removed before any of the steps that can skip a mate: the remote lock-path lookup, validate_secondmate_home, the local lock-path lookup, and the lock acquire. A skip at any of those dropped the flag, so the watcher lost track of the reread. - Now the flag is removed in one place only, when the send succeeds (rc 0). A skipped home, a busy lock, a deferred send (rc 4) or a failed send all leave it in place. The re-mark calls on a busy lock and on rc 4 were no longer needed, so I removed them. I updated the comment above the function to match. - Side effect: a send that keeps failing now stays flagged, so the watcher retries it on every poll and logs each failure. That follows your "don't clear until delivered" rule, but it replaces the old behaviour of leaving a failed send to the next config push or session start. - New test in tests/fm-secondmate-sync.test.sh: T8j "a deferred flag survives a skipped invalid home and is retried once it validates". It takes the home's marker away to make validation fail, checks that nothing is sent and the flag stays, then puts the marker back and checks that the nudge is delivered and both the flag and the retry marker are cleared. It fails on the old code and passes now. ci-3 (bin/fm-secondmate-restart.sh) - Rule that must hold: no automatic send wakes a mate that is waiting on its own open decision. - The two automatic sends in this script are the fallback reread nudge (fall_back_to_nudge) and the persist request. Both now pass --automatic. If a persist request is deferred, its correlation is discarded and the mate goes to the fallback nudge, which is also deferred, so the mate is reported as unreached. - New test in tests/fm-secondmate-restart.test.sh: T3b. It gives a mate an open needs-decision and runs a restart. It checks that both sends report as deferred, the mate's doorbell is never rung, its inbox gets no message, nothing is stopped, and the mate is reported as unreached with exit status 3. It fails on the old code and passes now. - The test marks the watcher as alive first. Without that, the watcher-down warning is printed first and becomes the reported reason instead of the deferral message. Verification - tests/fm-secondmate-sync.test.sh passes. - tests/fm-secondmate-restart.test.sh passes. - tests/fm-secondmate-harness.test.sh (the other test that exercises --retry-deferred) passes. - The fm-send-inbox test that covers automatic deferral passes. I only looked at the last lines of that run, not the whole file. - `shellcheck -x` on the four changed files is clean

* Pin autoarm supervision model in secondmate restart T3b

The fresh watcher beat the test writes proves a live watcher only under the
autoarm model; on CI hosts with no detected harness the persistent model
demands a lock-holding watcher, so the watcher-down banner became the
reported reason and the deferral assertion failed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Keep deferred secondmate nudges retryable under the inheritance lock.

A bootstrap instruction nudge could write its deferral flag outside the lock the watcher retry holds, so a concurrent retry could delete a flag that had just been set. A restart fallback that is deferred now records the same marker and flag, so the watcher delivers it once the decision closes.

* no-mistakes(document): Document watcher retry of deferred restart re-read nudges

* Send secondmate reread and restart nudges immediately again.

Deferring those nudges let a later config push drop an incomplete transfer once the decision closed. They now send as they do on main.

* Make the no-turn wait opt-in behind config/wait-no-turns.

Homes that do not create the file keep the previous briefs, drive text, and sends.

* no-mistakes(document): Document wait-no-turns inbox wording change in configuration

* no-mistakes(review): Keep checkpoint inbox check; forbid only polling while waiting

* no-mistakes(ci): Fixed ci-2 (Greptile: a concurrent retry marker gets lost). The rule that was broken: the watcher may remove only the `.retry-ring` mark for the record it just processed. A newer mark written in the meantime is owed its own retry. `fm_task_inbox_clear_retry` is the one shared function that removes the mark, and I fixed it there. In `bin/fm-task-inbox-lib.sh` it now takes the record path. It compares the mark's content with that record's name and removes the mark only when they match. When the mark names a different record it returns success and leaves the mark alone. It still fails only when the processed record's own mark can't be removed. Both callers in `bin/fm-watch.sh` now pass `"$rec"`: the dead or missing pane path and the path after a retry ring. So the fix holds at both removal sites. Tests, in `tests/fm-task-inbox.test.sh`: - I added an optional `FM_RING_MARKS_RETRY` hook to the fake tmux. It writes a newer record's mark while the doorbell is being typed, which reproduces the race deterministically. - I added `test_watcher_retry_keeps_a_newer_mark`. The owed retry rings once, the newer mark survives, and a later check rings the newer record once and then clears its mark. The test fails without the fix ("the spent retry removed a newer record's mark written during its ring") and passes with it. - I updated the direct `clear_retry` call in the existing unit test to pass the record. Results: `tests/fm-task-inbox.test.sh` passes in full and `tests/fm-send-inbox.test.sh` passes 15/15. Shellcheck reports only SC1091 "not following sourced file" notices. As instructed, I didn't change the brief inbox wording

* no-mistakes(document): Fix stale wait-no-turns inbox wording in inbox lib comment

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
…note (kunchenguid#6140)

* fix(bin): record Gerrit change URLs as close notes

Teardown's backlog_done_args hands every ship's recorded pr= URL to
fm_backlog_done as --pr, and tasks-axi refuses any --pr that is not a
canonical GitHub or Forgejo pull request. A Gerrit change URL therefore
left the item In flight after cleanup, and the pending backlog-close
record replayed into the same refusal at every session start.

fm_backlog_done now rewrites a --pr whose value fm_pr_url_parse reads
as a Gerrit change into --note "Gerrit change <url>". The mapping sits
at the tasks-axi call rather than in the pending-close record, so
records already written with --pr replay to a close unchanged. The
captain-held retain path records the URL in its deliverable line and
skips the update --pr it cannot make.

* no-mistakes(review): Note retained Gerrit change URL when captain answers early

* no-mistakes(document): Document Gerrit change URL handling in captain-hold retention
* perf(remote): separate active job sampling from dispatcher cadence

* no-mistakes(document): Link remote wait timing to its authoritative contract

* no-mistakes(ci): Fixed ci-1 with two narrowly scoped SC2030 annotations documenting intentional subshell-local legacy and active cadence overrides in tests/fm-remote-job.test.sh. Runtime behavior is unchanged. Reproduced the lint failure before the fix; afterward ShellCheck 0.11.0 with source following, Bash syntax validation, the complete remote-job behavior suite, and git diff --check all passed

* perf(supervision): reduce park, delta and dispatcher polling

* no-mistakes(document): Clarify poll latency contracts and authoritative documentation pointers
…henguid#6221)

* fix(bin): load backend sibling libraries under zsh

fm_backend_source kept each backend's sibling list in one space-separated
string and iterated it unquoted. zsh does not word-split an unquoted
expansion, so the readability check saw the whole list as one path and
refused every backend with more than one sibling. Hold the list in the
function's positional parameters instead, which needs no word splitting
in Bash 3.2, Bash 5, or zsh.

The existing zsh case in tests/fm-backend.test.sh covers it wherever zsh
is installed.

* test: run the Calm mod suite on stock Bash 3.2

The suite injected shell values into its generated Node scripts with the
${value@Q} transformation, which needs Bash 4.4. Stock macOS Bash 3.2
reports a bad substitution, so every case failed before it asserted
anything. Build each JavaScript string literal with JSON.stringify
through a small helper instead, which works on any Bash and is a valid
literal for any value.

* no-mistakes(review): fix(bin): rename zsh-special path local in fm_backend_source

* test: narrow the zsh backend claim to name matching

Under zsh the adapters locate their siblings through BASH_SOURCE, so a
successful fm_backend_source is not a full load. Assert only what the
contract states, and pass js_string values after -- so node never reads
a leading-dash value as its own option.

---------

Co-authored-by: Nova Agent B <novaagentb@gmail.com>
…tatus scans (kunchenguid#5263)

* fix(bin): exclude a remote mate's own parent channel from self-home scans

A remote secondmate home's outbound parent channel lives at state/parent-replies.status inside its own state dir, so the watcher's signal scan enumerated it as a task status file and the open-decisions fold classified it as a phantom task named parent-replies: every parent-channel append spun a spurious signal wake and a phantom open decision in the mate's own home.
fm-parent-channel-lib.sh gains fm_parent_channel_outbound_status, which resolves the channel into the mate's own state dir for the remote route only, and fm-classify-lib.sh's status_scan_parent_channel_exclude wraps it for the fleet-wide scans.
The watcher's scan_signals and heartbeat fail-safe backstop, the whole-file and incremental open-decisions folds, the presentation snapshot, and the unread-surface scan now skip exactly that resolved path.
The exclusion is home-shape-aware: a parent-replies.status in a main home or a local mate is an ordinary task log and keeps waking and folding, and every other status file is untouched.

* no-mistakes(review): exclude a remote mate's parent channel from the daemon heartbeat scan

* no-mistakes(document): Document remote mate parent-channel scan exclusion

* ci: retrigger portable serial 4

* no-mistakes(ci): CI check 'Behavior portable serial 7' failed in tests/fm-contributions.test.sh ('reservation poll failed'). CI stderr showed bin/fm-contributions.sh:345 arithmetic 'DEADLINE - 6\n90077104: syntax error in expression': the fixture's fake date returned a torn two-line clock value. Root cause: the fake forge wrapper in wrap_forge advances the shared controllable clock via a non-atomic read-modify-write ('$(cat $FORGE/clock) + 6' with truncate-in-place '> $FORGE/clock') while concurrent background gh calls run and the fake date reads the same file; an interleaved truncate+write publishes a half-written value (CI's torn '6\n90077104', tail of 1790077104) or an emptied-read value ('6'), which either breaks the poll's arithmetic (nonzero exit -> 'reservation poll failed') or defeats the 15-second reservation defer. This is a pre-existing test-fixture race, not caused by the PR's diff (base..target touches no contributions code; the same commit passed this shard in run 35711207830 earlier the same day). Fixed the flaky fixture at its root: clock_bump() now writes each new value to a per-process mktemp file in the same directory and publishes it with mv (atomic rename), so concurrent forge callers and the fake date always read one complete old-or-new clock; fault patterns and deltas are unchanged. Verified: minimal 3-way concurrency repro shows the old wrapper corrupting (12/32/38 outcomes incl. empty-read) while the rename-based wrapper never corrupts (20/20 clean); the full tests/fm-contributions.test.sh passes twice (all 38 assertions ok, incl. the reservation, budget-exhaustion, genuine-failure, shared-once, and latency tests); 10 isolated reservation runs pass; shellcheck rc=0; worktree contains only this one-file change

* no-mistakes(document): drop stale file-set copy in daemon catch-all comment
…6307)

* fix(bin): name the accepted verdict actors in fm-contributions help and refusal

* fix(ci): Updated tests/fm-contributions.test.sh to assert exactly captain, fleet, maintainer, and nobody in command-emitted help and refusal output. Three focused regressions passed; all three extra-actor mutations were rejected. ShellCheck, syntax, and diff checks passed. Production code remains unchanged
…chenguid#6306)

* fix(bin): recognise a clone root git names with different path spelling

fm-fleet-sync compared git's --show-toplevel with pwd -P as strings, so a clone
root that git recorded with different casing (case-insensitive volume) was
skipped as not a clone root and never refreshed. Compare filesystem identity
instead, which also covers symlink spelling.

* fix(document): Remove stale clone-root comparison comment
)

* test(calm): pin Pi's regular TUI mode where pane assertions read scrollback

Pi 1.0.0 defaults its TUI to a fullscreen alternate-screen mode whose
scrollable transcript is application-owned, so rows that leave the viewport
never enter terminal scrollback and tmux capture-pane -S can no longer see
them. The Pi Calm e2e launches now pass --tui-mode regular wherever the flag
exists so the transcript assertions keep reading real scrollback on both the
Pi 1.0.0 line and earlier Pi lines, which have no such flag and render
regular-only anyway.

* no-mistakes(document): Correct Pi TUI documentation and scrollback rationale
…ries (kunchenguid#6331)

* fix(bin): encode captain-hold reasons and reject self-inventory in complete

hold now stores a reason with parentheses, line breaks, or percent signs
through a reversible percent encoding that every reader decodes, instead of
refusing it. hold --origin records the origin on the held task, and complete
refuses the origin as its own inventory entry and an entry held for a
different origin; holds with no recorded origin are accepted and flagged.

* fix(review): Decode marked hold reasons consistently across readers

* fix(review): Remove unnecessary lifecycle test dispatch

* fix(review): Correct hold origin identity and inventory recovery

* fix(review): Record origins before placing backend holds

* fix(document): Clarify captain-hold validation and reason reader documentation

* fix(ci): Fixed both findings: failed backend holds restore the previous origin, and invalid base64/UTF-8 reasons remain verbatim. Added regressions and documented valid-literal ambiguity. Both failures were reproduced before fixes. Verification: 54 lifecycle tests and 9 wrapper tests passed; 7 Beads-specific cases skipped because tasks-axi is markdown-only. Focused lint and diff checks passed. No pipeline or publication actions performed
* fix(bin): take over the watcher cycle a main-only pass-through leaves

An attended main-only pass-through leaves a successor watcher cycle
running through main's handling turn. The session's next park attached
to that cycle instead of owning it, so the successor's arm, orphaned by
its host's exit, kept owning the watcher while the new park's arm polled
it twice a second until the next close or the park boundary, hours later
in a quiet second mate. A second-mate restart hit this every time, since
its persist request is a main-only close.

The host now records the successor it leaves for main, and the next
host's first cycle runs bin/fm-watch-arm.sh --take-over on it: when that
arm still owns the healthy watcher, the new arm stops it, reports a
reason the cycle delivered first, and otherwise owns a fresh cycle. The
stop's own downtime publication is undone over an acknowledged episode
when no wake was appended in between, so the handover wakes nobody.

* no-mistakes(review): Keep left-arm record until the orphaned arm is gone

* no-mistakes(review): Relinquish successor arm only after durably recording it

* no-mistakes(review): Relinquish successor only after its record reads back

* no-mistakes(document): Clarify successor takeover guarantees and authoritative documentation

* no-mistakes(ci): Fixed ci-1: acknowledgement restore now requires the exact taken-over arm/watcher ledger row with signal=TERM, awaited within a short bound. Otherwise takeover proceeds without erasing downtime. Added a self-exit regression confirmed failing before the fix and passing afterward; ordinary takeover tests and the full watcher-arm suite pass. Updated Generation reuse documentation. Syntax, diff checks, and ShellCheck pass with existing SC1091/SC2034 warnings excluded. ci-2 remains unchanged

* no-mistakes(document): Clarify watcher take-over recovery and restart limits
…cessor already closed (kunchenguid#6355)

* fix(bin): restore supervision host hand-back continuity

* no-mistakes(review): Scope host hand-back failure fallback to lost pending:handling

* no-mistakes(review): Remove stray scratch test copy tests/.tmp-rest.test.sh

* no-mistakes(test): Initialise successor globals so early hand-back survives set -u

* no-mistakes(document): Document host hand-back downtime failure and Claude lost-handback notice

* no-mistakes(ci): I fixed the Greptile finding. The rule that must hold: when the supervision host hands back an actionable wake, its rewake is refused, and no watcher is healthy, the hand-back still has to reach main as a delivered notice. That must be true whether the recovery marker is `pending:handling` or `announced:handling`. Only one place applies this check: the lost hand-back fallback in `bin/fm-claude-stop-autoarm.sh`. **Fix:** that check now accepts both `pending:handling:*` and `announced:handling:*` tokens (a one-line change). Nothing else in the fallback changed: - Refusals on any other marker, such as an already acknowledged one, still exit 0 silently and open no failure episode. - The notice is still sent once per episode, and repeats are recorded as `failed-suppressed`. **Tests:** - `tests/fm-claude-stop-autoarm.test.sh`: the lost hand-back test now runs as a shared helper with two variants, one writing a `pending:handling` marker and a new one writing `announced:handling` (`test_host_lost_announced_handback_notifies_once_per_episode`). - `tests/fm-supervision-host.test.sh`: the end-to-end test where downtime restoration fails is now a shared helper too, with a new `announced` variant (`test_claude_stop_hook_notifies_when_closed_announced_successor_downtime_restore_fails`). It moves the handling episode to `announced` before the host hands back. Without the fix the hook would exit 0 here; the test requires exit 2, `outcome=failed` and a delivered failure notice. **Verification (all under nice -n 10):** - The full `tests/fm-claude-stop-autoarm.test.sh` suite passed (rc=0), including both lost hand-back variants and the benign-refusal test. - In `tests/fm-supervision-host.test.sh` I ran only the four hand-back test functions, all passing (rc=0). The suite can't run single functions, so I used a temporary copy with a trimmed test list and deleted it afterwards; `git status` shows only the 3 intended files changed. - shellcheck is clean on all three changed files. I did not run `bin/fm-lint.sh`. - I did not run the new tests against the unfixed code; the claim that they fail without the fix comes from reading the old check
Freudator86 and others added 29 commits October 6, 2026 15:36
* Add daily startup growth check

* no-mistakes(review): watch printed startup memory, retain growth baselines, drop knobs

* no-mistakes(review): delegate budget verdict, report before publish, pin shim home

* no-mistakes(review): baseline first sightings silently, drop mtime, spare secondmates

* no-mistakes(review): cap the wake line, validate budget verdict fields

* no-mistakes(review): guard record schema, check appends, tighten assertions

* no-mistakes(review): drop arbitrary tracked library, make re-arm idempotent

* no-mistakes(review): correct tracked-set wording, restore shim guards, register artifacts

* no-mistakes(review): exit on signal instead of publishing partial record

* no-mistakes(document): correct startup-growth record removal cost in state registry

* no-mistakes(test): sweep orphaned empty startup-growth temp records on due evaluation

* no-mistakes(document): note watcher need for armed startup growth check

* no-mistakes(ci): Diagnosed "Behavior portable serial 5": the shard failed on tests/fm-contributions.test.sh in test_arm_plumbs_a_configured_budget_into_the_check_shim (inherited mode) with "generated check did not attempt a read" and a missing forge/calls file. Root cause: bin/fm-contributions.sh:379-382 computes DEADLINE=$(date +%s)+BUDGET and then gates each read on DEADLINE-$(date +%s) >= OBSERVATION_RESERVE (= min(BUDGET,15)). At the one-second budget this test uses, that gate demands zero elapsed whole seconds, so a wall-clock second boundary crossing between the two date calls makes the poll loop break before any forge read. The test calls wrap_forge (which installs a fake date reading $FORGE/clock only when that file exists) but never seeded forge/clock, so it ran against the real clock. The repo already documents this exact hazard at tests/fm-contributions.test.sh:675-677 and freezes the clock in every other budget-constrained test: test_budget_exhaustion_keeps_prior_record (both modes), test_genuine_failure_near_deadline_is_unavailable, test_reservation_defers_later_url_when_fifteen_seconds_do_not_remain, test_slow_read_deadline_kill_is_budget_refusal, test_budget_is_cut_down_to_the_watcher_check_bound. test_arm_plumbs was the only omission; its single loop body covers both the configured and inherited modes. The remaining wrap_forge users run the default 20-second budget (5 seconds of slack above the reserve) and are not reachable by this race. Fix (smallest, matching the sibling idiom): added the clock freeze `/bin/date +%s > "$home/forge/clock"` plus a two-line comment inside that test's loop, before the poll. Three added lines in tests/fm-contributions.test.sh; no production code and no other file touched. Verification: reproduced the exact CI failure message by running a scratch copy whose fake date advances one second per read (worst case of the unfrozen clock) - missing forge/calls, "not ok - generated check did not attempt a read". With the fix, the single test passes 10/10 standalone and the full tests/fm-contributions.test.sh passes all 46 assertions with exit 0. tests/fm-startup-growth-check.test.sh still exits 0. bash -n and shellcheck -S warning are clean on the edited file; scratch repro files were removed, leaving only the intended three-line change. Note for the author: this flake is not caused by this branch - tests/fm-contributions.test.sh is untouched by the PR and the race is pre-existing, surfacing only on a loaded runner (the shard's neighbouring assertions show 80+ second gaps). It was fixed because it is genuine nondeterminism with an established in-repo remedy and the check cannot otherwise go green. No work was done on the unselected greptile findings (ci-2..ci-5) or the cancelled "Behavior portable serial 2" check (ci-6)

* no-mistakes(ci): Diagnosed "Behavior portable serial 6": the omitted log section (fetched with gh api) shows the shard failed on tests/fm-calm-pi-extension.test.sh in test_hidden_block_geometry_e2e with "Pi Calm hidden-block geometry E2E did not complete the /reload viewport transition" (exit=1, duration_ms=23437; the family line pure-contract-unit count=3 duration_ms=51193 failed=1 matches devin-harness 3203 + calm-pi 23437 + task-delivery 24553). Root cause: tests/fm-calm-pi-extension.test.sh:2831 wait_for_geometry_transition detected the /reload by SAMPLING a single intermediate frame - it polled tmux capture-pane for Pi's transient "Reloading keybindings, extensions, skills, prompts, themes, and context files..." box and only accepted the rebuilt transcript afterwards (elif gated on saw_transient). Pi shows that box only while the reload runs and replaces it via dismissReloadBox when done, so on a loaded runner the box can live entirely between two polls; saw_transient then stays 0 forever and the 600-attempt wait times out although the reload fully succeeded. Reproduced locally under CPU oversubscription with CI's Pi 1.0.4: 1 failure in 8 runs, diagnostics proving the mechanism ("DIAG: transition timeout saw_transient=0 attempt=600") while the captured viewport already showed the completed reload status row and the intact collapsed transcript. Not a Pi-version regression: 1.0.4's handleReloadCommand and both banner strings are identical to the installed 0.85.1, and shard 6 of the previous pipeline run (job 112490351105) passed this same script on Pi 1.0.4. Invariant violated: a TUI E2E assertion must key off a durable state the UI retains, never off one intermediate frame polling may miss. Sibling enumeration: the helper was defined once and called once; grep for transient/saw_/two-stage waits across tests/, bin/ and .pi/ found no other one-shot-frame wait - every other wait in the file keys off durable text presence/absence or off continuously repeating animation frames (the working-ship motion checks) where sampling eventually connects; no doc or other test referenced the helper. Fix (smallest, one helper): renamed it wait_for_geometry_reload and made it wait for the durable status row Pi appends once the reload completed and the chat was rebuilt ("Reloaded keybindings, extensions, skills, prompts, themes, and context files", a persistent chat child via showStatus, present in both 0.85.1 and 1.0.4) together with CALM_GEOMETRY_FINAL. That marker is absent before the reload and never appears if the reload throws, so the check is strictly stronger than before (a failed reload now fails the test instead of passing on transient-then-final). One file changed: tests/fm-calm-pi-extension.test.sh, 2 hunks, no production code. Verification: deterministic before/after - with polls spaced 0.4s apart (the loaded-runner worst case) the pre-fix helper fails with exactly the CI message and the post-fix helper passes; 12/12 consecutive passes of the isolated test under load with Pi 1.0.4; full file 15/15 ok exit 0 under LANG=C.utf8 (matching the runner locale); bash -n, shellcheck -S warning and bin/fm-lint.sh clean; scratch repro files and the temp Pi 1.0.4 install removed, git status shows only the one intended modification. Notes: (1) the flake is not caused by this branch - that test file is untouched by the PR and the race is pre-existing; fixed because it is genuine nondeterminism and the check cannot otherwise go green. (2) Under far harsher load than CI applies a separate unrelated flake appeared in the same test: Pi's "Warning: tmux extended-keys is off..." row occasionally lands between the collapsed [skill] ahoy row and the final response, making assert_geometry_gap see 5 rows instead of 2. Different cause, never seen in CI, and its plausible remedies would change key delivery for the other TUI tests in the file, so it was left alone rather than growing this fix

---------

Co-authored-by: Quartermaster <quartermaster@users.noreply.github.com>
kunchenguid#6708)

* fix: reopen a remote-reply continuity break after repair

A later break for the same route and reason was swallowed after the
operator resolved the first one, because the status line matched for
the life of the log. The continuity ingest now appends again when the
cursor has moved or retirement has reset that episode, and an unchanged
re-read still appends nothing.

status_event_recorded is unchanged. Its other callers are the
pending-reply escalation, which already decides its own episode, the
parent-channel note append, and the remote document transfer note.

* no-mistakes(review): seed continuity episode for already recorded break line

* no-mistakes(document): document when a remote-reply continuity break reopens

* no-mistakes(ci): The adapter now stores the continuity episode record before it appends the `blocked` line, so a failed store appends nothing. The changes are uncommitted in the worktree, in `bin/fm-procevent-remote-reply.sh` and `tests/fm-remote-reply.test.sh`. **Invariant:** a continuity `blocked` line is on the parent status log only when the episode record for that break is already stored. Only the `continuity-broken` branch of `cmd_ingest` appends that line, so the fix is at that one place. **What changed in `cmd_ingest`:** - It decides first whether the break needs a line, then stores the episode record, then appends. - If storing the record fails, it stops with "cannot record continuity episode" and appends nothing. - If the line is already on the log and no episode record exists, it stores the record and does not append. **One addition you did not ask for:** if the append fails after the record is stored, the adapter puts the earlier record back (or deletes the new one when none existed). Without that, the retry would read as the unchanged repeat and append nothing, which is the issue 6701 failure again. **Deviation from your test instruction:** the test does not use `chmod` on the cursor directory. `write_continuity_episode` runs `chmod 700` on that directory before every write, so a read-only directory is made writable again. The test instead makes `mktemp` fail for the episode's temporary file in that directory, the same way the existing receipt-failure test does. **Tests added to `tests/fm-remote-reply.test.sh`:** - Store failure: the second break exits 1, appends no `blocked` line and opens no decision. One retry after storage recovers exits 3 and appends a single line. - Append failure: with the status log read-only, the third break exits 1 and appends nothing. One retry after the log is writable exits 3 and appends a single line. **Verification:** `bash tests/fm-remote-reply.test.sh` ends with "ALL TESTS PASSED" with the fix. Against the script at commit b750c828 the same test file fails at "a continuity break appended its line before its episode was stored". `shellcheck -S warning` reports only an unused loop variable at line 667 of the test file, which this change does not touch. I ran no other test files. The Greptile Review check log could not be retrieved, so I worked from the finding text alone

* fix: record a continuity break's reader position on its status line

A later break at another cursor is then a different line, so the existing
duplicate check appends it and reopens the decision. An unchanged re-read
builds the same line and appends nothing.

* fix: reopen a continuity break after an identical restore

A retirement that puts the same bytes back used to rebuild the recorded line, so the later break stayed closed. The retirement count on that line makes the later break distinct.

* no-mistakes(review): remove continuity match for full-prefix line without retirement count

* no-mistakes(document): clarify what a continuity break status line records

* fix: remove the reply cursor before recording retirement

A stop between those steps must leave the count unchanged, so an unchanged continuity break still builds the same line.
…is retired (kunchenguid#6733)

* fix(control): drop busy_gen when an incarnation is retired

A deliberate exit removed the busy sidecar and left busy_gen in the task record, so the two records disagreed about whether that incarnation was still observable.

* no-mistakes(review): drop GNU-only chmod and unreached sidecar-absent branch

* no-mistakes(review): correct lock comment to name the deadlock

* no-mistakes(ci): The test `test_exit_drops_meta_busy_gen_with_the_sidecar` in tests/fm-control.test.sh now compares the whole task record (the `state/<id>.meta` file), so the Greptile finding is fixed. Invariant: after `exit` retires an incarnation, the task record must equal the record from before `exit` with only the `busy_gen` line removed. This test is the only place in the change that asserts the record survives the rewrite, so it is the only site to fix. The other `busy_gen` tests assert that the line stays, and they do not go through the rewrite. What changed: before `exit`, the test writes the record without its `busy_gen` line to `expected.meta`. After `exit`, the test runs `diff` between that expected copy and the real record, and fails with the diff output if they differ. This one comparison replaces the two earlier checks (no `busy_gen` line left, and the `window` line present), because it covers both. I did not change bin/fm-control.sh or any other file. How I know it works: - I ran `bash tests/fm-control.test.sh`: exit code 0, 45 lines starting with `ok`, no other lines. - I temporarily changed the rewrite in bin/fm-control.sh to also drop the `harness` line. The test then failed with `not ok - exit should drop only busy_gen from the task record:` and the diff `< harness=codex`. The earlier `window`-only check would have passed that rewrite. I restored bin/fm-control.sh afterwards; `git status` shows only tests/fm-control.test.sh modified. - `bash -n` and `shellcheck` on the test file report no new warnings from the edit. The change is not committed; the working tree holds it
…#6484)

* test(secondmate-harness): scope fake ps -codex label for pid 5252 to the liveness probe

Closes kunchenguid#6456

* no-mistakes(ci): Updated the collision test to log and assert that PID 5252 was queried before selecting 4242. Full fm-secondmate harness suite passes

---------

Co-authored-by: YifuGu <ironerumi@users.noreply.github.com>
* fix(calm): share the standalone Pi Calm working-ship widget slot

Firstmate Calm and the user-global standalone Pi Calm both install an
animated working-ship widget during agent runs. Each claimed its own Pi
widget key, so a session loading both (the main Firstmate home) rendered
two boats. Pi replaces widgets under one key, so claiming the shared
"calm-working-ship" slot keeps dual-install sessions to a single boat
while a Firstmate-only session is unchanged.

Pins the shared slot contract in the working-ship module test so the key
cannot silently diverge again.

* test(calm): pin the shared working-ship widget key in CI, document dual-install

The key-parity assertion inside the Pi fixture only runs where the
@earendil-works/pi-coding-agent package is installed, so CI never
exercised it. Add a source-level twin that needs nothing but the
tracked file, and note in docs/calm.md that the boat shares the
standalone Pi Calm working-row widget slot.

* no-mistakes(review): Add executable dual-install widget replacement coverage

* no-mistakes(review): Guard shared widget cleanup with disposal ownership

* no-mistakes(document): Document shared Calm working-ship slot behavior

* test(calm): read the standalone Calm slot from its own module

The dual-install check registered both boats itself under the shared
slot, so it could only prove that Pi replaces a widget under one key: it
would still pass if the standalone Pi Calm extension installed its boat
under a different key, which is the two-boat regression the check exists
to prevent.

Read the standalone extension's own working-ship module when it is
installed - FM_STANDALONE_CALM_SHIP, else ~/.pi/agent/extensions/calm -
and drive the check with the key that module exports, so a rename on
either side registers two widgets and fails naming both keys. A pinned
shared-slot contract still covers a machine without the extension, and
the run reports which side it used instead of passing silently over an
absent extension.

Verified: the touched Pi Calm suite passes and reads the installed
standalone extension; with a copy of it whose key is renamed to
calm-working-ship-v2 the suite fails naming the drift.

* no-mistakes(review): Gate stock-row restoration by shared-widget ownership

* no-mistakes(review): Removed redundant widget-key source assertions

* no-mistakes(document): Document shared Calm working-ship widget ownership
…id#6649)

* fix(herdr): make exact-resume presentation-lock wait instead of a bounded timeout

The exact-resume path in bin/fm-spawn.sh used the same 50-attempt-then-
give-up lock acquire as the new-task-create path, but the two paths are
not equivalent on contention: a create has no prior state to strand and
can safely fall back to a flat layout, while a resume is recovering a
specific existing identity that a concurrent recovery may legitimately
be holding the lock for. Giving up there does not degrade gracefully,
it hard-fails the resume outright. The suite's own concurrent
cross-home recoveries test already asserts both concurrent recoveries
succeed with a genuine reclaim, and the file's header comment already
(inaccurately) claimed lock contention falls back to the ordinary flat
layout for both paths alike, so the intended contract was always that
recoveries serialize and both succeed, not that either one refuses
under a short bound.

Give spawn_herdr_presentation_order_lock_acquire a wait mode that uses
this file's own established fm_lock_acquire_wait idiom (already used
for its other fleet-shared locks) instead of the bounded loop, and use
it only at the exact-resume call site. The new-task-create call site
is unchanged and keeps its bounded-then-flat-fallback behavior, which
is already covered by its own passing test. Dead-owner PID-liveness
reclaim inside fm_lock_try_acquire still bounds the wait against a
holder that crashed mid-hold.

Adds a deterministic regression test that holds the shared session
lock from an unrelated process for well past the old bound, then
asserts the resume succeeds with a genuine reclaim and took close to
the full hold duration, so a fix that merely widens the bound rather
than genuinely waiting is still caught. The existing concurrent
cross-home recovery test exercises this under real timing but does not
reliably outlast a fixed bound on its own.

Corrects the header comment's claim that create and resume share one
bounded-then-flat-fallback behavior on lock contention; they no longer
do.

* no-mistakes(document): Document Herdr recovery waiting for presentation lock

* no-mistakes(document): Update stale hard-refusal claim in verification log

* no-mistakes(ci): Fixed the Greptile finding on tests/fm-backend-herdr-presentation-e2e.test.sh:1389 by bounding the resume lock-wait regression's spawn_task call. Added an optional 4th `deadline_seconds` arg to the `spawn_task` helper (defaults to empty, so all ~20 other existing call sites are unaffected and unwrapped by `timeout`). The lock-wait test now passes `LOCK_WAIT_HOLD_SECONDS + 60` (90s) as the deadline, and a dedicated check for exit code 124 emits a clear "hung for over Xs instead of waiting out a Ys lock hold" diagnostic before falling through to the existing pass/fail assertions, which are unchanged. No product code was touched. Verified with `bash -n`, `shellcheck -x` (no warnings), a standalone reproduction of the timeout/no-timeout/success paths, the project's `bin/fm-lint.sh --fast` on the file (clean), and the full `tests/fm-lint.test.sh` suite (all 46 assertions pass)

* no-mistakes(ci): Replaced the direct `timeout "$deadline_seconds"` call in `spawn_task()` (tests/fm-backend-herdr-presentation-e2e.test.sh) with the repo's portable bounded-execution helper: sourced `bin/fm-timeout-lib.sh` at the top of the file and changed `deadline_cmd=(timeout "$deadline_seconds")` to `deadline_cmd=(fm_run_timed "$deadline_seconds")`. This removes the GNU/BSD `timeout` dependency that would fail with exit 127 on a stock macOS host without coreutils, while preserving identical semantics (exit 124 on bound-hit, command's own exit otherwise), which the existing `[ "$LOCK_WAIT_STATUS" -eq 124 ]` diagnostic check already relies on. Verified: `bash -n` syntax check, `bin/fm-lint.sh --fast` clean, full `tests/fm-lint.test.sh` suite (46/46 pass), and a standalone repro confirming `fm_run_timed` returns 124 on timeout and 0 on success identically to the prior `timeout` call. No other direct `timeout` calls exist in this file or elsewhere in the PR's diff, so no sibling sites remain

* fix(herdr): gate exact-resume lock wait behind --herdr-resume-lock-wait

Keep refuse-by-default on presentation-order lock contention for Herdr
exact resume. Callers that need concurrent recoveries to serialize must
pass --herdr-resume-lock-wait; unbounded blocking on a third-party session
lock is never the default.

Update docs and the real-Herdr e2e suite so the default path asserts the
refusal and the opt-in path asserts the wait.

* no-mistakes(test): Fix e2e test's lost exit status after if/fi with no else branch

* docs(herdr): stop advertising --herdr-resume-lock-wait on --relaunch

The relaunch path reuses the recorded endpoint and never takes the
presentation-order lock, so the flag is inert there. Drop it from the
--relaunch usage line and state where the flag applies.

* no-mistakes(review): Clarify lock-wait docs; simplify bash-3.2-safe spawn_task helper

* no-mistakes(ci): Fixed ci-1 (Greptile P2). In tests/fm-backend-herdr-presentation-e2e.test.sh, the failure cleanup `cleanup_all` stopped only `LOCK_CONTENTION_OWNER_PID`. It now also stops `LOCK_REFUSE_HOLDER_PID` and `LOCK_WAIT_HOLDER_PID`, the holders of the two new contention cases, so a `fail` before their explicit `wait` no longer leaves them running. Both new PIDs are initialised empty next to the existing one, and each is cleared right after its successful `wait` so cleanup never touches a finished PID. I changed nothing else. `bash -n` passes. The real Herdr e2e run passed both new cases ("default resumed identity refuses session lock contention" and "--herdr-resume-lock-wait waits out session lock contention instead of refusing"). The full run hit my 550s timeout in a later, unrelated case, after the new cases passed
….2 (kunchenguid#6762)

* fix(bin): let TERM stop a watcher blocked in a pane capture on bash 3.2

Stock macOS bash 3.2 holds a HUP or TERM until a running command
substitution's child exits, and the watcher read every pane through
$(fm_backend_capture ...). A blocked backend read therefore held the
watcher's stop for as long as the read lasted, and a stopped watcher left
the hung read orphaned. tests/fm-watch-triage.test.sh
test_term_stops_a_watcher_blocked_inside_a_poll failed on /bin/bash 3.2
for this reason while passing on bash 5.

Pane captures now go through watcher_capture, which runs the read as a
waited background process group recorded like a check's, so the stop is
honored at once and watcher_cleanup stops a read still in flight along
with its per-call output file.

* no-mistakes(review): Run drain-ring idle capture in watcher shell, add regression test

* no-mistakes(document): Document watcher TERM handling for blocked checks and captures

* no-mistakes(test): Silence bash 3.2 setpgid race noise from watcher captures

* fix(bin): verify the capture group and scope the stop claim to pane reads

watcher_capture now confirms its background read leads its own process
group, as run_check_capture already does, so watcher_cleanup never relies
on a group that set -m failed to create. The comment and continuity doc
now say only fm_backend_capture pane reads go through watcher_capture;
agent-state and composer-state reads still run inside command
substitutions.
* feat(spawn): add per-home worker tool exclusions

Add an optional per-home config/crew-exclude-tools file listing tool names to hide from workers, one per line, with blank lines and # comments allowed.
It applies to every ship and scout launch and relaunch in that home, is never inherited by another home, and does not affect secondmate agents.
Pi and pi-signed apply it through --exclude-tools, which also covers MCP tool names.
Any other runtime, and a raw launch command, refuses the launch when the list is non-empty rather than ignoring it.
Malformed entries are refused before provisioning, and before a relaunch stops a running worker.
Exclusions that match no tool in the worker's loaded registry are reported as unverified warnings in its status record instead of refusing the worker.

Closes kunchenguid#6744

* no-mistakes(review): Preserve UTF-8 exclusion paths and verify Pi lifecycle behavior

* no-mistakes(document): Clarify worker tool exclusion documentation

* no-mistakes(ci): Fixed ci-1 in bin/fm-exclude-tools-lib.sh: a failed read now returns an error before printing names, so all shared launch and relaunch callers refuse rather than silently dropping exclusions. Added deterministic regression coverage for a file disappearing after readability checks across Pi/pi-signed ship and scout launches. Reproduced the original failure; verified 83 spawn checks, 77 relaunch checks, direct parser/runtime failure cases, full targeted lint, Bash syntax, and git diff --check. Relaunch tests passed with existing fixture-cleanup permission warnings. ci-2 remains unchanged per the user's decision; the outer executor owns the fresh CI run
kunchenguid#5343)

* refactor(bin): share the local Firstmate home walk from the wake library

Teardown's walk over the root home and its registered local secondmate homes
moves into bin/fm-wake-lib.sh as fm_local_firstmate_state_dirs, next to
fm_firstmate_root_home, so a second consumer can count task records across
this machine's homes without a copy. Teardown keeps its exact refusal wording
through a thin wrapper.

* feat(bin): defer spawns beyond a project's declared machine capacity

A project whose machine-local resource only serves a few workers at once had
no way to tell Firstmate so: every queued item was launched, and the surplus
workers spent full-context turns retrying the resource.

config/project-capacity in the root home now declares how many workers each
named project admits at once on this machine. bin/fm-spawn.sh counts the ship
and scout records on the same project origin across the root and its local
secondmate homes, skipping ones whose ready PR is recorded, while holding the
shared project lock through publication. A spawn with every place held exits 75
before any brief render, endpoint, worktree, record, or backlog move, so the
item stays queued; batches report it as deferred. Undeclared projects keep
today's uncapped dispatch, and an unreadable declaration refuses rather than
guessing the limit.

Refs kunchenguid#4237

* no-mistakes(review): Document that capacity matches the clone directory name

* no-mistakes(document): Rewrap stale fm-wake-lib root-home doc comment

* no-mistakes(review): Dedupe local state dirs by identity to avoid double-counting

* no-mistakes(document): Rewrap fm_local_firstmate_state_dirs error doc comment

* no-mistakes(ci): I fixed all four Greptile findings. All 14 tests in tests/fm-project-capacity.test.sh pass, and shellcheck at warning level is clean on the changed files. Each new test failed against the old code and passes now. - **ci-1 (spaced names):** a declaration line must give a name its capacity whenever the name is a valid clone directory name. `fm_project_capacity_lookup` now trims each line, skips blank lines and lines whose first non-blank character is `#`, and takes the last field as the capacity. Everything before that field is the name, so it may contain spaces. The old error cases still refuse: a single field is rejected, and trailing text leaves a last field that is not an integer. The library header and docs/configuration.md now say a name starting with `#` cannot be declared. New test `test_spaced_project_name_is_declared` declares `my heavy project 1` next to an indented comment line and gets a deferral. - **ci-2 (unreadable records):** the holder count must never silently leave out a holder. `fm_project_capacity_occupants` now refuses when a local home's state directory exists but cannot be read or listed, or when a `.meta` file cannot be read. The error names the path, and `fm-spawn.sh` shows it in its existing refusal message. New test `test_unreadable_holders_refuse_admission` covers an unreadable record in the root home and an unreadable state directory in a registered local secondmate home, then checks that the spawn is admitted once both are readable. The test is skipped when run as root. - **ci-3 (Orca lock):** any spawn that can become a holder for a capped project must take that project's lock. The lookup now also reports whether the declaration caps any project at all, and an Orca spawn takes the per-origin lock whenever it does. This covers every capped same-origin clone. It also covers some cases where no same-origin clone is capped, because a spawn cannot find clones under other directory names without searching for them. With no declaration file, Orca still skips the lock. The comments in the library and in the `fm-spawn.sh` header are updated. The Orca test now clones the origin as `project-2`, which has no declaration, and checks that its Orca spawn refuses while the lock is held and publishes no record. - **ci-4 (worktrees):** `assert_nothing_created` now also compares the project's `git worktree list` from before and after a deferred spawn. Both tests that call it take that snapshot first. Files changed: bin/fm-project-capacity-lib.sh, bin/fm-spawn.sh, docs/configuration.md, tests/fm-project-capacity.test.sh

* fix(bin): declare capacity for a project name that begins with #

A clone directory whose name begins with # was skipped as a comment, so that project stayed uncapped. A line is a declaration when the # is written against the rest of the name and the line ends with a capacity; a # followed by whitespace stays a comment.

* no-mistakes(document): Rewrap project-capacity library header comment

* no-mistakes(ci): Lint 2 fails because this PR's code pushes ShellCheck past its memory cap. ShellCheck ran out of memory analyzing bin/fm-teardown.sh in CI (reason=memory, rc=251, peak about 8.39 GB). On current main the same file passes at about 7.29 GB. **Cause:** the new `fm_local_firstmate_state_dirs` function in bin/fm-wake-lib.sh had a conditional `. fm-secondmate-registry-lib.sh` with a `# shellcheck source=` directive inside the function. ShellCheck followed that source again, inside a function scope, wherever fm-wake-lib.sh is sourced, and bin/fm-teardown.sh is the heaviest root that sources it. Measured locally with `shellcheck --norc --external-sources bin/fm-teardown.sh`: - current main (fd325b1): 7.29 GB - main merged with this PR: 7.86 GB - the same merge without the in-function source: 7.27 GB **Rule this restores:** this change must not make any lint root heavier than it is on main. That function holds the only new nested source in the change. **Fix:** I removed the in-function source, which no caller needs. Both callers already load the registry library at top level before calling the function: - bin/fm-teardown.sh sources it directly. - bin/fm-spawn.sh, the only user of bin/fm-project-capacity-lib.sh, gets it through bin/fm-ff-lib.sh. I also documented the requirement in the function's comment and in the "Requires" note in bin/fm-project-capacity-lib.sh. No behaviour changes. **Verification:** - ShellCheck on head: bin/fm-teardown.sh peaks at 7.12 GB and bin/fm-spawn.sh at 6.68 GB, both with rc=0. bin/fm-wake-lib.sh and bin/fm-project-capacity-lib.sh lint clean. - tests/fm-project-capacity.test.sh, tests/fm-teardown.test.sh (102 ok) and tests/fm-teardown-endpoint-safety.test.sh all pass. Files changed: bin/fm-wake-lib.sh, bin/fm-project-capacity-lib.sh

* fix(bin): release the Herdr session lock when reclaim finishes

A concurrent resume in another home waits five seconds for that lock.
Reclaim is the last presentation change on the recovery path, so holding
the lock through the launch tail made the waiter time out. The contributions
arm check also freezes its one-second clock, the same way the budget tests
do, because an unfrozen clock can tick past before the first forge read.

* no-mistakes(review): Keep Herdr session lock through launch handoff after reclaim

* no-mistakes(review): Skip the spawning task's own record in capacity count

* no-mistakes(review): Restore release test comment above its test

* docs: scope PR-ready re-evaluation to a declared project capacity

A ready pull request frees a place only when that project declares capacity, so the always-loaded backlog contract should re-evaluate on that handoff only in that case.
… section (kunchenguid#6785)

fm-procevent-lavish.sh read labels tag=message rows SESSION-ENDING MESSAGE
only when session_ended is true and CAPTAIN MESSAGE otherwise, but the
count line always said session_ending_message_count. Several composer
messages on a still-open board were therefore counted as session-ending.

The count line now follows the same session_ended switch:
session_ending_message_count once the session ended, captain_message_count
otherwise. Message rows stay out of the annotation count, per triage.

Fixes kunchenguid#6743
…henguid#6780)

The Herdr presentation lock namespace was the fixed machine-global
/tmp/firstmate-herdr-presentation, so on a host where two OS users run
Firstmate on Herdr the first account to create it owned it and every
teardown from the other account was refused with no way to clear it.

Suffix the namespace with the account uid. The owner-uid and mode-700
checks are unchanged, so a foreign-owned or wrong-mode name at this
account's path is still refused and never adopted, chowned, or removed.

Fixes kunchenguid#4716.
…te (kunchenguid#6809)

The OpenCode session plugin's shouldArm kept its own copy of the need
test that only looked for in-flight task records, while the turn-end
guard decides with fm_supervision_needed in bin/fm-supervision-lib.sh,
which also counts registered process-event sources and trusted custom
checks. With an empty fleet but any registered source or check, the
guard blocked every turn end while the plugin declined to arm - a loop
the guard's own repair line could not resolve because it names the
plugin as the fix.

The plugin now delegates the decision to the shared predicate through
bash, keeping the local away-record decline and the x-mode.env arm
override. OpenCode plugin test fixtures now carry the real predicate
their arming path sources, and the arm suite gains six cases asserting
the plugin's decision against the shared verdict over the same
synthetic state directories.

Co-authored-by: Mia Sun <mia@Bigs-Mac-mini.localdomain>
kunchenguid#6792)

* fix(bin): resolve a pending reply only from its own task's status line

Remote reply ingestion handed every corr= token in a mate's payload to
fm_pending_reply_try_resolve together with that mate's own status log,
so one mate echoing another mate's token resolved the other request.
Honor a status-file override only when it is the record's own
parent_status, and match the corr= token as a whole word.

Fixes kunchenguid#6538

* no-mistakes(document): docs: scope remote reply settlement to the asked mate
…nguid#6784)

agent-skill-trigger-index claims to be the complete agent-only trigger
index but omitted operational-home-layout, session-start-recovery,
validation-supervision, ship-landing, scout-completion, and
away-quiet-supervision. Add each with its own description's trigger,
placed beside the related entries.

The decision-hold-lifecycle redirect stub stays out, per triage.

Fixes kunchenguid#6503
…nchenguid#6814)

* test: share a rename-safe agent stand-in across liveness suites

On Ubuntu 26.04, `sleep` is the uutils multicall binary, which refuses to
run when invoked through a symlink named after another utility. The Herdr
descendant process-walk tests built their agent-named process as a `pi`
symlink to the host `sleep`, so the process exited at once, its parent shell
was gone before the walk ran, and both cases read `unknown unreadable` and
failed on that host. The suite stops at its first failure, so every later
case went unrun. The Herdr control smoke test's `claude` symlink has the
same construction.

The tmux liveness suite already solved this with a host-compiled spinner and
a survival-checked `sleep` fallback. That builder moves into tests/lib.sh as
fm_agent_standin, and the tmux suite, both Herdr descendant cases, and the
Herdr control smoke test now use it. When no stand-in can survive a foreign
name, a case skips with the reason instead of failing.

tests/fm-test-fixtures.test.sh gains a portable regression with a fake
multicall `sleep`, so it bites on hosts whose own `sleep` is single-purpose.

* no-mistakes(document): Correct Herdr verification fixture reference

* ci: retrigger cancelled shard
…he Stop hook's group is torn down (kunchenguid#6787)

* fix(bin): keep the supervision host's pass-through successor out of the hook's process group

The successor a main-only pass-through leaves for main shared the Stop hook's
process group, so the harness tearing that group down after the exit-2 rewake
stopped it. The stop published downtime and the next park's first cycle
announced an empty check: rearm-resurface, which woke main again in a loop.
Start that successor in a process group of its own, as the hook's own
handling successor already is.

* no-mistakes(review): Give the at-turn successor left for main its own group

* no-mistakes(document): Document own-group successor for turn-start hand-back too

* no-mistakes(ci): I made the change you asked for: both new teardown tests in tests/fm-supervision-host.test.sh now call the existing `stop_home_processes "$home"` just before `pass`. The tests are `test_successor_left_at_the_turn_survives_the_hook_process_group_teardown` and `test_pass_through_successor_survives_the_hook_process_group_teardown`. No production code and no other tests changed. The rule broken was that a test must not leave a home's watcher or arm processes running after it passes. These two were the only cases in the changed area that broke it. The other host+hook tests already stop their home, and `test_successor_close_during_main_turn_is_delivered_at_the_next_turn_end` leaves its watcher behind too, but it is an older test you said not to touch. The only reason anything was left over is that the successor's arm now sits in its own process group, outside the hook's teardown. `stop_home_processes` kills the watcher by the pid in its lock file, which stops it no matter which group it is in. **Checks run:** - I ran just these two tests from a scratch copy of the suite (since deleted). Both pass in about 13 seconds. - After each test, a process listing filtered to that test's home directory came back empty once the processes had about a second to exit after TERM. - `bash -n` on the test file passes. - `shellcheck` is not installed here, so I did not lint the file. - I did not run the full serial-2 suite locally. Whether it now finishes under its 30-minute limit will only show on the next CI run
…ters (kunchenguid#6823)

* test: use idle composer readiness for Claude tmux guards

* no-mistakes(test): Fix attended supervision test expectations and isolate worker state

* no-mistakes(document): Correct live guard coverage and readiness documentation

* no-mistakes(ci): Captain, fixed SC2100 by quoting the cursor-agent assignment in tests/fm-host-mirror-live-e2e.test.sh. Reproduced the failure before editing; pinned ShellCheck lint on both PR test files, bash syntax checks, and git diff --check now pass

* test: preserve attended successor close assertions

* no-mistakes(test): Fix attended live test watcher takeover expectations

* no-mistakes(document): Correct stale attended guard documentation

* Revert "no-mistakes(document): Correct stale attended guard documentation"

This reverts commit 8e59d89.

* Revert "no-mistakes(test): Fix attended live test watcher takeover expectations"

This reverts commit c0b8510.
* test: order stale watcher-lock fixture races explicitly

* no-mistakes(review): Removed duplicate stale-steal reap call
…ting server (kunchenguid#6920)

* fix(bin): keep herdr composer probes observable-only without autostarting server

- Refuse bare pane identifiers lacking session prefix in fm_backend_herdr_parse_target
- Add fm_backend_herdr_target_observable to passively verify server liveness without autostarting
- Route visible viewport captures through fm_backend_herdr_target_observable
- Detach stdin with </dev/null on background server launch in fm_backend_herdr_server_ensure
- Add unit and regression tests in tests/fm-backend-herdr.test.sh

* fix(bin): drop parse_target colon restriction to preserve existing fixtures

Revert the fm_backend_herdr_parse_target colon check so that single-colon
session targets like sess:p1 in secondmate liveness and other fixtures parse
properly. The probe path remains observable-only via fm_backend_herdr_target_observable.
…kunchenguid#6956)

* test: clear inherited Git repository location in fixture helper

An absolute GIT_DIR inherited from the caller (a linked-worktree hook
exports one) overrides the discovery that git -C implies, so fixture
setup moved the caller's branches and registered fixture worktrees in
it. tests/git-config-helpers.sh now unsets GIT_DIR, GIT_WORK_TREE,
GIT_INDEX_FILE and GIT_COMMON_DIR alongside the config isolation.

Fixes kunchenguid#6953

* no-mistakes(document): Document fixture Git location isolation in CONTRIBUTING
…shed in #43, fd9b8b0; real merge history on fm/fm-upstream-sync e24c85a, identical tree 00cbc3b)
, kunchenguid#6787, kunchenguid#6823, kunchenguid#6818, kunchenguid#6887, kunchenguid#6920, kunchenguid#6956)

Merges the seven upstream commits that follow 2ce57d0:
329ad4e, 3c58ec8, 0ec1c5a, b062eb9, 19fcbbd, b75658b, fb75c1f.

Conflict resolutions:

- tests/fm-supervision-host.test.sh: kept the fork's split layout.
  Upstream kunchenguid#6787's test changes move to the split files:
  turn_main_only_at_second_offer's quiet mode to
  tests/fm-supervision-host-helpers.sh, and start_hook_session's
  own-process-group hook plus the two new teardown tests to
  tests/fm-supervision-host-hook.test.sh, registered with run_host_case.
- tests/fm-control-herdr-smoke.test.sh: kept the fork's Python worker named
  claude, because the launch proof reads FM_SPAWN_GEN from the worker's
  environment and upstream's fm_agent_standin can fall back to the host
  sleep, whose environment macOS hides. Upstream's unused stand-in prelude
  is dropped here; its lab cleanup and exit-refusal assertions are kept.
- bin/backends/herdr.sh: kept both sides - the fork's COMPACT_ADVISER
  environment names and upstream's </dev/null server start.
- docs/verification/runtime-backends.md: kept the fork's sentence, which
  describes the stand-in the smoke guard still uses.
- CONTRIBUTING.md: union of both helper descriptions.
- bin/fm-control.sh: the first line of the picker comment before the
  dialog check was lost in a conflict resolution, leaving the comment to
  start mid-sentence.
- .agents/skills/afk/SKILL.md: the supervision-host bullet kept the old
  "with config/supervision-host" condition. Upstream changed it to
  "that runs the supervision host", because a Claude primary runs the
  host by default with no file.
Fork-local fix to the text of upstream kunchenguid#6887's
test_lock_stale_steal_single_winner_under_concurrency. A contender exit can
interrupt the parent's open of the first-winner FIFO (EINTR); bash does not
retry it, so the winner PID stayed empty, the parent waited on the winning
contender, and that contender waited on the hold FIFO the parent writes only
after the wait. Retry the read a bounded number of times, and if no contender
ever reports a win, stop every contender and the waiter before failing so
nothing is left blocked on the hold FIFO.
@MrGTV-love
MrGTV-love merged commit 7a53654 into main Oct 11, 2026
24 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.