Skip to content

feat: sync upstream Firstmate capabilities into the fork - #9

Merged
dscott98 merged 136 commits into
mainfrom
fm/fm-fork-sync
Oct 1, 2026
Merged

dscott98 merged 136 commits into
mainfrom
fm/fm-fork-sync

Conversation

@dscott98

@dscott98 dscott98 commented Oct 1, 2026 •

Copy link
Copy Markdown
Owner

Intent

Keep this Firstmate work in the captain's own fork (https://github.com/dscott98/firstmate), never upstream, while still being able to update from upstream (https://github.com/kunchenguid/firstmate).
Concretely: bring the fork's main current with upstream main so this install can run from the fork.
The fork's main is diverged: 10 fork-only commits ahead and about 133 upstream commits behind.

Merge method

Merge this PR into dscott98/firstmate main with a merge commit, not squash or rebase, so the upstream ancestry remains available for future synchronization.

What Changed

  • Merge upstream main into the fork, preserving fork-specific provider switching, worker startup verification, and capture-inbox guarantees.
  • Bring in supervision hosts, Calm presentation updates, Gerrit publishing and monitoring, configurable branch prefixes, and Claude/Pi worker account pins.
  • Incorporate upstream fixes for watcher continuity, remote secondmate recovery, task delivery, and merge checks, alongside expanded regression coverage, bounded CI linting, and operational documentation.

Risk Assessment

⚠️ Medium: The upstream merge preserves the fork customizations, but an opt-in delivery race can silently lose a worker’s promised retry.

Testing

Live Git synchronization and Firstmate inbox checks passed, as did the targeted inbox regression test. CLI transcripts were captured; an initial acknowledgement invocation was corrected to supply its required ID. No source changes or remote writes occurred.

  • Live validation: ✅ go - 4 of 5 scenarios driven live against the product
Scenario Result Live Evidence
Advance an isolated fork main to the candidate while retaining both fork and current upstream histories ✅ pass live Live fork synchronization transcript
Fetch upstream updates while keeping origin's push destination exclusively on dscott98/firstmate ✅ pass live Live fork synchronization transcript
Repeat upstream integration without rolling back or losing the merged fork commit ✅ pass live Live fork synchronization transcript
Queue, list, and acknowledge a note; acknowledgement without an ID refuses without losing the note ✅ pass live Live Firstmate inbox transcript
Publish and land the synchronization exclusively on the fork's main ⏸️ untested no This test phase lacks delivery authority. The outer executor must perform push, PR, CI, and landing and verify the fork-only destinations.
Evidence: Live fork synchronization transcript

Source: Live fork synchronization transcript

$ git ls-remote https://github.com/dscott98/firstmate.git refs/heads/main
22534f7857ad75e2b8b474d6ac6571fd5d580616	refs/heads/main
exit=0
$ git ls-remote https://github.com/kunchenguid/firstmate.git refs/heads/main
8f756bbc287c5bdfacc64a7cc09e8516c64fc919	refs/heads/main
exit=0
$ git merge-base --is-ancestor 8f756bbc287c5bdfacc64a7cc09e8516c64fc919 77b87efb050c2bcd9904aecc49153311a6b28961
exit=0
$ git merge-base --is-ancestor 22534f7857ad75e2b8b474d6ac6571fd5d580616 77b87efb050c2bcd9904aecc49153311a6b28961
exit=0
$ git remote get-url --push --all origin
https://github.com/dscott98/firstmate.git
exit=0
$ git bundle create ~/.no-mistakes/worktrees/a1211da81307/01M3W9CMQ1R2W9DRAK32K4F14H/.fork-sync-test-vhr878df/source.bundle HEAD
exit=0
$ git clone --quiet ~/.no-mistakes/worktrees/a1211da81307/01M3W9CMQ1R2W9DRAK32K4F14H/.fork-sync-test-vhr878df/source.bundle ~/.no-mistakes/worktrees/a1211da81307/01M3W9CMQ1R2W9DRAK32K4F14H/.fork-sync-test-vhr878df/install
Note: switching to '77b87efb050c2bcd9904aecc49153311a6b28961'.

You are in 'detached HEAD' state. You can look around, make experimental
changes and commit them, and you can discard any commits you make in this
state without impacting any branches by switching back to a branch.

If you want to create a new branch to retain commits you create, you may
do so (now or later) by using -c with the switch command. Example:

  git switch -c <new-branch-name>

Or undo this operation with:

  git switch -

Turn off this advice by setting config variable advice.detachedHead to false

exit=0
$ git remote set-url origin https://github.com/dscott98/firstmate.git
exit=0
$ git switch -C main 22534f7857ad75e2b8b474d6ac6571fd5d580616
Warning: you are leaving 134 commits behind, not connected to
any of your branches:

  77b87ef Merge upstream main into fork (2026-10-01 sync)
  8f756bb fix(bin): recognize clone roots across path spelling differences (#6306)
  b5d9061 fix(bin): document accepted contribution verdict actors (#6307)
  f593060 fix(bin): exclude a remote mate's own parent channel from self-home status scans (#5263)
 ... and 130 more.

If you want to keep them by creating a new branch, this may be a good time
to do so with:

 git branch <new-branch-name> 77b87ef

Switched to a new branch 'main'
exit=0
$ git merge --ff-only --stat 77b87efb050c2bcd9904aecc49153311a6b28961
Updating 22534f7..77b87ef
Fast-forward
 .agents/skills/afk/SKILL.md                        |   58 +-
 .agents/skills/agent-skill-trigger-index/SKILL.md  |   28 +
 .agents/skills/ahoy/SKILL.md                       |    3 +-
 .agents/skills/away-quiet-supervision/SKILL.md     |   24 +
 .agents/skills/bootstrap-diagnostics/SKILL.md      |    3 +-
 .agents/skills/captain-hold-lifecycle/SKILL.md     |    3 +-
 .agents/skills/firstmate-codexapp/SKILL.md         |    2 +-
 .../skills/firstmate-coding-guidelines/SKILL.md    |    7 +-
 .agents/skills/fmx-respond/SKILL.md                |   15 +
 .../references/common/control-and-recovery.md      |    7 +-
 .../references/common/primary-hooks.md             |    2 +-
 .../harness-adapters/references/harness/claude.md  |   11 +-
 .../harness-adapters/references/harness/codex.md   |    1 +
 .../harness-adapters/references/harness/cursor.md  |    2 +
 .../harness-adapters/references/harness/devin.md   |    2 +-
 .../harness-adapters/references/harness/grok.md    |   31 +-
 .../harness-adapters/references/harness/omp.md     |    2 +-
 .../references/harness/opencode.md                 |    3 +-
 .../harness-adapters/references/harness/pi.md      |    5 +-
 .agents/skills/operational-home-layout/SKILL.md    |  123 +
 .agents/skills/process-event-sources/SKILL.md      |   12 +-
 .agents/skills/project-management/SKILL.md         |   11 +-
 .agents/skills/quiet/SKILL.md                      |   48 +-
 .agents/skills/scout-completion/SKILL.md           |   16 +
 .agents/skills/secondmate-provisioning/SKILL.md    |    9 +-
 .agents/skills/session-start-recovery/SKILL.md     |   43 +
 .agents/skills/ship-landing/SKILL.md               |   30 +
 .agents/skills/stow/SKILL.md                       |   10 +-
 .agents/skills/validation-supervision/SKILL.md     |   32 +
 .../mods/firstmate-calm/.claude-plugin/plugin.json |    2 +-
 .claude/mods/firstmate-calm/hooks/register.ts      |  227 +-
 .claude/mods/firstmate-calm/lib/fm-branch-notes.ts |  166 ++
 .../firstmate-calm/lib/fm-calm-presentation.ts     |   23 +-
 .../firstmate-calm/lib/fm-operational-input.ts     |   34 +
 .../mods/firstmate-calm/tests/branch-notes.test.ts |  175 ++
 .claude/mods/firstmate-calm/tests/calm.test.ts     |   51 +-
 .claude/mods/firstmate-calm/tests/support.ts       |   37 +-
 .claude/settings.json                              |   16 +
 .cursor/hooks.json                                 |   14 +
 .github/workflows/ci.yml                           |   24 +-
 .no-mistakes.yaml                                  |    6 +-
 .omp/extensions/fm-primary-omp-watch.ts            |   93 +-
 .opencode/plugins/fm-primary-watch-arm.js          |   86 +-
 .pi/extensions/fm-branch-supervision.ts            |  135 +-
 .pi/extensions/fm-calm.ts                          |   14 +-
 .pi/extensions/fm-primary-pi-watch.ts              |   45 +-
 .pi/extensions/lib/fm-branch-dispatch.ts           |  321 ++-
 .../lib/fm-calm-operational-user-layout.ts         |   11 +-
 .../lib/fm-calm-pending-operational-layout.ts      |  310 +++
 .pi/extensions/lib/fm-operational-input.ts         |   16 +
 AGENTS.md                                          |  347 +--
 README.md                                          |   12 +-
 bin/backends/herdr.sh                              |  240 +-
 bin/fm-afk-contract.sh                             |  130 +-
 bin/fm-afk-launch.sh                               |  242 +-
 bin/fm-afk-return.sh                               |  197 +-
 bin/fm-afk-start.sh                                |    3 +-
 bin/fm-agent-process-lib.sh                        |    7 +-
 bin/fm-arm-command-policy.mjs                      |   15 +-
 bin/fm-backend.sh                                  |   42 +-
 bin/fm-backlog-transition-lib.sh                   |  114 +-
 bin/fm-bearings-snapshot.sh                        |   15 +-
 bin/fm-bootstrap.sh                                |  187 +-
 bin/fm-branch-dispatch.mjs                         |  131 +
 bin/fm-branch-outcome.sh                           |  258 +-
 bin/fm-branch-prompt.sh                            |   33 +-
 bin/fm-branch-report.sh                            |  153 ++
 bin/fm-brief-heading-lib.sh                        |   94 +
 bin/fm-brief.sh                                    |  264 +-
 bin/fm-captain-hold.sh                             |    7 +-
 bin/fm-classify-lib.sh                             |  352 ++-
 bin/fm-claude-stop-autoarm.sh                      |  216 +-
 bin/fm-claude-trust.sh                             |   53 +-
 bin/fm-composer-lib.sh                             |   33 +-
 bin/fm-config-inherit-lib.sh                       |  145 +-
 bin/fm-contributions.sh                            |  113 +-
 bin/fm-control-lib.sh                              |   50 +-
 bin/fm-control.sh                                  |   13 +
 bin/fm-crew-state.sh                               |   94 +-
 bin/fm-devin-config.sh                             |   13 +-
 bin/fm-dispatch-resolve.sh                         |  161 +-
 bin/fm-dod-lib.sh                                  |  470 +++-
 bin/fm-ensure-agents-md.sh                         |    8 +-
 bin/fm-ff-lib.sh                                   |   16 +
 bin/fm-fleet-snapshot.sh                           |    3 +
 bin/fm-fleet-sync.sh                               |   17 +-
 bin/fm-forge-detect.sh                             |   63 +
 bin/fm-gate-refuse-lib.sh                          |  106 +-
 bin/fm-git-strip-ai-trailers.sh                    |  265 ++
 bin/fm-guard.sh                                    |   15 +-
 bin/fm-harness.sh                                  |   34 +-
 bin/fm-herdr-lab.sh                                |   14 +-
 bin/fm-home-seed.sh                                |   33 +-
 bin/fm-host-mirr

... [9906 bytes truncated] ...

live-lab-up-mate.test.sh                  |   69 +
 tests/fm-live-lab.test.sh                          |  764 ++++++
 tests/fm-mail-check.test.sh                        |    1 +
 tests/fm-omp-harness.test.sh                       |  293 ++-
 tests/fm-operational-input.test.sh                 |   98 +
 tests/fm-parent-channel-scan-exclusion.test.sh     |  414 +++
 tests/fm-pending-reply.test.sh                     |  405 ++-
 tests/fm-pi-branch-extension.test.sh               |  581 ++++-
 tests/fm-pi-codex-native.test.sh                   |    2 +-
 tests/fm-pi-primary-live-e2e.test.sh               |    1 +
 tests/fm-pi-primary-types.test.sh                  |    1 +
 tests/fm-pi-watch-extension.test.sh                |   94 +
 tests/fm-pr-check-security.test.sh                 |  580 ++++-
 tests/fm-pr-merge.test.sh                          |  642 ++++-
 tests/fm-procevent.test.sh                         |  702 ++++-
 tests/fm-quota-choose.test.sh                      |   34 +
 tests/fm-remote-backlog-handoff.test.sh            |    2 +-
 tests/fm-remote-job.test.sh                        |  539 ++++
 tests/fm-remote-reply.test.sh                      |  245 +-
 tests/fm-remote-secondmate-lifecycle-e2e.test.sh   |  321 ++-
 tests/fm-remote-secondmate-relaunch.test.sh        |  192 ++
 tests/fm-remote-transport-lanes.test.sh            |    2 +-
 tests/fm-review-diff.test.sh                       |   49 +
 tests/fm-secondmate-harness.test.sh                |  100 +-
 tests/fm-secondmate-lifecycle-e2e.test.sh          |    8 +
 tests/fm-secondmate-liveness.test.sh               |  166 +-
 tests/fm-secondmate-reconcile.test.sh              |    1 +
 tests/fm-secondmate-restart.test.sh                |   12 +-
 tests/fm-secondmate-safety.test.sh                 |   64 +-
 tests/fm-secondmate-sync.test.sh                   |    2 +-
 tests/fm-send-inbox-doorbell-live-e2e.test.sh      |   21 +-
 tests/fm-send-inbox.test.sh                        |  100 +-
 tests/fm-session-lock-ancestry.test.sh             |   22 +-
 tests/fm-session-start.test.sh                     |  348 ++-
 tests/fm-sessionstart-nudge.test.sh                |   48 +
 tests/fm-shared-captain-inheritance.test.sh        |  212 +-
 tests/fm-spawn-compact-adviser-disable.test.sh     |   44 +-
 tests/fm-spawn-dispatch-profile.test.sh            |  397 ++-
 tests/fm-spawn-orca-worktree.test.sh               |  170 ++
 tests/fm-spawn-worktree-settle.test.sh             |    2 +-
 tests/fm-startup-memory-budget.test.sh             |    4 +-
 tests/fm-startup-network.test.sh                   |  101 +
 .../fm-supervision-host-attended-live-e2e.test.sh  |  443 ++++
 tests/fm-supervision-host-live-e2e.test.sh         |  140 +
 tests/fm-supervision-host.test.sh                  | 2689 ++++++++++++++++++++
 tests/fm-supervision-instructions.test.sh          |   88 +
 tests/fm-task-delivery.test.sh                     |  746 +++++-
 tests/fm-task-inbox.test.sh                        |  214 +-
 tests/fm-tasks-axi.test.sh                         |   24 +
 tests/fm-teardown-endpoint-safety.test.sh          |   55 +
 tests/fm-teardown.test.sh                          |  417 +++
 tests/fm-test-run.test.sh                          |  100 +-
 tests/fm-timeout-lib.test.sh                       |  344 +++
 tests/fm-tool-update-check.test.sh                 |   52 +
 tests/fm-trace-context-spawn.test.sh               |    3 +
 tests/fm-turnend-guard.test.sh                     |   12 +-
 tests/fm-voice-relay.test.sh                       |   44 +
 tests/fm-wake-drain-outcome-backstop.test.sh       |    8 +
 tests/fm-wake-drain-unread-status.test.sh          |    8 +
 tests/fm-wake-queue.test.sh                        |  992 +++++++-
 tests/fm-watch-arm.test.sh                         |  408 ++-
 tests/fm-watch-checkpoint.test.sh                  |  119 +
 tests/fm-watch-triage.test.sh                      |  528 +++-
 tests/fm-watcher-lock.test.sh                      |  393 +++
 tests/fm-worker-account-live-e2e.test.sh           |  132 +
 tests/fm-worker-account.test.sh                    |  401 +++
 tests/fm-x-mode.test.sh                            |  179 +-
 tests/lib.sh                                       |  151 +-
 tests/wake-helpers.sh                              |    9 +
 321 files changed, 47360 insertions(+), 4134 deletions(-)
 create mode 100644 .agents/skills/agent-skill-trigger-index/SKILL.md
 create mode 100644 .agents/skills/away-quiet-supervision/SKILL.md
 create mode 100644 .agents/skills/operational-home-layout/SKILL.md
 create mode 100644 .agents/skills/scout-completion/SKILL.md
 create mode 100644 .agents/skills/session-start-recovery/SKILL.md
 create mode 100644 .agents/skills/ship-landing/SKILL.md
 create mode 100644 .agents/skills/validation-supervision/SKILL.md
 create mode 100644 .claude/mods/firstmate-calm/lib/fm-branch-notes.ts
 create mode 100644 .claude/mods/firstmate-calm/tests/branch-notes.test.ts
 create mode 100644 .pi/extensions/lib/fm-calm-pending-operational-layout.ts
 create mode 100755 bin/fm-branch-dispatch.mjs
 create mode 100755 bin/fm-branch-report.sh
 create mode 100644 bin/fm-brief-heading-lib.sh
 create mode 100755 bin/fm-forge-detect.sh
 create mode 100755 bin/fm-git-strip-ai-trailers.sh
 create mode 100755 bin/fm-host-mirror.sh
 create mode 100755 bin/fm-jev-mem-guard.py
 create mode 100755 bin/fm-jev-mem-guard.sh
 create mode 100755 bin/fm-lab-home.sh
 create mode 100755 bin/fm-live-lab.sh
 create mode 100644 bin/fm-path-lib.sh
 create mode 100755 bin/fm-remote-secondmate-relaunch.sh
 create mode 100644 bin/fm-secondmate-liveness-lib.sh
 create mode 100644 bin/fm-supervision-engine-lib.sh
 create mode 100755 bin/fm-supervision-host.sh
 create mode 100644 bin/fm-worker-account-lib.sh
 create mode 100644 docs/gerrit-change-watch.md
 create mode 100644 docs/gerrit-forge-integration.md
 create mode 100644 docs/jev-guards.md
 create mode 100644 docs/supervision-host.md
 create mode 100644 docs/supervision-protocols/supervision-host.md
 create mode 100755 tests/fm-calm-pi-queue-retention-live-e2e.test.sh
 create mode 100755 tests/fm-forge-detect.test.sh
 create mode 100755 tests/fm-fork-free-helpers.test.sh
 create mode 100644 tests/fm-git-strip-ai-trailers.test.sh
 create mode 100755 tests/fm-host-mirror-live-e2e.test.sh
 create mode 100755 tests/fm-host-mirror.test.sh
 create mode 100755 tests/fm-jev-mem-guard.test.sh
 create mode 100644 tests/fm-live-lab-up-mate.test.sh
 create mode 100755 tests/fm-live-lab.test.sh
 create mode 100755 tests/fm-parent-channel-scan-exclusion.test.sh
 create mode 100755 tests/fm-remote-secondmate-relaunch.test.sh
 create mode 100755 tests/fm-spawn-orca-worktree.test.sh
 create mode 100755 tests/fm-supervision-host-attended-live-e2e.test.sh
 create mode 100755 tests/fm-supervision-host-live-e2e.test.sh
 create mode 100755 tests/fm-supervision-host.test.sh
 create mode 100755 tests/fm-timeout-lib.test.sh
 create mode 100755 tests/fm-worker-account-live-e2e.test.sh
 create mode 100755 tests/fm-worker-account.test.sh
exit=0
$ git rev-parse HEAD
77b87efb050c2bcd9904aecc49153311a6b28961
exit=0
$ git merge-base --is-ancestor 22534f7857ad75e2b8b474d6ac6571fd5d580616 HEAD
exit=0
$ git merge-base --is-ancestor 8f756bbc287c5bdfacc64a7cc09e8516c64fc919 HEAD
exit=0
PASS: fork main fast-forwards to candidate retaining fork and upstream histories.
$ git fetch https://github.com/kunchenguid/firstmate.git refs/heads/main:refs/remotes/upstream/main
From https://github.com/kunchenguid/firstmate
 * [new branch]      main       -> upstream/main
exit=0
$ git remote get-url --push --all origin
https://github.com/dscott98/firstmate.git
exit=0
$ git remote
origin
exit=0
$ git merge-base --is-ancestor refs/remotes/upstream/main HEAD
exit=0
PASS: documented upstream fetch works without adding an upstream push remote.
$ git merge --ff-only refs/remotes/upstream/main
Already up to date.
exit=0
$ git rev-parse HEAD
77b87efb050c2bcd9904aecc49153311a6b28961
exit=0
$ git status --porcelain
exit=0
PASS: repeat upstream integration leaves merged fork commit and clean files intact.
Removed isolated installation and bundle. No remote writes performed.
Evidence: Live Firstmate inbox transcript

Source: Live Firstmate inbox transcript

$ bin/fm-inbox.sh note Fork synchronization live smoke
queued 1790877525-wSqb0m
  Fork synchronization live smoke
  firstmate will pick this up at its next check.
exit=0
$ bin/fm-inbox.sh list
1790877525-wSqb0m
    Fork synchronization live smoke
exit=0
$ bin/fm-inbox.sh drain --ack
fm-inbox: usage: fm-inbox.sh drain --ack <id>...
exit=1

Re-drive with required acknowledgement ID in a fresh isolated home:
$ bin/fm-inbox.sh note Fork synchronization live smoke
queued 1790877545-O0bb3X
  Fork synchronization live smoke
  firstmate will pick this up at its next check.
exit=0
$ bin/fm-inbox.sh list
1790877545-O0bb3X
    Fork synchronization live smoke
exit=0
$ bin/fm-inbox.sh drain --ack
fm-inbox: usage: fm-inbox.sh drain --ack <id>...
exit=1
$ bin/fm-inbox.sh list
1790877545-O0bb3X
    Fork synchronization live smoke
exit=0
$ bin/fm-inbox.sh drain --ack 1790877545-O0bb3X
acked 1790877545-O0bb3X
exit=0
$ bin/fm-inbox.sh list
(inbox empty)
exit=0
PASS: queue/list/ack works; missing ID is refused without losing the note.

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

⚠️ **Rebase** - 1 warning

Confirm these commits belong in this PR before approving, or manually separate the intended work onto origin/main before gating.

⚠️ **Review** - 1 warning
  • ⚠️ bin/fm-task-inbox-lib.sh:402 - With wait-no-turns enabled, the watcher can read retry marker A here, a concurrent fm-send can write marker B after its doorbell fails, and line 403 then deletes B. Fire-and-forget messages bypass the ordinary retry ladder, leaving B without its promised retry. The cleanup at lines 422–424 has the same race. Serialize marker publication and conditional removal through the shared inbox helpers.
✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 4 of 5 scenarios driven live against the product
Scenario Result Live Evidence
Advance an isolated fork main to the candidate while retaining both fork and current upstream histories ✅ pass live Live fork synchronization transcript
Fetch upstream updates while keeping origin's push destination exclusively on dscott98/firstmate ✅ pass live Live fork synchronization transcript
Repeat upstream integration without rolling back or losing the merged fork commit ✅ pass live Live fork synchronization transcript
Queue, list, and acknowledge a note; acknowledgement without an ID refuses without losing the note ✅ pass live Live Firstmate inbox transcript
Publish and land the synchronization exclusively on the fork's main ⏸️ untested no This test phase lacks delivery authority. The outer executor must perform push, PR, CI, and landing and verify the fork-only destinations.
  • git ls-remote for both repositories' main branches
  • Isolated bundle-backed installation: git merge --ff-only and ancestry checks for both parents
  • Documented upstream fetch followed by push-URL verification and repeat merge
  • TMPDIR=<worktree-local temporary directory> bash tests/fm-inbox.test.sh
  • Real bin/fm-inbox.sh note, list, and drain --ack commands in an isolated home
  • Removed temporary installations and homes; verified clean git status --short
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

kunchenguid and others added 30 commits September 23, 2026 11:31
…ns (kunchenguid#5389)

The sibling secondmate stall cases in tests/fm-wake-queue.test.sh now wait for the watcher's recorded observation instead of a one-second wall-clock checkpoint, so they can neither fail nor pass vacuously under load. Deterministic proof with a 5s watcher launch delay: before the fix 4 cases passed vacuously and 6 failed; after it all 10 pass on the recorded observation.

Also includes a CI flake fix from validation: fm_control_harness_supported in bin/fm-control-lib.sh finishes reading the harness allowlist before returning, removing intermittent broken-pipe diagnostics. Behavior is unchanged.
… tail (kunchenguid#5336)

* fix(bin): refuse a Herdr submit that would send only a message tail

A long typed payload can sit in the composer as a suffix, or as a paste placeholder plus a remainder, and the following Enter was still reported as delivered. Prove the selected composer holds the payload before Enter, and report failure when it does not.

* no-mistakes(review): Scope Herdr payload proof to Claude, clear composer on refusal

* no-mistakes(test): Clear refused Herdr composer drafts one wrapped row per press

* no-mistakes(test): Accept Claude's multi-line paste placeholder in Herdr submit proof

* no-mistakes(review): Accept Claude read-back that drops U+2063 in Herdr proof

* no-mistakes(document): Document Herdr proof ignoring U+2063 operational mark

* no-mistakes(ci): I made a one-line test change. The failing check comes from a timing race in an existing test that this PR doesn't touch. **What failed:** `tests/fm-procevent.test.sh` failed at "the superseded paced runner invoked its stale command" (line ~3313). The PR only changes the Herdr files and their tests, and the same shard passed on main at the base commit. **Why it can fail:** the fixture starts a second runner with a 3-second launch floor (the minimum wait since the source's last launch). That runner sleeps for the rest of the floor and only then checks whether its registration was replaced (`fm_procevent_launch_floor_wait` in `bin/fm-procevent-lib.sh`). The test then waits for the claim and re-registers the source. If that takes longer than about 3 seconds after the first launch, the old runner wakes up, finds its registration still current, and runs the stale command. That produces the second log line the test reports. The CI shard was slow (this one test took 160 s). **Fix:** in `tests/fm-procevent.test.sh` I raised the superseded runner's floor from 3 to 15 seconds and added a comment explaining why. The floor now outlasts the fixture setup even on a loaded runner. Nothing else changed: the first launch and the later fresh-registration start still use a 3-second floor, and no product code changed. **Verification:** - The full test file can't give a reliable result on this machine (load average about 64 on 8 cores). It failed earlier, at the reconcile assertion around line 1680, before it reached this section. - I ran the changed section by itself (file setup plus the pacing-race block) five times with the fix: all passed, in about 9-13 s each. - The original code also passed five out of five, so the race didn't reproduce locally. The diagnosis rests on the code path and the CI log. - I haven't seen the full file or the CI shard pass with the fix yet

* no-mistakes(test): Accept Claude folder-trust prompt via down+enter in live e2e

* no-mistakes(document): Note unreadable Claude composer refusal in Herdr docs
…unchenguid#5427)

Speaking as Kun's firstmate: squash-merging — opt-in (forge=gerrit registry-gated; default project-mode stdout restored to two words), attestation MATCH, CI+NM green, safe review, MERGEABLE.
…uid#5358)

* feat(bin): add an opt-in per-home worker account pin

A home that mixes work and personal accounts for one runner had no way to
say which account its workers launch on: Claude workers inherited whatever
CLAUDE_CONFIG_DIR the supervising process had, Pi workers the pane's ambient
root, and an ambient API key outranked both, with no signal at launch.

config/claude-account and config/pi-account now pin that choice per home.
With neither file every launch is unchanged. With one, every launch of that
runner from the home (ship, scout, local secondmate, raw Claude command, and
relaunch) runs under the declared root, and the spawn refuses before any
endpoint exists when the file is malformed or the runner's own check
(claude auth status, pi auth check with a model-listing fallback) says the
pinned account is not signed in. The check runs in a cleared environment so
an ambient credential cannot answer for an empty root. A pinned Claude launch
sheds the environment credentials Claude ranks above a stored login; a pinned
Pi launch needs an explicit <provider>/<id> model for a declared provider and
also carries --provider. The chosen account is printed on the spawned line and
recorded in the task record, and relaunch checks the pin before stopping the
running agent.

* test(secondmate): give the concurrent config-push wait room for a slow host

test_config_reread_serializes_concurrent_pushes waited about two seconds for
the first fm-config-push.sh to reach its first send-keys. On a slower host
that push takes four to five seconds, so the test failed on main before the
push ever got there. The loop still leaves as soon as the marker appears, so
the larger bound costs nothing where the push is fast.

* no-mistakes(review): Refuse raw Claude account overrides under a pin
…ts (kunchenguid#5470)

* fix(bin): keep the Herdr lab session option before a -- delimiter

fm-herdr-lab.sh run appended --session <lab> after every argument, so a
command with a passthrough delimiter such as agent start ... -- <agent args>
handed the session flag to the agent and Herdr routed the call by the
caller's ambient socket instead of the lab.
The helper now inserts --session <lab> immediately before the first --
delimiter and keeps the trailing form otherwise.

* no-mistakes(document): Clarify Herdr lab session option placement
* feat(bin): guard the partition, harness pin, and bounded exec for a non-Pi supervision host

Lease liveness is now the pure record test in every calling context, so an
unmarked main honors a live branch lease held by a separate process, and a
lease file engages the guard's claim serialization for any caller; a home
with no lease files still takes no lock.

bin/fm-harness.sh honors FM_SUPERVISION_PRIMARY_HARNESS while
FM_SUPERVISION_ACTOR=branch, so a supervision branch running under another
harness resolves own, crew, and secondmate to the primary's harness.

fm_tasks_axi's watchdog moves into bin/fm-timeout-lib.sh as fm_exec_timed with
a separate grace: the perl watchdog is preferred, runs the command in its own
process group against wall-clock deadlines, forwards TERM/INT/HUP, and reaps
the group, so a descendant holding captured output can no longer keep the
caller waiting past the bound on a host without timeout.

The Claude Stop auto-arm header records that Claude drops the exit 2 of a hook
it terminated at the configured timeout, re-measured on Claude Code 2.1.281.

* fix(bin): state that fm_exec_timed cannot reach a descendant in its own process group

Live runs of real Claude and Pi engine turns under the bound showed both CLIs
start every tool command in a process group of its own, so those processes end
through the engine's own TERM handling rather than the group signal or reap.
Also clears the new timeout test's ShellCheck findings.

* no-mistakes(document): Clarify cross-harness lease documentation
…henguid#2648)

* feat(bin): make the ship-branch prefix configurable per project

fm-brief.sh hardcoded every generated ship branch to fm/<task-id>, which
leaks that firstmate produced the branch/PR - unwanted for a third-party
public repo that does not use this tooling.

Add an optional --branch-prefix flag to fm-brief.sh (default "fm/", so
existing installs are unaffected) and teach fm-project-mode.sh - the
registry's single-owner parser - to resolve a project's optional
"branch=<prefix>" data/projects.md annotation via a new --branch-prefix
query, order-independent with the existing mode/+yolo tokens. Firstmate
resolves the override at task intake and passes it explicitly, mirroring
how --mode already works; fm-brief.sh itself never reads the registry.

An empty override resolves to a bare "<task-id>" branch rather than a
leading slash. All five previously hardcoded fm/$ID sites (branch
creation, never-push rule text, definition-of-done text, and the status
message) now render the resolved prefix consistently.

* no-mistakes(review): Wire branch-prefix intake in AGENTS.md; fix fm-merge-local.sh hardcoded fm/ prefix

* no-mistakes(document): docs: document configurable ship-branch prefix in architecture.md

* no-mistakes(review): Persist immutable branch contracts

* no-mistakes(document): Document configurable ship branch prefixes

* no-mistakes(lint): Captain: fix ShellCheck test warnings

* fix(bin): map bearings PR rows to their recorded ship branch (kunchenguid#1887)

fm-bearings-snapshot.sh keyed a PR back to its task by string-matching
the headRefName against the fm/ prefix, so any project whose branch
prefix was overridden (e.g. via kunchenguid#2648's branch=<prefix> registry
annotation) had its PRs silently drop to task "-" in the bearings
view, exactly the third fm/-assumption issue kunchenguid#1887 named alongside
fm-merge-local.sh and fm-bearings-snapshot.sh itself.

fm-fleet-snapshot.sh now surfaces each task's recorded branch=
metadata field in its JSON task rows, and fm-bearings-snapshot.sh
cross-references a PR's headRefName against those recorded branches
before falling back to the legacy fm/ prefix heuristic, so a custom
branch prefix maps a PR back to its real task.

Adds a regression test proving a PR opened against a fix/<task-id>
branch resolves to that task instead of "-"; confirmed it fails on
the prior startswith("fm/") logic and passes with this change.
ShellCheck clean; full fm-bearings-snapshot.test.sh and
fm-fleet-snapshot-view.test.sh suites pass.

* fix(ci): align lint arithmetic-looking assignment and stale Bearings snapshot count

- Quote the --branch-prefix want_value assignment in fm-brief.sh, fm-promote.sh,
  and fm-spawn.sh so ShellCheck SC2100 no longer misreads the plain string
  'branch-prefix' as arithmetic shorthand.
- Bump the Stock macOS Bash snapshot job's hardcoded Bearings test-count
  assertion from 59 to 60: this PR added a Bearings test, so the count was
  stale, not the feature.

* no-mistakes(review): fix(bin): honor recorded ship branch in relaunch and review-diff

* no-mistakes(document): docs: complete branch-prefix flag in brief and promote headers

* fix(lint): quote branch-prefix parser token; drop unused BRANCH_Q after rebase

* no-mistakes(review): Restore %q branch escaping in promotion instructions with regression test

* no-mistakes(document): document recorded ship branch and prefix flag

fm-review-diff.sh's header is the owner of its branch-resolution
contract; it still described only the legacy local-branch behavior
after the change made review-diff honor state/<id>.meta's recorded
ship branch. README's feature bullet enumerates the registry's
optional flags and was missing the new branch=<prefix> override.

* no-mistakes(lint): Silence SC2016 on intentional single-quoted sed expression

* no-mistakes(review): address branch-prefix review findings in DoD and project-mode

* no-mistakes(test): branch-prefix suites pass under tasks-axi 0.2.6; environment-only failure

* no-mistakes(document): purge stale fm/ branch naming from docs and headers

* fix(test): assert the merged epoch status wording in the branch-prefix override test

The rebase resolution of tests/fm-brief.test.sh kept the branch's
pre-merge \`done: ready in branch ...\` assertion while the merged
fm-dod-lib.sh (carrying main's epoch-stamped status line) renders
\`done [at=<epoch>]: ready in branch ...\`. Align the assertion so the
override-consistency test matches the behavior it verifies.

* no-mistakes(review): Address remaining branch-prefix findings in four bin scripts

* no-mistakes(test): skip real-tasks-axi tests below the repo's 0.2.6 floor

* no-mistakes(document): document spawn's branch-prefix registry deviation notice
…d per-rule confidence floors (kunchenguid#5478)

* feat(bin): send dispatch resolver only the brief's task sections

* Sent Jev only the scaffolded Captain's intent and Firstmate spec
  sections, falling back to the whole brief when neither heading is
  present, so the identical setup, rules, and definition-of-done
  boilerplate no longer reads as a signal about the task
* Added an optional per-rule min_confidence that replaces the global 0.6
  floor for that rule; a picked rule below its own floor falls to the most
  probable other option that clears its floor, or returns ambiguous
* Kept files with no declared floor on the exact previous behavior and
  kept the model blind to the new field
* Recorded the live old-versus-new comparison over scaffolded fixtures

* no-mistakes(review): share brief heading parser, add kind line, fix floors

* no-mistakes(test): stop sending ship delivery mode to jev, keep scout tag

* no-mistakes(document): docs: list shared brief heading lib in scripts inventory
* feat(bin): supervision host core behind config/supervision-host

Add the supervision host (bin/fm-supervision-host.sh): beside a Claude
primary it owns the watcher cycle for the Stop auto-arm and, while the
away-posture record exists, hands each wake to a bounded headless Claude
engine session that runs the supervision branch's contract - the same
generated prompt, row eligibility, wake grant, per-actor drain, outcome
store, leases, and away relocation the Pi branch uses. Attended wakes pass
straight to main. Every path that cannot finish a wake hands it to main
with a supervision-host line; the park ends itself before the Stop hook
timeout with a cycle-boundary wake.

- bin/fm-supervision-engine-lib.sh: opt-in parse, verified engines
  (claude, default sonnet), one bounded engine turn, and a reap of engine
  tool processes that sit in their own process groups.
- bin/fm-branch-report.sh: the command twin of fm_branch_report, scoped
  to the tasks the current host turn claimed.
- bin/fm-branch-dispatch.mjs: command entry to the Pi dispatch module, so
  eligibility and the wake prompt have one owner.
- bin/fm-claude-stop-autoarm.sh runs the host in the arm's place when
  config/supervision-host exists; nothing changes without the file.
- bin/fm-watch-arm.sh --stop: home-scoped stop without a re-arm.
- bin/fm-lease-lib.sh: an opted-in home takes the lease-command lock for
  unmarked main too, closing the first-claim race; the refusal tells the
  caller to leave the lease alone and retry.
- /afk launches no away daemon on an opted-in Claude home; /quiet still
  does. Session start renders the host's main-side protocol there.

* fix(bin): relay a host turn's outcomes when the captain returns mid-turn, and log per-turn engine cost

Live validation found two supervision host gaps. A captain who returns while
an engine turn is running gets a return brief rendered before that turn's
outcomes exist, so the host now hands the close to main with those outcomes.
Claude reports a resumed conversation's running cost, so the engine lib now
derives each turn's cost from the total the host records, and the host log
records every close's destination.

* docs(verification): record the supervision host's live evidence

The dated live results behind docs/supervision-host.md: the Claude engine's
live guard, the away-wake cases against real workers, the engine's cost
reporting, and the flag-off before-and-after regression.

* docs: describe the supervision host ledger as covering every close

* no-mistakes(review): Harden supervision host ownership, boundary, ack, and late outcomes

* no-mistakes(review): Recheck park boundary just before starting an engine turn

* no-mistakes(review): Cap park boundary, deliver all host lines, reject incomplete results

* no-mistakes(document): Correct supervision host documentation and stale pointers
…t-in, and delivery (kunchenguid#5506)

Attestation MATCH; contract-class restore; CI/NM green. Squash-merged by Kun's firstmate.
…kunchenguid#5528)

* fix(bin): bound the startup-network worker's lock waits by its budget

Fixes kunchenguid#5377

The deferred startup network worker bounded its sweeps with a stage budget but
took the publish lock and the fleet-lock lease with an unbounded wait, so a live
holder of that lock kept the detached worker alive for hours past its timeout
with its output discarded at the end. Every wait now goes through the bounded
acquire and shares the remaining stage or delivery budget; a lock a live process
still holds at the deadline ends the worker with a failed record naming the
holder and the rerun command, and a wake so the result surfaces.

* no-mistakes(review): propagate publish exit code from cmd_run terminal paths
…dth (kunchenguid#5517)

* fix(bin): keep the ps fallback identity independent of terminal width

fm_pid_identity's portable fallback read the command column at the
ambient COLUMNS width, so an identity recorded from a wide shell never
matched the one recomputed inside a narrow hook and the continuity guard
denied every fleet command. Pass -ww so the column is never cut.

Fixes kunchenguid#799

* no-mistakes(ci): Fixed CI failure in Behavior portable serial 4. Root cause: the -ww flag added in commit ac7ab5d to fm_pid_identity (bin/fm-wake-lib.sh) shifted the ps argv so $1 became -ww instead of -p, breaking the positional fake-ps fixtures in tests/fm-procevent.test.sh (lines 2736, 3571) which then fell through to real ps and failed the fm-procevent test. Fix (already applied in the worktree, matching the authoritative user instruction exactly): replaced -ww with a COLUMNS=10000 environment pin so the call is COLUMNS=10000 LC_ALL=C ps -p "$pid" -o lstart= -o command=, mirroring fm_pending_reply_pid_identity in bin/fm-pending-reply-lib.sh:982. argv is back to -p PID -o lstart= -o command=, so the fixtures match again with no fixture edits. Comments above the call in bin/fm-wake-lib.sh and in test_pid_identity_is_terminal_width_invariant (tests/fm-watcher-lock.test.sh) now describe the COLUMNS pin instead of -ww; the regression test still asserts narrow-vs-wide byte equality and the full command. Verified: the terminal-width-invariant regression test passes. The only local not-ok results were flaky, run-varying timing tests (procevent launch/claim confirmation, listener reparenting) that differ each run and are unrelated to the ps argv change
…thorized intent (kunchenguid#5526)

Fixes kunchenguid#3608

When a scout is promoted to a ship, the captain's authorized intent is
extracted from a legacy `# Task` body by matching `Captain:` and `[captain]`
lines anywhere in the body, including inside fenced code blocks and indented
examples, while the heading reader already tracks fences. A fenced `Captain:`
example therefore passed the provenance gate and became the ship contract's
intent while the real ask was dropped.

Make the captain-words extractor fence-aware like the heading reader: a line
inside a ``` or ~~~ fenced block, or indented four spaces or a tab, is never a
marked line. The promotion and spawn callers need no change. The regression
test covers both the extractor and the promotion provenance gate refusing a
brief whose only Captain lines are fenced or indented examples.
…riefs (kunchenguid#2868)

* fix(bin): forbid administering the shared worktree pool in crewmate briefs

A crewmate ran a `git worktree remove` loop over the treehouse pool its own
worktree came from, destroying five worktrees - four belonging to tasks that
were running mid-pipeline. The generated brief's rule 2, "stay inside this
worktree; modify nothing outside it", is a rule about files: removing a
worktree is administration of shared state, not an edit outside a directory,
so the sentence never reached the act. The worker satisfied its brief
completely.

Rule 7 already named one piece of shared infrastructure - the no-mistakes
daemon, one instance serving every lane - with the reason stated plainly. The
worktree pool is the same class of thing and was unnamed.

Fold the pool into that existing rule rather than adding a second warning:
state the constraint around the act (create, remove, return, prune, move,
reassign a worktree or pool slot; write into a sibling slot), keep concrete
commands as examples rather than as the definition so no single provider is
pinned, and give the prohibition a real exit through `blocked:`.

The rule is emitted from one shared string interpolated into both crewmate
scaffolds, so the ship and scout copies cannot drift apart. The secondmate
charter deliberately omits it: that home runs its own fleet and legitimately
allocates and returns slots for its own crewmates.

Contract text only; no runtime enforcement layer.

* no-mistakes(document): Distill pool-safety comment rationale

* no-mistakes(review): align pool-rule test grep patterns with emitted [at=<epoch>] text
… a dispatch record (kunchenguid#5524)

* fix(bin): refuse tasks-axi add --start so In flight always has a dispatch record

Fixes kunchenguid#4753

Dispatch (bin/fm-spawn.sh) is the only path that moves a backlog row to
In flight, because it creates the task record, status file, and inbox
that go with the row. A row hand-placed there through the wrapper's
`add --start` had none of those, and nothing later noticed, so the
live-task count included work nobody was doing. The wrapper now refuses
`add --start` (exit 2) and names the dispatch path; plain `add` and
`start <id>` pass through unchanged, and the lifecycle transitions
address tasks-axi directly so dispatch is unaffected.

The issue's other half, a reconcile sweep in bin/fm-inactive-reconcile.sh
that notices an In flight row with no task record, is left as is; this
change closes the only path that creates such a row.

* no-mistakes(review): refuse create --start alias, not just add --start

* no-mistakes(review): reword add --start guard docs to drop only-path overclaim

* no-mistakes(review): scope add/create --start guard docs, drop universal claim
…#5503)

* feat(bin): run the supervision host beside the other non-Pi primaries while away

Cursor's stop-hook park, the OpenCode plugin, the omp watch extension, Grok's
model-owned background arm, and Codex's foreground checkpoint now run
bin/fm-supervision-host.sh in the watcher arm's place when the home opted in
with config/supervision-host, so the host's Claude engine takes away-posture
wakes beside those primaries exactly as it does beside Claude. Without the
file nothing changes.

- The host streams its first cycle's status line, accepts --restart and the
  owner's predecessor arm for its first cycle, and prints each exit in one
  write, so owners that wait for arm readiness and restart their own
  successor (OpenCode, omp) keep their handling handoff.
- Codex's checkpoint passes its bound to the host as the park boundary,
  raises it to FM_CODEX_WATCH_CHECKPOINT_AWAY (3600 s) while the away record
  exists, and lets an engine turn that starts before the bound finish after
  it (FM_SUPERVISION_HOST_PARK_LIMIT).
- /afk launches no away daemon on an opted-in home of those harnesses and
  says so at entry when the file selects no engine for that primary.
- Session start renders the host protocol for each arm owner, and Grok's
  arm command becomes the host.

* fix(bin): keep the watcher-down banner away from the supervision branch actor

A supervision host's engine turn runs guarded commands after its successor
watcher cycle may already have closed on a newer wake, so the guard showed it
the watcher-down banner with the primary's repair line. Under a Codex primary
pin that line is the checkpoint, and a live Codex lab run showed the away
session running it mid-turn (the nested host stood down on its ownership
check). The branch actor never owns watcher continuity, so the banner, its
reminder, and the episode state now leave that actor out, as the queued-wake
warning already does.

The lint telemetry fixture counts bin/fm-afk-launch.sh's source directives,
which the host engine note raised from four to five.

* fix(bin): queue away-session outcomes recorded after the return for main

A Cursor park superseded by the captain's return stops its host as the
engine turn ends, so the host's own handoff of that turn's outcomes was
never printed and the outcomes never reached main. The report surface now
queues every outcome it records after the away record is gone as a durable
check wake; the return owner archives the record before it reads the store,
so each outcome is in the return brief, queued, or both. A host stopped
mid-turn also removes its turn's result and error files.

The stream test now acknowledges its first close and accepts a restarted
cycle that closes on its resurface before the arm confirms it.

* fix(bin): clear a hard-killed host's turn at the next activation

A Cursor park superseded mid-turn can kill its host outright, which runs no
cleanup, so the turn's result, error, and descendant files stayed behind and
any tool process the engine started was left running. The next host's
activation now reaps the descendants that turn recorded and removes its
files. The host suite also registers its homes in a file, because
make_home runs in a command substitution, so its cleanup now stops every
host a case leaves running.

* fix(bin): leave rows that arrive after main's drain unclaimed at its acknowledgement

Main's acknowledgement re-claimed every unreserved queued row, including one
that arrived after the drain above the acknowledged cutoff. That row stayed
main's without ever being shown to it, so while away the supervision host
refused every later wake that included it and handed each back to main until
main drained again. The acknowledgement now claims only unreserved rows at or
below its cutoff.

* docs: name the killed turn's engine and files in the host's failure direction

* docs: record live supervision host runs on the non-Pi primaries

* no-mistakes(review): Replay host-only supervision boundaries across omp session replacement

* no-mistakes(review): Deliver omp supervision-host wakes only at the host's close

* no-mistakes(document): Correct supervision host documentation for non-Pi primaries

* no-mistakes(ci): Fixed the CI failure by naming FM_CODEX_WATCH_CHECKPOINT_AWAY in the rendered Codex host instructions. The focused instruction and checkpoint suites pass
)

* feat(bin): auto-relaunch dead persistent secondmates during ordinary supervision

A persistent secondmate whose primary agent exits mid-session previously
stayed down until the next session-start liveness sweep. Extract the
sweep's probe/classify/relaunch mechanics into a shared library and drive
the same contract from a cadence-gated watcher tick, so a positively dead
or missing endpoint is relaunched through the guarded spawn path within a
poll cycle instead of an hour later.

Only the recovery-grade `dead` and `missing` verdicts authorize relaunch;
ambiguous, unreadable, unverified, and unreachable-remote reads stay
fail-closed and a remote route is never replaced by a local endpoint.
Each relaunch emits exactly one `check` wake and appends to a durable
per-mate ledger; a mate exceeding the bounded attempt budget is parked
behind a marker until a live probe rearms it. A per-mate liveness lock
serializes the tick against a concurrent session-start sweep.

* no-mistakes(review): Fail closed on relaunch ledger errors; clear state on remote teardown

* no-mistakes(review): Share ledger read guard; retire relaunch state under liveness lock

* no-mistakes(review): Lazy-load wake lib; live rearm restores full relaunch budget

* no-mistakes(review): Finish liveness tick for every mate before waking once

* no-mistakes(review): Keep liveness tick scanning past per-mate errors, then wake

* no-mistakes(review): Wake only on queued rows; teardown holds liveness lock

* no-mistakes(review): Queue liveness outcome wake before releasing mate lock

* no-mistakes(document): Update secondmate liveness documentation for mid-session recovery

* no-mistakes(lint): Fix empty assignments flagged by ShellCheck

* no-mistakes(ci): Added ShellCheck analysis boundaries for the shared liveness library in both callers and marked its result globals as intentional library outputs. Changed-file lint passed; full CI partitions were not run locally

* no-mistakes(ci): Fixed Lint 2 by removing an unused test variable in tests/fm-wake-queue.test.sh. ShellCheck, bash syntax, and the full wake-queue test script pass

* no-mistakes(ci): Fixed the CI wake-queue fixture: stall-only watcher legs now seed the liveness cadence marker, preventing the new endpoint probe from interfering with their assertions. The full wake-queue test, ShellCheck, and diff checks pass locally
…cycle ends (kunchenguid#5550)

* fix(bin): start a successor when the Claude Stop-hook arm's attached cycle ends

Fixes kunchenguid#2381

When the Claude Stop hook's foreground arm attached to a peer watcher cycle
and that cycle ended, the arm reported the delivered wake and the hook exited
2 without starting a successor, so the handling turn ran with no watcher.
Pi, omp, and OpenCode start the next arm before delivering the wake and pass
the closed arm's pid as FM_WATCH_PREDECESSOR_ARM_PID; the Claude hook never
passed that predecessor identity.

The hook now runs its arm as a tracked child it waits on, so it holds that
arm's pid, and after any actionable close starts one handling-successor
bin/fm-watch-arm.sh with the closed arm's pid as FM_WATCH_PREDECESSOR_ARM_PID.
The successor is launched the one way a process outlives a Claude hook's
exit-2 rewake (nohup, detached stdio, own process group, the shape
bin/fm-startup-network.sh already uses); the hook waits for its status line
and adds one banner line when no live watcher was confirmed, never withholding
the wake. The supervision-host path is unchanged, as is the arm wrapper.

The regression test drives the real hook against an arm fixture whose attached
peer cycle ends: it fails on the previous tip because no successor starts, and
now asserts the successor names the closed arm as its predecessor and outlives
the rewake. A second case pins the unconfirmed-successor banner line.
docs/watcher-continuity.md no longer records the Claude asymmetry.

* no-mistakes(ci): Serial-4 failure was a real regression: tests/fm-session-lock-ancestry.test.sh asserts exact cumulative arm-invocation counts while driving the real fm-claude-stop-autoarm.sh hook against a stubbed fm-watch-arm.sh. This PR makes the hook start a handling successor after an actionable close, so every owned actionable phase now records TWO arm invocations (foreground arm + successor) instead of one, breaking "healthy chain: expected 1 arm(s), got 2". Fixed by updating the cumulative expectations to match the new behavior: owned phases 1/2/6 -> 2/4/6, foreign carry phases 3/4/5 -> 4, plus a comment explaining the +2-per-owned-phase model. Verified: phase-1 (the CI failure point) now passes on every run, syntax checks clean, and sibling arm-count tests (fm-claude-stop-autoarm.test.sh, fm-cursor-primary, fm-turnend-guard) pass unchanged. The only remaining local not-ok is a WSL-only environmental artifact (orphan reparents to a subreaper, not PID 1) that passes on the CI runner. Parallel-1 failure is an unrelated flake: its 11 tests (fm-lint, fm-pr-merge, fm-test-run, fm-cd-pretool-check, fm-pi-primary-types, fm-grok-harness, fm-composer-lib, fm-review-diff, fm-tmux-submit-busy, fm-composer-ghost, fm-brief) do not include fm-session-lock-ancestry and none reads any file this PR touches; all pass locally. It should clear on CI re-run. Made the smallest root-cause fix (one test file, 6 count updates + a clarifying comment). Validation of the branch continues through the no-mistakes pipeline, which owns re-running CI
…ning (kunchenguid#5566)

* fix(bin): report a Lavish source armed only after its listener is running

Registration alone was treated as ready, so arm could succeed before anything was collecting from the board.

* no-mistakes(review): Guard Lavish arm launches, keep retire refusals, report live prior listener

* no-mistakes(review): Keep polling through window before reporting a still-live prior listener

* test: wait for a capture's claim to drop before the next arm

The result is stored before the runner exits, so a re-arm in that gap was meeting a live claim.

* no-mistakes(document): Record Lavish arm readiness evidence in verification doc

* no-mistakes(ci): Both failures were caused by this PR, and both are fixed with test-only edits. Lint 2 (ShellCheck SC2034): this branch removed the only use of `reply_id` (a `start "$reply_id"` call) from tests/fm-procevent.test.sh, which left the assignment at line 1450 unused. I deleted that assignment. It was the only `reply_id` in the file. ShellCheck is now clean on both test files. Behavior portable serial 4: the failing test was tests/fm-bearings-board.test.sh, in the check "registration consumed its answer before the any-origin binding existed". I reproduced it locally: the hold was still `state: queued` when the test checked it. - What must hold: the test's check that the hold is closed must run after the listener has captured the answer. - Why it broke: the test used a stand-in adapter that ran `fm-procevent.sh start` in the foreground after `arm`, so capture finished before build returned. On this branch, `arm` starts the listener itself in the background, so the real listener captures the answer and closes the hold a moment after build returns. - Fix: removed the now-redundant stand-in adapter, the copied runtime directory, and its extra environment variables. The test now runs the real build through the existing `run_board` helper and waits up to about 10s for the hold to reach `state: done`. The checks that follow are unchanged: `Resolution mode: answered` and the any-origin binding. - Other tests: this was the only test in the file that stood in for the adapter this way. The shard's other pure-contract-unit test (tests/fm-trace-context-lib.test.sh) passed unchanged. Verification: - tests/fm-bearings-board.test.sh passed 3 times in a row via bin/fm-test-run.sh, all 18 checks, about 53s per run. - tests/fm-procevent.test.sh was not rerun, because the lint fix only removed an unused assignment
…elivered (kunchenguid#5599)

* fix(bin): acknowledge a delivered unknown-wake escalation

The same unrecognized wake was escalated again after it had already been handled, because delivery never recorded that identity.

* no-mistakes(review): Scope unknown-wake acknowledgements to one away session

* no-mistakes(review): Clear delivered digest when unknown-wake ack write fails

* no-mistakes(review): Limit unknown-wake suppression to acknowledged lines

* no-mistakes(document): List unknown-wake ack file among away-session artifacts
…ess wait (kunchenguid#5587)

* fix(bin): keep a stated default retraction from cancelling a keyless wait

A resolved line that names the shared default decision bucket was closing the keyless live wait that only prints as that same key. Keyless self-retraction still closes the keyless wait.

* no-mistakes(review): Keep declared waits standing past foreign-key resolved lines

* no-mistakes(review): Bound declared-wait read and share one decision-key parser

* no-mistakes(document): Document supervisors' key-aware declared-wait read
…unchenguid#5544)

* fix(bin): terminate a remote job worker that lost ownership when it receives TERM

A serving worker whose lock directory is gone can no longer quarantine shutdown, and resuming service publishes a false ready heartbeat. Exit after stopping only that worker's own command tree, without removing a replacement owner's lock.

* no-mistakes(review): Check worker lock ownership before publishing shutdown quarantine

* no-mistakes(document): Correct worker shutdown comment on replacement-owned lock

* fix(bin): keep an ousted remote job worker off the replacement quarantine

Shutdown can lose the lock after the first ownership check and before it
writes or clears quarantine. Bind both operations to the directory object
this process still owns so a replacement's quarantine stays untouched.

* no-mistakes(review): Make ousted-worker shutdown test reliably reach quarantine clear

* no-mistakes(document): Reattach worker_shutdown doc comment to its function

* no-mistakes(ci): Fixed the failing check (Behavior portable serial 7) with a test-only change to the stall test in tests/fm-remote-job.test.sh. Product code is unchanged; no other test changed. Cause: after the decoy dies, both workers run the same check-exists, read, delete sequence on the job records. On the CI runner the replacement deleted a record between the ousted worker's check and its read. The ousted worker exited 125, and because the file runs under set -e the unguarded `wait` ended the test with 125. The exit trap then killed the replacement, which produced the "Killed" line. Reproduction: a temporary 0.3 s delay between the check and the read, applied to the ousted worker only, made the committed test fail exactly as in CI (exit 125 and the "Killed" line). The new test passed with the same delay. The delay is reverted, along with a similar debug hook that the timed-out attempt had left in bin/fm-remote-job-worker.sh. Test changes: - The replacement is frozen (and confirmed stopped) before the decoy is killed and resumed only after the ousted worker exits, so only one worker touches the job records at a time. - The ousted worker is stopped only once its quarantine exists and its lane is reaped, which places it inside its stop loop. - Every fixed poll loop is now a wait on a named condition with a 30 s deadline and an explicit failure message. Exit detection also handles zombies. - The exit trap kills and waits for the decoy and both workers on every path. - A non-zero exit from the ousted worker now fails with its exit code and stderr instead of silently ending the file. The test still proves that the resumed ousted worker exits 0 and leaves the replacement's lock, quarantine contents and quarantine inode unchanged. Verification: the full test file passed four times on its own and three times under nice -n 10 with four busy-loop CPU hogs; bin/fm-lint.sh passes. Changes are not committed

* no-mistakes(ci): I fixed the failing check (Behavior portable serial 7) by changing only the stall test in tests/fm-remote-job.test.sh. Product code is unchanged. **What failed:** "an ousted worker in shutdown leaves the replacement quarantine untouched" failed on CI with the ousted worker exiting 125 ("could not stop the active command tree"). **Why:** during shutdown, the worker retries the still-running decoy command group a fixed 100 times, 0.01 s apart, then gives up and exits 125. The test tried to freeze the worker partway through those retries by sending SIGSTOP from outside. On a slow runner the retries ran out before the stop arrived, so the worker had already given up. The invariant is that the test must hold the ousted worker inside that retry loop until the replacement owns the lock. That was the only place the test depended on timing. The other waits already watch for a named state change with a 30 s deadline. **Fix:** - The ousted worker now starts with a small `sleep` wrapper at the front of its PATH, and the SIGSTOP race is gone. - The wrapper only holds a `sleep` called directly by that worker's own process (it checks its parent pid against a hold file) while its quarantine file exists. - The only such `sleep` is the first retry in the shutdown stop loop, so the worker waits there as long as needed. - The wrapper writes a marker when it starts holding. The test waits for that marker, then hands the lock to the replacement, freezes the replacement, and kills the decoy. - The test releases the worker by deleting the hold file. Deleting the whole temp directory also releases it, so a failed run cannot leave the wrapper looping. - A process leak: the test overwrites the job's command-group record with the decoy, so no worker ever stopped the job's real command. `fm-hold-job.sh` and its `sleep 30` stayed running for up to 30 s after the test. The test now records that group before overwriting it and kills it at the end of the test and in the exit cleanup. - The test still asserts the same things: the ousted worker exits 0, and the replacement's lock, quarantine contents and quarantine inode are unchanged. **Verification:** - The full file passed twice on its own, twice under `nice -n 10` with six busy-loop CPU hogs, and twice more after the leak fix. - `pgrep` found no leftover processes afterwards. - With the worker from just before the fix commit (cf45cb6^), the test still fails with "the ousted worker wrote or cleared the replacement quarantine during shutdown", so it still proves the fix. - `bin/fm-lint.sh` passes. - I did not reproduce the CI failure locally. The cause comes from the fixed retry limit and the CI error message. The changes are not committed

* no-mistakes(ci): I changed only the stall test ("an ousted worker in shutdown leaves the replacement quarantine untouched") in tests/fm-remote-job.test.sh. Product code is unchanged, and so is every other test. **Invariant:** the pid written to the job's group record must be a process-group leader whose group dies when that one process is killed. Otherwise the worker's bounded stop loop never sees the group die, gives up, and exits 125 ("could not stop the active command tree") before it reaches the lost-ownership exit. The decoy is the only place in this test that depends on this. **Fix:** - The decoy used to be `set -m; sleep 30 &`. It now starts as `perl -MPOSIX=setsid -e 'setsid() >= 0 or exit 1; exec @argv' sleep 30 &`, which gets its own session and group without shell job control. tests/fm-procevent.test.sh already uses the same idiom. - The test now waits, with the file's usual 30 s deadline and a named failure, until `ps -o pgid=` of the decoy equals its pid before writing it into the group record. This way the worker can never read the record before `setsid` has run. - The existing steps are unchanged: the test kills the decoy, reaps it with `wait` before releasing the hold file, and the exit trap still kills and reaps the decoy and both workers. - The assertions are unchanged: the ousted worker exits 0, and the replacement's lock pid, quarantine text and quarantine inode stay the same. **Cleanup:** I reverted a debug `printf` hook that the timed-out previous attempt had left in bin/fm-remote-job-worker.sh, and deleted its untracked `.tmp-repro/` directory. Neither was committed. **Verification:** - The full tests/fm-remote-job.test.sh passed twice normally and once under `setsid -w` with stdin from /dev/null (no controlling terminal). - `bin/fm-lint.sh` passes. - No leftover `sleep 30` processes afterwards. **Not reproduced:** I could not reproduce the CI failure locally. On this host `set -m` made the decoy its own group leader even without a controlling terminal, so the cause on the runner is not confirmed. The change removes the test's reliance on shell job control, as the user asked. Changes are not committed

* no-mistakes(ci): I changed only the stall test ("an ousted worker in shutdown leaves the replacement quarantine untouched") in tests/fm-remote-job.test.sh. Product code is unchanged, and so is every other test. **Invariant:** the group record the ousted worker checks in its stop loop must stay the job's own command group, and the test must stop that group before it releases the hold. Otherwise the bounded retry keeps seeing a live group, gives up, and exits 125 ("could not stop the active command tree") before it reaches the lost-ownership exit. The test overwrote this record in one place (the decoy) and stopped the group in one place (killing the decoy); both are changed. **Fix:** - I removed the setsid decoy and the overwrite of `.claim/group`. The record keeps the job's real command group, which the test still saves as `STALL_JOB_GROUP`. - The two-line `group_start` stays. It is still needed: without it the worker kills the real group on its first pass, before the replacement takes over, so the hold would never matter. - The `sleep` wrapper that holds the worker at its first stop-loop retry is unchanged. - After the replacement owns the lock, its quarantine is planted and it is frozen, the test runs `kill -KILL -- -$STALL_JOB_GROUP`. It then waits, with the file's usual 30 s deadline and a named failure, until `kill -0` on the group fails. Only then does it remove the hold file. The worker therefore always sees its own command already stopped and never races its retry budget. - The exit trap still kills the saved command group if the test fails. It can't `wait` on that group because the group is not a child of the test shell. The decoy variable and its cleanup entry are gone. - The assertions are unchanged: the ousted worker exits 0, and the replacement's lock pid, quarantine text and quarantine inode stay the same. **Verification:** - The full tests/fm-remote-job.test.sh passed twice normally. - It passed once under `setsid -w` with stdin from /dev/null (no controlling terminal). - It passed once under `nice -n 10` with six busy-loop CPU hogs. - With the worker from before the fix (cf45cb6^), the test still fails with "the ousted worker wrote or cleared the replacement quarantine during shutdown", so it still proves the fix. - No `fm-hold-job` or `sleep 30` processes were left afterwards. - `bin/fm-lint.sh` passes. **Not reproduced:** I couldn't reproduce the CI failure locally; the decoy version also passed on this host. So I can't confirm why the decoy group stayed alive on the runner. The new wait turns any leftover live group into a clear named failure instead of an exit 125. The changes are not committed

* fix(bin): keep a dead command group dead on bash 5.2

A bare return inside the liveness check drops the failing kill status when the check runs in a conditional, so shutdown keeps treating a stopped group as alive and exits 125.

* no-mistakes(review): Use bash 3.2 fd syntax and fix trap return comments
…henguid#5589)

* docs: make configuration settings easier to find and understand

* no-mistakes(review): Restore dropped qualifiers and fix misplaced config doc labels

* no-mistakes(review): Restore three dropped qualifiers in configuration reference
)

* fix: bound worker edits of project AGENTS.md/CLAUDE.md to factual corrections

These files are loaded into every agent session of a project, so additions
should be a deliberate human choice rather than automated task output. The
ship brief's project-memory section and AGENTS.md section 6 previously invited
workers to record durable knowledge, which let project AGENTS.md files accrete
detail the codebase or README already carries. Workers now edit only to fix
factually wrong content - including content their own change made wrong - and
fm-ensure-agents-md.sh runs only alongside such a correction. Stow no longer
routes project-memory additions through ship tasks, and the generated skeleton
no longer invites discovery-driven additions.

* no-mistakes(review): Stop running fm-ensure-agents-md.sh on memory-file corrections

* no-mistakes(document): Clarify manual project-memory initialization and remove duplicate guidance
…enguid#5635)

* fix(bin): let gate agents drive lifecycle against marked lab homes

Part 2 of the kunchenguid#5615 split. A no-mistakes gate agent runs inside a
checkout carrying the fleet-captain identity, so fm-gate-refuse-lib
refuses fleet mutation on the gate signal. That refusal was absolute,
which kept gate validation from ever exercising the real lifecycle.

Stamp a disposable lab FM_HOME with a .fm-lab-home marker file that
only bin/fm-lab-home.sh writes, and only onto a fresh empty dir, so
no call path can mark a populated real home. fm_refuse_if_gate_agent
then permits lifecycle only when FM_HOME carries the marker and is
driven through its stock layout - any FM_*_OVERRIDE relocation stays
refused so part of the "lab" cannot be split back onto the real fleet.
The threat model is a confused agent touching the real fleet, not
deliberate forgery, so the marker is a plain token file rather than a
bound record. FM_GATE_REFUSE_BYPASS is unchanged: it still serves the
test harness, which cannot mark hundreds of temp homes.

Teardown's slot-ownership scan compared state-dir paths textually
while fm_firstmate_root_home canonicalizes, so a lab home under a
symlinked TMPDIR scanned its own record twice and self-collided;
compare file identity (-ef) instead.

* no-mistakes(review): Refuse unlistable lab homes and hardlinked slot records

* no-mistakes(review): Mint lab markers only on verified-empty fresh dirs

* no-mistakes(document): Clarify lab-home gate documentation and comment contracts

* no-mistakes(document): Clarify lab-home gate documentation and remove stale claims

* no-mistakes(document): Clarify gate lab-home documentation and boundary wording
…ad of refusing every re-arm (kunchenguid#5594)

* fix(bin): replace a watcher whose beacon stalls past a hard bound instead of refusing every re-arm

A fleet watcher that is alive but whose liveness beacon has gone stale could
never be replaced: every re-arm was refused because the lock holder was a live
pid, and the holder was never evicted because it was not dead. Add
FM_WATCHER_STALL_BOUND (default 3x the stale grace): below it the refusal is
unchanged; at or past it the arm re-verifies the holder against the lock's
recorded identity, sends TERM, waits boundedly, and takes the lock the normal
way, ledgering a stalled-holder-replaced row. A holder that survives TERM keeps
the old refusal.

Fixes kunchenguid#4400

* no-mistakes(test): poll for replacement message to fix watcher-lock test flake

* no-mistakes(document): document FM_WATCHER_STALL_BOUND in config inventory
…n can keep them (kunchenguid#5563)

* fix(pi): hide queued Firstmate notifications under Calm only when the session can keep them

Calm now keeps authenticated Firstmate operational inputs out of Pi's queued-message
listing, but only after proving the live session exposes every member needed to keep
them across Escape. A session missing any of them keeps stock rows and Escape and shows
one generic warning. Escape and the dequeue key return only captain-authored messages to
the editor and re-queue hidden notifications in order; after an abort that kept any in
Pi's agent queue, the adapter starts the delivery turn itself because Pi 0.87.1 does not
continue an aborted run. Compaction-held notifications stay with Pi's compaction flush and
never start or announce a turn.

Fixes kunchenguid#1588

* docs(calm): record Pi 0.87.1 queued-row retention verification

* no-mistakes(review): Deliver kept Calm notifications after tree-navigation aborts too

* no-mistakes(review): Defer Calm notification turn until tree navigation finishes

* no-mistakes(lint): Silence SC2016 for literal JavaScript in queue-retention e2e test
…uid#5548)

* fix(bin): refuse teardown when a required source disappears

A missing sibling was sourced after cleanup had started, so Bash 3.2
exited 0 from the EXIT trap and Bash 5 continued and reported success.

* no-mistakes(review): Remove unused FM_TEST_ONLY hook from teardown tests

* no-mistakes(review): Check task backend sources before any teardown cleanup

* test(gotmp): give teardown fixtures every tmux adapter sibling

Teardown now refuses when a sibling the recorded backend's adapter sources
is missing, so the fake bin must carry fm-session-lock-lib.sh,
fm-agent-process-lib.sh and fm-gemini-lib.sh.
Restructure the supervision host doc's prose into shorter sections, lists,
and tables without changing documented behavior. Every original heading,
anchor, identifier, number, quoted string, and link target is preserved.
* docs: make herdr-backend easier to read

Restructure the Herdr backend doc's prose into shorter sections, lists, numbered procedures, and tables without changing documented behavior. Every original heading, anchor, fenced code block, link target, and documented fact is kept.

* no-mistakes(document): Restore composer-proof reason and complete Herdr topic table
kunchenguid and others added 29 commits September 28, 2026 17:19
…6064)

* fix(bin): read a live quiet record as a present captain at the host and watcher

A quiet record left without its daemon (a quiet start that never ran or was
interrupted) was read as away by the supervision host, so it parked a present
captain's main and held captain outcomes for a return that never comes, and
the watcher and daemon silenced captain-held rechecks on record presence.

The host's posture checks, the watcher's and daemon's captain-held silencing,
and the host's outcome path (branch report, drain BRANCH OUTCOMES, relocated
branch authority, the owners' away wake note, and the Codex checkpoint bound)
now ask the record owner's away-or-quiet reading, so only an away record is
away. A live away record keeps today's behavior.

* no-mistakes(document): Correct quiet-record documentation and supervision guidance

* no-mistakes(document): Clarify quiet-record posture and captain-held rechecks

* no-mistakes(document): Clarify quiet-record posture in documentation
…kunchenguid#6043)

* fix(bin): name an in-window engine latch in the return brief and drop the false handling GAP line

The away return brief said nothing had failed after the supervision host
latched on engine errors during the window, and printed a GAP: watcher
downtime line whenever a wake was merely being handled or queued at return.

The failures section now reads the host ledger and latch record and names
the latch time, the window's engine-error count, and whether the session
is still paused or recovered. An open recovery episode is reported as
information, and as a gap only when a queued episode outlived the return
grace or the marker cannot be read.

* no-mistakes(review): Fix latch trip time, drop marker-age grace, bound error count

* no-mistakes(review): Report paused latch without ledger trip row; bound errors

* no-mistakes(review): Never report a failed probe's latch row as trip time

* no-mistakes(review): Only a retained trip row marks a pre-window latch

* no-mistakes(document): Clarify return-brief latch and watcher-gap documentation

* no-mistakes(ci): Fixed Lint 1 by marking the shared cooldown constant as used by sourcing scripts. The repository lint command and diff check pass; the return test run was stopped by a 180-second timeout after its completed cases passed

* no-mistakes(ci): Fixed the return brief so the trip time and error count come from the same initial latch row, and ledger rows before the current session’s lock boundary cannot affect its latch report. Added real-script regressions for both findings. The return test suite, repository lint, and diff check pass

* no-mistakes(ci): Fixed the return brief’s restart cutoff so it retains in-window failures, prints one line per initial-trip row, and omits zero-error count wording. Added real-script restart regressions. The return test suite, ShellCheck, and diff check pass

* no-mistakes(ci): Fixed the return brief so a recorded trip followed by recovery stays recovered, while a later pause with a lost trip append gets a separate “trip time unavailable” line. Added a real-script regression that failed before the fix. The return test suite, ShellCheck, syntax checks, and diff check pass

* no-mistakes(ci): Fixed the false second latch during recovery. A real-script regression failed before the fix and passes now; the lost-second-trip test still passes. The return test suite, ShellCheck, syntax checks, and diff check pass
* fix(calm): name the Claude Code Calm plugin fm so supervision notes read "fm: "

Claude Code labels every mod transcript line with the plugin name, so the
notes rendered as "firstmate-calm: ⚓ ...". Rename the plugin to fm, update
the live guard to assert the fm: label, and document the one-time replay for
sessions resumed across the rename.

* no-mistakes(document): Clarify Calm plugin rename in documentation
…#6037)

* feat(bin): add fm-live-lab.sh, a one-command live supervision lab builder

* fix(bin): exact lab windows, per-lab task ids, self-safe teardown

* fix(bin): target lab windows by id, stop lab descendants, add readiness tests

* fix(bin): keep Claude's auto-updater off in live labs; list fm-live-lab.sh

* fix(bin): start the lab tmux server without user config

* no-mistakes(review): Scope lab teardown to its store, root, and task ids

* no-mistakes(review): Record selected user stores at up for check and down

* no-mistakes(document): Clarify live lab documentation and remove stale narratives

* no-mistakes(ci): Fixed the CI failure by checking for an existing lab root before looking up the harness executable. The affected behavioral test and shell syntax check pass; the refusal also works with Claude absent from PATH

* no-mistakes(ci): Fixed all four Greptile findings: teardown signals only recorded lab processes and their descendants; the worker gate is in its granted task directory and its path is exposed; readiness uses current crew state; and mate and worker IDs use 12 nonce hex digits. The CLI behavior tests pass, as do shell syntax, ShellCheck, and diff checks. The Claude no-host path is unchanged

* no-mistakes(ci): Fixed the CI test’s dependence on an installed Claude binary by supplying a test-local stub. The full fm-live-lab test, shell syntax check, and diff check pass

* no-mistakes(ci): Fixed all three selected findings in bin/fm-live-lab.sh: down waits for recorded processes and escalates before cleanup, PID roots are checked against recorded start times, and Claude primary trust is rechecked after mate/worker readiness. Added behavioral tests in tests/fm-live-lab.test.sh. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass

* no-mistakes(ci): Fixed the pre-primary settle wait, worker gate instructions, unused retry variable, and teardown PID revalidation in bin/fm-live-lab.sh. Added behavioral tests in tests/fm-live-lab.test.sh. Both requested commands pass: tests/fm-live-lab.test.sh and bin/fm-lint.sh

* no-mistakes(ci): Fixed teardown to track pre-kill lab processes by PID and start time, including children orphaned when a root exits. Up now rejects an empty pane PID before calling ps. Added regression tests and a Linux-safe worker fixture. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass

* no-mistakes(ci): Fixed teardown tracking for children spawned during shutdown and made the worker fixture verify its exact window with a Linux-available shell. Both requested checks pass. The lab test takes about 66 seconds locally, so the under-one-minute target remains unmet

* no-mistakes(ci): Fixed ci-2 and ci-4 in bin/fm-live-lab.sh and tests/fm-live-lab.test.sh. Teardown now tracks identity-checked members of captured lab process groups, including children orphaned during shutdown, without signaling the caller’s group or unrelated processes. Lint passed, and the lab test passed four times

* no-mistakes(ci): Fixed teardown so an observed-empty process group is permanently dropped, preventing a reused group ID from signalling unrelated work. Added a ps-shim regression test. The lab test, lint, and diff checks pass

* no-mistakes(ci): Fixed ci-1 in bin/fm-live-lab.sh and tests/fm-live-lab.test.sh. The TERM-born-child fixture now waits until its handler is installed before calling down. Down sends SIGKILL to identity-valid survivors on every pass from pass 20 onward and includes survivor process details if it must refuse cleanup. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass locally; Linux CI remains to be verified

* no-mistakes(ci): Fixed down’s teardown wait to require two empty identity-checked scans separated by 0.5 seconds, and removed the unused test loop variable without changing the TERM-born-child test. The lab test, lint, and diff check pass locally
…unchenguid#6103)

* fix(bin): keep slow watcher cycles and preempted reply polls from breaking supervision

- fm_pending_reply_tick selects the records it has work for in one awk pass,
  so settled records cost no lock or fork and the walk no longer grows with
  the never-pruned store.
- An attached arm keeps following a live, identity-matched holder whose beacon
  went stale until the lock changes or the shared stall bound
  (fm_watcher_stall_bound), then fails with a typed stalled-holder line so the
  retry replaces the holder.
- The remote-reply adapter reports the job worker's preemption (exit 76) as a
  closed window, so the listener keeps its claim and polls again instead of
  being relaunched every watcher cycle.

* no-mistakes(document): Clarify watcher grace and attached-arm documentation
…geable is UNKNOWN (kunchenguid#6110)

* fix(bin): retry a bounded number of times when GitHub mergeable is UNKNOWN

Fixes kunchenguid#6020

bin/fm-pr-merge.sh refused a GitHub merge whenever the pull request's
mergeable field was not literally MERGEABLE. GitHub reports UNKNOWN for
a short while after a push or a base-branch change while it recomputes
mergeability, so a green, conflict-free pull request was refused as if
it could not be merged.

github_verify_mergeable now returns a distinct status when mergeable is
the only failing condition and reads UNKNOWN. The caller retries up to
5 times, 3 seconds apart (overridable in tests), re-reading and
re-checking every live condition on each attempt. Once the bound is
spent it reports mergeability as still being computed rather than
unmergeable, with the same nonzero exit as before. Every other refusal
(closed, draft, conflicting, red or missing checks, away authority,
queue protection) is unchanged and never retried.

* no-mistakes(ci): I fixed both review findings the way you asked. The full suite (`bash tests/fm-pr-merge.test.sh`) ran to completion. Its last lines showed all `ok`, and any failure would have stopped the run early. I watched the output through `tail`, so I didn't see the new test's own `ok` line directly. **ci-2 (`bin/fm-pr-merge.sh`), retry delay not validated.** What must hold: the retry wait is always a short, valid `sleep` argument, so a bad `FM_PR_GITHUB_MERGEABLE_RETRY_DELAY` can never trip `set -e` or hold the task lock for a long time. The retry loop is the only place that reads this variable. The script now reads the value once before the loop and accepts only whole numbers from 0 to 10. Anything else (empty, `abc`, `-1`, `1.5`, `11`, a huge number, leading spaces) falls back to 3. I ran those values through the check by hand and each came out as expected. The loop now sleeps on that checked value. **ci-1 (`tests/fm-pr-merge.test.sh`), no test for a check changing between UNKNOWN reads.** What must hold: every retry re-checks all live conditions, not just mergeable. The fake `gh pr view` in the test can now take an optional second word on each line of the mergeable sequence, which sets the first check's result. The new test `test_github_mergeable_unknown_retry_rechecks_checks` feeds `UNKNOWN`, then `UNKNOWN FAILURE`. It asserts: - exit code 1 after exactly 2 reads, - the refusal names `check 'ci' is not green`, - the message does not say mergeability is still being computed, - `pr merge` was never called. If a later change made the retry look only at mergeable, the loop would read UNKNOWN 5 times, end with the "still being computed" message, and this test would fail. I didn't run it against a deliberately broken script to confirm that. `bash -n` passes. `shellcheck` reports only the existing info-level notes about files it can't follow. Only `bin/fm-pr-merge.sh` and `tests/fm-pr-merge.test.sh` changed
kunchenguid#6112)

* fix(bin): converge every open owner onto a known terminal contribution

settle_final only cleared a stale error on retry, so an owner whose saved
row still said open kept projecting a merged or closed pull request as
open after another owner's row had already recorded the terminal
observation. Copy the known terminal observation to every owner whose
saved row is not itself terminal, keeping that owner's own pending and
notified state, and clear its error.

* no-mistakes(review): Carry terminal checked_at when converging existing owner rows

* no-mistakes(ci): I fixed Greptile finding ci-2 as you asked, with a change to tests/fm-contributions.test.sh only. The rule it enforces: when a retry converges an owner onto a URL that is already merged or closed, that owner gets the terminal owner's whole observation, not just its state. The same weak check appeared twice in test_interrupted_multi_owner_poll_settles_every_owner, so I fixed both: - **Open owner (line 784):** the check now also requires `.observation == $terminal[0].records[0].observation`. The existing checks for error, checked_at, pending and notified are unchanged. - **Errored owner (just below):** it only checked state and error before. It now reads the terminal owner's file and makes the same full-observation comparison. Adding the comparison alone would not have caught anything. The test fixtures gave both owners identical observations apart from `state`, so copying only the state would still have passed. In both cases I also set the terminal owner's observation head to HEAD_B, so the two observations now really differ. Verification: - The focused test passes against the current bin/fm-contributions.sh. - I temporarily changed `settle_final` so it copied only the state. The test then failed, reporting the owner still on the old head (HEAD_A). I restored the file afterwards, and `git status` shows only the test file modified. - The full tests/fm-contributions.test.sh suite exits 0. No product code changed. The other CI finding (ci-1, "Behavior portable serial 9") was left alone because you chose to ignore it
…nguid#6124)

* feat: run the supervision host by default on a Claude primary

An absent config/supervision-host on a Claude primary now reads as on with
the default engine, and a file holding `off` opts any home out. Cursor,
OpenCode, omp, Grok, and Codex stay file-gated, with `off` read as disabled
there too. Every reader asks fm_supervision_host_enabled instead of testing
the file, and non-bash readers query it through the lib's `enabled` entry.
A primary's `off` is not inherited by secondmates: each home keeps its own
supervision posture.

* test: pin the watcher-path posture in fixtures that assume no supervision host

Fixtures that drive the watcher arm or assert a non-host drain now write
an explicit off file, and fixtures that copy the Stop auto-arm or the
supervision instructions carry the engine lib they now source. The two
drain suites also stop reading the code root's config.

* fix: name the opt-out when an off home passes an attended wake to main

A host parked when the home writes off now logs that the home does not run
the supervision host, rather than claiming it has no engine.

* no-mistakes(document): Clarify Claude supervision defaults and historical evidence

* no-mistakes(ci): Fixed process leaks in the two added host tests. Each case now stops its recorded watcher and host/arm processes; fake hook sessions exit through session.stop. The full host suite passed before the final cleanup refinement, and both affected cases, bash syntax, ShellCheck, and diff checks passed afterward. CI runtime still needs confirmation
…start scope check (kunchenguid#6125)

* fix(bin): create the state dir on a fresh primary before the session-start scope check

fm_primary_scope_matches required an already-existing state directory, so
bin/fm-sessionstart-run.sh stood down on a fresh clone before anything could
create it. Split out fm_primary_root_matches so the run wrapper can confirm
primary-home identity first, create the gitignored state dir when it is
missing, and only then run the unchanged scope check.

* no-mistakes(document): Document session-start state dir creation on fresh clones

* no-mistakes(ci): I fixed the Greptile P1 the way you asked. When a fresh primary can't create `state/`, the run wrapper no longer stands down silently. **Invariant:** when an otherwise eligible fresh primary cannot create `state/`, startup must never fail silently. This path has only one site: the mkdir in `bin/fm-sessionstart-run.sh`. Other hooks and the nudge wrapper never create `state/`, so they have no equivalent failure. **What changed:** - **Run wrapper** (`bin/fm-sessionstart-run.sh`): it captures mkdir's error and prints one line to stderr before standing down as before (exit 0, or 3 for the Pi prerequisite). The line looks like `fm-sessionstart-run: startup could not create the state directory <path>: <reason>`. - **Test** (`tests/fm-sessionstart-nudge.test.sh`): the new case `test_run_reports_a_state_dir_it_cannot_create` uses a fresh primary with no `state/` and a read-only (0500) root. It checks four things: exit 0, no digest on stdout, no state dir created, and exactly one stderr line ending in "Permission denied". It fails without the fix and passes with it. - **Docs** (`docs/sessionstart-nudge.md`): I added one sentence describing the stderr line and one describing what the new test proves. **Verification:** I ran `tests/fm-sessionstart-nudge.test.sh`, and every test passes. `bin/fm-lint.sh` on the changed scripts (pinned ShellCheck 0.11.0) and `tests/fm-documentation-audiences.test.sh` also pass. As you asked, the wrapper still stands down with the ineligible-checkout status afterwards. It does not report this as a failed eligible startup, which is what the bot suggested
…ery (kunchenguid#6126)

* fix(bin): measure pending-reply grace from turn completion, not delivery

Fixes kunchenguid#6057

The pending-reply guard demanded a repost ("REPOST REQUIRED: previous
marked request had no correlated parent report") while the second
mate's correlated reply was already on its way.
fm_pending_reply_send_recovery measured its grace window from delivery
instead of from the request turn's completion, so any turn longer than
the grace fired the demand the moment the turn ended, before the reply
could have landed. The missed-report escalation had the same gap: it
fired the instant the recovery turn's completion was observed, with no
grace at all.

Both now measure grace from the relevant turn's completion (request
turn for the recovery repost, recovery turn for the escalation), and
both take one fresh, uncached read of the parent status file
immediately before firing, accepting a correlated line regardless of
its verb. Transport-failure escalations stay immediate, and the
one-repost limit is unchanged.

* no-mistakes(review): Document grace window as measured from turn completion

* no-mistakes(ci): Both Greptile findings were real and caused by this PR, so I fixed them. The full `tests/fm-pending-reply.test.sh` suite passes. **ci-1 (a reply could be overwritten by a repost).** The rule that must hold: a recovery send is recorded only if the record is still unresolved, checked under the same per-correlation lock that resolution uses. The escalation path already did this (`_fm_pending_reply_maybe_escalate_locked` reads fresh and publishes under one lock). The recovery path did not: `fm_pending_reply_send_recovery` did its fresh read through `fm_pending_reply_try_resolve`, which let go of the lock before the send was recorded. A reply landing in that gap could be overwritten, and the repost would go out anyway. Now `send_recovery` takes the lock once and, while holding it, re-checks that the phase is still `awaiting_report`, runs the fresh uncached read, and records the send (sender pid and identity, attempt time, phase `recovery_sending`). It releases the lock before actually sending, so the lock is not held during the send. It uses the same lock helpers the other lock wrappers use. Grace timing, the one-repost limit and the escalation path are unchanged. **ci-2 (the test would pass even without the fix).** In `test_recovery_fresh_status_read_resolves_before_firing`, the reply is still appended to the status file, but the stored file signature is then set to the file's new signature. That stands in for a same-size rewrite that the signature cache cannot see. The test first checks that a normal cached read misses the reply, then that the fresh read before sending catches it. I also added the same check for the fresh read before escalation, which the review said was uncovered. The test now sets its own send hook, so it no longer depends on one left over from an earlier test (that leftover had made failures exit silently). **Checks:** - I removed the fresh-read bypass at each site in turn and reran the suite. With it gone from recovery, the test fails with "recovery must not fire once a correlated reply has landed". With it gone from escalation, it fails with "the fresh pre-escalation read should have resolved the record, got escalated". With both in place, all tests pass. - Shellcheck with `-x` timed out locally. Without `-x` and ignoring SC1091, the only warnings are SC2034 on the existing `maybe_escalate` lock wrapper, which is not part of this change. The new code adds no warnings. Changes are in `bin/fm-pending-reply-lib.sh` and `tests/fm-pending-reply.test.sh`. Nothing is committed yet; a plain commit message such as "fix(bin): record the pending-reply recovery send under the fresh-read lock" fits the instruction

* no-mistakes(ci): ci-1 was real and caused by this PR. The same bug was also in the escalation path, so both are fixed. The full tests/fm-pending-reply.test.sh suite passes. The rule that must hold: a recovery repost or an escalation goes out only if the record's phase, read after the fresh-read resolve, is still what it was before. The resolver writes phase=resolved first and only then writes the other resolution fields. If one of those later writes fails, it returns an error even though the record is already resolved. Places this rule applies, both fixed: - Recovery (fm_pending_reply_send_recovery): the fresh-read resolve now runs first, and the phase is re-read right after it, whatever it returned. The send is recorded and made only if the phase is still exactly awaiting_report. This replaces the earlier phase check rather than adding a second one. - Escalation (_fm_pending_reply_maybe_escalate_locked): same bug. After a failed resolve it went on to publish the blocked line and set phase=escalated. One added line after the resolve call returns 1 without publishing if the phase has changed. Test: added test_partial_resolve_write_blocks_firing. It forces a failure on the resolved_epoch write after a correlated reply has landed. It checks that the recovery send hook is never called, that no escalation line is published, and that the phase stays resolved. The forced failure runs in a subshell so it can't affect later tests. Checks: - With the recovery fix reverted, the new test fails with "recovery must not fire after a partial resolve". - With the escalation fix reverted, it fails with "partial resolve should block escalation, got escalated". - With both fixes in, every test passes. - Shellcheck was run with SC1091 excluded and without -x, not through the repo's lint script. The only new message is one SC2329 info on the test's override function; other test overrides in the same file already get that same info, unsuppressed. Changed files: bin/fm-pending-reply-lib.sh and tests/fm-pending-reply.test.sh. Nothing is committed. Suggested plain commit message: "fix(bin): recheck pending-reply phase after the fresh read before sending
…to stderr (kunchenguid#6001)

* fix: provider-table lookup never writes a broken-pipe error to stderr

Fixes kunchenguid#5956

fm_quota_single_provider_for_harness returned from its while read loop
as soon as it found a match, closing the pipe while
fm_quota_single_provider_table's printf could still be writing.
Where SIGPIPE is ignored, as on GitHub Actions runners, bash then
prints "printf: write error: Broken pipe" on the resolver's stderr,
which intermittently broke the one-diagnostic-line assertions in
tests/fm-dispatch-resolve.test.sh.

Read the whole table before answering, the way
fm_control_harness_supported already does, so the writer always
finishes. Return values and output are unchanged.

Reproduced by running tests/fm-dispatch-resolve.test.sh with SIGPIPE
ignored on a single pinned core under CPU contention: 30 of 30 runs
failed before the fix, 0 of 30 after. Note: reproducing requires
setting the trap inside the tested shell because nice(1) resets an
inherited SIGPIPE ignore to SIG_DFL. tests/fm-quota-choose.test.sh
passes and bin/fm-lint.sh is clean.

* no-mistakes(ci): Fixed both Greptile findings the user chose to address. ci-1 (bin/fm-quota-axi-lib.sh:154). Invariant: looking up a harness must always end with status 0 and print the provider, even when the caller runs under `set -e`. The loop body `[ -z "$found" ] && [ "$harness" = "$1" ] && found=$provider` now ends in `|| :`. Every iteration succeeds and the whole table is still read. Only `fm_quota_single_provider_for_harness` loops over the table this way, so this is the one place the fix was needed. One caveat: on bash 5.3 the old code did not actually exit under `set -e`, because the `while` loop is not the function's last command, so the new `set -e` test would have passed before this fix too. The change makes the loop's success explicit, as the user asked. ci-2 (regression coverage). I added three cases to the existing `tests/fm-quota-choose.test.sh`, all calling the public lookup function after sourcing the library: 1. With SIGPIPE ignored (`trap "" PIPE`), it looks up every harness 200 times and checks that nothing reaches stderr. 2. A deterministic version of the race: the table function is wrapped so it writes the first row, pauses 0.2 s, then writes the rest. With SIGPIPE ignored, it checks that looking up `claude` prints `claude` and writes nothing to stderr. The stress loop alone reproduced the bug in only about 1 of 5 local runs, which is why this case exists. 3. A direct call under `set -e` prints `claude`. Verification: - `bash tests/fm-quota-choose.test.sh`: all pass. - Same test against the pre-PR library (fa48367, via `FM_ROOT_OVERRIDE`): fails with `printf: write error: Broken pipe`. The deterministic case failed in one run and the stress loop caught it in another. - `shellcheck` on both files: clean. - `tests/fm-dispatch-resolve.test.sh`: passes
…isioning (kunchenguid#6162)

* fix: survive Pi 0.99 rendering and Git 2.55 local-clone races

Pi 0.99 puts arguments on the stock tool header and leaves hidden custom messages in the export conversation column. Match that header, and keep Calm's boundary on the visible column. Clone a remote home with --no-local so a prune during Git's loose-object copy cannot fail the seed.

* no-mistakes(review): Stop SIGPIPE write errors; cover older Pi export and project clones

* no-mistakes(document): Clarify Calm export visibility and tool rendering

* no-mistakes(ci): Fixed the dispatch diagnostic to list every provider-less use/default profile in one line and added a multi-profile behavior test. Shortened supervision fixtures using the existing engine-grace and park-clock knobs; removed stray scratch files. Dispatch tests, syntax checks, and three targeted supervision cases passed. CI’s prior supervision duration was 751s; the single permitted local full-suite run timed out at 1200s, so an after-duration is not established. The cancelled serial check had no failure verdict. The outer executor should record the measured before/after duration in the PR body when available

* no-mistakes(review): Gate Pi 0.99 call headers by version; drop hidden-row assertion

* no-mistakes(review): Test stock call headers under Pi 0.87 and 0.99 stubs

* no-mistakes(test): Fix older-Pi queued-row test and verify park-boundary behavior

* no-mistakes(document): Clarify Pi Calm export and queued-turn documentation

* no-mistakes(ci): Fixed the stock macOS Bash 3.2 parse failure in tests/fm-calm-pi-extension.test.sh; its parse check passes. The watcher CI failure is in unchanged code: the isolated five-minute/66-minute case passes locally, but the CI log omits the drain error needed to establish its cause. No speculative watcher fix was made. The full local watcher suite timed out after 500 seconds
…6169)

* Prevent premature Lavish board handoffs

* Prove Lavish arm lacks reply acknowledgement

* Confirm Lavish replies before arming worker boards

* no-mistakes(review): Post Lavish reply only after locked arm eligibility checks

* no-mistakes(review): Fail Lavish reply closed on unknown version

* no-mistakes(document): Correct Lavish reply documentation and remove stale guidance

* no-mistakes(document): Clarify Lavish reply routing and remove duplicate version guidance
…#6154)

* feat: inherit the supervision-host opt-out from the primary

Move the supervision host's off opt-out out of config/supervision-host into
its own presence flag, config/supervision-host-off, and add that flag to the
primary-authoritative inherited config set. A primary that opts out now opts
every secondmate home out at spawn and convergence, and clearing it converges
them back. config/supervision-host stays the home-local engine choice.

Shape: config/supervision-host mixed two things, a fleet posture (off) and a
per-home engine and model. Only the posture should follow the primary, so it
becomes a separate presence flag that rides the existing inherited-config
mechanism (FM_INHERITABLE_CONFIG in bin/fm-config-inherit-lib.sh) with no new
machinery, while the engine line stays local. The parse stays in its one
owner, fm_supervision_host_enabled. There is no migration or compatibility
handling for a home that still holds off in config/supervision-host.

Primary off, mate on: inherited material is primary-authoritative by design,
so a mate cannot keep the host while the primary is opted out, and a mate's
own opt-out is removed at the next convergence while the primary has none.
Running the host on a mate is the primary's choice for the fleet; no override
mechanism is added.

Live validation (disposable bin/fm-live-lab.sh lab, Claude primary with a
real seeded secondmate, --supervision-host off):
- up: every readiness check ok, including "host: none running, as expected"
  and a live mate session; the spawned mate home held the inherited
  config/supervision-host-off and the gate read primary OFF, mate OFF.
- primary removed its opt-out, then bin/fm-config-push.sh reported
  "supervision-host-off: pushed - mirrored primary absence" and a config
  reread sent; the gate read primary ON, mate ON, and the live mate handled
  the reread.
- primary opted out again and pushed: "supervision-host-off: pushed", mate
  gate OFF.
- down stopped every lab process and left no lab process running.

Out of scope, follow-up: default-on for the other harnesses, away-daemon
retirement, rollout.

* no-mistakes(document): Document inherited supervision-host opt-out ownership

* no-mistakes(ci): Fixed ci-4: with `--supervision-host off --mate`, lab readiness now requires the inherited flag in the mate home and a disabled mate supervision-host gate. The focused behavior test, shellcheck, and diff checks pass. Left ci-1–ci-3 untouched as directed

* no-mistakes(test): Fix mate readiness HOST_OFF initialization in lab up

* no-mistakes(ci): Fixed Lint 2 by making the new test’s fixtures source resolvable to ShellCheck; its off/on readiness test and ShellCheck now pass locally. Behavior portable serial 5 failed in the unchanged remote-reply test at generation 7. That test passes locally, and no PR-caused defect was identified, so no remote-reply code was changed
…nguid#6179)

* fix(tests): cut the fixed sleeps in supervision-host cycles

The serial CI lane keeps brushing its 30-minute cap because
fm-supervision-host.test.sh spends ~903s of the job, and per the
run-36635306527 case profile the top nine cases are all multi-cycle
ones (3-10 park/close/turn cycles each): every close waits out the
host's sleep $POLL in await_close plus a watcher sleep $FM_POLL scan
cycle, and every engine turn waits out the fixed sleep 1 descendant
snapshot. That is ~3s of pure sleep per cycle before any real work.

The host poll now accepts positive decimal seconds through a new
seconds_or validator (FM_SUPERVISION_HOST_POLL), and the engine turn's
snapshot loop takes FM_SUPERVISION_ENGINE_SNAPSHOT_SECONDS, also a
positive decimal defaulting to one second - the smallest seam at each
wait's single owner. The suite drives them at 0.2 alongside the
existing FM_POLL=0.5 and FM_ARM_ATTACH_POLL=0.2 knobs, so the real
poll loops still run. The park-boundary case moves onto the injected
test clock instead of a real 3s wait, per-case cleanup polls the host
pid rather than sleeping a full second, and the proof-by-absence
windows (flood re-escalation, successor re-announce, watcher
persistence, recovery staying off main) shrink from 2-3s to 1s, which
still spans two watcher polls at the test cadence.

Every assertion, process lifecycle, and reaping path is unchanged;
production defaults stay at one second. Isolated case timings on a
contended host, base vs branch: attended-latch 54.3->34.6s,
undelivered-dialog 67.7->59.1s, away-latch 46.5->30.5s, held-cadence
47.9->21.6s, unreadable-mirror 39.2->38.5s, park-limit 18.2->12.3s,
registration-fallback 14.1->10.0s, first-cycle-status 12.6->8.4s,
latch-scope 16.7->16.3s. Full suite: 65/65 pass. fm-lint and
shellcheck clean.

* no-mistakes(review): Wait for scan lock release before duplicate check

* no-mistakes(document): Correct supervision snapshot cadence documentation

* fix(tests): keep production poll cadence, probe exits at 0.1s

The fractional poll cadences multiplied the cost of each loop body:
full process-table scans in the engine turn and process refreshes in
await_close ran five times more often, which swamped the thin CI runner
and nearly doubled every multi-cycle case (serial 5 was cancelled at its
30-minute limit on run 36635306527's successor). Restore the production
cadence and notice arm/engine exits with a cheap kill -0 probe at a
tenth of a second between the one-second bodies instead: strictly less
dead time than baseline with no added CPU.

Also hold each injected-clock park bound well past its case's
wall-clock checks so a host that ignored the test clock fails instead
of silently passing at a real-time boundary, and restore the shortened
proof windows (watcher liveness, recovery-off-main absence, first-cycle
stream) to their baseline depth.

* no-mistakes(document): Clarify supervision engine snapshot documentation
…henguid#6192)

* fix: rebalance portable CI from current duration measurements

* no-mistakes(test): Test serial packing boundary and verify endpoint timeout cleanup

* no-mistakes(document): Clarify timeout guidance and remove duplicated packing estimates
…nguid#6216)

* fix(bin): run no repository hook when core.hooksPath is empty

The per-task hook wrapper refused every commit in a repository whose own
config sets core.hooksPath to the empty string, because git rev-parse
--git-path hooks fails on it. Plain git reads that setting as no hooks, so
the wrapper now runs none; every other lookup failure still refuses and
shows git's error.

Fixes kunchenguid#6171

* no-mistakes(review): Refuse commits when core.hooksPath is a valueless key

* no-mistakes(document): Document empty core.hooksPath handling in commit attribution docs

* no-mistakes(ci): When the wrapper refuses a commit, Git's hook-lookup error now shows up once instead of twice. That required changing one line in the wrapper, and the tests were extended so both bad-config cases would catch the duplicate. Invariant: when the wrapper refuses, Git's lookup error must appear exactly once. In the failure path, the only Git call besides the deliberate second lookup is the `git config --get --type=path core.hooksPath` check in `runtime_chain_body` (`bin/fm-git-strip-ai-trailers.sh:168`). That check prints the same error, so it was the one place to fix. I added `2>/dev/null` to it. Its exit status still decides the outcome: an empty value still runs no hook, and anything else goes on to the second lookup, which prints Git's error once, and the commit is refused. Tests (`tests/fm-git-strip-ai-trailers.test.sh`): - The unresolvable-path test (`~fm-no-such-user-6171/hooks`) now requires `failed to expand user dir` to appear exactly once in the refused commit's output. - The valueless-key test now requires `missing value for 'core.hookspath'` to appear exactly once. - Pre-existing bug in the unresolvable-path test: its `git add` ran after the bad config was set, so it failed silently (exit 128) and the "refused commit" had nothing staged. The test now stages the file before writing the config, the same way the valueless test does, so a real commit gets refused. - The empty-string test is unchanged and still passes, so an empty `core.hooksPath` still runs no hook. Verification: - With the wrapper change reverted, both new checks fail with `expected '1', got '2'`. With the change in place, the whole suite passes. - `bash -n` passes. shellcheck shows only an info-level SC1091 note about sourcing `lib.sh`, which was already there before this change. - `git status` lists only the two intended files
…kunchenguid#6213)

* fix(bin): let a stale record on a reassigned slot retire records-only

When a pool slot's owner claim names another task, the stale record's
teardown touches nothing under the slot, so the exclusive-slot record scan
no longer refuses it. Full teardowns of a slot this task still claims, or
one with no claim, keep the refusal.

Fixes kunchenguid#6184

* no-mistakes(document): Note claim-over-record precedence for reassigned teardown slots
…uid#6240)

* fix(bin): keep the steering doorbell short under deep homes

The doorbell printed the task inbox's absolute path twice, so under a deep
home it grew to about 290 characters and a Herdr submit reported it never
reached the pane on every re-ring. It now names the inbox once by its short
<task>.inbox name and points at the full path the worker's brief already
gives, so its length no longer depends on the home's depth.

Fixes kunchenguid#6120

* no-mistakes(review): Export FM_TASK_INBOX at launch and name it in doorbell

* no-mistakes(ci): ci-1 (Behavior portable serial 9) was caused by this PR, and I fixed it in the test. tests/fm-claude-trust.test.sh failed with "the launch command did not carry a brief doorbell". Its claude_launch_doorbell helper stripped exactly two leading `export ...;` statements before reading the final prompt argument. This PR adds a third one (`export FM_TASK_INBOX=...`) to every launch, so the helper was reading the wrong command. The invariant: a test that parses the launch command must skip every leading export statement, however many there are. I checked every test that parses the launch this way. The only other ones are the two helpers in tests/fm-spawn-dispatch-profile.test.sh, and they already loop over all exports. The kimi and dispatch-profile exact-string checks were updated earlier in this PR. The fix makes claude_launch_doorbell use the same loop (`while [[ "$command" == export\ *\;* ]]; do command=${command#*; }; done`) and then take the last argument. The ordinary path still works: the claude spawn test and the secondmate-clone spawn test both resolve the brief record through the same helper. Verified locally: `bash tests/fm-claude-trust.test.sh` exits 0 with no failing cases. ci-2 (Behavior tests (Herdr)) was not caused by this change, and I made no code change for it. In tests/fm-backend-herdr-presentation-e2e.test.sh, the concurrent secondmate recovery failed with "herdr presentation recovery could not acquire its session lock; refusing a concurrent resume". Two reasons it is not this PR: - The same failure, in the same test and case, happened on run 36655209015 for the unrelated branch fm/fm-contributions-old-gh-compat about 14 hours earlier. - This PR's change cannot lengthen how long the lock is held. The launch is written to a file and sent to the pane as `. launch.N.sh`, so the extra export changes neither the pane submit nor the lock hold time. The cause is a race that was already there: spawn_herdr_presentation_order_lock_acquire gives up after 5 seconds, and a concurrent real-Herdr recovery can hold the lock longer. Fixing that means changing the product's lock timeout, which is outside this PR. It should be tracked separately, and a rerun of the Herdr job is expected to pass. The only file changed is tests/fm-claude-trust.test.sh
… no turns (kunchenguid#4859)

* fix(dod): drive no-mistakes with one foreground call, not a background poll

The brief told workers to background the drive call and poll `axi status`
because one call "routinely outlives what your harness lets a single
command run". That advice contradicts the tool it drives: `no-mistakes
axi run --help` documents `--wait` with an 8m default, existing precisely
"so an agent harness with a 10-minute tool cap gets a structured return
instead of an unbounded hang".

Following the old text, a worker could never idle - a backgrounded call
returns in milliseconds, so it does not wait at all - and each attempt
leaked a live timer that later fired as a paid wake. Tell workers to make
one foreground call, let it block, and repeat it when it returns on
elapsed wait rather than on a gate or outcome.

Also drops the generalisation that told workers on any unestablished
harness to assume a command cap and use the same shape, which exported
the defect to harnesses with no such cap.

* fix(bin): let a waiting worker spend no turns until it is answered

A worker waiting on a decision, a pipeline gate, CI, or a heavy-test slot
kept taking model turns: the brief told it to list its inbox at any natural
checkpoint, and six automatic senders nudged secondmates whatever their open
decisions.

- The ship and scout briefs gain one Waiting section: end the turn after
  needs-decision or blocked, and hold an external wait inside ONE blocking
  command bounded by the harness's own command ceiling. The checkpoint clause
  is deleted. Forbidding the wrong shapes is not enough on its own, so the
  section also names the blocking foreground `until` loop as the wait a Claude
  Code worker may use, because that harness can refuse a sleep-then-check
  command while pointing at backgrounding, which is the one shape a waiting
  worker must not take.
- fm-send --automatic defers (exit 4, nothing written or rung) while the
  target has an open decision or blocker of its own; every automatic sender
  passes it and keeps its retry state, and the pending-reply recovery waits
  the same way.
- The two senders that report the result classified it by matching the text of
  the send's captured output against `deferred:*`. fm-send runs bin/fm-guard.sh
  as a supervision warning, and that guard prints its worktree-tangle banner
  whenever the primary checkout is on a feature branch, which is exactly what a
  CI pull-request checkout is. The banner lands ahead of the `deferred:` line,
  so the match fell through and a waiting mate was reported as a failed send,
  with the banner as the reason. Both senders now classify on fm-send's exit
  status, which is the contract the deferral is actually stated in, and select
  the `deferred:` line out of the output rather than assuming it came first.

The third root cause, a no-mistakes definition of done that backgrounded the
drive call and polled axi status, is fixed by this branch's parent commit
"drive no-mistakes with one foreground call, not a background poll"; this
commit takes that text as is and adds the regression test.

Upstream's spawn abort path no longer calls the lease-return helper at all,
so the fork's missing-helper guard and its pin-feature test line are moot
here and are not ported.

The command ceilings each harness enforces, and the probes behind the named
Claude Code wait, are recorded in docs/verification/runtime-backends.md.

* no-mistakes(review): Exempt captain holds, quiet deferred reconcile, clarify worker pauses

* no-mistakes(document): Document deferred automatic nudges, rereads, and reply recovery

* no-mistakes(document): Ring unlanded fire-and-forget steers exactly once more

* no-mistakes(ci): The failing check, "PR must be raised via no-mistakes", reads the pipeline's attestation record, which says document=skipped. No file in the repository can change that record, so I did not touch the check or the PR body. As you said, the no-mistakes rerun after this run finishes will re-execute the document step and record document=completed. The one change is the documentation sentence you ordered. It adds a line to docs/remote-secondmates.md, right after the line saying the remote host runs no re-ring ladder of its own: "A fire-and-forget record, such as a reconcile ask, gets its single retry ring only on the local plane: the remote steer leg owes no re-ring, so a swallowed remote doorbell for one waits for the next ring into that inbox, and a remote-side retry is known follow-up scope." No behavior changed. Checks: tests/fm-documentation-audiences.test.sh passes (4/4) and bin/fm-lint.sh is clean. The change is left uncommitted in the working tree for the pipeline to pick up

* no-mistakes(review): Hold automatic wakes until a mate's own decision closes

* no-mistakes(document): Document watcher delivery of deferred remote re-read nudges

* no-mistakes(review): Merge duplicate elapsed-wait reattach instructions in DOD

* no-mistakes(test): Resolve merged default decision in remote-reply recovery fixture

* no-mistakes(test): Source classify lib so config-push retry-deferred honors open decisions

* no-mistakes(ci): Fixed a flaky test that also fails on main. Neither this PR's bin/fm-brief.sh nor its bin/fm-dod-lib.sh change is involved: bin/fm-dispatch-resolve.sh sources neither file. Another branch (fm-attended-cutover-smoothing-s1, run 36343879084) failed the same shard 8 check the same way, on a different case ("a rule-criterion match prints one diagnostic line, got 2"). Root cause: `fm_quota_single_provider_for_harness` in bin/fm-quota-axi-lib.sh returned from its `while read` loop as soon as it found a match. That closed the pipe while `fm_quota_single_provider_table`'s `printf` was sometimes still writing. GitHub Actions runners ignore SIGPIPE, so bash printed `fm-quota-axi-lib.sh: line 138: printf: write error: Broken pipe` to the resolver's stderr. That is the extra line. I reproduced it locally by running the test with SIGPIPE ignored: 2 of 20 runs failed, one with the resolver's diagnostic line plus two broken-pipe lines. Invariant: looking up a harness in the provider table must never make the table writer fail. The only reader of that table is this function, and all of the resolver's lookups (line 208 without stderr redirected, line 222 with it) go through it. So the fix is in that one place: read the whole table, then print the match. The same file now shows it reads the full table first, like `fm_control_harness_supported` does. Return values and output are unchanged. Verification: with SIGPIPE ignored, tests/fm-dispatch-resolve.test.sh failed 0 of 30 runs after the fix (2 of 20 before). tests/fm-dispatch-resolve.test.sh, tests/fm-brief.test.sh, tests/fm-send-inbox.test.sh, tests/fm-quota-choose.test.sh and tests/fm-quota-array-dispatch-live-e2e.test.sh all pass, and shellcheck is clean. tests/fm-procevent-quota.test.sh fails locally with or without the change ("process-event state root is not a private directory"), so that failure comes from the local environment, not from this fix. No new test was added: the existing one-diagnostic-line assertions already catch this whenever SIGPIPE is ignored, as it is in CI

* Revert "no-mistakes(ci): Fixed a flaky test that also fails on main. Neither this PR's bin/fm-brief.sh nor its bin/fm-dod-lib.sh change is involved: bin/fm-dispatch-resolve.sh sources neither file. Another branch (fm-attended-cutover-smoothing-s1, run 36343879084) failed the same shard 8 check the same way, on a different case ("a rule-criterion match prints one diagnostic line, got 2"). Root cause: `fm_quota_single_provider_for_harness` in bin/fm-quota-axi-lib.sh returned from its `while read` loop as soon as it found a match. That closed the pipe while `fm_quota_single_provider_table`'s `printf` was sometimes still writing. GitHub Actions runners ignore SIGPIPE, so bash printed `fm-quota-axi-lib.sh: line 138: printf: write error: Broken pipe` to the resolver's stderr. That is the extra line. I reproduced it locally by running the test with SIGPIPE ignored: 2 of 20 runs failed, one with the resolver's diagnostic line plus two broken-pipe lines. Invariant: looking up a harness in the provider table must never make the table writer fail. The only reader of that table is this function, and all of the resolver's lookups (line 208 without stderr redirected, line 222 with it) go through it. So the fix is in that one place: read the whole table, then print the match. The same file now shows it reads the full table first, like `fm_control_harness_supported` does. Return values and output are unchanged. Verification: with SIGPIPE ignored, tests/fm-dispatch-resolve.test.sh failed 0 of 30 runs after the fix (2 of 20 before). tests/fm-dispatch-resolve.test.sh, tests/fm-brief.test.sh, tests/fm-send-inbox.test.sh, tests/fm-quota-choose.test.sh and tests/fm-quota-array-dispatch-live-e2e.test.sh all pass, and shellcheck is clean. tests/fm-procevent-quota.test.sh fails locally with or without the change ("process-event state root is not a private directory"), so that failure comes from the local environment, not from this fix. No new test was added: the existing one-diagnostic-line assertions already catch this whenever SIGPIPE is ignored, as it is in CI"

This reverts commit c719928.

* no-mistakes(review): Retry deferred local instruction nudges via the watcher

* no-mistakes(review): Document watcher retry for deferred local instruction nudges

* no-mistakes(ci): I fixed both review findings you selected (ci-1 and ci-3). I did not touch the deferral check in bin/fm-send.sh. ci-1 (bin/fm-config-push.sh, retry_deferred_rereads) - Rule that must hold: a deferred reread stays flagged until it is actually delivered. - Before the fix, the flag was removed before any of the steps that can skip a mate: the remote lock-path lookup, validate_secondmate_home, the local lock-path lookup, and the lock acquire. A skip at any of those dropped the flag, so the watcher lost track of the reread. - Now the flag is removed in one place only, when the send succeeds (rc 0). A skipped home, a busy lock, a deferred send (rc 4) or a failed send all leave it in place. The re-mark calls on a busy lock and on rc 4 were no longer needed, so I removed them. I updated the comment above the function to match. - Side effect: a send that keeps failing now stays flagged, so the watcher retries it on every poll and logs each failure. That follows your "don't clear until delivered" rule, but it replaces the old behaviour of leaving a failed send to the next config push or session start. - New test in tests/fm-secondmate-sync.test.sh: T8j "a deferred flag survives a skipped invalid home and is retried once it validates". It takes the home's marker away to make validation fail, checks that nothing is sent and the flag stays, then puts the marker back and checks that the nudge is delivered and both the flag and the retry marker are cleared. It fails on the old code and passes now. ci-3 (bin/fm-secondmate-restart.sh) - Rule that must hold: no automatic send wakes a mate that is waiting on its own open decision. - The two automatic sends in this script are the fallback reread nudge (fall_back_to_nudge) and the persist request. Both now pass --automatic. If a persist request is deferred, its correlation is discarded and the mate goes to the fallback nudge, which is also deferred, so the mate is reported as unreached. - New test in tests/fm-secondmate-restart.test.sh: T3b. It gives a mate an open needs-decision and runs a restart. It checks that both sends report as deferred, the mate's doorbell is never rung, its inbox gets no message, nothing is stopped, and the mate is reported as unreached with exit status 3. It fails on the old code and passes now. - The test marks the watcher as alive first. Without that, the watcher-down warning is printed first and becomes the reported reason instead of the deferral message. Verification - tests/fm-secondmate-sync.test.sh passes. - tests/fm-secondmate-restart.test.sh passes. - tests/fm-secondmate-harness.test.sh (the other test that exercises --retry-deferred) passes. - The fm-send-inbox test that covers automatic deferral passes. I only looked at the last lines of that run, not the whole file. - `shellcheck -x` on the four changed files is clean

* Pin autoarm supervision model in secondmate restart T3b

The fresh watcher beat the test writes proves a live watcher only under the
autoarm model; on CI hosts with no detected harness the persistent model
demands a lock-holding watcher, so the watcher-down banner became the
reported reason and the deferral assertion failed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Keep deferred secondmate nudges retryable under the inheritance lock.

A bootstrap instruction nudge could write its deferral flag outside the lock the watcher retry holds, so a concurrent retry could delete a flag that had just been set. A restart fallback that is deferred now records the same marker and flag, so the watcher delivers it once the decision closes.

* no-mistakes(document): Document watcher retry of deferred restart re-read nudges

* Send secondmate reread and restart nudges immediately again.

Deferring those nudges let a later config push drop an incomplete transfer once the decision closed. They now send as they do on main.

* Make the no-turn wait opt-in behind config/wait-no-turns.

Homes that do not create the file keep the previous briefs, drive text, and sends.

* no-mistakes(document): Document wait-no-turns inbox wording change in configuration

* no-mistakes(review): Keep checkpoint inbox check; forbid only polling while waiting

* no-mistakes(ci): Fixed ci-2 (Greptile: a concurrent retry marker gets lost). The rule that was broken: the watcher may remove only the `.retry-ring` mark for the record it just processed. A newer mark written in the meantime is owed its own retry. `fm_task_inbox_clear_retry` is the one shared function that removes the mark, and I fixed it there. In `bin/fm-task-inbox-lib.sh` it now takes the record path. It compares the mark's content with that record's name and removes the mark only when they match. When the mark names a different record it returns success and leaves the mark alone. It still fails only when the processed record's own mark can't be removed. Both callers in `bin/fm-watch.sh` now pass `"$rec"`: the dead or missing pane path and the path after a retry ring. So the fix holds at both removal sites. Tests, in `tests/fm-task-inbox.test.sh`: - I added an optional `FM_RING_MARKS_RETRY` hook to the fake tmux. It writes a newer record's mark while the doorbell is being typed, which reproduces the race deterministically. - I added `test_watcher_retry_keeps_a_newer_mark`. The owed retry rings once, the newer mark survives, and a later check rings the newer record once and then clears its mark. The test fails without the fix ("the spent retry removed a newer record's mark written during its ring") and passes with it. - I updated the direct `clear_retry` call in the existing unit test to pass the record. Results: `tests/fm-task-inbox.test.sh` passes in full and `tests/fm-send-inbox.test.sh` passes 15/15. Shellcheck reports only SC1091 "not following sourced file" notices. As instructed, I didn't change the brief inbox wording

* no-mistakes(document): Fix stale wait-no-turns inbox wording in inbox lib comment

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
…note (kunchenguid#6140)

* fix(bin): record Gerrit change URLs as close notes

Teardown's backlog_done_args hands every ship's recorded pr= URL to
fm_backlog_done as --pr, and tasks-axi refuses any --pr that is not a
canonical GitHub or Forgejo pull request. A Gerrit change URL therefore
left the item In flight after cleanup, and the pending backlog-close
record replayed into the same refusal at every session start.

fm_backlog_done now rewrites a --pr whose value fm_pr_url_parse reads
as a Gerrit change into --note "Gerrit change <url>". The mapping sits
at the tasks-axi call rather than in the pending-close record, so
records already written with --pr replay to a close unchanged. The
captain-held retain path records the URL in its deliverable line and
skips the update --pr it cannot make.

* no-mistakes(review): Note retained Gerrit change URL when captain answers early

* no-mistakes(document): Document Gerrit change URL handling in captain-hold retention
* perf(remote): separate active job sampling from dispatcher cadence

* no-mistakes(document): Link remote wait timing to its authoritative contract

* no-mistakes(ci): Fixed ci-1 with two narrowly scoped SC2030 annotations documenting intentional subshell-local legacy and active cadence overrides in tests/fm-remote-job.test.sh. Runtime behavior is unchanged. Reproduced the lint failure before the fix; afterward ShellCheck 0.11.0 with source following, Bash syntax validation, the complete remote-job behavior suite, and git diff --check all passed

* perf(supervision): reduce park, delta and dispatcher polling

* no-mistakes(document): Clarify poll latency contracts and authoritative documentation pointers
…henguid#6221)

* fix(bin): load backend sibling libraries under zsh

fm_backend_source kept each backend's sibling list in one space-separated
string and iterated it unquoted. zsh does not word-split an unquoted
expansion, so the readability check saw the whole list as one path and
refused every backend with more than one sibling. Hold the list in the
function's positional parameters instead, which needs no word splitting
in Bash 3.2, Bash 5, or zsh.

The existing zsh case in tests/fm-backend.test.sh covers it wherever zsh
is installed.

* test: run the Calm mod suite on stock Bash 3.2

The suite injected shell values into its generated Node scripts with the
${value@Q} transformation, which needs Bash 4.4. Stock macOS Bash 3.2
reports a bad substitution, so every case failed before it asserted
anything. Build each JavaScript string literal with JSON.stringify
through a small helper instead, which works on any Bash and is a valid
literal for any value.

* no-mistakes(review): fix(bin): rename zsh-special path local in fm_backend_source

* test: narrow the zsh backend claim to name matching

Under zsh the adapters locate their siblings through BASH_SOURCE, so a
successful fm_backend_source is not a full load. Assert only what the
contract states, and pass js_string values after -- so node never reads
a leading-dash value as its own option.

---------

Co-authored-by: Nova Agent B <novaagentb@gmail.com>
…tatus scans (kunchenguid#5263)

* fix(bin): exclude a remote mate's own parent channel from self-home scans

A remote secondmate home's outbound parent channel lives at state/parent-replies.status inside its own state dir, so the watcher's signal scan enumerated it as a task status file and the open-decisions fold classified it as a phantom task named parent-replies: every parent-channel append spun a spurious signal wake and a phantom open decision in the mate's own home.
fm-parent-channel-lib.sh gains fm_parent_channel_outbound_status, which resolves the channel into the mate's own state dir for the remote route only, and fm-classify-lib.sh's status_scan_parent_channel_exclude wraps it for the fleet-wide scans.
The watcher's scan_signals and heartbeat fail-safe backstop, the whole-file and incremental open-decisions folds, the presentation snapshot, and the unread-surface scan now skip exactly that resolved path.
The exclusion is home-shape-aware: a parent-replies.status in a main home or a local mate is an ordinary task log and keeps waking and folding, and every other status file is untouched.

* no-mistakes(review): exclude a remote mate's parent channel from the daemon heartbeat scan

* no-mistakes(document): Document remote mate parent-channel scan exclusion

* ci: retrigger portable serial 4

* no-mistakes(ci): CI check 'Behavior portable serial 7' failed in tests/fm-contributions.test.sh ('reservation poll failed'). CI stderr showed bin/fm-contributions.sh:345 arithmetic 'DEADLINE - 6\n90077104: syntax error in expression': the fixture's fake date returned a torn two-line clock value. Root cause: the fake forge wrapper in wrap_forge advances the shared controllable clock via a non-atomic read-modify-write ('$(cat $FORGE/clock) + 6' with truncate-in-place '> $FORGE/clock') while concurrent background gh calls run and the fake date reads the same file; an interleaved truncate+write publishes a half-written value (CI's torn '6\n90077104', tail of 1790077104) or an emptied-read value ('6'), which either breaks the poll's arithmetic (nonzero exit -> 'reservation poll failed') or defeats the 15-second reservation defer. This is a pre-existing test-fixture race, not caused by the PR's diff (base..target touches no contributions code; the same commit passed this shard in run 35711207830 earlier the same day). Fixed the flaky fixture at its root: clock_bump() now writes each new value to a per-process mktemp file in the same directory and publishes it with mv (atomic rename), so concurrent forge callers and the fake date always read one complete old-or-new clock; fault patterns and deltas are unchanged. Verified: minimal 3-way concurrency repro shows the old wrapper corrupting (12/32/38 outcomes incl. empty-read) while the rename-based wrapper never corrupts (20/20 clean); the full tests/fm-contributions.test.sh passes twice (all 38 assertions ok, incl. the reservation, budget-exhaustion, genuine-failure, shared-once, and latency tests); 10 isolated reservation runs pass; shellcheck rc=0; worktree contains only this one-file change

* no-mistakes(document): drop stale file-set copy in daemon catch-all comment
…6307)

* fix(bin): name the accepted verdict actors in fm-contributions help and refusal

* fix(ci): Updated tests/fm-contributions.test.sh to assert exactly captain, fleet, maintainer, and nobody in command-emitted help and refusal output. Three focused regressions passed; all three extra-actor mutations were rejected. ShellCheck, syntax, and diff checks passed. Production code remains unchanged
…chenguid#6306)

* fix(bin): recognise a clone root git names with different path spelling

fm-fleet-sync compared git's --show-toplevel with pwd -P as strings, so a clone
root that git recorded with different casing (case-insensitive volume) was
skipped as not a clone root and never refreshed. Compare filesystem identity
instead, which also covers symlink spelling.

* fix(document): Remove stale clone-root comparison comment
…rror expectations, Herdr startup-event simulation, and secondmate Git initialization. All three affected test scripts and targeted lint pass locally. Live verification used Pi 0.99.1; Pi-signed was unavailable
@dscott98
dscott98 merged commit 1f523ef into main Oct 1, 2026
19 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.