Skip to content

feat(bin): defer spawns beyond a declared per-project capacity, opt-in - #5343

Merged
kunchenguid merged 20 commits into
kunchenguid:mainfrom
tiago-peixoto:fm/upstream-4237-project-capacity-admission
Oct 7, 2026
Merged

kunchenguid merged 20 commits into
kunchenguid:mainfrom
tiago-peixoto:fm/upstream-4237-project-capacity-admission

Conversation

@tiago-peixoto

@tiago-peixoto tiago-peixoto commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Intent

Fixes #4237.
Keep #5343 (issue #4237) mergeable: it turned CONFLICTING against upstream main on 2026-09-24 after being green. Accepted behavior to preserve: project-capacity admission as accepted in the existing pull request and its maintainer triage (opt-in class).

The accepted admission stays opt-in. Without config/project-capacity, dispatch stays uncapped. With a declaration, a spawn beyond the declared number prints one deferred line, exits 75, creates nothing, and leaves the backlog item queued. Relaunches and secondmates are not counted. Holders are counted across the root home and local secondmate homes, including other clones of the same origin, until a ready pull request is recorded or the task is cleaned up. A malformed or unreadable declaration refuses the spawn. Matching is by clone directory name. A project name may begin with # when that # is written against the rest of the name. A line that is only #, or that begins with # followed by whitespace, stays a comment.

Maintainer triage treats the change as opt-in once the always-loaded backlog contract re-evaluates queued work after every teardown and heartbeat, and also after a recorded PR-ready handoff only when config/project-capacity caps that project. The conditional dispatch line in section 7 stays as it is.

Added 2026-10-04: "on opensource, some of my PRs need resync with main. make sure that is done, and CI is fully green for maintainer, on all of our work." "make sure the work is moving, nudging maintainer, making sure PRs are green and ready for maintainer"

What Changed

  • bin/fm-spawn.sh now reads an optional config/project-capacity from the root Firstmate home through the new bin/fm-project-capacity-lib.sh. With no declaration, dispatch stays uncapped. When every declared place for the project is held, a fresh ship or scout spawn prints one deferred: line and exits 75 before it creates a brief, endpoint, worktree, record, or backlog move. A malformed or unreadable declaration, or a task record that cannot be read while counting, refuses the spawn with exit 1. A batch reports such a pair as batch: DEFERRED. Relaunches and --secondmate spawns are not counted.
  • Holders are counted under the shared project lock, which the spawn now also takes on the Orca backend whenever the declaration caps any project. The count covers the root home and every local secondmate home, including other clones of the same origin, until a ready PR is recorded or the task is cleaned up. The walk over local homes moved from bin/fm-teardown.sh into fm_local_firstmate_state_dirs in bin/fm-wake-lib.sh, so teardown and capacity admission share it.
  • AGENTS.md, docs/configuration.md ("Project capacity"), and the operational-home-layout skill document the file format, the # comment rules, and matching by clone directory name. The backlog contract now also re-evaluates queued work after a recorded PR-ready handoff when config/project-capacity caps that project. The branch adds tests/fm-project-capacity.test.sh, registers it in bin/fm-test-run.sh, and freezes the clock in one tests/fm-contributions.test.sh case so its one-second budget cannot expire before the first forge call.

Risk Assessment

✅ Low: The admission is opt-in, runs before any brief, worktree, record or backlog change, releases its lock through the existing exit trap and publication release, and matches every stated intent criterion; the traces I ran (declaration parsing, own-record skip, symlinked-home dedupe, batch exit codes, Orca lock path, teardown refactor) found no wrong result.

Testing

No baseline command ran before this step. I ran the targeted behavior test tests/fm-project-capacity.test.sh (16 cases, all passed, real fm-spawn with fake tmux and treehouse). I then built a disposable lab home with bin/fm-lab-home.sh, a throwaway project with a local bare origin, a second clone named '#repo', a registered local secondmate home, and a markdown backlog, and drove ten script-level scenarios through a real tmux pane on the private fm-lab socket. Worker launches used the documented raw launch command 'sleep 3600', so no paid worker agent ran; tmux windows, treehouse worktrees, task records and backlog moves were all real. Finally I started a real Claude primary (normal login) in the lab and drove three turns covering deferral, re-dispatch after teardown, and re-dispatch after a recorded PR-ready handoff. Two scenarios were not driven live and are reported as untested: a secondmate-kind record not holding a place, and two concurrent spawns racing for the last place. The automated behavior test covers both (cases only this project's ship and scout records hold places and concurrent spawns can never both take the last place in capacity-behavior-test.log), but that is not a live result. This is a CLI change with no rendered UI, so the evidence is CLI transcripts and a pane capture rather than screenshots. One fixture problem (the bare origin's default branch) was mine and fixed before the recorded run. Overall result: all thirteen live-driven scenarios passed, and two scenarios are untested live.

  • Live validation: ✅ go - 13 of 15 scenarios driven live against the product
Scenario Result Live Evidence
With no config/project-capacity, the captain spawns two workers on one project and both launch ✅ pass live live-lab-transcript.txt section S1: cap-a and cap-b both print spawned ... with exit 0; two task records, two tmux windows, two treehouse worktrees, both items under In flight
With heavy 2 declared and two holders, a third spawn prints one deferred line, exits 75, creates nothing, and its backlog item stays queued ✅ pass live live-lab-transcript.txt section S2: exit 75, exactly one line starting deferred: naming holders cap-a, cap-b; data/cap-c holds only brief.md; records, windows and worktrees unchanged; cap-c still un…
Adversarial: a malformed or unreadable declaration refuses the spawn instead of guessing a limit ✅ pass live live-lab-transcript.txt section S3: non-integer capacity, zero, duplicate project, nameless line, mode 000 file and a directory each give `error: spawn refused: the project capacity declaration is unr…
A line that is only #, or # followed by whitespace, stays a comment and does not change the cap ✅ pass live live-lab-transcript.txt section S4: with #, # heavy 9, an indented # ... line, #note-with-no-number and a blank line above heavy 2, the spawn is deferred at 2, not 9, and not refused
A project whose clone directory is named #repo is capped by the line #repo 2, matched by clone directory name ✅ pass live live-lab-transcript.txt section S5: spawn from <lab>/projects/#repo prints deferred: project #repo admits 2 worker(s) with exit 75; after changing the line to #repo 3 the same spawn is admitted
Holders in another clone of the same origin count against the cap ✅ pass live live-lab-transcript.txt sections S5 and S6: the '#repo' spawn is deferred by holders cap-a and cap-b running in the 'heavy' clone; with heavy 3, cap-c is deferred with holders cap-a, cap-b, cap-d…
A relaunch of a worker that already holds a place is not counted while the project is full ✅ pass live live-lab-transcript.txt section S7: with 3 of 3 places held, bin/fm-spawn.sh cap-a --relaunch exits 0 and the cap-a pane runs the agent command again
Cleaning up a task frees exactly one place ✅ pass live live-lab-transcript.txt section S8: after bin/fm-teardown.sh cap-d, cap-c is admitted and the next spawn cap-e is deferred with holders cap-a, cap-b, cap-c
Recording a ready pull request frees that worker's place ✅ pass live live-lab-transcript.txt section S9: the real bin/fm-pr-check.sh cap-b &lt;url&gt; writes pr= to cap-b's record, and the next spawn cap-e is admitted although cap-b's record and window still exist
The cap is shared between the root home and a local secondmate home, in both directions ✅ pass live live-lab-transcript.txt section S10: a spawn from the secondmate home, whose own config/ is empty, is deferred on the root's declaration with holders `cap-a in <root>, cap-c in <root>, cap-e in <root>…
A secondmate-kind record does not hold a place ⏸️ untested no The prior payload did not establish a live result for this scenario. It was not driven against the running product: the run did not provision and launch a real persistent secondmate in the lab home, w…
Two concurrent spawns cannot both take the last place ⏸️ untested no The prior payload did not establish a live result for this scenario. It was not driven against the running product: the race needs one spawn held inside the admission lock at a precise point, and real…
A real primary that gets a deferral leaves the item queued, not blocked, and says it will retry when a place frees ✅ pass live live-primary-transcript.txt first turn: the Claude primary runs bin/fm-spawn.sh, sees exit 75, reports the item still under Queued, not blocked, and names cleanup or a ready PR as the retry trigger
A real primary re-evaluates queued work after a teardown and dispatches the deferred item ✅ pass live live-primary-transcript.txt second turn: asked only to clean up cap-c and carry on, the primary runs teardown, cites its standing instruction to re-evaluate queued work after every cleanup, and spawns…
A real primary re-evaluates queued work after a recorded PR-ready handoff on a capped project ✅ pass live live-primary-transcript.txt third turn: after bin/fm-pr-check.sh cap-a the primary states Heavy has a declared capacity, so my standing instructions call for re-evaluating the queue now and spawns c…
Evidence: Live lab transcript: real fm-spawn, fm-teardown and fm-pr-check in a disposable lab home (scenarios S1-S10)

Source: Live lab transcript: real fm-spawn, fm-teardown and fm-pr-check in a disposable lab home (scenarios S1-S10)

## Lab: FM_HOME=/tmp/fm-lab.1AKLOV, real bin/fm-spawn.sh typed into a real tmux pane on the private fm-lab socket, real treehouse pool, real tasks-axi backlog.
## Worker launch command is the documented raw-launch escape hatch ('sleep 3600') so no paid agent runs.

\### S1 - no config/project-capacity: dispatch is uncapped
(config/ listed above; no project-capacity file)
(cap-a was already spawned the same way; its output is in live-s1a below)
fm-gate-refuse: gate agent lifecycle permitted only against lab home /tmp/fm-lab.1AKLOV
warning: /tmp/fm-lab.1AKLOV/data/cap-a/launch-brief.md records no ship branch; defaulting to legacy branch fm/cap-a
notice: cap-a ships mode=local-only while the standing posture for heavy is no-mistakes - less rigor than the captain's standing posture; proceed only on a current explicit captain instruction or an intake judgment you can state
spawned cap-a harness=sleep kind=ship mode=local-only yolo=off window=primary:fm-cap-a worktree=/tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/1/heavy
[exit 0]
$ bin/fm-spawn.sh cap-b /tmp/fm-lab.1AKLOV/projects/heavy --mode local-only --yolo off --harness 'sleep 3600'
fm-gate-refuse: gate agent lifecycle permitted only against lab home /tmp/fm-lab.1AKLOV
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
●  WATCHER DOWN - SUPERVISION IS OFF
●  1 task(s) in flight, but no watcher has a fresh beacon (last beat: never, grace 300s).
●  Trust the emitted supervision protocol for this harness; do not use shell & for watcher repair.
●  This is a supervision warning only; the guarded operation WILL still run.
●  watcher supervision needs Stop-owned automatic recovery; inspect the hook registration and startup status before ending the turn.
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
warning: /tmp/fm-lab.1AKLOV/data/cap-b/launch-brief.md records no ship branch; defaulting to legacy branch fm/cap-b
notice: cap-b ships mode=local-only while the standing posture for heavy is no-mistakes - less rigor than the captain's standing posture; proceed only on a current explicit captain instruction or an intake judgment you can state
spawned cap-b harness=sleep kind=ship mode=local-only yolo=off window=primary:fm-cap-b worktree=/tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/2/heavy
[exit 0]
--- state records:
cap-a.meta
cap-b.meta
--- tmux windows:
primary:bash bash ~/.no-mistakes/worktrees/5284051b2355/01M468TAJ5H4QA2XF5A86T06ZW
primary:fm-cap-a sleep /tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/1/heavy
primary:fm-cap-b sleep /tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/2/heavy
--- git worktrees:
/tmp/fm-lab.1AKLOV/projects/heavy                            47eeee2 [main]
/tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/1/heavy 47eeee2 (detached HEAD)
/tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/2/heavy 47eeee2 (detached HEAD)
--- backlog:
# Backlog

## In flight
- [ ] cap-b - item for cap-b (kind: ship) (since 2026-10-05)
- [ ] cap-a - item for cap-a (kind: ship) (since 2026-10-05)

## Queued
- [ ] cap-c - item for cap-c (kind: ship) (since 2026-10-05)
## Done

\### S2 - declare 'heavy 2' with both places held: spawn cap-c is deferred
config/project-capacity:
# heavy suite serves two workers
#
heavy 2
$ bin/fm-spawn.sh cap-c /tmp/fm-lab.1AKLOV/projects/heavy --mode local-only --yolo off --harness 'sleep 3600'
fm-gate-refuse: gate agent lifecycle permitted only against lab home /tmp/fm-lab.1AKLOV
WARNING: watcher still down (same stale episode; last beat: never, grace 300s) - full banner already printed this episode.
deferred: project heavy admits 2 worker(s) at once on this machine (/tmp/fm-lab.1AKLOV/config/project-capacity) and 2 already hold a place (cap-a, cap-b); task cap-c was not launched and its backlog item stays queued - dispatch it again once one of them records its ready PR or is cleaned up
[exit 75]
stderr line count: 3 ; lines starting 'deferred:': 1
brief.md
--- state records:
cap-a.meta
cap-b.meta
--- tmux windows:
primary:bash bash ~/.no-mistakes/worktrees/5284051b2355/01M468TAJ5H4QA2XF5A86T06ZW
primary:fm-cap-a sleep /tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/1/heavy
primary:fm-cap-b sleep /tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/2/heavy
--- git worktrees:
/tmp/fm-lab.1AKLOV/projects/heavy                            47eeee2 [main]
/tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/1/heavy 47eeee2 (detached HEAD)
/tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/2/heavy 47eeee2 (detached HEAD)
--- backlog:
# Backlog

## In flight
- [ ] cap-b - item for cap-b (kind: ship) (since 2026-10-05)
- [ ] cap-a - item for cap-a (kind: ship) (since 2026-10-05)

## Queued
- [ ] cap-c - item for cap-c (kind: ship) (since 2026-10-05)
## Done

\### S3 - a malformed or unreadable declaration refuses the spawn (both places still held by cap-a, cap-b)

#### capacity is not an integer
config/project-capacity:
  | heavy two
$ bin/fm-spawn.sh cap-c /tmp/fm-lab.1AKLOV/projects/heavy --mode local-only --yolo off --harness 'sleep 3600'
error: spawn refused: the project capacity declaration is unreadable (/tmp/fm-lab.1AKLOV/config/project-capacity line 1 gives heavy a capacity that is not a positive integer); fix it so the captain's worker limits are known (docs/configuration.md "Project capacity")
[exit 1]
after: cap-c record: 0, cap-c windows: 0, backlog row: - [ ] cap-c - item for cap-c (kind: ship) (since 2026-10-05) [section: ## Queued]

#### capacity of zero
config/project-capacity:
  | heavy 0
$ bin/fm-spawn.sh cap-c /tmp/fm-lab.1AKLOV/projects/heavy --mode local-only --yolo off --harness 'sleep 3600'
error: spawn refused: the project capacity declaration is unreadable (/tmp/fm-lab.1AKLOV/config/project-capacity line 1 gives heavy a capacity that is not a positive integer); fix it so the captain's worker limits are known (docs/configuration.md "Project capacity")
[exit 1]
after: cap-c record: 0, cap-c windows: 0, backlog row: - [ ] cap-c - item for cap-c (kind: ship) (since 2026-10-05) [section: ## Queued]

#### project named twice
config/project-capacity:
  | heavy 2
  | heavy 3
$ bin/fm-spawn.sh cap-c /tmp/fm-lab.1AKLOV/projects/heavy --mode local-only --yolo off --harness 'sleep 3600'
error: spawn refused: the project capacity declaration is unreadable (/tmp/fm-lab.1AKLOV/config/project-capacity line 2 names heavy a second time); fix it so the captain's worker limits are known (docs/configuration.md "Project capacity")
[exit 1]
after: cap-c record: 0, cap-c windows: 0, backlog row: - [ ] cap-c - item for cap-c (kind: ship) (since 2026-10-05) [section: ## Queued]

#### another project's line is malformed (line with no name)
config/project-capacity:
  | other 2
  | 7
$ bin/fm-spawn.sh cap-c /tmp/fm-lab.1AKLOV/projects/heavy --mode local-only --yolo off --harness 'sleep 3600'
error: spawn refused: the project capacity declaration is unreadable (/tmp/fm-lab.1AKLOV/config/project-capacity line 2 is not '<project-name> <capacity>'); fix it so the captain's worker limits are known (docs/configuration.md "Project capacity")
[exit 1]
after: cap-c record: 0, cap-c windows: 0, backlog row: - [ ] cap-c - item for cap-c (kind: ship) (since 2026-10-05) [section: ## Queued]

#### file is unreadable (mode 000)
config/project-capacity:
---------- 1 firstmate firstmate 8 Oct  5 16:57 /tmp/fm-lab.1AKLOV/config/project-capacity
$ bin/fm-spawn.sh cap-c /tmp/fm-lab.1AKLOV/projects/heavy --mode local-only --yolo off --harness 'sleep 3600'
error: spawn refused: the project capacity declaration is unreadable (/tmp/fm-lab.1AKLOV/config/project-capacity is not a readable regular file); fix it so the captain's worker limits are known (docs/configuration.md "Project capacity")
[exit 1]
after: cap-c record: 0, cap-c windows: 0, backlog row: - [ ] cap-c - item for cap-c (kind: ship) (sin

... [3878 bytes truncated] ...

] cap-d - item for cap-d (kind: ship) (since 2026-10-05)
- [ ] cap-b - item for cap-b (kind: ship) (since 2026-10-05)
- [ ] cap-a - item for cap-a (kind: ship) (since 2026-10-05)

## Queued
- [ ] cap-c - item for cap-c (kind: ship) (since 2026-10-05)
- [ ] cap-e - item for cap-e (kind: ship) (since 2026-10-05)
## Done

\### S6 - 'heavy 3' with three holders (cap-a, cap-b in 'heavy'; cap-d in the '#repo' clone of the same origin)
config/project-capacity:
  | heavy 3
$ bin/fm-spawn.sh cap-c /tmp/fm-lab.1AKLOV/projects/heavy --mode local-only --yolo off --harness 'sleep 3600'
deferred: project heavy admits 3 worker(s) at once on this machine (/tmp/fm-lab.1AKLOV/config/project-capacity) and 3 already hold a place (cap-a, cap-b, cap-d); task cap-c was not launched and its backlog item stays queued - dispatch it again once one of them records its ready PR or is cleaned up
[exit 75]

\### S7 - a relaunch of a worker that already holds a place is not counted (project still full)
cap-a pane after stopping its agent: fm-cap-a bash
$ bin/fm-spawn.sh cap-a --relaunch --harness 'sleep 3600'
warning: /tmp/fm-lab.1AKLOV/data/cap-a/launch-brief.md records no ship branch; defaulting to legacy branch fm/cap-a
notice: cap-a ships mode=local-only while the standing posture for heavy is no-mistakes - less rigor than the captain's standing posture; proceed only on a current explicit captain instruction or an intake judgment you can state
spawned cap-a harness=sleep kind=ship mode=local-only yolo=off window=primary:fm-cap-a worktree=/tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/1/heavy
[exit 0]
cap-a pane after relaunch: fm-cap-a sleep

\### S8 - cleanup frees exactly one place
$ bin/fm-teardown.sh cap-d
teardown: reaping leaked worktree process(es) for cap-d: 493073 499040
teardown: force-killing leaked worktree process(es) for cap-d: 493073
A new version of treehouse is available: v2.3.0 → v3.1.2
Run "treehouse update" to update

🌳 Worktree returned to pool.
teardown cap-d complete (window primary:fm-cap-d, worktree /tmp/fm-lab.1AKLOV/treehouse/.treehouse/#repo-1ac586/1/#repo)
Backlog: cap-d is closed in /tmp/fm-lab.1AKLOV/data/backlog.md. Run bin/fm-tasks-axi.sh ready for dependency-cleared candidates, check date gates, and dispatch only work whose blockers are gone and date is due.
[exit 0]
records now: cap-a.meta cap-b.meta 
$ bin/fm-spawn.sh cap-c /tmp/fm-lab.1AKLOV/projects/heavy --mode local-only --yolo off --harness 'sleep 3600'
warning: /tmp/fm-lab.1AKLOV/data/cap-c/launch-brief.md records no ship branch; defaulting to legacy branch fm/cap-c
notice: cap-c ships mode=local-only while the standing posture for heavy is no-mistakes - less rigor than the captain's standing posture; proceed only on a current explicit captain instruction or an intake judgment you can state
spawned cap-c harness=sleep kind=ship mode=local-only yolo=off window=primary:fm-cap-c worktree=/tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/3/heavy
[exit 0]
$ bin/fm-spawn.sh cap-e /tmp/fm-lab.1AKLOV/projects/heavy --mode local-only --yolo off --harness 'sleep 3600'
deferred: project heavy admits 3 worker(s) at once on this machine (/tmp/fm-lab.1AKLOV/config/project-capacity) and 3 already hold a place (cap-a, cap-b, cap-c); task cap-e was not launched and its backlog item stays queued - dispatch it again once one of them records its ready PR or is cleaned up
[exit 75]

\### S9 - a recorded ready PR frees exactly one place
$ bin/fm-pr-check.sh cap-b https://github.com/o/r/pull/7
armed: state/cap-b.check.sh
[exit 0]
cap-b record pr line: pr=https://github.com/o/r/pull/7
$ bin/fm-spawn.sh cap-e /tmp/fm-lab.1AKLOV/projects/heavy --mode local-only --yolo off --harness 'sleep 3600'
warning: /tmp/fm-lab.1AKLOV/data/cap-e/launch-brief.md records no ship branch; defaulting to legacy branch fm/cap-e
notice: cap-e ships mode=local-only while the standing posture for heavy is no-mistakes - less rigor than the captain's standing posture; proceed only on a current explicit captain instruction or an intake judgment you can state
spawned cap-e harness=sleep kind=ship mode=local-only yolo=off window=primary:fm-cap-e worktree=/tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/4/heavy
[exit 0]

\### S10 - capacity is shared with a local secondmate home (/tmp/fm-lab.1AKLOV/mate, registered in the root's data/secondmates.md; its own config/ has no declaration)
mate config dir: (empty)
$ env FM_HOME=/tmp/fm-lab.1AKLOV/mate bin/fm-spawn.sh mate-1 /tmp/fm-lab.1AKLOV/mate/projects/heavy --mode local-only --yolo off --harness 'sleep 3600'
deferred: project heavy admits 3 worker(s) at once on this machine (/tmp/fm-lab.1AKLOV/config/project-capacity) and 3 already hold a place (cap-a in /tmp/fm-lab.1AKLOV, cap-c in /tmp/fm-lab.1AKLOV, cap-e in /tmp/fm-lab.1AKLOV); task mate-1 was not launched and its backlog item stays queued - dispatch it again once one of them records its ready PR or is cleaned up
[exit 75]
mate records: 0; mate backlog section of mate-1: ## Queued
$ bin/fm-teardown.sh cap-e
teardown: reaping leaked worktree process(es) for cap-e: 586093 595840
teardown: force-killing leaked worktree process(es) for cap-e: 586093

🌳 Worktree returned to pool.
teardown cap-e complete (window primary:fm-cap-e, worktree /tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/4/heavy)
Backlog: cap-e is closed in /tmp/fm-lab.1AKLOV/data/backlog.md. Run bin/fm-tasks-axi.sh ready for dependency-cleared candidates, check date gates, and dispatch only work whose blockers are gone and date is due.
[exit 0]
$ env FM_HOME=/tmp/fm-lab.1AKLOV/mate bin/fm-spawn.sh mate-1 /tmp/fm-lab.1AKLOV/mate/projects/heavy --mode local-only --yolo off --harness 'sleep 3600'
warning: /tmp/fm-lab.1AKLOV/mate/data/mate-1/launch-brief.md records no ship branch; defaulting to legacy branch fm/mate-1
notice: mate-1 ships mode=local-only while the standing posture for heavy is no-mistakes - less rigor than the captain's standing posture; proceed only on a current explicit captain instruction or an intake judgment you can state
spawned mate-1 harness=sleep kind=ship mode=local-only yolo=off window=primary:fm-mate-1 worktree=/tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/4/heavy
[exit 0]
(cap-e was closed by its teardown; a new queued item cap-f is dispatched from the root home)
$ bin/fm-spawn.sh cap-f /tmp/fm-lab.1AKLOV/projects/heavy --mode local-only --yolo off --harness 'sleep 3600'
deferred: project heavy admits 3 worker(s) at once on this machine (/tmp/fm-lab.1AKLOV/config/project-capacity) and 3 already hold a place (cap-a, cap-c, mate-1 in /tmp/fm-lab.1AKLOV/mate); task cap-f was not launched and its backlog item stays queued - dispatch it again once one of them records its ready PR or is cleaned up
[exit 75]

\### Final lab state
--- state records:
cap-a.meta
cap-b.meta
cap-c.meta
--- tmux windows:
primary:bash bash ~/.no-mistakes/worktrees/5284051b2355/01M468TAJ5H4QA2XF5A86T06ZW
primary:fm-cap-a sleep /tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/1/heavy
primary:fm-cap-b sleep /tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/2/heavy
primary:fm-cap-c sleep /tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/3/heavy
primary:fm-mate-1 sleep /tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/4/heavy
--- git worktrees:
/tmp/fm-lab.1AKLOV/projects/heavy                            47eeee2 [main]
/tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/1/heavy 47eeee2 (detached HEAD)
/tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/2/heavy 47eeee2 (detached HEAD)
/tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/3/heavy 47eeee2 (detached HEAD)
/tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/4/heavy 47eeee2 (detached HEAD)
--- backlog:
# Backlog

## In flight
- [ ] cap-c - item for cap-c (kind: ship) (since 2026-10-05)
- [ ] cap-b - item for cap-b (kind: ship) (since 2026-10-05)
- [ ] cap-a - item for cap-a (kind: ship) (since 2026-10-05)

## Queued
- [ ] cap-f - item for cap-f (kind: ship) (since 2026-10-05)
## Done
- [x] cap-e - item for cap-e (kind: ship) (done 2026-10-05)
  local main
- [x] cap-d - item for cap-d (kind: ship) (done 2026-10-05)
  local main
--- mate records:
mate-1.meta
Evidence: Real Claude primary in the lab: pane capture of three turns plus lab state

Source: Real Claude primary in the lab: pane capture of three turns plus lab state

## Real Claude primary (claude CLI, normal login) in the lab home on the private fm-lab tmux socket, session 'captain'. Pane capture:
 ▐▛███▛█   Claude Code v2.1.289
▝▜██████▀  Opus 5.5 · Claude Max
 ▝▝   ▝▝   ~/.no-mistakes/worktrees/5284051b2355/01M468TAJ5H4QA2XF5A86T06ZW
❯ Captain here. This is a disposable lab home (FM_HOME=/tmp/fm-lab.1AKLOV). Lab rule for this whole session: every worker launch uses --harness 'sleep 3600' with --mode local-only
  --yolo off, briefs are already written and must not be edited, and skip watcher repair. Item cap-f is queued in the backlog for project /tmp/fm-lab.1AKLOV/projects/heavy.
  Dispatch it now with bin/fm-spawn.sh, then tell me in three lines what happened, what state the backlog item is in, and when you will try it again.
● Bash(export FM_HOME=/tmp/fm-lab.1AKLOV; ls -la $FM_HOME $FM_HOME/data $FM_HOME/config $FM_HOME/state 2>&1 | head -60; cat $FM_HOME/data/backlog.md 2>&1 | head -40; c…)
  ⎿  /tmp/fm-lab.1AKLOV:
     total 12
     drwx------   11 firstmate firstmate    280 Oct  5 16:58 .
     … +126 lines (ctrl+o to expand)
  ⎿  Allowed by auto mode classifier
● Bash(export FM_HOME=/tmp/fm-lab.1AKLOV; bin/fm-spawn.sh cap-f /tmp/fm-lab.1AKLOV/projects/heavy --mode local-only --yolo off --harness 'sleep 3600'; echo "EXIT=$?";…)
  ⎿  fm-gate-refuse: gate agent lifecycle permitted only against lab home /tmp/fm-lab.1AKLOV
     WARNING: watcher still down (same stale episode; last beat: never, grace 300s) - full banner already printed this episode.
     deferred: project heavy admits 3 worker(s) at once on this machine (/tmp/fm-lab.1AKLOV/config/project-capacity) and 3 already hold a place (cap-a, cap-c, mate-1 in /tmp/f
     … +17 lines (ctrl+o to expand)
  ⎿  Allowed by auto mode classifier
● Captain, the worker for cap-f did not start: project heavy allows 3 workers at once on this machine, and 3 already hold a place (cap-a, cap-c, and the second mate's worker
  mate-1), so bin/fm-spawn.sh deferred it with exit code 75.
  The cap-f backlog item is unchanged and still under Queued, not blocked - I read data/backlog.md after the run, and no cap-f worker record was created.
  I will dispatch cap-f again as soon as one of those three reports a ready branch or is cleaned up, and I re-check at every cleanup and fleet review; monitoring is down in this
  lab and I left it unrepaired per your rule, so that re-check happens on my next turn rather than by automatic notification.
✻ Cogitated for 17s · done 5:00 PM
❯ cap-c is finished and nothing in it needs keeping. Clean up cap-c with bin/fm-teardown.sh, then carry on as your standing instructions say.
● Skill(ship-landing)
  ⎿  Successfully loaded skill
● Bash(export FM_HOME=/tmp/fm-lab.1AKLOV; bin/fm-teardown.sh --help 2>&1 | head -40; echo ---; cat $FM_HOME/state/cap-c.meta)
  ⎿  fm-gate-refuse: gate agent lifecycle permitted only against lab home /tmp/fm-lab.1AKLOV
     error: teardown refused: task record is not a regular file at /tmp/fm-lab.1AKLOV/state/--help.meta
     ---
     … +13 lines (ctrl+o to expand)
  ⎿  Allowed by auto mode classifier
● Bash(export FM_HOME=/tmp/fm-lab.1AKLOV; bin/fm-teardown.sh cap-c; echo "EXIT=$?"; echo ---; cat $FM_HOME/data/backlog.md; ls $FM_HOME/state | grep cap-c)
  ⎿  fm-gate-refuse: gate agent lifecycle permitted only against lab home /tmp/fm-lab.1AKLOV
     WARNING: watcher still down (same stale episode; last beat: never, grace 300s) - full banner already printed this episode.
     teardown: reaping leaked worktree process(es) for cap-c: 549336 556179
     … +25 lines (ctrl+o to expand)
  ⎿  Allowed by auto mode classifier
● Standing instructions: re-evaluate queued work after every teardown. A place on heavy is now free, so cap-f goes again.
● Bash(export FM_HOME=/tmp/fm-lab.1AKLOV; bin/fm-spawn.sh cap-f /tmp/fm-lab.1AKLOV/projects/heavy --mode local-only --yolo off --harness 'sleep 3600'; echo "EXIT=$?";…)
  ⎿  fm-gate-refuse: gate agent lifecycle permitted only against lab home /tmp/fm-lab.1AKLOV
     WARNING: watcher still down (same stale episode; last beat: never, grace 300s) - full banner already printed this episode.
     warning: /tmp/fm-lab.1AKLOV/data/cap-f/launch-brief.md records no ship branch; defaulting to legacy branch fm/cap-f
     … +34 lines (ctrl+o to expand)
  ⎿  Allowed by auto mode classifier
● Bash(export FM_HOME=/tmp/fm-lab.1AKLOV; bin/fm-tasks-axi.sh --help 2>&1 | head -30)
  ⎿  fm-tasks-axi.sh - run tasks-axi against THIS home's backlog from any working directory.
     Usage: fm-tasks-axi.sh [<tasks-axi command> [args...]]
     … +27 lines (ctrl+o to expand)
  ⎿  Allowed by auto mode classifier
● Bash(export FM_HOME=/tmp/fm-lab.1AKLOV; bin/fm-tasks-axi.sh update --help 2>&1 | head -25; bin/fm-tasks-axi.sh show cap-f 2>&1 | head -20)
  ⎿  usage: tasks-axi update <id> [flags]
     aliases: edit
     flags:
     … +25 lines (ctrl+o to expand)
  ⎿  Allowed by auto mode classifier
● Bash(export FM_HOME=/tmp/fm-lab.1AKLOV; bin/fm-tasks-axi.sh update cap-f --body "Delivery: local-only, yolo off. Deviation from the standing no-mistakes posture for…)
  ⎿  ok: updated cap-f (body)
     task:
       id: cap-f
     … +23 lines (ctrl+o to expand)
  ⎿  Allowed by auto mode classifier
● Captain, cap-c is cleaned up and the cap-f worker is now running on project heavy.
  - cap-c cleanup: bin/fm-teardown.sh cap-c exited 0 without any force flag, so its own landed-work check passed. It stopped two leftover processes, returned the isolated copy to
    the pool, and moved cap-c to Done in the backlog.
  - cap-f dispatch: my standing instructions say to re-evaluate queued work after every cleanup, and cap-c's place on heavy was now free. bin/fm-spawn.sh exited 0 this time and
    moved cap-f from Queued to In flight. The Queued section is now empty.
  - Delivery note: the launch warned that local-only is less rigor than heavy's standing no-mistakes setting. I proceeded on your lab rule and recorded that reason on the cap-f
    backlog item.
  - Branch warning: the launch also warned that cap-f's instructions name no branch, so it defaulted to fm/cap-f. I left the instructions unedited per your rule.
  Heavy is at its limit of 3 again, and nothing else is queued. Monitoring is still down and unrepaired per your lab rule, so I will not get automatic notifications from cap-a,
  cap-b or cap-f.
✻ Churned for 45s · done 5:01 PM
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
❯ Show me the current backlog and fleet state
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
  ⏵⏵ auto mode on (shift+tab to cycle) · ← for agents

## Lab state after the primary's second turn
--- state records:
cap-a.meta
cap-b.meta
cap-f.meta
--- tmux windows:
captain:claude claude ~/.no-mistakes/worktrees/5284051b2355/01M468TAJ5H4QA2XF5A86T06ZW
captain:fm-cap-f sleep /tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/3/heavy
primary:bash bash ~/.no-mistakes/worktrees/5284051b2355/01M468TAJ5H4QA2XF5A86T06ZW
primary:fm-cap-a sleep /tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/1/heavy
primary:fm-cap-b sleep /tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/2/heavy
primary:fm-mate-1 sleep /tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/4/heavy
--- git worktrees:
/tmp/fm-lab.1AKLOV/projects/heavy                            47eeee2 [main]
/tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/1/heavy 47eeee2 (detached HEAD)
/tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/2/heavy 47eeee2 (detached HEAD)
/tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/3/heavy 47eeee2 (detached HEAD)
/tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/4/heavy 47eeee2 (detached HEAD)
--- backlog:
# Backlog

## In flight
- [ ] cap-f - item for cap-f (kind: ship) (since 2026-10-05)
  Delivery: local-only, yolo off. Deviation from the standing no-mistakes posture for heavy: explicit lab-session instruction that every worker launch uses local-only with yolo off.
- [ ] cap-b - item for cap-b (kind: ship) (since 2026-10-05)
- [ ] cap-a - item for cap-a (kind: ship) (since 2026-10-05)

## Queued
## Done
- [x] cap-c - item for cap-c (kind: ship) (done 2026-10-05)
  local main
- [x] cap-e - item for cap-e (kind: ship) (done 2026-10-05)
  local main
- [x] cap-d - item for cap-d (kind: ship) (done 2026-10-05)
  local main

## Third turn: a recorded PR-ready handoff on a capped project. Pane capture:
❯ Two things. I queued a new item cap-g for project heavy (brief already written). And worker cap-a just reported its pull request is ready: https://github.com/o/r/pull/9 - record
  that handoff with bin/fm-pr-check.sh cap-a and that URL (skip any forge verification, this is a lab). Then carry on as your standing instructions say.
● Bash(export FM_HOME=/tmp/fm-lab.1AKLOV; sed -n 1,60p bin/fm-pr-check.sh; echo ---; cat $FM_HOME/data/backlog.md; cat $FM_HOME/state/cap-a.meta; ls $FM_HOME/data/cap-…)
  ⎿  #!/usr/bin/env bash
     # Record a PR-ready task: store one validated canonical pr=<url> and the forge's
     # exact pr_head=<sha> when available, then atomically arm a static merge poll.
     … +89 lines (ctrl+o to expand)
  ⎿  Allowed by auto mode classifier
● Bash(export FM_HOME=/tmp/fm-lab.1AKLOV; bin/fm-pr-check.sh cap-a https://github.com/o/r/pull/9; echo "EXIT=$?"; cat $FM_HOME/state/cap-a.meta | grep -E '^pr')
  ⎿  WARNING: watcher still down (same stale episode; last beat: never, grace 300s) - full banner already printed this episode.
     armed: state/cap-a.check.sh
     EXIT=0
     … +2 lines (ctrl+o to expand)
  ⎿  Allowed by auto mode classifier
● The handoff is recorded. Heavy has a declared capacity, so my standing instructions call for re-evaluating the queue now: cap-a's recorded PR frees its place, and cap-g is the
  queued candidate.
● Bash(export FM_HOME=/tmp/fm-lab.1AKLOV; bin/fm-spawn.sh cap-g /tmp/fm-lab.1AKLOV/projects/heavy --mode local-only --yolo off --harness 'sleep 3600'; rc=$?; echo "EXI…)
  ⎿  fm-gate-refuse: gate agent lifecycle permitted only against lab home /tmp/fm-lab.1AKLOV
     WARNING: watcher still down (same stale episode; last beat: never, grace 300s) - full banner already printed this episode.
     warning: /tmp/fm-lab.1AKLOV/data/cap-g/launch-brief.md records no ship branch; defaulting to legacy branch fm/cap-g
     … +26 lines (ctrl+o to expand)
  ⎿  Allowed by auto mode classifier
● Captain, cap-a's ready PR is recorded and the cap-g worker is now running on project heavy.
  - cap-a handoff: bin/fm-pr-check.sh cap-a exited 0 and stored https://github.com/o/r/pull/9 on cap-a's record, with merge monitoring registered for it. cap-a stays In flight in
    the backlog until that PR merges, and merging it needs your word.
  - cap-g dispatch: heavy has a declared limit of 3 workers, so my standing instructions call for re-evaluating the queue after a recorded ready PR. That record freed cap-a's
    place, bin/fm-spawn.sh exited 0, and cap-g moved from Queued to In flight. The Queued section is empty.
  - Same two warnings as cap-f: local-only is less rigor than heavy's standing no-mistakes setting, which I proceeded on under your lab rule and recorded on the cap-g item. The
    instructions name no branch, so it defaulted to fm/cap-g, and I left them unedited.
  Heavy is at its limit again. Monitoring is still down and unrepaired per your lab rule, so a merge of PR 9 will not notify me automatically.
✻ Worked for 24s · done 5:02 PM
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
❯ I queued cap-h for heavy, dispatch it now
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
  ⏵⏵ auto mode on (shift+tab to cycle) · ← for agents

## Lab state after the third turn
--- state records:
cap-a.meta
cap-b.meta
cap-f.meta
cap-g.meta
--- tmux windows:
captain:claude claude ~/.no-mistakes/worktrees/5284051b2355/01M468TAJ5H4QA2XF5A86T06ZW
captain:fm-cap-f sleep /tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/3/heavy
captain:fm-cap-g sleep /tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/5/heavy
primary:bash bash ~/.no-mistakes/worktrees/5284051b2355/01M468TAJ5H4QA2XF5A86T06ZW
primary:fm-cap-a sleep /tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/1/heavy
primary:fm-cap-b sleep /tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/2/heavy
primary:fm-mate-1 sleep /tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/4/heavy
--- git worktrees:
/tmp/fm-lab.1AKLOV/projects/heavy                            47eeee2 [main]
/tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/1/heavy 47eeee2 (detached HEAD)
/tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/2/heavy 47eeee2 (detached HEAD)
/tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/3/heavy 47eeee2 (detached HEAD)
/tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/4/heavy 47eeee2 (detached HEAD)
/tmp/fm-lab.1AKLOV/treehouse/.treehouse/heavy-1ac586/5/heavy 47eeee2 (detached HEAD)
--- backlog:
# Backlog

## In flight
- [ ] cap-g - item for cap-g (kind: ship) (since 2026-10-05)
  Delivery: local-only, yolo off. Deviation from the standing no-mistakes posture for heavy: explicit lab-session instruction that every worker launch uses local-only with yolo off.
- [ ] cap-f - item for cap-f (kind: ship) (since 2026-10-05)
  Delivery: local-only, yolo off. Deviation from the standing no-mistakes posture for heavy: explicit lab-session instruction that every worker launch uses local-only with yolo off.
- [ ] cap-b - item for cap-b (kind: ship) (since 2026-10-05)
- [ ] cap-a - item for cap-a (kind: ship) (since 2026-10-05)

## Queued
## Done
- [x] cap-c - item for cap-c (kind: ship) (done 2026-10-05)
  local main
- [x] cap-e - item for cap-e (kind: ship) (done 2026-10-05)
  local main
- [x] cap-d - item for cap-d (kind: ship) (done 2026-10-05)
  local main
cap-a pr line: pr=https://github.com/o/r/pull/9
Evidence: Deferred spawn at capacity (live)
$ bin/fm-spawn.sh cap-c <lab>/projects/heavy --mode local-only --yolo off --harness 'sleep 3600'
deferred: project heavy admits 2 worker(s) at once on this machine (<lab>/config/project-capacity) and 2 already hold a place (cap-a, cap-b); task cap-c was not launched and its backlog item stays queued - dispatch it again once one of them records its ready PR or is cleaned up
[exit 75]
state records: cap-a.meta cap-b.meta | data/cap-c: brief.md only | backlog: cap-c under ## Queued
Evidence: Malformed declaration refuses (live)
config/project-capacity: heavy two
$ bin/fm-spawn.sh cap-c <lab>/projects/heavy --mode local-only --yolo off --harness 'sleep 3600'
error: spawn refused: the project capacity declaration is unreadable (<lab>/config/project-capacity line 1 gives heavy a capacity that is not a positive integer); fix it so the captain's worker limits are known (docs/configuration.md "Project capacity")
[exit 1]
Evidence: Shared cap across the root home and a local secondmate home (live)
$ env FM_HOME=<lab>/mate bin/fm-spawn.sh mate-1 <lab>/mate/projects/heavy --mode local-only --yolo off --harness 'sleep 3600'
deferred: project heavy admits 3 worker(s) at once on this machine (<lab>/config/project-capacity) and 3 already hold a place (cap-a in <lab>, cap-c in <lab>, cap-e in <lab>); task mate-1 was not launched and its backlog item stays queued - ...
[exit 75]

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - 1 info
  • ℹ️ bin/fm-spawn.sh:3046 - Every fresh ship or scout spawn now resolves the root Firstmate home through fm_project_capacity_config_dir before it knows whether a declaration exists. On non-Orca backends this adds no new failure, because fm_treehouse_project_lock_path already resolved the same root at line 3059. On Orca it is new: an Orca spawn from a secondmate home whose parent binding is malformed or unreachable now exits 1 with "could not resolve the root Firstmate home that declares project capacity", even when no config/project-capacity exists anywhere. This is fail-closed and matches the intent that an unreadable declaration refuses the spawn, since the spawn cannot tell whether the root declares a cap. No change is needed; the note records the one path where a home without a declaration behaves differently from before.
✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 13 of 15 scenarios driven live against the product
Scenario Result Live Evidence
With no config/project-capacity, the captain spawns two workers on one project and both launch ✅ pass live live-lab-transcript.txt section S1: cap-a and cap-b both print spawned ... with exit 0; two task records, two tmux windows, two treehouse worktrees, both items under In flight
With heavy 2 declared and two holders, a third spawn prints one deferred line, exits 75, creates nothing, and its backlog item stays queued ✅ pass live live-lab-transcript.txt section S2: exit 75, exactly one line starting deferred: naming holders cap-a, cap-b; data/cap-c holds only brief.md; records, windows and worktrees unchanged; cap-c still un…
Adversarial: a malformed or unreadable declaration refuses the spawn instead of guessing a limit ✅ pass live live-lab-transcript.txt section S3: non-integer capacity, zero, duplicate project, nameless line, mode 000 file and a directory each give `error: spawn refused: the project capacity declaration is unr…
A line that is only #, or # followed by whitespace, stays a comment and does not change the cap ✅ pass live live-lab-transcript.txt section S4: with #, # heavy 9, an indented # ... line, #note-with-no-number and a blank line above heavy 2, the spawn is deferred at 2, not 9, and not refused
A project whose clone directory is named #repo is capped by the line #repo 2, matched by clone directory name ✅ pass live live-lab-transcript.txt section S5: spawn from <lab>/projects/#repo prints deferred: project #repo admits 2 worker(s) with exit 75; after changing the line to #repo 3 the same spawn is admitted
Holders in another clone of the same origin count against the cap ✅ pass live live-lab-transcript.txt sections S5 and S6: the '#repo' spawn is deferred by holders cap-a and cap-b running in the 'heavy' clone; with heavy 3, cap-c is deferred with holders cap-a, cap-b, cap-d…
A relaunch of a worker that already holds a place is not counted while the project is full ✅ pass live live-lab-transcript.txt section S7: with 3 of 3 places held, bin/fm-spawn.sh cap-a --relaunch exits 0 and the cap-a pane runs the agent command again
Cleaning up a task frees exactly one place ✅ pass live live-lab-transcript.txt section S8: after bin/fm-teardown.sh cap-d, cap-c is admitted and the next spawn cap-e is deferred with holders cap-a, cap-b, cap-c
Recording a ready pull request frees that worker's place ✅ pass live live-lab-transcript.txt section S9: the real bin/fm-pr-check.sh cap-b &lt;url&gt; writes pr= to cap-b's record, and the next spawn cap-e is admitted although cap-b's record and window still exist
The cap is shared between the root home and a local secondmate home, in both directions ✅ pass live live-lab-transcript.txt section S10: a spawn from the secondmate home, whose own config/ is empty, is deferred on the root's declaration with holders `cap-a in <root>, cap-c in <root>, cap-e in <root>…
A secondmate-kind record does not hold a place ⏸️ untested no The prior payload did not establish a live result for this scenario. It was not driven against the running product: the run did not provision and launch a real persistent secondmate in the lab home, w…
Two concurrent spawns cannot both take the last place ⏸️ untested no The prior payload did not establish a live result for this scenario. It was not driven against the running product: the race needs one spawn held inside the admission lock at a precise point, and real…
A real primary that gets a deferral leaves the item queued, not blocked, and says it will retry when a place frees ✅ pass live live-primary-transcript.txt first turn: the Claude primary runs bin/fm-spawn.sh, sees exit 75, reports the item still under Queued, not blocked, and names cleanup or a ready PR as the retry trigger
A real primary re-evaluates queued work after a teardown and dispatches the deferred item ✅ pass live live-primary-transcript.txt second turn: asked only to clean up cap-c and carry on, the primary runs teardown, cites its standing instruction to re-evaluate queued work after every cleanup, and spawns…
A real primary re-evaluates queued work after a recorded PR-ready handoff on a capped project ✅ pass live live-primary-transcript.txt third turn: after bin/fm-pr-check.sh cap-a the primary states Heavy has a declared capacity, so my standing instructions call for re-evaluating the queue now and spawns c…
  • bash tests/fm-project-capacity.test.sh - 16 behavior cases passed, exit 0
  • bin/fm-lab-home.sh create $LAB plus a private tmux server tmux -L fm-lab with FM_HOME=$LAB; all commands below were typed into that pane with send-keys
  • bin/fm-spawn.sh cap-a|cap-b &lt;lab&gt;/projects/heavy --mode local-only --yolo off --harness &#39;sleep 3600&#39; with no config/project-capacity - both launched, exit 0
  • bin/fm-spawn.sh cap-c ... with heavy 2 declared and two holders - one deferred: line, exit 75, no record, no window, no worktree, no launch brief, item still under Queued
  • bin/fm-spawn.sh cap-c ... against six bad declarations (heavy two, heavy 0, a project named twice, a line with no name, mode 000 file, a directory) - each refused with exit 1 and created nothing
  • bin/fm-spawn.sh cap-c ... with #, # heavy 9, #note-with-no-number, an indented # ... line and a blank line before heavy 2 - still deferred at 2
  • bin/fm-spawn.sh cap-d &#39;&lt;lab&gt;/projects/#repo&#39; ... with #repo 2 - deferred, naming holders cap-a and cap-b from the other clone; with #repo 3 - admitted
  • bin/fm-spawn.sh cap-c ... with heavy 3 - deferred, counting cap-d from the '#repo' clone as the third holder
  • bin/fm-spawn.sh cap-a --relaunch --harness &#39;sleep 3600&#39; while the project was full - relaunched, exit 0
  • bin/fm-teardown.sh cap-d then bin/fm-spawn.sh cap-c ... (admitted) then bin/fm-spawn.sh cap-e ... (deferred)
  • bin/fm-pr-check.sh cap-b https://github.com/o/r/pull/7 (recorded pr=) then bin/fm-spawn.sh cap-e ... (admitted)
  • env FM_HOME=&lt;lab&gt;/mate bin/fm-spawn.sh mate-1 &lt;lab&gt;/mate/projects/heavy ... from a registered local secondmate home - deferred on the root home's declaration and holders; admitted after bin/fm-teardown.sh cap-e; then root bin/fm-spawn.sh cap-f ... deferred naming mate-1 in &lt;lab&gt;/mate
  • Real claude primary in the lab (tmux session captain): asked to dispatch cap-f at full capacity, then to clean up cap-c, then to record cap-a's ready PR with cap-g queued
  • Teardown: tmux -L fm-lab kill-server, removed the lab directory, confirmed git status --short is empty
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

Teardown's walk over the root home and its registered local secondmate homes
moves into bin/fm-wake-lib.sh as fm_local_firstmate_state_dirs, next to
fm_firstmate_root_home, so a second consumer can count task records across
this machine's homes without a copy. Teardown keeps its exact refusal wording
through a thin wrapper.
A project whose machine-local resource only serves a few workers at once had
no way to tell Firstmate so: every queued item was launched, and the surplus
workers spent full-context turns retrying the resource.

config/project-capacity in the root home now declares how many workers each
named project admits at once on this machine. bin/fm-spawn.sh counts the ship
and scout records on the same project origin across the root and its local
secondmate homes, skipping ones whose ready PR is recorded, while holding the
shared project lock through publication. A spawn with every place held exits 75
before any brief render, endpoint, worktree, record, or backlog move, so the
item stays queued; batches report it as deferred. Undeclared projects keep
today's uncapped dispatch, and an unreadable declaration refuses rather than
guessing the limit.

Refs kunchenguid#4237
Keep project-capacity admission beside main's worker-account pin so the existing pull request can merge.
@tiago-peixoto

Copy link
Copy Markdown
Contributor Author

Friendly nudge: head 9a63c565 is mergeable against current main and all CI checks pass. Happy to adjust anything if needed.

@greptile-apps

greptile-apps Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

[High risk] Adds machine-local worker spawn admission control by project capacity.

The PR appears safe to merge; no new actionable issue or outstanding previous finding remains.

Reviews (12) · Last reviewed commit: "docs: scope PR-ready re-evaluation to a ..."

Comment thread bin/fm-project-capacity-lib.sh Outdated
Comment thread bin/fm-project-capacity-lib.sh
Comment thread bin/fm-spawn.sh Outdated
Comment thread tests/fm-project-capacity.test.sh Outdated
…tests/fm-project-capacity.test.sh pass, and shellcheck at warning level is clean on the changed files. Each new test failed against the old code and passes now. - **ci-1 (spaced names):** a declaration line must give a name its capacity whenever the name is a valid clone directory name. `fm_project_capacity_lookup` now trims each line, skips blank lines and lines whose first non-blank character is `#`, and takes the last field as the capacity. Everything before that field is the name, so it may contain spaces. The old error cases still refuse: a single field is rejected, and trailing text leaves a last field that is not an integer. The library header and docs/configuration.md now say a name starting with `#` cannot be declared. New test `test_spaced_project_name_is_declared` declares `my heavy project 1` next to an indented comment line and gets a deferral. - **ci-2 (unreadable records):** the holder count must never silently leave out a holder. `fm_project_capacity_occupants` now refuses when a local home's state directory exists but cannot be read or listed, or when a `.meta` file cannot be read. The error names the path, and `fm-spawn.sh` shows it in its existing refusal message. New test `test_unreadable_holders_refuse_admission` covers an unreadable record in the root home and an unreadable state directory in a registered local secondmate home, then checks that the spawn is admitted once both are readable. The test is skipped when run as root. - **ci-3 (Orca lock):** any spawn that can become a holder for a capped project must take that project's lock. The lookup now also reports whether the declaration caps any project at all, and an Orca spawn takes the per-origin lock whenever it does. This covers every capped same-origin clone. It also covers some cases where no same-origin clone is capped, because a spawn cannot find clones under other directory names without searching for them. With no declaration file, Orca still skips the lock. The comments in the library and in the `fm-spawn.sh` header are updated. The Orca test now clones the origin as `project-2`, which has no declaration, and checks that its Orca spawn refuses while the lock is held and publishes no record. - **ci-4 (worktrees):** `assert_nothing_created` now also compares the project's `git worktree list` from before and after a deferred spawn. Both tests that call it take that snapshot first. Files changed: bin/fm-project-capacity-lib.sh, bin/fm-spawn.sh, docs/configuration.md, tests/fm-project-capacity.test.sh
A clone directory whose name begins with # was skipped as a comment, so that project stayed uncapped. A line is a declaration when the # is written against the rest of the name and the line ends with a capacity; a # followed by whitespace stays a comment.
…k past its memory cap. ShellCheck ran out of memory analyzing bin/fm-teardown.sh in CI (reason=memory, rc=251, peak about 8.39 GB). On current main the same file passes at about 7.29 GB. **Cause:** the new `fm_local_firstmate_state_dirs` function in bin/fm-wake-lib.sh had a conditional `. fm-secondmate-registry-lib.sh` with a `# shellcheck source=` directive inside the function. ShellCheck followed that source again, inside a function scope, wherever fm-wake-lib.sh is sourced, and bin/fm-teardown.sh is the heaviest root that sources it. Measured locally with `shellcheck --norc --external-sources bin/fm-teardown.sh`: - current main (fd325b1): 7.29 GB - main merged with this PR: 7.86 GB - the same merge without the in-function source: 7.27 GB **Rule this restores:** this change must not make any lint root heavier than it is on main. That function holds the only new nested source in the change. **Fix:** I removed the in-function source, which no caller needs. Both callers already load the registry library at top level before calling the function: - bin/fm-teardown.sh sources it directly. - bin/fm-spawn.sh, the only user of bin/fm-project-capacity-lib.sh, gets it through bin/fm-ff-lib.sh. I also documented the requirement in the function's comment and in the "Requires" note in bin/fm-project-capacity-lib.sh. No behaviour changes. **Verification:** - ShellCheck on head: bin/fm-teardown.sh peaks at 7.12 GB and bin/fm-spawn.sh at 6.68 GB, both with rc=0. bin/fm-wake-lib.sh and bin/fm-project-capacity-lib.sh lint clean. - tests/fm-project-capacity.test.sh, tests/fm-teardown.test.sh (102 ok) and tests/fm-teardown-endpoint-safety.test.sh all pass. Files changed: bin/fm-wake-lib.sh, bin/fm-project-capacity-lib.sh
A concurrent resume in another home waits five seconds for that lock.
Reclaim is the last presentation change on the recovery path, so holding
the lock through the launch tail made the waiter time out. The contributions
arm check also freezes its one-second clock, the same way the budget tests
do, because an unfrozen clock can tick past before the first forge read.
@tiago-peixoto

Copy link
Copy Markdown
Contributor Author

Friendly follow-up for review when you have time: head effe03d6 is mergeable and 19 of 20 checks pass, including Greptile Review, the no-mistakes contribution check and the Herdr behavior tests. The one failing check, Lint 1, is the ShellCheck out-of-memory that currently fails the same way on main (tracked in #5620), not something this change introduces. Project-capacity admission stays opt-in: without the declaration file nothing is capped. Happy to adjust anything you would like changed.

@tiago-peixoto tiago-peixoto changed the title feat(bin): defer spawns beyond a project's declared machine capacity feat(bin): defer worker spawns beyond a declared project capacity Oct 4, 2026
@tiago-peixoto

Copy link
Copy Markdown
Contributor Author

Friendly follow-up for review when you have time: the current head is f65549a2, GitHub reports it mergeable with a clean merge against main, and every check on that head is passing, including the no-mistakes contribution check. This change defers worker spawns past a declared project capacity. Happy to address any remaining concerns.

@kunchenguid

kunchenguid commented Oct 5, 2026 •

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate:

First look at head f65549a2a8de vs main f470a01c. Whole thread re-read, plus #4237 (ready-for-pr, opt-in). Attestation MATCH. NM SUCCESS. Greptile 5/5, prior findings resolved. CI all SUCCESS (run 37176844304). MERGEABLE / CLEAN. Fixes #4237 verified.

Tip vs main, unconfigured path. The scripted admission is properly gated: with no config/project-capacity, fm_project_capacity_lookup returns no cap, non-Orca spawns take the same project lock they already took, Orca still takes none, and nothing is deferred. The teardown home walk moved into fm_local_firstmate_state_dirs with equivalent refusals. (Minor note, not a blocker: an unconfigured fresh spawn now also resolves the root home through fm_project_capacity_config_dir, so a malformed root binding would refuse an Orca spawn that main allowed.)

The one piece that is not gated is in the always-loaded contract. AGENTS.md section 10 (Backlog contract) changes "Re-evaluate queued work after every teardown and heartbeat" to "after every teardown, recorded PR-ready handoff, and heartbeat" for every home. Without a declaration a PR-ready handoff frees nothing, so that is a new default re-evaluation turn on every PR handoff that only costs tokens.

contract-class: new-default (narrowly, because of that unconditional re-evaluation trigger). It becomes opt-in once that sentence is scoped to homes that declare project capacity, for example "...after every teardown and heartbeat, and also after a recorded PR-ready handoff when config/project-capacity caps that project...". The new section 7 line already reads as conditional and is fine.

VISION (per rule):

  1. One captain, one interface: aligns. Surplus workers stop burning turns retrying a saturated resource.
  2. Authority is explicit: aligns for the script (the declaration is the explicit grant; absent means uncapped). The unconditional section 10 trigger is the part that applies without that grant.
  3. Scripts own the mechanics: aligns. Counting and deferral are deterministic; exit 75 is a clean script verdict.
  4. A restart is a non-event: aligns. Holders are read from durable task records; a deferred item stays queued.
  5. Delegation with a spine: aligns. Composes with the existing dispatch rule; no new task shape.
  6. The fleet outlives any vendor: aligns. Backend-agnostic, Orca included.
  7. Scope: aligns. Command-layer dispatch admission. The token-efficiency clause is why the extra default turn matters.

Waiting on you, @tiago-peixoto: scope that AGENTS.md re-evaluation sentence to declared capacity. With that change, green checks, and a matching attestation, this is eligible to land as opt-in. Security FYI: none. Fork workflows already ran on this head, so there is nothing to approve.

A ready pull request frees a place only when that project declares capacity, so the always-loaded backlog contract should re-evaluate on that handoff only in that case.
@tiago-peixoto tiago-peixoto changed the title feat(bin): defer worker spawns beyond a declared project capacity feat(bin): defer spawns beyond a declared per-project capacity, opt-in Oct 5, 2026
@tiago-peixoto

Copy link
Copy Markdown
Contributor Author

The backlog re-evaluation sentence is now scoped to declared capacity. Queued work is re-evaluated after every teardown and heartbeat, and also after a recorded pull-request handoff only when config/project-capacity caps that project. Head 7ea7862a is current with main, mergeable, and every check is passing.

@tiago-peixoto

Copy link
Copy Markdown
Contributor Author

Friendly follow-up for review when you have time: the current head is 7ea7862a, GitHub reports it mergeable with a clean merge, and every check on that head is passing, including the no-mistakes contribution check. The always-loaded backlog sentence is now scoped to a declared project capacity, as requested on 2026-10-05. Happy to address any remaining concerns.

@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate: thanks, @tiago-peixoto, and sorry this sat after you made the change. I re-reviewed head 7ea7862a against current main 47aff866.

What I checked since my last stamp on f65549a2: the one thing I asked for is done. AGENTS.md section 10 now reads "Re-evaluate queued work after every teardown and heartbeat, and also after a recorded PR-ready handoff when config/project-capacity caps that project…". A home with no declaration gets no new re-evaluation turn. The new section 7 line describes conditional behavior. The rest of the tip matches what I reviewed before:

  • With no config/project-capacity, fm_project_capacity_lookup returns no cap and nothing is deferred.
  • Non-Orca spawns take the same project lock they already took. Orca takes it only when the declaration caps a project.
  • The teardown home walk moved into fm_local_firstmate_state_dirs with the same refusals.

All Greptile threads are resolved (5/5 on this head). The no-mistakes attestation matches 7ea7862a, "PR must be raised via no-mistakes" passed, all 22 checks are green, and GitHub reports MERGEABLE/CLEAN. Main has moved since CI ran, so I also merged this head into 47aff866 locally. The merge was clean, bash -n passes, and tests/fm-project-capacity.test.sh passed 16/16 on the merged tree. No workflow, network, or secret changes. No fork runs needed approval.

Contract-class: opt-in. The admission and its contract trigger stay off unless the captain writes config/project-capacity. The one unconfigured-path difference I found is minor: an Orca spawn now also resolves the root home, so a broken root binding refuses it, the same way non-Orca spawns already refuse. That's a fail-closed on corrupt state, not a new behavior on a healthy home.

VISION.md, rule by rule

  • One captain, one interface: aligns. Surplus workers stop burning turns against a saturated machine, and nothing new reaches the captain.
  • Authority is explicit and never inferred: aligns. The declaration is the explicit grant, and without it dispatch stays uncapped.
  • Scripts own the mechanics, agents own the judgment: aligns. Counting and deferral are deterministic, and exit 75 is a clean script verdict.
  • A restart is a non-event: aligns. Holders come from durable task records, and a deferred item stays queued.
  • Delegation with a spine: aligns. It composes with the existing dispatch rule, with no new task shape.
  • The fleet outlives any vendor: aligns. It's backend-agnostic, Orca included.
  • Scope: aligns. This is command-layer dispatch admission, and the contract trigger no longer spends tokens in homes that don't use it.

Next step: it's restore-or-opt-in, fully green, and attested, so I'm landing it now (closes #4237).

@kunchenguid
kunchenguid merged commit ac0811c into kunchenguid:main Oct 7, 2026
22 checks passed
@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate: this is merged. Thank you @tiago-peixoto — really appreciate you taking the time on this.

yelenplays added a commit to yelenplays/firstmate that referenced this pull request Oct 8, 2026
* fix: prevent unnecessary remote worker turnover (#6431)

* fix: prevent healthy remote job worker turnover

* no-mistakes(review): Serialize full LaunchAgent repair and verify launchd-tracked owners

* no-mistakes(review): Let launchd-tracked unpublished spawns start before reloading

* no-mistakes(document): Document remote worker heartbeat and serialized LaunchAgent recovery

* no-mistakes(ci): Full CI log showed the idle-worker regression exceeded its outdated command budget (82 versus 80) after independent heartbeat ownership checks were added. Raised the budget to 120 while retaining the separate busy-poll sleep limit. Remote-job and LaunchAgent executable tests passed, as did bash syntax validation and git diff --check. No production behavior changed

* no-mistakes(ci): Fixed missing readiness recovery under verified live ownership, preserving the serving PID and private file mode. Added executable regressions for deletion during a blocked sweep and stale readiness diagnostics without LaunchAgent reload. Deletion regression failed before the fix. Both remote-job test suites, bash syntax validation, and git diff --check passed

* no-mistakes(ci): Full CI log identified a flaky ownership-loss test racing an already-authorized heartbeat refresh. Replaced backdating and a fixed sleep with bounded observation of readiness expiry through the public probe. Production behavior unchanged. Remote-job and LaunchAgent executable suites passed; bash syntax validation and git diff --check passed

* fix: wait for launchd bootout cleanup

* no-mistakes(document): Document remote worker recovery and read-only turnover verification

* no-mistakes(review): Publish worker identity before lock owner records

* no-mistakes(document): Document worker identity publication safety invariant

* no-mistakes(ci): Fixed ci-1: replacement workers discard predecessor readiness before publishing identity and roll back identity if lock-owner recording fails. Added executable regressions reproducing both defects. LaunchAgent and orphan-reap suites, shell syntax checks, and git diff --check passed. No live service state was modified

* no-mistakes(ci): Restored lock-owner-before-identity publication and removed the identity rollback and reordering-only tests. Retained an executable regression proving predecessor readiness is rejected until replacement startup completes. LaunchAgent and orphan-reap suites, shell syntax checks, and git diff --check passed. Other changes remain intact; the retained-identity interrupted-repair edge remains out of scope. No live service state was modified

* fix: speed up remote-job sequence claim cleanup (#6575)

* fix(remote-job): reap expired seq claims with one directory walk

The hourly claim sweep forked uname+stat per .seq-claims entry and blocked
serving for ~85s at ~17k dirs. Delete expired empty claim dirs with a single
find -exec rmdir batch and cache the host uname for remaining mtime reads.

* no-mistakes(review): Restore original path mtime helper and drop uname cache

* no-mistakes(test): Restore claim retention eligibility; focused retention and serving tests pass

* no-mistakes(document): Document single-walk claim cleanup and regression entrypoints

* no-mistakes(ci): Fixed Lint 2’s reproduced SC1091 by adding the tests/lib.sh ShellCheck source directive to the retention test. Runtime behavior is unchanged. ShellCheck passed for both new claim tests; Bash syntax, the retention behavior test, and git diff --check passed

* fix(tests): keep fixture registries out of git worktree roots

A TMPDIR pointed at a repository root placed live .fm-test-* registries
beside tracked files, and a concurrent git add during the claim-walk CI
fix round committed three of them. Route registries and fixture roots
through a TMPDIR that refuses git worktree roots, remove the stray files,
and pin the escape with a behavioral cleanup test.

* no-mistakes(review): Preserve whole-second claim expiry in single-walk sweep

* no-mistakes(document): Correct temporary-directory resolution documentation

* no-mistakes(ci): Fixed ci-3 by changing only the stale, fresh, and read-only orphan fixture paths in tests/fm-test-fixture-cleanup.test.sh to use $FM_TEST_TMPDIR. All seven tests passed both normally and with TMPDIR set to the worktree root. Shell syntax and git diff --check passed

* feat(bin): record no-mistakes pipeline spend per task at teardown, opt-in (#5354)

* feat(bin): record each task's no-mistakes pipeline spend at cleanup

no-mistakes keeps every pipeline agent invocation's token usage only in its
local agent_invocations records, and cold pipeline agents (review, test,
document) leave no session log, so a task's review-loop cost never reached
Firstmate's records and could not be attributed after cleanup.

bin/fm-pipeline-spend.sh attributes a task's runs by the repository
no-mistakes resolves for the task copy, the task branch, and the branch's
creation (a relaunch mints a new spawn_gen while the branch keeps
validating), then sums every invocation, failed and cancelled included.
Token fields use no-mistakes' per-round deltas so resumed review rounds are
not counted twice, and unrecorded values stay unknown rather than zero.
show prints the record read-only; record appends it once per task
incarnation to data/pipeline-spend.jsonl. Teardown records it for every
ship task it cleans up, before deleting the branch and the task record.

fm_nm_state_db becomes the one owner of where no-mistakes' state database
lives, shared with the capped run-inventory reader.

* no-mistakes(review): Drop spend show command and timeout override

* style(bin): rewrap fm-pipeline-spend header comment

* Make pipeline spend recording opt-in

* feat(bin): record each task's no-mistakes pipeline spend at cleanup

no-mistakes keeps every pipeline agent invocation's token usage only in its
local agent_invocations records, and cold pipeline agents (review, test,
document) leave no session log, so a task's review-loop cost never reached
Firstmate's records and could not be attributed after cleanup.

bin/fm-pipeline-spend.sh attributes a task's runs by the repository
no-mistakes resolves for the task copy, the task branch, and the branch's
creation (a relaunch mints a new spawn_gen while the branch keeps
validating), then sums every invocation, failed and cancelled included.
Token fields use no-mistakes' per-round deltas so resumed review rounds are
not counted twice, and unrecorded values stay unknown rather than zero.
show prints the record read-only; record appends it once per task
incarnation to data/pipeline-spend.jsonl. Teardown records it for every
ship task it cleans up, before deleting the branch and the task record.

fm_nm_state_db becomes the one owner of where no-mistakes' state database
lives, shared with the capped run-inventory reader.

* no-mistakes(review): Drop spend show command and timeout override

* style(bin): rewrap fm-pipeline-spend header comment

* Make pipeline spend recording opt-in

* no-mistakes(document): Note pipeline-spend opt-in gate in teardown comments

* no-mistakes(ci): I fixed all three Greptile findings following your decision. Both test suites pass: tests/fm-teardown.test.sh (105 ok, exit 0) and tests/fm-pipeline-spend.test.sh (8 ok). I didn't run either new test against the old code, so the claim that they fail before the fix is from reading the code, not a run. Nothing was committed or pushed. **ci-1 / ci-2 (bin/fm-teardown.sh):** The rule is that every owned ship task in a home that has opted in gets a durable spend record before its task record is deleted. The `[ -d "$WT" ]` check skipped the recorder when the worktree folder was already gone, so no record was written. I removed that check and kept the other conditions: the task must be a ship task, teardown must own its worktree, and `config/pipeline-spend` must exist. `bin/fm-pipeline-spend.sh` already writes an unavailable-source line when the worktree is missing. The new test `test_teardown_records_unavailable_spend_for_a_gone_worktree` in tests/fm-teardown.test.sh sets up an owned ship task whose worktree is missing, with recording turned on. It checks that teardown succeeds, writes a line with `source=="unavailable"`, `total==null` and a reason saying the task copy is gone, and still removes the task record. Under the old check no ledger would have been written. **ci-3 (bin/fm-nm-run-lib.sh):** When `NM_HOME` and `HOME` were both unset, `fm_nm_state_db` fell back to `${HOME:-}/.no-mistakes`, which resolves to `/.no-mistakes`. It now uses `root=~/.no-mistakes`. Bash expands `~` to the account's home directory even when `HOME` is unset, which matches the previous Python `Path.home()` lookup. The new test `test_state_db_without_nm_home_or_home_uses_the_account_home` in tests/fm-pipeline-spend.test.sh runs the function with both variables unset and compares the result to the home directory from the system's password database. The old code would have returned `/.no-mistakes/state.sqlite`. shellcheck reports nothing new on the changed files. I added one `disable=SC2016` comment in the new test, where the single-quoted `$1` is meant to expand in the child shell

* feat(bin): start tasks from a named base branch (#6442)

* feat(bin): start tasks from a named base branch

Spawns always reset a task's pooled copy to origin's default branch,
so work that belongs on a feature, integration, or release branch
started from the wrong code and opened its PR against the default.

fm-brief.sh --base-branch records the base in the brief, fm-spawn.sh
resets the copy to origin/<base> and records base_branch= in task
meta, and the worker targets its PR at that branch. Review diffs,
the cleanup content check, and scout promotion read the recorded
base. local-only and Gerrit deliveries refuse a named base.

* no-mistakes(review): Anchor brief base-branch parse on scaffold Setup line

* no-mistakes(review): Require explicit --base-branch on spawn, matching brief lines

* no-mistakes(document): Note base-branch reset in fm-spawn freshness header

* no-mistakes(ci): I fixed all four Greptile findings the way you specified. The affected suites pass locally: fm-brief, fm-dod-lib, fm-spawn-pool-base-freshen, fm-review-diff, fm-teardown and fm-task-delivery. shellcheck on the changed files reports only info-level notes, no warnings or errors. I didn't revert the code to watch each new test fail before the fix. - **ci-1 (brief text blocked or redirected spawns):** In bin/fm-dod-lib.sh, `fm_brief_base_branches` now only accepts a `Base branch: <value>` line that comes directly after the base-variant Setup sentence fm-brief.sh writes. Any other `Base branch:` line is ignored. `--base-branch` is still the only authority, and the existing checks in fm-spawn.sh are unchanged. With the flag, the spawn is refused unless a matching pair exists and no pair names a different branch. Without the flag, it is refused only if such a pair exists. The header comment in fm-spawn.sh is updated to match. - The `brief_with_base` test helper now writes the real two-line pair. - The decoy refusal case now uses a full decoy pair in the captain's intent. - New test `test_prose_base_branch_line_is_ignored`: a brief with no base whose intent quotes `Base branch: release/1.2` spawns normally with no `base_branch=` recorded. A spawn with `--base-branch feature/hub` whose intent names release/1.2 starts from `origin/feature/hub` and records it. - **ci-2 (scout base not checked against the forge):** fm-spawn.sh now looks up the project's registered forge with `bin/fm-project-mode.sh --forge`, like the ship path does, before `fm_base_branch_valid` runs. Previously it passed `none`. New test `test_scout_base_branch_refused_on_gerrit_forge`: a scout with a base on a forge=gerrit project is refused at spawn and no meta file is written. - **ci-3 (base not shell-quoted in worker commands):** `fm_dod_block` now quotes the base with `printf %q` in the `--base` and `--base-branch` arguments, and the branch name in the prose stays readable. The test in fm-brief.test.sh uses the base `release/$HOTFIX` and checks that both direct-PR and no-mistakes briefs render `release/\$HOTFIX` in the commands. - **ci-4 (ambiguous PR sentence):** The direct-PR sentence now reads "open a PR with `gh-axi` that is ready for review, not a draft, against the base branch `X` (`--base X`), not the repository default." A brief with no base renders exactly as before. The existing assertion in fm-brief.test.sh is updated

* no-mistakes(ci): Both real failures on PR 6442 (run 37064756768) come from CI tooling and a Pi version bump, not from the base-branch feature. Fixes for both already exist on upstream main, so I applied those two commits to the working tree with `git cherry-pick --no-commit`. Nothing is committed or pushed, and no feature file changed. - **Lint 2:** ShellCheck ran out of memory on `bin/fm-teardown.sh` (rss about 8 GB, exit 251, "shellcheck: out of memory"). The rule that broke is that the lint gate must finish on any valid shell root without exceeding the runner's memory ceiling. Upstream commit d719ef3 (#6443) fixes this in `bin/fm-lint.sh`: a root that hits the memory ceiling is retried without `--external-sources`. It also updates `tests/fm-lint.test.sh` and `docs/fm-test-portable-shards.md`. - **Behavior portable serial 6:** `tests/fm-calm-pi-extension.test.sh` failed with "grep disappeared from /export calm.html HTML while calm mode was on". Pi 1.0.1 renamed its HTML export renderer lookup, and upstream commit fede619 (#6530) updates the test to accept the new name. This change is test-only. Both commits applied cleanly. The four touched files are `bin/fm-lint.sh`, `docs/fm-test-portable-shards.md`, `tests/fm-lint.test.sh` and `tests/fm-calm-pi-extension.test.sh`. None of them is in the feature diff, so the feature contract is unchanged. **Local checks:** - `tests/fm-lint.test.sh` passes. - `bin/fm-lint.sh` on `bin/fm-teardown.sh`, `bin/fm-spawn.sh` and `bin/fm-dod-lib.sh` exits 0 with full ShellCheck analysis. - `tests/fm-calm-pi-extension.test.sh` exits 0 with 7 ok and 0 not ok. One case skipped because the Pi package isn't installed locally, so the Pi 1.0.1 export path that failed in CI couldn't be reproduced here. CI needs to confirm it. - I didn't reproduce the 8 GB OOM locally. The PR has to be updated through the pipeline with a normal push: no force push, no replacement PR, no merge

* fix(bin): name the Firstmate skill file as a fallback in the worker role (#6647)

* fix(bin): point project workers at the Firstmate skill file

The Skill tool cannot resolve a Firstmate skill from another project's worktree, so the launch role names the readable skill file instead.

* no-mistakes(review): fix(bin): name Firstmate skill file as fallback only

* feat(bin): retire a contribution whose forge object is permanently gone (#6655)

* feat(bin): retire a contribution whose forge object is permanently gone

Add fm-contributions.sh retire <task> <url> <captain|fleet> <reason>,
which records actor, reason and time on the saved record and removes
that task/url pair from known, poll rotation and coverage even while a
backlog link remains. Repeating a retire keeps the first provenance;
an unrecorded pair, unknown actor, empty reason or unacknowledged
pending signal is refused.

* no-mistakes(review): Keep retirement per task when settling final owners

* no-mistakes(ci): Made the three Greptile fixes the captain chose. ci-1, retirement needs the captain's word: `retire` now accepts only the actor `captain`. Any other actor is refused, including `fleet`. I removed `fleet` from the usage line, the script header (`bin/fm-contributions.sh`), the check in `bin/fm-contributions.sh` and the retired-record validation in `bin/fm-contributions.jq`, which now requires `.actor == "captain"`. The script header now says a retirement always records the captain's word, because the script cannot verify who runs it and the authority to retire is the captain's. The Bearings skill line in `.agents/skills/bearings/SKILL.md` now says `retire` records the captain's word. In the tests, a `fleet` retire is now a refusal case and must leave the record unretired. Every other retire call in the tests uses `captain`. The PR description has not been changed yet: the ruling asks for the same captain's-word statement there, and that belongs to the PR phase. ci-2, blank reasons: the reason check is now `[ -n "${5//[[:space:]]/}" ]`, so a reason made only of spaces or tabs is refused. I added a refusal test that passes a space-tab-space reason. The header also lists a blank reason among the refusals. ci-3, the gone-object test: the test forge wrapper has a new `not-found` fault that returns HTTP 404 for every `api repos/o/r/...` read. The happy-path test `test_retire_ends_observation_of_a_gone_contribution` now uses it in place of the 502 `down` fault, so it models a permanently gone repository rather than an outage. Verification: `bash tests/fm-contributions.test.sh` passed with exit 0. The last lines of its output include the three retire tests passing and no failures. `shellcheck` flagged only SC2034 (`FM_WAKE_QUEUE` unused) and SC1091 (an unfollowed source of `bin/fm-path-lib.sh`), neither on a line this round changed

* test: extend remote-reply whole-log recapture waits (#6639)

* test: give the remote-reply whole-log recapture a longer wait

* no-mistakes(review): Extend both recapture waits and simplify retry handling

* fix(bin): reopen a pending-reply escalation after its resolve (#6654)

* fix(bin): reopen a pending-reply escalation after its resolve

A retry of an open decision still appends nothing. A later same-kind escalation after a resolved line for that key appends again, so the next loss is visible.

* no-mistakes(ci): Both Greptile findings are fixed in the working tree; the four directly affected test suites pass, and a wider run of related suites had not finished when I returned this result. Invariant: a line already in the status file is recorded, and only a caller with its own evidence of a new episode may append it again. It must hold for every caller of status_event_recorded: the continuity break and the other call in bin/fm-procevent-remote-reply.sh, fm_parent_channel_append_once, and the pending-reply escalation. Changes: - bin/fm-classify-lib.sh: status_event_recorded is back to the idempotent retry check. It returns at the first matching line, and a later resolved line no longer makes that line look new. This fixes ci-1 for every caller and restores the early exit for ci-2, with no separate scan optimization. - bin/fm-pending-reply-lib.sh: _fm_pending_reply_maybe_escalate_locked now owns the reopen. It appends a new blocked line when the record already has an escalated_epoch (it escalated before and was reset) and the keyed decision is no longer open. Otherwise it uses status_event_recorded, so a retry or a resend that leaves the decision open appends nothing. - Comments on status_event_recorded and the pending-reply header describe this scope. Tests: - tests/fm-pending-reply.test.sh: the regression for escalate, operator resolve, second escalate is unchanged and passes. - tests/fm-remote-reply.test.sh: new regression drives the real reader. After an operator resolve of remote-reply-continuity-ios, a repeated handle and ingest of the same break appends nothing and leaves the decision closed. I placed it beside the existing continuity-break test, not in the pending-reply file, because only that file has the reader harness. - tests/fm-classify-corr-token.test.sh and tests/fm-classify-decision-key.test.sh: the cases that expected a blocked line to append again after a resolve now assert it stays recorded. Verification: - fm-classify-decision-key, fm-classify-corr-token, fm-pending-reply and fm-remote-reply test suites all exit 0 with the fix. - With the two bin files reverted to the commit under review, the new continuity regression and the updated parent-publisher case both fail, so they reproduce ci-1. - shellcheck -x on the changed files reports nothing. - Not finished: a background run of every other suite that mentions the parent channel or pending replies had produced no output, pass or fail, when I returned. One known gap: the escalation writes escalated_epoch before phase. If the process dies between those two writes and an operator resolves the key before the next tick, that tick appends a blocked line for the same episode. I left it, because closing it needs new state. Issue 4755 and other escalation behavior are untouched. Nothing is committed

* fix: refuse a confirming Enter on the Claude background-task exit picker (#6666)

* fix: refuse a confirming Enter on the Claude exit picker

The background-task picker still classifies as pending, so a retried Enter confirms "Exit and stop tasks".
Stop after the Enter that opened it, and raise the existing stale wake with the dialog name.

A non-paused secondmate still skips that wake, so a secondmate parked on the picker still looks idle.
Model-downgrade confirmation, MCP approval, and any other Claude exit confirmation are not covered, because there is no recorded screen for them.

* no-mistakes(review): fix: anchor exit picker match and wake second mates

* fix: refuse a typed submit while the Claude exit picker is open

A pane that already shows the picker must not receive the message or a confirming Enter.
The match requires the recorded heading and footer lines, and it is skipped when no dialog name is being recorded.

* no-mistakes(review): restore watcher to main behaviour, drop dialog wake

* no-mistakes(test): test: align Herdr picker fixtures with the preflight read

* no-mistakes(document): document exit refusal on a recognised dialog

* fix: remove the dialog file when exit runs in a subshell

do_exit is invoked from a command substitution, so the parent cleanup never saw the path it is supposed to delete.

* no-mistakes(review): fix: set the dialog file path after the control lock

* no-mistakes(document): document why the dialog file path follows the lock

* no-mistakes(ci): The exit cleanup in bin/fm-control.sh now deletes the dialog file first and releases the control lock second. The changes are in the worktree and are not committed. Invariant: only the process that holds the control lock for a task may write or delete that task's dialog file ($STATE/<id>.composer-dialog, the file that carries the picker name from the composer read to the Enter checks). The old cleanup released the lock and then deleted the file, so the delete ran outside the lock. Places where this invariant must hold: control_cleanup is the only place that deletes the file, and the single assignment after CONTROL_LOCK_HELD=1 is the only place that names it. do_exit only truncates the file, and it runs while the parent holds the lock. A process that loses the lock never sets the path, so it deletes nothing (earlier fix, unchanged). No sibling site needed a change. Fix: in control_cleanup, the rm of the dialog file moved above the fm_lock_release call, with a two-line comment that gives the reason. No dialog recognition or supervision code changed. Test: added test_exit_removes_the_dialog_file_before_releasing_the_lock in tests/fm-control-relaunch.test.sh. It runs one `fm-control exit` with a recording rm on PATH. The lock release removes paths at or under the control lock with rm, so the recording rm writes whether the dialog file exists at that moment. The test asserts the last record is "absent". It starts no second command. Verification, all run locally: - Before the fix, the new test failed with "the dialog file must be gone when the control lock is released, got: present". - After the fix, tests/fm-control-relaunch.test.sh exits 0 with 75 ok lines and no "not ok" line, including the new test and the existing test that the file is gone after exit and after relaunch. - tests/fm-control.test.sh exits 0 with 44 ok lines. - shellcheck -x on both changed files exits 0 with no output. One thing I noticed and did not change: tests/fm-control-relaunch.test.sh prints 615 "rm: cannot remove ... Permission denied" lines during its temp-directory teardown. The count is the same with my change reverted, so this change does not cause it, and the suite still exits 0. I could not read the Greptile check log (the check run was not found), so I worked from the finding text and the user instructions

* fix: restore timeout watchdog compatibility with macOS Bash 3.2 (#6028)

* fix: support stock macOS Bash in timeout watchdog

* no-mistakes(ci): Fixed ci-1 in tests/fm-timeout-lib.test.sh: both startup-owner failure paths now kill the bounded command group before killing the watchdog. Fault injection reproduced both leaks before the fix and confirmed cleanup afterward. Bash 3.2 suite, syntax check, ShellCheck, and diff checks passed; GNU timeout coverage skipped because the binary is unavailable. Production code unchanged

* fix(project-management): use subshell form for Initialize command (#6699)

Wrap the cd command in a subshell to comply with the cd-guard policy
that blocks persistent top-level directory changes in the primary
firstmate checkout. The subshell form (cd projects/<name> && ...) is
accepted by the policy as documented in issue #6502.

Fixes #6502

* feat(bin): add armable daily startup growth check (#6725)

* Add daily startup growth check

* no-mistakes(review): watch printed startup memory, retain growth baselines, drop knobs

* no-mistakes(review): delegate budget verdict, report before publish, pin shim home

* no-mistakes(review): baseline first sightings silently, drop mtime, spare secondmates

* no-mistakes(review): cap the wake line, validate budget verdict fields

* no-mistakes(review): guard record schema, check appends, tighten assertions

* no-mistakes(review): drop arbitrary tracked library, make re-arm idempotent

* no-mistakes(review): correct tracked-set wording, restore shim guards, register artifacts

* no-mistakes(review): exit on signal instead of publishing partial record

* no-mistakes(document): correct startup-growth record removal cost in state registry

* no-mistakes(test): sweep orphaned empty startup-growth temp records on due evaluation

* no-mistakes(document): note watcher need for armed startup growth check

* no-mistakes(ci): Diagnosed "Behavior portable serial 5": the shard failed on tests/fm-contributions.test.sh in test_arm_plumbs_a_configured_budget_into_the_check_shim (inherited mode) with "generated check did not attempt a read" and a missing forge/calls file. Root cause: bin/fm-contributions.sh:379-382 computes DEADLINE=$(date +%s)+BUDGET and then gates each read on DEADLINE-$(date +%s) >= OBSERVATION_RESERVE (= min(BUDGET,15)). At the one-second budget this test uses, that gate demands zero elapsed whole seconds, so a wall-clock second boundary crossing between the two date calls makes the poll loop break before any forge read. The test calls wrap_forge (which installs a fake date reading $FORGE/clock only when that file exists) but never seeded forge/clock, so it ran against the real clock. The repo already documents this exact hazard at tests/fm-contributions.test.sh:675-677 and freezes the clock in every other budget-constrained test: test_budget_exhaustion_keeps_prior_record (both modes), test_genuine_failure_near_deadline_is_unavailable, test_reservation_defers_later_url_when_fifteen_seconds_do_not_remain, test_slow_read_deadline_kill_is_budget_refusal, test_budget_is_cut_down_to_the_watcher_check_bound. test_arm_plumbs was the only omission; its single loop body covers both the configured and inherited modes. The remaining wrap_forge users run the default 20-second budget (5 seconds of slack above the reserve) and are not reachable by this race. Fix (smallest, matching the sibling idiom): added the clock freeze `/bin/date +%s > "$home/forge/clock"` plus a two-line comment inside that test's loop, before the poll. Three added lines in tests/fm-contributions.test.sh; no production code and no other file touched. Verification: reproduced the exact CI failure message by running a scratch copy whose fake date advances one second per read (worst case of the unfrozen clock) - missing forge/calls, "not ok - generated check did not attempt a read". With the fix, the single test passes 10/10 standalone and the full tests/fm-contributions.test.sh passes all 46 assertions with exit 0. tests/fm-startup-growth-check.test.sh still exits 0. bash -n and shellcheck -S warning are clean on the edited file; scratch repro files were removed, leaving only the intended three-line change. Note for the author: this flake is not caused by this branch - tests/fm-contributions.test.sh is untouched by the PR and the race is pre-existing, surfacing only on a loaded runner (the shard's neighbouring assertions show 80+ second gaps). It was fixed because it is genuine nondeterminism with an established in-repo remedy and the check cannot otherwise go green. No work was done on the unselected greptile findings (ci-2..ci-5) or the cancelled "Behavior portable serial 2" check (ci-6)

* no-mistakes(ci): Diagnosed "Behavior portable serial 6": the omitted log section (fetched with gh api) shows the shard failed on tests/fm-calm-pi-extension.test.sh in test_hidden_block_geometry_e2e with "Pi Calm hidden-block geometry E2E did not complete the /reload viewport transition" (exit=1, duration_ms=23437; the family line pure-contract-unit count=3 duration_ms=51193 failed=1 matches devin-harness 3203 + calm-pi 23437 + task-delivery 24553). Root cause: tests/fm-calm-pi-extension.test.sh:2831 wait_for_geometry_transition detected the /reload by SAMPLING a single intermediate frame - it polled tmux capture-pane for Pi's transient "Reloading keybindings, extensions, skills, prompts, themes, and context files..." box and only accepted the rebuilt transcript afterwards (elif gated on saw_transient). Pi shows that box only while the reload runs and replaces it via dismissReloadBox when done, so on a loaded runner the box can live entirely between two polls; saw_transient then stays 0 forever and the 600-attempt wait times out although the reload fully succeeded. Reproduced locally under CPU oversubscription with CI's Pi 1.0.4: 1 failure in 8 runs, diagnostics proving the mechanism ("DIAG: transition timeout saw_transient=0 attempt=600") while the captured viewport already showed the completed reload status row and the intact collapsed transcript. Not a Pi-version regression: 1.0.4's handleReloadCommand and both banner strings are identical to the installed 0.85.1, and shard 6 of the previous pipeline run (job 112490351105) passed this same script on Pi 1.0.4. Invariant violated: a TUI E2E assertion must key off a durable state the UI retains, never off one intermediate frame polling may miss. Sibling enumeration: the helper was defined once and called once; grep for transient/saw_/two-stage waits across tests/, bin/ and .pi/ found no other one-shot-frame wait - every other wait in the file keys off durable text presence/absence or off continuously repeating animation frames (the working-ship motion checks) where sampling eventually connects; no doc or other test referenced the helper. Fix (smallest, one helper): renamed it wait_for_geometry_reload and made it wait for the durable status row Pi appends once the reload completed and the chat was rebuilt ("Reloaded keybindings, extensions, skills, prompts, themes, and context files", a persistent chat child via showStatus, present in both 0.85.1 and 1.0.4) together with CALM_GEOMETRY_FINAL. That marker is absent before the reload and never appears if the reload throws, so the check is strictly stronger than before (a failed reload now fails the test instead of passing on transient-then-final). One file changed: tests/fm-calm-pi-extension.test.sh, 2 hunks, no production code. Verification: deterministic before/after - with polls spaced 0.4s apart (the loaded-runner worst case) the pre-fix helper fails with exactly the CI message and the post-fix helper passes; 12/12 consecutive passes of the isolated test under load with Pi 1.0.4; full file 15/15 ok exit 0 under LANG=C.utf8 (matching the runner locale); bash -n, shellcheck -S warning and bin/fm-lint.sh clean; scratch repro files and the temp Pi 1.0.4 install removed, git status shows only the one intended modification. Notes: (1) the flake is not caused by this branch - that test file is untouched by the PR and the race is pre-existing; fixed because it is genuine nondeterminism and the check cannot otherwise go green. (2) Under far harsher load than CI applies a separate unrelated flake appeared in the same test: Pi's "Warning: tmux extended-keys is off..." row occasionally lands between the collapsed [skill] ahoy row and the final response, making assert_geometry_gap see 5 rows instead of 2. Different cause, never seen in CI, and its plausible remedies would change key delivery for the other TUI tests in the file, so it was left alone rather than growing this fix

---------

Co-authored-by: Quartermaster <quartermaster@users.noreply.github.com>

* fix(bin): reopen the remote-reply continuity decision on a later break (#6708)

* fix: reopen a remote-reply continuity break after repair

A later break for the same route and reason was swallowed after the
operator resolved the first one, because the status line matched for
the life of the log. The continuity ingest now appends again when the
cursor has moved or retirement has reset that episode, and an unchanged
re-read still appends nothing.

status_event_recorded is unchanged. Its other callers are the
pending-reply escalation, which already decides its own episode, the
parent-channel note append, and the remote document transfer note.

* no-mistakes(review): seed continuity episode for already recorded break line

* no-mistakes(document): document when a remote-reply continuity break reopens

* no-mistakes(ci): The adapter now stores the continuity episode record before it appends the `blocked` line, so a failed store appends nothing. The changes are uncommitted in the worktree, in `bin/fm-procevent-remote-reply.sh` and `tests/fm-remote-reply.test.sh`. **Invariant:** a continuity `blocked` line is on the parent status log only when the episode record for that break is already stored. Only the `continuity-broken` branch of `cmd_ingest` appends that line, so the fix is at that one place. **What changed in `cmd_ingest`:** - It decides first whether the break needs a line, then stores the episode record, then appends. - If storing the record fails, it stops with "cannot record continuity episode" and appends nothing. - If the line is already on the log and no episode record exists, it stores the record and does not append. **One addition you did not ask for:** if the append fails after the record is stored, the adapter puts the earlier record back (or deletes the new one when none existed). Without that, the retry would read as the unchanged repeat and append nothing, which is the issue 6701 failure again. **Deviation from your test instruction:** the test does not use `chmod` on the cursor directory. `write_continuity_episode` runs `chmod 700` on that directory before every write, so a read-only directory is made writable again. The test instead makes `mktemp` fail for the episode's temporary file in that directory, the same way the existing receipt-failure test does. **Tests added to `tests/fm-remote-reply.test.sh`:** - Store failure: the second break exits 1, appends no `blocked` line and opens no decision. One retry after storage recovers exits 3 and appends a single line. - Append failure: with the status log read-only, the third break exits 1 and appends nothing. One retry after the log is writable exits 3 and appends a single line. **Verification:** `bash tests/fm-remote-reply.test.sh` ends with "ALL TESTS PASSED" with the fix. Against the script at commit b750c828 the same test file fails at "a continuity break appended its line before its episode was stored". `shellcheck -S warning` reports only an unused loop variable at line 667 of the test file, which this change does not touch. I ran no other test files. The Greptile Review check log could not be retrieved, so I worked from the finding text alone

* fix: record a continuity break's reader position on its status line

A later break at another cursor is then a different line, so the existing
duplicate check appends it and reopens the decision. An unchanged re-read
builds the same line and appends nothing.

* fix: reopen a continuity break after an identical restore

A retirement that puts the same bytes back used to rebuild the recorded line, so the later break stayed closed. The retirement count on that line makes the later break distinct.

* no-mistakes(review): remove continuity match for full-prefix line without retirement count

* no-mistakes(document): clarify what a continuity break status line records

* fix: remove the reply cursor before recording retirement

A stop between those steps must leave the count unchanged, so an unchanged continuity break still builds the same line.

* fix(control): drop busy_gen from the task record when an incarnation is retired (#6733)

* fix(control): drop busy_gen when an incarnation is retired

A deliberate exit removed the busy sidecar and left busy_gen in the task record, so the two records disagreed about whether that incarnation was still observable.

* no-mistakes(review): drop GNU-only chmod and unreached sidecar-absent branch

* no-mistakes(review): correct lock comment to name the deadlock

* no-mistakes(ci): The test `test_exit_drops_meta_busy_gen_with_the_sidecar` in tests/fm-control.test.sh now compares the whole task record (the `state/<id>.meta` file), so the Greptile finding is fixed. Invariant: after `exit` retires an incarnation, the task record must equal the record from before `exit` with only the `busy_gen` line removed. This test is the only place in the change that asserts the record survives the rewrite, so it is the only site to fix. The other `busy_gen` tests assert that the line stays, and they do not go through the rewrite. What changed: before `exit`, the test writes the record without its `busy_gen` line to `expected.meta`. After `exit`, the test runs `diff` between that expected copy and the real record, and fails with the diff output if they differ. This one comparison replaces the two earlier checks (no `busy_gen` line left, and the `window` line present), because it covers both. I did not change bin/fm-control.sh or any other file. How I know it works: - I ran `bash tests/fm-control.test.sh`: exit code 0, 45 lines starting with `ok`, no other lines. - I temporarily changed the rewrite in bin/fm-control.sh to also drop the `harness` line. The test then failed with `not ok - exit should drop only busy_gen from the task record:` and the diff `< harness=codex`. The earlier `window`-only check would have passed that rewrite. I restored bin/fm-control.sh afterwards; `git status` shows only tests/fm-control.test.sh modified. - `bash -n` and `shellcheck` on the test file report no new warnings from the edit. The change is not committed; the working tree holds it

* test: cover PID collisions in harness ancestry detection (#6484)

* test(secondmate-harness): scope fake ps -codex label for pid 5252 to the liveness probe

Closes #6456

* no-mistakes(ci): Updated the collision test to log and assert that PID 5252 was queried before selecting 4242. Full fm-secondmate harness suite passes

---------

Co-authored-by: YifuGu <ironerumi@users.noreply.github.com>

* fix: share Pi Calm's working-ship widget slot (#1854)

* fix(calm): share the standalone Pi Calm working-ship widget slot

Firstmate Calm and the user-global standalone Pi Calm both install an
animated working-ship widget during agent runs. Each claimed its own Pi
widget key, so a session loading both (the main Firstmate home) rendered
two boats. Pi replaces widgets under one key, so claiming the shared
"calm-working-ship" slot keeps dual-install sessions to a single boat
while a Firstmate-only session is unchanged.

Pins the shared slot contract in the working-ship module test so the key
cannot silently diverge again.

* test(calm): pin the shared working-ship widget key in CI, document dual-install

The key-parity assertion inside the Pi fixture only runs where the
@earendil-works/pi-coding-agent package is installed, so CI never
exercised it. Add a source-level twin that needs nothing but the
tracked file, and note in docs/calm.md that the boat shares the
standalone Pi Calm working-row widget slot.

* no-mistakes(review): Add executable dual-install widget replacement coverage

* no-mistakes(review): Guard shared widget cleanup with disposal ownership

* no-mistakes(document): Document shared Calm working-ship slot behavior

* test(calm): read the standalone Calm slot from its own module

The dual-install check registered both boats itself under the shared
slot, so it could only prove that Pi replaces a widget under one key: it
would still pass if the standalone Pi Calm extension installed its boat
under a different key, which is the two-boat regression the check exists
to prevent.

Read the standalone extension's own working-ship module when it is
installed - FM_STANDALONE_CALM_SHIP, else ~/.pi/agent/extensions/calm -
and drive the check with the key that module exports, so a rename on
either side registers two widgets and fails naming both keys. A pinned
shared-slot contract still covers a machine without the extension, and
the run reports which side it used instead of passing silently over an
absent extension.

Verified: the touched Pi Calm suite passes and reads the installed
standalone extension; with a copy of it whose key is renamed to
calm-working-ship-v2 the suite fails naming the drift.

* no-mistakes(review): Gate stock-row restoration by shared-widget ownership

* no-mistakes(review): Removed redundant widget-key source assertions

* no-mistakes(document): Document shared Calm working-ship widget ownership

* feat(bin): add opt-in --herdr-resume-lock-wait to fm-spawn (#6649)

* fix(herdr): make exact-resume presentation-lock wait instead of a bounded timeout

The exact-resume path in bin/fm-spawn.sh used the same 50-attempt-then-
give-up lock acquire as the new-task-create path, but the two paths are
not equivalent on contention: a create has no prior state to strand and
can safely fall back to a flat layout, while a resume is recovering a
specific existing identity that a concurrent recovery may legitimately
be holding the lock for. Giving up there does not degrade gracefully,
it hard-fails the resume outright. The suite's own concurrent
cross-home recoveries test already asserts both concurrent recoveries
succeed with a genuine reclaim, and the file's header comment already
(inaccurately) claimed lock contention falls back to the ordinary flat
layout for both paths alike, so the intended contract was always that
recoveries serialize and both succeed, not that either one refuses
under a short bound.

Give spawn_herdr_presentation_order_lock_acquire a wait mode that uses
this file's own established fm_lock_acquire_wait idiom (already used
for its other fleet-shared locks) instead of the bounded loop, and use
it only at the exact-resume call site. The new-task-create call site
is unchanged and keeps its bounded-then-flat-fallback behavior, which
is already covered by its own passing test. Dead-owner PID-liveness
reclaim inside fm_lock_try_acquire still bounds the wait against a
holder that crashed mid-hold.

Adds a deterministic regression test that holds the shared session
lock from an unrelated process for well past the old bound, then
asserts the resume succeeds with a genuine reclaim and took close to
the full hold duration, so a fix that merely widens the bound rather
than genuinely waiting is still caught. The existing concurrent
cross-home recovery test exercises this under real timing but does not
reliably outlast a fixed bound on its own.

Corrects the header comment's claim that create and resume share one
bounded-then-flat-fallback behavior on lock contention; they no longer
do.

* no-mistakes(document): Document Herdr recovery waiting for presentation lock

* no-mistakes(document): Update stale hard-refusal claim in verification log

* no-mistakes(ci): Fixed the Greptile finding on tests/fm-backend-herdr-presentation-e2e.test.sh:1389 by bounding the resume lock-wait regression's spawn_task call. Added an optional 4th `deadline_seconds` arg to the `spawn_task` helper (defaults to empty, so all ~20 other existing call sites are unaffected and unwrapped by `timeout`). The lock-wait test now passes `LOCK_WAIT_HOLD_SECONDS + 60` (90s) as the deadline, and a dedicated check for exit code 124 emits a clear "hung for over Xs instead of waiting out a Ys lock hold" diagnostic before falling through to the existing pass/fail assertions, which are unchanged. No product code was touched. Verified with `bash -n`, `shellcheck -x` (no warnings), a standalone reproduction of the timeout/no-timeout/success paths, the project's `bin/fm-lint.sh --fast` on the file (clean), and the full `tests/fm-lint.test.sh` suite (all 46 assertions pass)

* no-mistakes(ci): Replaced the direct `timeout "$deadline_seconds"` call in `spawn_task()` (tests/fm-backend-herdr-presentation-e2e.test.sh) with the repo's portable bounded-execution helper: sourced `bin/fm-timeout-lib.sh` at the top of the file and changed `deadline_cmd=(timeout "$deadline_seconds")` to `deadline_cmd=(fm_run_timed "$deadline_seconds")`. This removes the GNU/BSD `timeout` dependency that would fail with exit 127 on a stock macOS host without coreutils, while preserving identical semantics (exit 124 on bound-hit, command's own exit otherwise), which the existing `[ "$LOCK_WAIT_STATUS" -eq 124 ]` diagnostic check already relies on. Verified: `bash -n` syntax check, `bin/fm-lint.sh --fast` clean, full `tests/fm-lint.test.sh` suite (46/46 pass), and a standalone repro confirming `fm_run_timed` returns 124 on timeout and 0 on success identically to the prior `timeout` call. No other direct `timeout` calls exist in this file or elsewhere in the PR's diff, so no sibling sites remain

* fix(herdr): gate exact-resume lock wait behind --herdr-resume-lock-wait

Keep refuse-by-default on presentation-order lock contention for Herdr
exact resume. Callers that need concurrent recoveries to serialize must
pass --herdr-resume-lock-wait; unbounded blocking on a third-party session
lock is never the default.

Update docs and the real-Herdr e2e suite so the default path asserts the
refusal and the opt-in path asserts the wait.

* no-mistakes(test): Fix e2e test's lost exit status after if/fi with no else branch

* docs(herdr): stop advertising --herdr-resume-lock-wait on --relaunch

The relaunch path reuses the recorded endpoint and never takes the
presentation-order lock, so the flag is inert there. Drop it from the
--relaunch usage line and state where the flag applies.

* no-mistakes(review): Clarify lock-wait docs; simplify bash-3.2-safe spawn_task helper

* no-mistakes(ci): Fixed ci-1 (Greptile P2). In tests/fm-backend-herdr-presentation-e2e.test.sh, the failure cleanup `cleanup_all` stopped only `LOCK_CONTENTION_OWNER_PID`. It now also stops `LOCK_REFUSE_HOLDER_PID` and `LOCK_WAIT_HOLDER_PID`, the holders of the two new contention cases, so a `fail` before their explicit `wait` no longer leaves them running. Both new PIDs are initialised empty next to the existing one, and each is cleared right after its successful `wait` so cleanup never touches a finished PID. I changed nothing else. `bash -n` passes. The real Herdr e2e run passed both new cases ("default resumed identity refuses session lock contention" and "--herdr-resume-lock-wait waits out session lock contention instead of refusing"). The full run hit my 550s timeout in a later, unrelated case, after the new cases passed

* fix(bin): let TERM stop a watcher blocked in a pane capture on bash 3.2 (#6762)

* fix(bin): let TERM stop a watcher blocked in a pane capture on bash 3.2

Stock macOS bash 3.2 holds a HUP or TERM until a running command
substitution's child exits, and the watcher read every pane through
$(fm_backend_capture ...). A blocked backend read therefore held the
watcher's stop for as long as the read lasted, and a stopped watcher left
the hung read orphaned. tests/fm-watch-triage.test.sh
test_term_stops_a_watcher_blocked_inside_a_poll failed on /bin/bash 3.2
for this reason while passing on bash 5.

Pane captures now go through watcher_capture, which runs the read as a
waited background process group recorded like a check's, so the stop is
honored at once and watcher_cleanup stops a read still in flight along
with its per-call output file.

* no-mistakes(review): Run drain-ring idle capture in watcher shell, add regression test

* no-mistakes(document): Document watcher TERM handling for blocked checks and captures

* no-mistakes(test): Silence bash 3.2 setpgid race noise from watcher captures

* fix(bin): verify the capture group and scope the stop claim to pane reads

watcher_capture now confirms its background read leads its own process
group, as run_check_capture already does, so watcher_cleanup never relies
on a group that set -m failed to create. The comment and continuity doc
now say only fm_backend_capture pane reads go through watcher_capture;
agent-state and composer-state reads still run inside command
substitutions.

* feat: add per-home worker tool exclusions (#6750)

* feat(spawn): add per-home worker tool exclusions

Add an optional per-home config/crew-exclude-tools file listing tool names to hide from workers, one per line, with blank lines and # comments allowed.
It applies to every ship and scout launch and relaunch in that home, is never inherited by another home, and does not affect secondmate agents.
Pi and pi-signed apply it through --exclude-tools, which also covers MCP tool names.
Any other runtime, and a raw launch command, refuses the launch when the list is non-empty rather than ignoring it.
Malformed entries are refused before provisioning, and before a relaunch stops a running worker.
Exclusions that match no tool in the worker's loaded registry are reported as unverified warnings in its status record instead of refusing the worker.

Closes #6744

* no-mistakes(review): Preserve UTF-8 exclusion paths and verify Pi lifecycle behavior

* no-mistakes(document): Clarify worker tool exclusion documentation

* no-mistakes(ci): Fixed ci-1 in bin/fm-exclude-tools-lib.sh: a failed read now returns an error before printing names, so all shared launch and relaunch callers refuse rather than silently dropping exclusions. Added deterministic regression coverage for a file disappearing after readability checks across Pi/pi-signed ship and scout launches. Reproduced the original failure; verified 83 spawn checks, 77 relaunch checks, direct parser/runtime failure cases, full targeted lint, Bash syntax, and git diff --check. Relaunch tests passed with existing fixture-cleanup permission warnings. ci-2 remains unchanged per the user's decision; the outer executor owns the fresh CI run

* feat(bin): defer spawns beyond a declared per-project capacity, opt-in (#5343)

* refactor(bin): share the local Firstmate home walk from the wake library

Teardown's walk over the root home and its registered local secondmate homes
moves into bin/fm-wake-lib.sh as fm_local_firstmate_state_dirs, next to
fm_firstmate_root_home, so a second consumer can count task records across
this machine's homes without a copy. Teardown keeps its exact refusal wording
through a thin wrapper.

* feat(bin): defer spawns beyond a project's declared machine capacity

A project whose machine-local resource only serves a few workers at once had
no way to tell Firstmate so: every queued item was launched, and the surplus
workers spent full-context turns retrying the resource.

config/project-capacity in the root home now declares how many workers each
named project admits at once on this machine. bin/fm-spawn.sh counts the ship
and scout records on the same project origin across the root and its local
secondmate homes, skipping ones whose ready PR is recorded, while holding the
shared project lock through publication. A spawn with every place held exits 75
before any brief render, endpoint, worktree, record, or backlog move, so the
item stays queued; batches report it as deferred. Undeclared projects keep
today's uncapped dispatch, and an unreadable declaration refuses rather than
guessing the limit.

Refs #4237

* no-mistakes(review): Document that capacity matches the clone directory name

* no-mistakes(document): Rewrap stale fm-wake-lib root-home doc comment

* no-mistakes(review): Dedupe local state dirs by identity to avoid double-counting

* no-mistakes(document): Rewrap fm_local_firstmate_state_dirs error doc comment

* no-mistakes(ci): I fixed all four Greptile findings. All 14 tests in tests/fm-project-capacity.test.sh pass, and shellcheck at warning level is clean on the changed files. Each new test failed against the old code and passes now. - **ci-1 (spaced names):** a declaration line must give a name its capacity whenever the name is a valid clone directory name. `fm_project_capacity_lookup` now trims each line, skips blank lines and lines whose first non-blank character is `#`, and takes the last field as the capacity. Everything before that field is the name, so it may contain spaces. The old error cases still refuse: a single field is rejected, and trailing text leaves a last field that is not an integer. The library header and docs/configuration.md now say a name starting with `#` cannot be declared. New test `test_spaced_project_name_is_declared` declares `my heavy project 1` next to an indented comment line and gets a deferral. - **ci-2 (unreadable records):** the holder count must never silently leave out a holder. `fm_project_capacity_occupants` now refuses when a local home's state directory exists but cannot be read or listed, or when a `.meta` file cannot be read. The error names the path, and `fm-spawn.sh` shows it in its existing refusal message. New test `test_unreadable_holders_refuse_admission` covers an unreadable record in the root home and an unreadable state directory in a registered local secondmate home, then checks that the spawn is admitted once both are readable. The test is skipped when run as root. - **ci-3 (Orca lock):** any spawn that can become a holder for a capped project must take that project's lock. The lookup now also reports whether the declaration caps any project at all, and an Orca spawn takes the per-origin lock whenever it does. This covers every capped same-origin clone. It also covers some cases where no same-origin clone is capped, because a spawn cannot find clones under other directory names without searching for them. With no declaration file, Orca still skips the lock. The comments in the library and in the `fm-spawn.sh` header are updated. The Orca test now clones the origin as `project-2`, which has no declaration, and checks that its Orca spawn refuses while the lock is held and publishes no record. - **ci-4 (worktrees):** `assert_nothing_created` now also compares the project's `git worktree list` from before and after a deferred spawn. Both tests that call it take that snapshot first. Files changed: bin/fm-project-capacity-lib.sh, bin/fm-spawn.sh, docs/configuration.md, tests/fm-project-capacity.test.sh

* fix(bin): declare capacity for a project name that begins with #

A clone directory whose name begins with # was skipped as a comment, so that project stayed uncapped. A line is a declaration when the # is written against the rest of the name and the line ends with a capacity; a # followed by whitespace stays a comment.

* no-mistakes(document): Rewrap project-capacity library header comment

* no-mistakes(ci): Lint 2 fails because this PR's code pushes ShellCheck past its memory cap. ShellCheck ran out of memory analyzing bin/fm-teardown.sh in CI (reason=memory, rc=251, peak about 8.39 GB). On current main the same file passes at about 7.29 GB. **Cause:** the new `fm_local_firstmate_state_dirs` function in bin/fm-wake-lib.sh had a conditional `. fm-secondmate-registry-lib.sh` with a `# shellcheck source=` directive inside the function. ShellCheck followed that source again, inside a function scope, wherever fm-wake-lib.sh is sourced, and bin/fm-teardown.sh is the heaviest root that sources it. Measured locally with `shellcheck --norc --external-sources bin/fm-teardown.sh`: - current main (fd325b1b): 7.29 GB - main merged with this PR: 7.86 GB - the same merge without the in-function source: 7.27 GB **Rule this restores:** this change must not make any lint root heavier than it is on main. That function holds the only new nested source in the change. **Fix:** I removed the in-function source, which no caller needs. Both callers already load the registry library at top level before calling the function: - bin/fm-teardown.sh sources it directly. - bin/fm-spawn.sh, the only user of bin/fm-project-capacity-lib.sh, gets it through bin/fm-ff-lib.sh. I also documented the requirement in the function's comment and in the "Requires" note in bin/fm-project-capacity-lib.sh. No behaviour changes. **Verification:** - ShellCheck on head: bin/fm-teardown.sh peaks at 7.12 GB and bin/fm-spawn.sh at 6.68 GB, both with rc=0. bin/fm-wake-lib.sh and bin/fm-project-capacity-lib.sh lint clean. - tests/fm-project-capacity.test.sh, tests/fm-teardown.test.sh (102 ok) and tests/fm-teardown-endpoint-safety.test.sh all pass. Files changed: bin/fm-wake-lib.sh, bin/fm-project-capacity-lib.sh

* fix(bin): release the Herdr session lock when reclaim finishes

A concurrent resume in another home waits five seconds for that lock.
Reclaim is the last presentation change on the recovery path, so holding
the lock through the launch tail made the waiter time out. The contributions
arm check also freezes its one-second clock, the same way the budget tests
do, because an unfrozen clock can tick past before the first forge read.

* no-mistakes(review): Keep Herdr session lock through launch handoff after reclaim

* no-mistakes(review): Skip the spawning task's own record in capacity count

* no-mistakes(review): Restore release test comment above its test

* docs: scope PR-ready re-evaluation to a declared project capacity

A ready pull request frees a place only when that project declares capacity, so the always-loaded backlog contract should re-evaluate on that handoff only in that case.

* no-mistakes(document): Document new worker controls and runtime behavior

* no-mistakes(ci): Fixed both CI failures. The Pi geometry test now waits for a fresh session-start event before checking the reloaded transcript, instead of mistaking stale content for a completed reload. The Herdr contention tests now hold the shared lock longer so resume setup cannot outlast the hold. The Pi test passed locally, but its live UI cases skipped because Pi and tmux are unavailable; shell syntax and diff checks passed. The Herdr E2E could not be verified locally against CI’s pinned Herdr 0.7.4

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: Tiago <tiagop@hey.com>
Co-authored-by: lakshyads05 <lakshya@solutions21.ae>
Co-authored-by: Stephen Haptonstahl <srh@haptonstahl.org>
Co-authored-by: cloud-practitioner <113355863+cloud-practitioner@users.noreply.github.com>
Co-authored-by: M00NLIG7 <57321738+M00NLIG7@users.noreply.github.com>
Co-authored-by: Symphony <otakugamerza1@gmail.com>
Co-authored-by: Freudator86 <94322668+Freudator86@users.noreply.github.com>
Co-authored-by: Quartermaster <quartermaster@users.noreply.github.com>
Co-authored-by: YifuGu <39033099+ironerumi@users.noreply.github.com>
Co-authored-by: YifuGu <ironerumi@users.noreply.github.com>
Co-authored-by: Thoughts-One <119157221+Thoughts-One@users.noreply.github.com>
Co-authored-by: Vytautas Stankus <svycka@users.noreply.github.com>
ashkonmousavi added a commit to ashkonmousavi/firstmate that referenced this pull request Oct 8, 2026
#85)

* fix(pi): restore watcher continuity across successor gaps and make extension log opt-in (#5489)

* Fix Pi watcher successor-gap confirmations and add extension log

Accept an already-acknowledged handling confirmation as a no-op when the
generation matches, confirm the restoration's own recovery token with a
superseded (not rejected) outcome on generation mismatch, retire an arm on
confirm failure only when the failed token names that exact pid, and record
restore attempts, readiness timeouts, and confirm results in the bounded
state/.watch-extension.log. Regression tests: already-acked no-op plus
mismatch/dead-pid/lock-mismatch rejections and the manual-restart churn
contract in fm-watch-arm.test.sh, and a mid-restore marker advance with no
rejection appendix in fm-pi-watch-extension.test.sh.

* Treat a dead arm child as an empty slot so repair and retry recover

startArm and scheduleRetry answered unchanged while holding a ChildProcess
whose OS process was already gone but whose close had not fired, so neither
the repair tool nor the retry timer started anything until that close fired.
Gate slot occupancy on a liveness check (exit/signal codes plus pid probe)
and start a fresh arm instead, with a regression test driving the repair
tool against a dead-but-unclosed child.

* no-mistakes(document): Document new Pi extension log knob

* no-mistakes(review): Fix confirm-failure retire token match, add distinct-pid test

* no-mistakes(document): Clarify retire guard needs pid and generation

* Make the Pi extension diagnostic log opt-in and default-off

Only a positive FM_WATCH_EXTENSION_LOG_KEEP_LINES enables
state/.watch-extension.log. Unset, empty, non-numeric, zero, and
negative values disable logging entirely, so the default run writes
nothing and never creates the file. The shared positiveInteger
fallback semantics stay untouched for the retry and timeout knobs.
docs/configuration.md owns the knob contract and
docs/watcher-continuity.md points at it. Tests: the superseded-delivery
case runs opted in, and a new case proves unset, zero, and non-numeric
values create no log file while delivery still succeeds.

* no-mistakes(document): Qualify extension-log coverage bullet as opt-in

* no-mistakes(ci): The two reported checks (CI run 36372002913, Require no-mistakes run 36372069208) show conclusion action_required with 0 jobs and no logs because this is a fork PR (RibatTRW/firstmate) and GitHub is holding the workflow runs awaiting maintainer approval; that gate is external to the code and needs a maintainer to approve the runs. While verifying the change locally I found a real defect the approved CI run would hit: the PR's new test test_handling_delivered_rejects_a_superseded_generation failed deterministically. Invariant violated: reopen-announced (a non-successor/manual arm start) mints a fresh recovery generation only when the durable wake queue holds unrecovered work; an announced episode with an empty queue must be left untouched so idle arm starts never churn generations (the [ -s queue ] guard from #4819, relied on by bin/fm-watch.sh:2413 and covered by the append-reopens and announcement-bound sibling tests). The test called reopen with an empty queue and expected churn, so the fix establishes the queued-work precondition (append one wake, re-announce, re-read the generation) before asserting the reopen mints and the old confirmation mismatches. Test-only change, 14 lines in tests/fm-watch-arm.test.sh. Verified: fm-watch-arm.test.sh 25/25 ok on two consecutive runs, fm-pi-watch-extension.test.sh 55/55 ok, and shellcheck reports only one pre-existing warning outside the edited region

* no-mistakes(document): Restore blank line in watcher-continuity docs

* Route superseded Pi deliveries like confirmed ones and cover the retire guard

A superseded handling confirmation now falls through to the normal delivery
path, so an accepting supervision branch owns the wake instead of main.
Scope the watcher-continuity token-pinned confirmation and narrowed retire
rule to Pi, since omp and OpenCode still confirm against the current
successor. Add tests that fail when the retire guard, the scheduled-retry
gate, or the deferred-close gate is reverted, relabel the churned-generation
characterization test, and use a reaped pid for the dead-pid rejection.

* fix(pi): hide duplicate assistant finals from hidden processing retries (#5863)

* fix(pi): silence unacknowledged processing retry replies

Suppress autonomous processing prose before persistence and during streaming while retaining tool calls, signed reasoning, and retryable outcomes. Restore ordinary output after acknowledgement or a user message.

Fixes #4954

* no-mistakes(review): Silence only processing retries, keep first presentation visible

* fix(pi): preserve differing processing retry replies

* test(pi): accept Pi 1.0.1's renamed HTML export renderer lookup (#6530)

Pi 1.0.1's createToolHtmlRenderer reads getToolRenderers and ignores
getToolDefinition. Calm /export still includes stock grep HTML; the
fixture has to pass the lookup key the installed Pi actually reads.

* fix(bin): tolerate transient quota read failures (#6490)

* fix(procevent-quota): tolerate consecutive slow quota-axi reads

The quota allowance poll treated any quota_json failure as terminal, so one
slow quota-axi --json (measured max ~29s under a 48s derived bound) shut the
watch down until someone re-armed it, and the detail always said
"missing/incompatible". Tolerate three consecutive failed or timed-out reads
before going terminal, reset the streak on any good read, and report a
timeout distinctly from a missing or incompatible tool. Each timed poll runs
exactly one bounded --version and one bounded --json: validate the captured
version text through fm_quota_axi_version_compatible rather than launching a
second probe, and describe a mixed failure streak by count plus last cause.

* no-mistakes(document): Document quota polling failure tolerance

* no-mistakes(ci): Fixed ci-2 and ci-3. Permanent quota read failures (rc 2 missing, rc 3 incompatible) now report on the first poll, while transient rc 1/4 failures retain the existing three-failure retry behavior. The missing-binary test now uses an isolated PATH without quota-axi and asserts both permanent failures stop at condition_polls: 1. Verification passed: tests/fm-procevent-quota.test.sh, canonical fast lint for both changed files, bash syntax checks, and git diff --check

* no-mistakes(review): Classify untimed quota version failures as transient

* no-mistakes(document): Clarify quota polling failure budget

* fix(bin): clarify scratch guidance and dirty teardown refusals (#6505)

* fix(teardown): clarify scratch guidance and dirty worktree refusals

Keep ship proof material outside the task worktree and distinguish untracked-only leftovers from tracked edits without changing cleanup guards.

Fixes #6319

* fix(ci): Fixed ci-3 only. Both promotion outputs now replace the scout restriction and require external scratch storage and a clean worktree before done. Verification: 27 delivery tests passed, five mutations caught, restored test passed, pinned ShellCheck and syntax/whitespace checks passed. ci-1, ci-2, and ci-4 remain untouched

* fix(bin): escalate inbox instructions blocked by busy workers (#6518)

* Escalate inbox instructions stuck behind a busy worker

Count consecutive busy-deferred due doorbells durably and escalate at the configured bound without typing into the worker pane.

Fixes #6445

* fix(review): Fix inbox escalation deduplication and busy streak resets

* fix(review): Preserve busy inbox escalations through daemon supervision

* fix(document): Correct busy-inbox escalation documentation

* fix(ci): Fixed SC2034 in tests/fm-task-inbox.test.sh by including the loop counter in the failure diagnostic. Source-aware lint, all 34 inbox tests, and git diff --check pass. Behavior portable serial 6 reproduces identically on base 1f3e7696 and target 78156b86 with Pi 1.0.1: an unrelated renderer API change breaks the unchanged Calm test. No Calm changes made; that failure is addressed separately by https://github.com/kunchenguid/firstmate/pull/6516. Logs retained in scratchpad-ci/

* fix(ci): Fixed ci-2, ci-3, and ci-4: successor failures surface, reset alerts deduplicate, and oversized busy limits fall back to two. Passed 47 inbox tests, 7 focused daemon checks, all 13 mutation checks, lint, documentation checks, and diff checks. Evidence: scratchpad-ci-selected/summary.json. ci-1 remains unchanged and unwaived. Fresh live Herdr proof remains with the outer driver

* fix: prevent unnecessary remote worker turnover (#6431)

* fix: prevent healthy remote job worker turnover

* no-mistakes(review): Serialize full LaunchAgent repair and verify launchd-tracked owners

* no-mistakes(review): Let launchd-tracked unpublished spawns start before reloading

* no-mistakes(document): Document remote worker heartbeat and serialized LaunchAgent recovery

* no-mistakes(ci): Full CI log showed the idle-worker regression exceeded its outdated command budget (82 versus 80) after independent heartbeat ownership checks were added. Raised the budget to 120 while retaining the separate busy-poll sleep limit. Remote-job and LaunchAgent executable tests passed, as did bash syntax validation and git diff --check. No production behavior changed

* no-mistakes(ci): Fixed missing readiness recovery under verified live ownership, preserving the serving PID and private file mode. Added executable regressions for deletion during a blocked sweep and stale readiness diagnostics without LaunchAgent reload. Deletion regression failed before the fix. Both remote-job test suites, bash syntax validation, and git diff --check passed

* no-mistakes(ci): Full CI log identified a flaky ownership-loss test racing an already-authorized heartbeat refresh. Replaced backdating and a fixed sleep with bounded observation of readiness expiry through the public probe. Production behavior unchanged. Remote-job and LaunchAgent executable suites passed; bash syntax validation and git diff --check passed

* fix: wait for launchd bootout cleanup

* no-mistakes(document): Document remote worker recovery and read-only turnover verification

* no-mistakes(review): Publish worker identity before lock owner records

* no-mistakes(document): Document worker identity publication safety invariant

* no-mistakes(ci): Fixed ci-1: replacement workers discard predecessor readiness before publishing identity and roll back identity if lock-owner recording fails. Added executable regressions reproducing both defects. LaunchAgent and orphan-reap suites, shell syntax checks, and git diff --check passed. No live service state was modified

* no-mistakes(ci): Restored lock-owner-before-identity publication and removed the identity rollback and reordering-only tests. Retained an executable regression proving predecessor readiness is rejected until replacement startup completes. LaunchAgent and orphan-reap suites, shell syntax checks, and git diff --check passed. Other changes remain intact; the retained-identity interrupted-repair edge remains out of scope. No live service state was modified

* fix: speed up remote-job sequence claim cleanup (#6575)

* fix(remote-job): reap expired seq claims with one directory walk

The hourly claim sweep forked uname+stat per .seq-claims entry and blocked
serving for ~85s at ~17k dirs. Delete expired empty claim dirs with a single
find -exec rmdir batch and cache the host uname for remaining mtime reads.

* no-mistakes(review): Restore original path mtime helper and drop uname cache

* no-mistakes(test): Restore claim retention eligibility; focused retention and serving tests pass

* no-mistakes(document): Document single-walk claim cleanup and regression entrypoints

* no-mistakes(ci): Fixed Lint 2’s reproduced SC1091 by adding the tests/lib.sh ShellCheck source directive to the retention test. Runtime behavior is unchanged. ShellCheck passed for both new claim tests; Bash syntax, the retention behavior test, and git diff --check passed

* fix(tests): keep fixture registries out of git worktree roots

A TMPDIR pointed at a repository root placed live .fm-test-* registries
beside tracked files, and a concurrent git add during the claim-walk CI
fix round committed three of them. Route registries and fixture roots
through a TMPDIR that refuses git worktree roots, remove the stray files,
and pin the escape with a behavioral cleanup test.

* no-mistakes(review): Preserve whole-second claim expiry in single-walk sweep

* no-mistakes(document): Correct temporary-directory resolution documentation

* no-mistakes(ci): Fixed ci-3 by changing only the stale, fresh, and read-only orphan fixture paths in tests/fm-test-fixture-cleanup.test.sh to use $FM_TEST_TMPDIR. All seven tests passed both normally and with TMPDIR set to the worktree root. Shell syntax and git diff --check passed

* feat(bin): record no-mistakes pipeline spend per task at teardown, opt-in (#5354)

* feat(bin): record each task's no-mistakes pipeline spend at cleanup

no-mistakes keeps every pipeline agent invocation's token usage only in its
local agent_invocations records, and cold pipeline agents (review, test,
document) leave no session log, so a task's review-loop cost never reached
Firstmate's records and could not be attributed after cleanup.

bin/fm-pipeline-spend.sh attributes a task's runs by the repository
no-mistakes resolves for the task copy, the task branch, and the branch's
creation (a relaunch mints a new spawn_gen while the branch keeps
validating), then sums every invocation, failed and cancelled included.
Token fields use no-mistakes' per-round deltas so resumed review rounds are
not counted twice, and unrecorded values stay unknown rather than zero.
show prints the record read-only; record appends it once per task
incarnation to data/pipeline-spend.jsonl. Teardown records it for every
ship task it cleans up, before deleting the branch and the task record.

fm_nm_state_db becomes the one owner of where no-mistakes' state database
lives, shared with the capped run-inventory reader.

* no-mistakes(review): Drop spend show command and timeout override

* style(bin): rewrap fm-pipeline-spend header comment

* Make pipeline spend recording opt-in

* feat(bin): record each task's no-mistakes pipeline spend at cleanup

no-mistakes keeps every pipeline agent invocation's token usage only in its
local agent_invocations records, and cold pipeline agents (review, test,
document) leave no session log, so a task's review-loop cost never reached
Firstmate's records and could not be attributed after cleanup.

bin/fm-pipeline-spend.sh attributes a task's runs by the repository
no-mistakes resolves for the task copy, the task branch, and the branch's
creation (a relaunch mints a new spawn_gen while the branch keeps
validating), then sums every invocation, failed and cancelled included.
Token fields use no-mistakes' per-round deltas so resumed review rounds are
not counted twice, and unrecorded values stay unknown rather than zero.
show prints the record read-only; record appends it once per task
incarnation to data/pipeline-spend.jsonl. Teardown records it for every
ship task it cleans up, before deleting the branch and the task record.

fm_nm_state_db becomes the one owner of where no-mistakes' state database
lives, shared with the capped run-inventory reader.

* no-mistakes(review): Drop spend show command and timeout override

* style(bin): rewrap fm-pipeline-spend header comment

* Make pipeline spend recording opt-in

* no-mistakes(document): Note pipeline-spend opt-in gate in teardown comments

* no-mistakes(ci): I fixed all three Greptile findings following your decision. Both test suites pass: tests/fm-teardown.test.sh (105 ok, exit 0) and tests/fm-pipeline-spend.test.sh (8 ok). I didn't run either new test against the old code, so the claim that they fail before the fix is from reading the code, not a run. Nothing was committed or pushed. **ci-1 / ci-2 (bin/fm-teardown.sh):** The rule is that every owned ship task in a home that has opted in gets a durable spend record before its task record is deleted. The `[ -d "$WT" ]` check skipped the recorder when the worktree folder was already gone, so no record was written. I removed that check and kept the other conditions: the task must be a ship task, teardown must own its worktree, and `config/pipeline-spend` must exist. `bin/fm-pipeline-spend.sh` already writes an unavailable-source line when the worktree is missing. The new test `test_teardown_records_unavailable_spend_for_a_gone_worktree` in tests/fm-teardown.test.sh sets up an owned ship task whose worktree is missing, with recording turned on. It checks that teardown succeeds, writes a line with `source=="unavailable"`, `total==null` and a reason saying the task copy is gone, and still removes the task record. Under the old check no ledger would have been written. **ci-3 (bin/fm-nm-run-lib.sh):** When `NM_HOME` and `HOME` were both unset, `fm_nm_state_db` fell back to `${HOME:-}/.no-mistakes`, which resolves to `/.no-mistakes`. It now uses `root=~/.no-mistakes`. Bash expands `~` to the account's home directory even when `HOME` is unset, which matches the previous Python `Path.home()` lookup. The new test `test_state_db_without_nm_home_or_home_uses_the_account_home` in tests/fm-pipeline-spend.test.sh runs the function with both variables unset and compares the result to the home directory from the system's password database. The old code would have returned `/.no-mistakes/state.sqlite`. shellcheck reports nothing new on the changed files. I added one `disable=SC2016` comment in the new test, where the single-quoted `$1` is meant to expand in the child shell

* feat(bin): start tasks from a named base branch (#6442)

* feat(bin): start tasks from a named base branch

Spawns always reset a task's pooled copy to origin's default branch,
so work that belongs on a feature, integration, or release branch
started from the wrong code and opened its PR against the default.

fm-brief.sh --base-branch records the base in the brief, fm-spawn.sh
resets the copy to origin/<base> and records base_branch= in task
meta, and the worker targets its PR at that branch. Review diffs,
the cleanup content check, and scout promotion read the recorded
base. local-only and Gerrit deliveries refuse a named base.

* no-mistakes(review): Anchor brief base-branch parse on scaffold Setup line

* no-mistakes(review): Require explicit --base-branch on spawn, matching brief lines

* no-mistakes(document): Note base-branch reset in fm-spawn freshness header

* no-mistakes(ci): I fixed all four Greptile findings the way you specified. The affected suites pass locally: fm-brief, fm-dod-lib, fm-spawn-pool-base-freshen, fm-review-diff, fm-teardown and fm-task-delivery. shellcheck on the changed files reports only info-level notes, no warnings or errors. I didn't revert the code to watch each new test fail before the fix. - **ci-1 (brief text blocked or redirected spawns):** In bin/fm-dod-lib.sh, `fm_brief_base_branches` now only accepts a `Base branch: <value>` line that comes directly after the base-variant Setup sentence fm-brief.sh writes. Any other `Base branch:` line is ignored. `--base-branch` is still the only authority, and the existing checks in fm-spawn.sh are unchanged. With the flag, the spawn is refused unless a matching pair exists and no pair names a different branch. Without the flag, it is refused only if such a pair exists. The header comment in fm-spawn.sh is updated to match. - The `brief_with_base` test helper now writes the real two-line pair. - The decoy refusal case now uses a full decoy pair in the captain's intent. - New test `test_prose_base_branch_line_is_ignored`: a brief with no base whose intent quotes `Base branch: release/1.2` spawns normally with no `base_branch=` recorded. A spawn with `--base-branch feature/hub` whose intent names release/1.2 starts from `origin/feature/hub` and records it. - **ci-2 (scout base not checked against the forge):** fm-spawn.sh now looks up the project's registered forge with `bin/fm-project-mode.sh --forge`, like the ship path does, before `fm_base_branch_valid` runs. Previously it passed `none`. New test `test_scout_base_branch_refused_on_gerrit_forge`: a scout with a base on a forge=gerrit project is refused at spawn and no meta file is written. - **ci-3 (base not shell-quoted in worker commands):** `fm_dod_block` now quotes the base with `printf %q` in the `--base` and `--base-branch` arguments, and the branch name in the prose stays readable. The test in fm-brief.test.sh uses the base `release/$HOTFIX` and checks that both direct-PR and no-mistakes briefs render `release/\$HOTFIX` in the commands. - **ci-4 (ambiguous PR sentence):** The direct-PR sentence now reads "open a PR with `gh-axi` that is ready for review, not a draft, against the base branch `X` (`--base X`), not the repository default." A brief with no base renders exactly as before. The existing assertion in fm-brief.test.sh is updated

* no-mistakes(ci): Both real failures on PR 6442 (run 37064756768) come from CI tooling and a Pi version bump, not from the base-branch feature. Fixes for both already exist on upstream main, so I applied those two commits to the working tree with `git cherry-pick --no-commit`. Nothing is committed or pushed, and no feature file changed. - **Lint 2:** ShellCheck ran out of memory on `bin/fm-teardown.sh` (rss about 8 GB, exit 251, "shellcheck: out of memory"). The rule that broke is that the lint gate must finish on any valid shell root without exceeding the runner's memory ceiling. Upstream commit d719ef3 (#6443) fixes this in `bin/fm-lint.sh`: a root that hits the memory ceiling is retried without `--external-sources`. It also updates `tests/fm-lint.test.sh` and `docs/fm-test-portable-shards.md`. - **Behavior portable serial 6:** `tests/fm-calm-pi-extension.test.sh` failed with "grep disappeared from /export calm.html HTML while calm mode was on". Pi 1.0.1 renamed its HTML export renderer lookup, and upstream commit fede619 (#6530) updates the test to accept the new name. This change is test-only. Both commits applied cleanly. The four touched files are `bin/fm-lint.sh`, `docs/fm-test-portable-shards.md`, `tests/fm-lint.test.sh` and `tests/fm-calm-pi-extension.test.sh`. None of them is in the feature diff, so the feature contract is unchanged. **Local checks:** - `tests/fm-lint.test.sh` passes. - `bin/fm-lint.sh` on `bin/fm-teardown.sh`, `bin/fm-spawn.sh` and `bin/fm-dod-lib.sh` exits 0 with full ShellCheck analysis. - `tests/fm-calm-pi-extension.test.sh` exits 0 with 7 ok and 0 not ok. One case skipped because the Pi package isn't installed locally, so the Pi 1.0.1 export path that failed in CI couldn't be reproduced here. CI needs to confirm it. - I didn't reproduce the 8 GB OOM locally. The PR has to be updated through the pipeline with a normal push: no force push, no replacement PR, no merge

* fix(bin): name the Firstmate skill file as a fallback in the worker role (#6647)

* fix(bin): point project workers at the Firstmate skill file

The Skill tool cannot resolve a Firstmate skill from another project's worktree, so the launch role names the readable skill file instead.

* no-mistakes(review): fix(bin): name Firstmate skill file as fallback only

* feat(bin): retire a contribution whose forge object is permanently gone (#6655)

* feat(bin): retire a contribution whose forge object is permanently gone

Add fm-contributions.sh retire <task> <url> <captain|fleet> <reason>,
which records actor, reason and time on the saved record and removes
that task/url pair from known, poll rotation and coverage even while a
backlog link remains. Repeating a retire keeps the first provenance;
an unrecorded pair, unknown actor, empty reason or unacknowledged
pending signal is refused.

* no-mistakes(review): Keep retirement per task when settling final owners

* no-mistakes(ci): Made the three Greptile fixes the captain chose. ci-1, retirement needs the captain's word: `retire` now accepts only the actor `captain`. Any other actor is refused, including `fleet`. I removed `fleet` from the usage line, the script header (`bin/fm-contributions.sh`), the check in `bin/fm-contributions.sh` and the retired-record validation in `bin/fm-contributions.jq`, which now requires `.actor == "captain"`. The script header now says a retirement always records the captain's word, because the script cannot verify who runs it and the authority to retire is the captain's. The Bearings skill line in `.agents/skills/bearings/SKILL.md` now says `retire` records the captain's word. In the tests, a `fleet` retire is now a refusal case and must leave the record unretired. Every other retire call in the tests uses `captain`. The PR description has not been changed yet: the ruling asks for the same captain's-word statement there, and that belongs to the PR phase. ci-2, blank reasons: the reason check is now `[ -n "${5//[[:space:]]/}" ]`, so a reason made only of spaces or tabs is refused. I added a refusal test that passes a space-tab-space reason. The header also lists a blank reason among the refusals. ci-3, the gone-object test: the test forge wrapper has a new `not-found` fault that returns HTTP 404 for every `api repos/o/r/...` read. The happy-path test `test_retire_ends_observation_of_a_gone_contribution` now uses it in place of the 502 `down` fault, so it models a permanently gone repository rather than an outage. Verification: `bash tests/fm-contributions.test.sh` passed with exit 0. The last lines of its output include the three retire tests passing and no failures. `shellcheck` flagged only SC2034 (`FM_WAKE_QUEUE` unused) and SC1091 (an unfollowed source of `bin/fm-path-lib.sh`), neither on a line this round changed

* test: extend remote-reply whole-log recapture waits (#6639)

* test: give the remote-reply whole-log recapture a longer wait

* no-mistakes(review): Extend both recapture waits and simplify retry handling

* fix(bin): reopen a pending-reply escalation after its resolve (#6654)

* fix(bin): reopen a pending-reply escalation after its resolve

A retry of an open decision still appends nothing. A later same-kind escalation after a resolved line for that key appends again, so the next loss is visible.

* no-mistakes(ci): Both Greptile findings are fixed in the working tree; the four directly affected test suites pass, and a wider run of related suites had not finished when I returned this result. Invariant: a line already in the status file is recorded, and only a caller with its own evidence of a new episode may append it again. It must hold for every caller of status_event_recorded: the continuity break and the other call in bin/fm-procevent-remote-reply.sh, fm_parent_channel_append_once, and the pending-reply escalation. Changes: - bin/fm-classify-lib.sh: status_event_recorded is back to the idempotent retry check. It returns at the first matching line, and a later resolved line no longer makes that line look new. This fixes ci-1 for every caller and restores the early exit for ci-2, with no separate scan optimization. - bin/fm-pending-reply-lib.sh: _fm_pending_reply_maybe_escalate_locked now owns the reopen. It appends a new blocked line when the record already has an escalated_epoch (it escalated before and was reset) and the keyed decision is no longer open. Otherwise it uses status_event_recorded, so a retry or a resend that leaves the decision open appends nothing. - Comments on status_event_recorded and the pending-reply header describe this scope. Tests: - tests/fm-pending-reply.test.sh: the regression for escalate, operator resolve, second escalate is unchanged and passes. - tests/fm-remote-reply.test.sh: new regression drives the real reader. After an operator resolve of remote-reply-continuity-ios, a repeated handle and ingest of the same break appends nothing and leaves the decision closed. I placed it beside the existing continuity-break test, not in the pending-reply file, because only that file has the reader harness. - tests/fm-classify-corr-token.test.sh and tests/fm-classify-decision-key.test.sh: the cases that expected a blocked line to append again after a resolve now assert it stays recorded. Verification: - fm-classify-decision-key, fm-classify-corr-token, fm-pending-reply and fm-remote-reply test suites all exit 0 with the fix. - With the two bin files reverted to the commit under review, the new continuity regression and the updated parent-publisher case both fail, so they reproduce ci-1. - shellcheck -x on the changed files reports nothing. - Not finished: a background run of every other suite that mentions the parent channel or pending replies had produced no output, pass or fail, when I returned. One known gap: the escalation writes escalated_epoch before phase. If the process dies between those two writes and an operator resolves the key before the next tick, that tick appends a blocked line for the same episode. I left it, because closing it needs new state. Issue 4755 and other escalation behavior are untouched. Nothing is committed

* fix: refuse a confirming Enter on the Claude background-task exit picker (#6666)

* fix: refuse a confirming Enter on the Claude exit picker

The background-task picker still classifies as pending, so a retried Enter confirms "Exit and stop tasks".
Stop after the Enter that opened it, and raise the existing stale wake with the dialog name.

A non-paused secondmate still skips that wake, so a secondmate parked on the picker still looks idle.
Model-downgrade confirmation, MCP approval, and any other Claude exit confirmation are not covered, because there is no recorded screen for them.

* no-mistakes(review): fix: anchor exit picker match and wake second mates

* fix: refuse a typed submit while the Claude exit picker is open

A pane that already shows the picker must not receive the message or a confirming Enter.
The match requires the recorded heading and footer lines, and it is skipped when no dialog name is being recorded.

* no-mistakes(review): restore watcher to main behaviour, drop dialog wake

* no-mistakes(test): test: align Herdr picker fixtures with the preflight read

* no-mistakes(document): document exit refusal on a recognised dialog

* fix: remove the dialog file when exit runs in a subshell

do_exit is invoked from a command substitution, so the parent cleanup never saw the path it is supposed to delete.

* no-mistakes(review): fix: set the dialog file path after the control lock

* no-mistakes(document): document why the dialog file path follows the lock

* no-mistakes(ci): The exit cleanup in bin/fm-control.sh now deletes the dialog file first and releases the control lock second. The changes are in the worktree and are not committed. Invariant: only the process that holds the control lock for a task may write or delete that task's dialog file ($STATE/<id>.composer-dialog, the file that carries the picker name from the composer read to the Enter checks). The old cleanup released the lock and then deleted the file, so the delete ran outside the lock. Places where this invariant must hold: control_cleanup is the only place that deletes the file, and the single assignment after CONTROL_LOCK_HELD=1 is the only place that names it. do_exit only truncates the file, and it runs while the parent holds the lock. A process that loses the lock never sets the path, so it deletes nothing (earlier fix, unchanged). No sibling site needed a change. Fix: in control_cleanup, the rm of the dialog file moved above the fm_lock_release call, with a two-line comment that gives the reason. No dialog recognition or supervision code changed. Test: added test_exit_removes_the_dialog_file_before_releasing_the_lock in tests/fm-control-relaunch.test.sh. It runs one `fm-control exit` with a recording rm on PATH. The lock release removes paths at or under the control lock with rm, so the recording rm writes whether the dialog file exists at that moment. The test asserts the last record is "absent". It starts no second command. Verification, all run locally: - Before the fix, the new test failed with "the dialog file must be gone when the control lock is released, got: present". - After the fix, tests/fm-control-relaunch.test.sh exits 0 with 75 ok lines and no "not ok" line, including the new test and the existing test that the file is gone after exit and after relaunch. - tests/fm-control.test.sh exits 0 with 44 ok lines. - shellcheck -x on both changed files exits 0 with no output. One thing I noticed and did not change: tests/fm-control-relaunch.test.sh prints 615 "rm: cannot remove ... Permission denied" lines during its temp-directory teardown. The count is the same with my change reverted, so this change does not cause it, and the suite still exits 0. I could not read the Greptile check log (the check run was not found), so I worked from the finding text and the user instructions

* fix: restore timeout watchdog compatibility with macOS Bash 3.2 (#6028)

* fix: support stock macOS Bash in timeout watchdog

* no-mistakes(ci): Fixed ci-1 in tests/fm-timeout-lib.test.sh: both startup-owner failure paths now kill the bounded command group before killing the watchdog. Fault injection reproduced both leaks before the fix and confirmed cleanup afterward. Bash 3.2 suite, syntax check, ShellCheck, and diff checks passed; GNU timeout coverage skipped because the binary is unavailable. Production code unchanged

* fix(project-management): use subshell form for Initialize command (#6699)

Wrap the cd command in a subshell to comply with the cd-guard policy
that blocks persistent top-level directory changes in the primary
firstmate checkout. The subshell form (cd projects/<name> && ...) is
accepted by the policy as documented in issue #6502.

Fixes #6502

* feat(bin): add armable daily startup growth check (#6725)

* Add daily startup growth check

* no-mistakes(review): watch printed startup memory, retain growth baselines, drop knobs

* no-mistakes(review): delegate budget verdict, report before publish, pin shim home

* no-mistakes(review): baseline first sightings silently, drop mtime, spare secondmates

* no-mistakes(review): cap the wake line, validate budget verdict fields

* no-mistakes(review): guard record schema, check appends, tighten assertions

* no-mistakes(review): drop arbitrary tracked library, make re-arm idempotent

* no-mistakes(review): correct tracked-set wording, restore shim guards, register artifacts

* no-mistakes(review): exit on signal instead of publishing partial record

* no-mistakes(document): correct startup-growth record removal cost in state registry

* no-mistakes(test): sweep orphaned empty startup-growth temp records on due evaluation

* no-mistakes(document): note watcher need for armed startup growth check

* no-mistakes(ci): Diagnosed "Behavior portable serial 5": the shard failed on tests/fm-contributions.test.sh in test_arm_plumbs_a_configured_budget_into_the_check_shim (inherited mode) with "generated check did not attempt a read" and a missing forge/calls file. Root cause: bin/fm-contributions.sh:379-382 computes DEADLINE=$(date +%s)+BUDGET and then gates each read on DEADLINE-$(date +%s) >= OBSERVATION_RESERVE (= min(BUDGET,15)). At the one-second budget this test uses, that gate demands zero elapsed whole seconds, so a wall-clock second boundary crossing between the two date calls makes the poll loop break before any forge read. The test calls wrap_forge (which installs a fake date reading $FORGE/clock only when that file exists) but never seeded forge/clock, so it ran against the real clock. The repo already documents this exact hazard at tests/fm-contributions.test.sh:675-677 and freezes the clock in every other budget-constrained test: test_budget_exhaustion_keeps_prior_record (both modes), test_genuine_failure_near_deadline_is_unavailable, test_reservation_defers_later_url_when_fifteen_seconds_do_not_remain, test_slow_read_deadline_kill_is_budget_refusal, test_budget_is_cut_down_to_the_watcher_check_bound. test_arm_plumbs was the only omission; its single loop body covers both the configured and inherited modes. The remaining wrap_forge users run the default 20-second budget (5 seconds of slack above the reserve) and are not reachable by this race. Fix (smallest, matching the sibling idiom): added the clock freeze `/bin/date +%s > "$home/forge/clock"` plus a two-line comment inside that test's loop, before the poll. Three added lines in tests/fm-contributions.test.sh; no production code and no other file touched. Verification: reproduced the exact CI failure message by running a scratch copy whose fake date advances one second per read (worst case of the unfrozen clock) - missing forge/calls, "not ok - generated check did not attempt a read". With the fix, the single test passes 10/10 standalone and the full tests/fm-contributions.test.sh passes all 46 assertions with exit 0. tests/fm-startup-growth-check.test.sh still exits 0. bash -n and shellcheck -S warning are clean on the edited file; scratch repro files were removed, leaving only the intended three-line change. Note for the author: this flake is not caused by this branch - tests/fm-contributions.test.sh is untouched by the PR and the race is pre-existing, surfacing only on a loaded runner (the shard's neighbouring assertions show 80+ second gaps). It was fixed because it is genuine nondeterminism with an established in-repo remedy and the check cannot otherwise go green. No work was done on the unselected greptile findings (ci-2..ci-5) or the cancelled "Behavior portable serial 2" check (ci-6)

* no-mistakes(ci): Diagnosed "Behavior portable serial 6": the omitted log section (fetched with gh api) shows the shard failed on tests/fm-calm-pi-extension.test.sh in test_hidden_block_geometry_e2e with "Pi Calm hidden-block geometry E2E did not complete the /reload viewport transition" (exit=1, duration_ms=23437; the family line pure-contract-unit count=3 duration_ms=51193 failed=1 matches devin-harness 3203 + calm-pi 23437 + task-delivery 24553). Root cause: tests/fm-calm-pi-extension.test.sh:2831 wait_for_geometry_transition detected the /reload by SAMPLING a single intermediate frame - it polled tmux capture-pane for Pi's transient "Reloading keybindings, extensions, skills, prompts, themes, and context files..." box and only accepted the rebuilt transcript afterwards (elif gated on saw_transient). Pi shows that box only while the reload runs and replaces it via dismissReloadBox when done, so on a loaded runner the box can live entirely between two polls; saw_transient then stays 0 forever and the 600-attempt wait times out although the reload fully succeeded. Reproduced locally under CPU oversubscription with CI's Pi 1.0.4: 1 failure in 8 runs, diagnostics proving the mechanism ("DIAG: transition timeout saw_transient=0 attempt=600") while the captured viewport already showed the completed reload status row and the intact collapsed transcript. Not a Pi-version regression: 1.0.4's handleReloadCommand and both banner strings are identical to the installed 0.85.1, and shard 6 of the previous pipeline run (job 112490351105) passed this same script on Pi 1.0.4. Invariant violated: a TUI E2E assertion must key off a durable state the UI retains, never off one intermediate frame polling may miss. Sibling enumeration: the helper was defined once and called once; grep for transient/saw_/two-stage waits across tests/, bin/ and .pi/ found no other one-shot-frame wait - every other wait in the file keys off durable text presence/absence or off continuously repeating animation frames (the working-ship motion checks) where sampling eventually connects; no doc or other test referenced the helper. Fix (smallest, one helper): renamed it wait_for_geometry_reload and made it wait for the durable status row Pi appends once the reload completed and the chat was rebuilt ("Reloaded keybindings, extensions, skills, prompts, themes, and context files", a persistent chat child via showStatus, present in both 0.85.1 and 1.0.4) together with CALM_GEOMETRY_FINAL. That marker is absent before the reload and never appears if the reload throws, so the check is strictly stronger than before (a failed reload now fails the test instead of passing on transient-then-final). One file changed: tests/fm-calm-pi-extension.test.sh, 2 hunks, no production code. Verification: deterministic before/after - with polls spaced 0.4s apart (the loaded-runner worst case) the pre-fix helper fails with exactly the CI message and the post-fix helper passes; 12/12 consecutive passes of the isolated test under load with Pi 1.0.4; full file 15/15 ok exit 0 under LANG=C.utf8 (matching the runner locale); bash -n, shellcheck -S warning and bin/fm-lint.sh clean; scratch repro files and the temp Pi 1.0.4 install removed, git status shows only the one intended modification. Notes: (1) the flake is not caused by this branch - that test file is untouched by the PR and the race is pre-existing; fixed because it is genuine nondeterminism and the check cannot otherwise go green. (2) Under far harsher load than CI applies a separate unrelated flake appeared in the same test: Pi's "Warning: tmux extended-keys is off..." row occasionally lands between the collapsed [skill] ahoy row and the final response, making assert_geometry_gap see 5 rows instead of 2. Different cause, never seen in CI, and its plausible remedies would change key delivery for the other TUI tests in the file, so it was left alone rather than growing this fix

---------

Co-authored-by: Quartermaster <quartermaster@users.noreply.github.com>

* fix(bin): reopen the remote-reply continuity decision on a later break (#6708)

* fix: reopen a remote-reply continuity break after repair

A later break for the same route and reason was swallowed after the
operator resolved the first one, because the status line matched for
the life of the log. The continuity ingest now appends again when the
cursor has moved or retirement has reset that episode, and an unchanged
re-read still appends nothing.

status_event_recorded is unchanged. Its other callers are the
pending-reply escalation, which already decides its own episode, the
parent-channel note append, and the remote document transfer note.

* no-mistakes(review): seed continuity episode for already recorded break line

* no-mistakes(document): document when a remote-reply continuity break reopens

* no-mistakes(ci): The adapter now stores the continuity episode record before it appends the `blocked` line, so a failed store appends nothing. The changes are uncommitted in the worktree, in `bin/fm-procevent-remote-reply.sh` and `tests/fm-remote-reply.test.sh`. **Invariant:** a continuity `blocked` line is on the parent status log only when the episode record for that break is already stored. Only the `continuity-broken` branch of `cmd_ingest` appends that line, so the fix is at that one place. **What changed in `cmd_ingest`:** - It decides first whether the break needs a line, then stores the episode record, then appends. - If storing the record fails, it stops with "cannot record continuity episode" and appends nothing. - If the line is already on the log and no episode record exists, it stores the record and does not append. **One addition you did not ask for:** if the append fails after the record is stored, the adapter puts the earlier record back (or deletes the new one when none existed). Without that, the retry would read as the unchanged repeat and append nothing, which is the issue 6701 failure again. **Deviation from your test instruction:** the test does not use `chmod` on the cursor directory. `write_continuity_episode` runs `chmod 700` on that directory before every write, so a read-only directory is made writable again. The test instead makes `mktemp` fail for the episode's temporary file in that directory, the same way the existing receipt-failure test does. **Tests added to `tests/fm-remote-reply.test.sh`:** - Store failure: the second break exits 1, appends no `blocked` line and opens no decision. One retry after storage recovers exits 3 and appends a single line. - Append failure: with the status log read-only, the third break exits 1 and appends nothing. One retry after the log is writable exits 3 and appends a single line. **Verification:** `bash tests/fm-remote-reply.test.sh` ends with "ALL TESTS PASSED" with the fix. Against the script at commit b750c828 the same test file fails at "a continuity break appended its line before its episode was stored". `shellcheck -S warning` reports only an unused loop variable at line 667 of the test file, which this change does not touch. I ran no other test files. The Greptile Review check log could not be retrieved, so I worked from the finding text alone

* fix: record a continuity break's reader position on its status line

A later break at another cursor is then a different line, so the existing
duplicate check appends it and reopens the decision. An unchanged re-read
builds the same line and appends nothing.

* fix: reopen a continuity break after an identical restore

A retirement that puts the same bytes back used to rebuild the recorded line, so the later break stayed closed. The retirement count on that line makes the later break distinct.

* no-mistakes(review): remove continuity match for full-prefix line without retirement count

* no-mistakes(document): clarify what a continuity break status line records

* fix: remove the reply cursor before recording retirement

A stop between those steps must leave the count unchanged, so an unchanged continuity break still builds the same line.

* fix(control): drop busy_gen from the task record when an incarnation is retired (#6733)

* fix(control): drop busy_gen when an incarnation is retired

A deliberate exit removed the busy sidecar and left busy_gen in the task record, so the two records disagreed about whether that incarnation was still observable.

* no-mistakes(review): drop GNU-only chmod and unreached sidecar-absent branch

* no-mistakes(review): correct lock comment to name the deadlock

* no-mistakes(ci): The test `test_exit_drops_meta_busy_gen_with_the_sidecar` in tests/fm-control.test.sh now compares the whole task record (the `state/<id>.meta` file), so the Greptile finding is fixed. Invariant: after `exit` retires an incarnation, the task record must equal the record from before `exit` with only the `busy_gen` line removed. This test is the only place in the change that asserts the record survives the rewrite, so it is the only site to fix. The other `busy_gen` tests assert that the line stays, and they do not go through the rewrite. What changed: before `exit`, the test writes the record without its `busy_gen` line to `expected.meta`. After `exit`, the test runs `diff` between that expected copy and the real record, and fails with the diff output if they differ. This one comparison replaces the two earlier checks (no `busy_gen` line left, and the `window` line present), because it covers both. I did not change bin/fm-control.sh or any other file. How I know it works: - I ran `bash tests/fm-control.test.sh`: exit code 0, 45 lines starting with `ok`, no other lines. - I temporarily changed the rewrite in bin/fm-control.sh to also drop the `harness` line. The test then failed with `not ok - exit should drop only busy_gen from the task record:` and the diff `< harness=codex`. The earlier `window`-only check would have passed that rewrite. I restored bin/fm-control.sh afterwards; `git status` shows only tests/fm-control.test.sh modified. - `bash -n` and `shellcheck` on the test file report no new warnings from the edit. The change is not committed; the working tree holds it

* test: cover PID collisions in harness ancestry detection (#6484)

* test(secondmate-harness): scope fake ps -codex label for pid 5252 to the liveness probe

Closes #6456

* no-mistakes(ci): Updated the collision test to log and assert that PID 5252 was queried before selecting 4242. Full fm-secondmate harness suite passes

---------

Co-authored-by: YifuGu <ironerumi@users.noreply.github.com>

* fix: share Pi Calm's working-ship widget slot (#1854)

* fix(calm): share the standalone Pi Calm working-ship widget slot

Firstmate Calm and the user-global standalone Pi Calm both install an
animated working-ship widget during agent runs. Each claimed its own Pi
widget key, so a session loading both (the main Firstmate home) rendered
two boats. Pi replaces widgets under one key, so claiming the shared
"calm-working-ship" slot keeps dual-install sessions to a single boat
while a Firstmate-only session is unchanged.

Pins the shared slot contract in the working-ship module test so the key
cannot silently diverge again.

* test(calm): pin the shared working-ship widget key in CI, document dual-install

The key-parity assertion inside the Pi fixture only runs where the
@earendil-works/pi-coding-agent package is installed, so CI never
exercised it. Add a source-level twin that needs nothing but the
tracked file, and note in docs/calm.md that the boat shares the
standalone Pi Calm working-row widget slot.

* no-mistakes(review): Add executable dual-install widget replacement coverage

* no-mistakes(review): Guard shared widget cleanup with disposal ownership

* no-mistakes(document): Document shared Calm working-ship slot behavior

* test(calm): read the standalone Calm slot from its own module

The dual-install check registered both boats itself under the shared
slot, so it could only prove that Pi replaces a widget under one key: it
would still pass if the standalone Pi Calm extension installed its boat
under a different key, which is the two-boat regression the check exists
to prevent.

Read the standalone extension's own working-ship module when it is
installed - FM_STANDALONE_CALM_SHIP, else ~/.pi/agent/extensions/calm -
and drive the check with the key that module exports, so a rename on
either side registers two widgets and fails naming both keys. A pinned
shared-slot contract still covers a machine without the extension, and
the run reports which side it used instead of passing silently over an
absent extension.

Verified: the touched Pi Calm suite passes and reads the installed
standalone extension; with a copy of it whose key is renamed to
calm-working-ship-v2 the suite fails naming the drift.

* no-mistakes(review): Gate stock-row restoration by shared-widget ownership

* no-mistakes(review): Removed redundant widget-key source assertions

* no-mistakes(document): Document shared Calm working-ship widget ownership

* feat(bin): add opt-in --herdr-resume-lock-wait to fm-spawn (#6649)

* fix(herdr): make exact-resume presentation-lock wait instead of a bounded timeout

The exact-resume path in bin/fm-spawn.sh used the same 50-attempt-then-
give-up lock acquire as the new-task-create path, but the two paths are
not equivalent on contention: a create has no prior state to strand and
can safely fall back to a flat layout, while a resume is recovering a
specific existing identity that a concurrent recovery may legitimately
be holding the lock for. Giving up there does not degrade gracefully,
it hard-fails the resume outright. The suite's own concurrent
cross-home recoveries test already asserts both concurrent recoveries
succeed with a genuine reclaim, and the file's header comment already
(inaccurately) claimed lock contention falls back to the ordinary flat
layout for both paths alike, so the intended contract was always that
recoveries serialize and both succeed, not that either one refuses
under a short bound.

Give spawn_herdr_presentation_order_lock_acquire a wait mode that uses
this file's own established fm_lock_acquire_wait idiom (already used
for its other fleet-shared locks) instead of the bounded loop, and use
it only at the exact-resume call site. The new-task-create call site
is unchanged and keeps its bounded-then-flat-fallback behavior, which
is already covered by its own passing test. Dead-owner PID-liveness
reclaim inside fm_lock_try_acquire still bounds the wait against a
holder that crashed mid-hold.

Adds a deterministic regression test that holds the shared session
lock from an unrelated process for well past the old bound, then
asserts the resume succeeds with a genuine reclaim and took close to
the full hold duration, so a fix that merely widens the bound rather
than genuinely waiting is still caught. The existing concurrent
cross-home recovery test exercises this under real timing but does not
reliably outlast a fixed bound on its own.

Corrects the header comment's claim that create and resume share one
bounded-then-flat-fallback behavior on lock contention; they no longer
do.

* no-mistakes(document): Document Herdr recovery waiting for presentation lock

* no-mistakes(document): Update stale hard-refusal claim in verification log

* no-mistakes(ci): Fixed the Greptile finding on tests/fm-backend-herdr-presentation-e2e.test.sh:1389 by bounding the resume lock-wait regression's spawn_task call. Added an optional 4th `deadline_seconds` arg to the `spawn_task` helper (defaults to empty, so all ~20 other existing call sites are unaffected and unwrapped by `timeout`). The lock-wait test now passes `LOCK_WAIT_HOLD_SECONDS + 60` (90s) as the deadline, and a dedicated check for exit code 124 emits a clear "hung for over Xs instead of waiting out a Ys lock hold" diagnostic before falling through to the existing pass/fail assertions, which are unchanged. No product code was touched. Verified with `bash -n`, `shellcheck -x` (no warnings), a standalone reproduction of the timeout/no-timeout/success paths, the project's `bin/fm-lint.sh --fast` on the file (clean), and the full `tests/fm-lint.test.sh` suite (all 46 assertions pass)

* no-mistakes(ci): Replaced the direct `timeout "$deadline_seconds"` call in `spawn_task()` (tests/fm-backend-herdr-presentation-e2e.test.sh) with the repo's portable bounded-execution helper: sourced `bin/fm-timeout-lib.sh` at the top of the file and changed `deadline_cmd=(timeout "$deadline_seconds")` to `deadline_cmd=(fm_run_timed "$deadline_seconds")`. This removes the GNU/BSD `timeout` dependency that would fail with exit 127 on a stock macOS host without coreutils, while preserving identical semantics (exit 124 on bound-hit, command's own exit otherwise), which the existing `[ "$LOCK_WAIT_STATUS" -eq 124 ]` diagnostic check already relies on. Verified: `bash -n` syntax check, `bin/fm-lint.sh --fast` clean, full `tests/fm-lint.test.sh` suite (46/46 pass), and a standalone repro confirming `fm_run_timed` returns 124 on timeout and 0 on success identically to the prior `timeout` call. No other direct `timeout` calls exist in this file or elsewhere in the PR's diff, so no sibling sites remain

* fix(herdr): gate exact-resume lock wait behind --herdr-resume-lock-wait

Keep refuse-by-default on presentation-order lock contention for Herdr
exact resume. Callers that need concurrent recoveries to serialize must
pass --herdr-resume-lock-wait; unbounded blocking on a third-party session
lock is never the default.

Update docs and the real-Herdr e2e suite so the default path asserts the
refusal and the opt-in path asserts the wait.

* no-mistakes(test): Fix e2e test's lost exit status after if/fi with no else branch

* docs(herdr): stop advertising --herdr-resume-lock-wait on --relaunch

The relaunch path reuses the recorded endpoint and never takes the
presentation-order lock, so the flag is inert there. Drop it from the
--relaunch usage line and state where the flag applies.

* no-mistakes(review): Clarify lock-wait docs; simplify bash-3.2-safe spawn_task helper

* no-mistakes(ci): Fixed ci-1 (Greptile P2). In tests/fm-backend-herdr-presentation-e2e.test.sh, the failure cleanup `cleanup_all` stopped only `LOCK_CONTENTION_OWNER_PID`. It now also stops `LOCK_REFUSE_HOLDER_PID` and `LOCK_WAIT_HOLDER_PID`, the holders of the two new contention cases, so a `fail` before their explicit `wait` no longer leaves them running. Both new PIDs are initialised empty next to the existing one, and each is cleared right after its successful `wait` so cleanup never touches a finished PID. I changed nothing else. `bash -n` passes. The real Herdr e2e run passed both new cases ("default resumed identity refuses session lock contention" and "--herdr-resume-lock-wait waits out session lock contention instead of refusing"). The full run hit my 550s timeout in a later, unrelated case, after the new cases passed

* fix(bin): let TERM stop a watcher blocked in a pane capture on bash 3.2 (#6762)

* fix(bin): let TERM stop a watcher blocked in a pane capture on bash 3.2

Stock macOS bash 3.2 holds a HUP or TERM until a running command
substitution's child exits, and the watcher read every pane through
$(fm_backend_capture ...). A blocked backend read therefore held the
watcher's stop for as long as the read lasted, and a stopped watcher left
the hung read orphaned. tests/fm-watch-triage.test.sh
test_term_stops_a_watcher_blocked_inside_a_poll failed on /bin/bash 3.2
for this reason while passing on bash 5.

Pane captures now go through watcher_capture, which runs the read as a
waited background process group recorded like a check's, so the stop is
honored at once and watcher_cleanup stops a read still in flight along
with its per-call output file.

* no-mistakes(review): Run drain-ring idle capture in watcher shell, add regression test

* no-mistakes(document): Document watcher TERM handling for blocked checks and captures

* no-mistakes(test): Silence bash 3.2 setpgid race noise from watcher captures

* fix(bin): verify the capture group and scope the stop claim to pane reads

watcher_capture now confirms its background read leads its own process
group, as run_check_capture already does, so watcher_cleanup never relies
on a group that set -m failed to create. The comment and continuity doc
now say only fm_backend_capture pane reads go through watcher_capture;
agent-state and composer-state reads still run inside command
substitutions.

* feat: add per-home worker tool exclusions (#6750)

* feat(spawn): add per-home worker tool exclusions

Add an optional per-home config/crew-exclude-tools file listing tool names to hide from workers, one per line, with blank lines and # comments allowed.
It applies to every ship and scout launch and relaunch in that home, is never inherited by another home, and does not affect secondmate agents.
Pi and pi-signed apply it through --exclude-tools, which also covers MCP tool names.
Any other runtime, and a raw launch command, refuses the launch when the list is non-empty rather than ignoring it.
Malformed entries are refused before provisioning, and before a relaunch stops a running worker.
Exclusions that match no tool in the worker's loaded registry are reported as unverified warnings in its status record instead of refusing the worker.

Closes #6744

* no-mistakes(review): Preserve UTF-8 exclusion paths and verify Pi lifecycle behavior

* no-mistakes(document): Clarify worker tool exclusion documentation

* no-mistakes(ci): Fixed ci-1 in bin/fm-exclude-tools-lib.sh: a failed read now returns an error before printing names, so all shared launch and relaunch callers refuse rather than silently dropping exclusions. Added deterministic regression coverage for a file disappearing after readability checks across Pi/pi-signed ship and scout launches. Reproduced the original failure; verified 83 spawn checks, 77 relaunch checks, direct parser/runtime failure cases, full targeted lint, Bash syntax, and git diff --check. Relaunch tests passed with existing fixture-cleanup permission warnings. ci-2 remains unchanged per the user's decision; the outer executor owns the fresh CI run

* feat(bin): defer spawns beyond a declared per-project capacity, opt-in (#5343)

* refactor(bin): share the local Firstmate home walk from the wake library

Teardown's walk over the root home and its registered local secondmate homes
moves into bin/fm-wake-lib.sh as fm_local_firstmate_state_dirs, next to
fm_firstmate_root_home, so a second consumer can count task records across
this machine's homes without a copy. Teardown keeps its exact refusal wording
through a thin wrapper.

* feat(bin): defer spawns beyond a project's declared machine capacity

A project whose machine-local resource only serves a few workers at once had
no way to tell Firstmate so: every queued item was launched, and the surplus
workers spent full-context turns retrying the resource.

config/project-capacity in the root home now declares how many workers each
named project admits at once on this machine. bin/fm-spawn.sh counts the ship
and scout records on the same project origin across the root and its local
secondmate homes, skipping ones whose ready PR is recorded, while holding the
shared project lock through publication. A spawn with every place held exits 75
before any brief render, endpoint, worktree, record, or backlog move, so the
item stays queued; batches report it as deferred. Undeclared projects keep
today's uncapped dispatch, and an unreadable declaration refuses rather than
guessing the limit.

Refs #4237

* no-mistakes(review): Document that capacity matches the clone directory name

* no-mistakes(document): Rewrap stale fm-wake-lib root-home doc comment

* no-mistakes(review): Dedupe local state dirs by identity to avoid double-counting

* no-mistakes(document): Rewrap fm_local_firstmate_state_dirs error doc comment

* no-mistakes(ci): I fixed all four Greptile findings. All 14 tests in tests/fm-project-capacity.test.sh pass, and shellcheck at warning level is clean on the changed files. Each new test failed against the old code and passes now. - **ci-1 (spaced names):** a declaration line must give a name its capacity whenever the name is a valid clone directory name. `fm_project_capacity_lookup` now trims each line, skips blank lines and lines whose first non-blank character is `#`, and takes the last field as the capacity. Everything before that field is the name, so it may contain spaces. The old error cases still refuse: a single field is rejected, and trailing text leaves a last field that is not an integer. The library header and docs/configuration.md now say a name starting with `#` cannot be declared. New test `test_spaced_project_name_is_declared` declares `my heavy project 1` next to an indented comment line and gets a deferral. - **ci-2 (unreadable records):** the holder count must never silently leave out a holder. `fm_project_capacity_occupants` now refuses when a local home's state directory exists but cannot be read or listed, or when a `.meta` file cannot be read. The error names the path, and `fm-spawn.sh` shows it in its existing refusal message. New test `test_unreadable_holders_refuse_admission` covers an unreadable record in the root home and an unreadable state directory in a registered local secondmate home, then checks that the spawn is admitted once both are readable. The test is skipped when run as root. - **ci-3 (Orca lock):** any spawn that can become a holder for a capped project must take that project's lock. The lookup now also reports whether the declaration caps any project at all, and an Orca spawn takes the per-origin lock whenever it does. This covers every capped same-origin clone. It also covers some cases where no same-origin clone is capped, because a spawn cannot find clones under other directory names without searching for them. With no declaration file, Orca still skips the lock. The comments in the library and in the `fm-spawn.sh` header are updated. The Orca test now clones the origin as `project-2`, which has no declaration, and checks that its Orca spawn refuses while the lock is held and publishes no record. - **ci-4 (worktrees):** `assert_nothing_created` now also compares the project's `git worktree list` from before and after a deferred spawn. Both tests that call it take that snapshot first. Files changed: bin/fm-project-capacity-lib.sh, bin/fm-spawn.sh, docs/configuration.md, tests/fm-project-capacity.test.sh

* fix(bin): declare capacity for a project name that begins with #

A clone directory whose name begins with # was skipped as a comment, so that project stayed uncapped. A line is a declaration when the # is written against the rest of the name and the line ends with a capacity; a # followed by whitespace stays a comment.

* no-mistakes(document): Rewrap project-capacity library header comment

* no-mistakes(ci): Lint 2 fails because this PR's code pushes ShellCheck past its memory cap. ShellCheck ran out of memory analyzing bin/fm-teardown.sh in CI (reason=memory, rc=251, peak about 8.39 GB). On current main the same file passes at about 7.29 GB. **Cause:** the new `fm_local_firstmate_state_dirs` function in bin/fm-wake-lib.sh had a conditional `. fm-secondmate-registry-lib.sh` with a `# shellcheck source=` directive inside the function. ShellCheck followed that source again, inside a function scope, wherever fm-wake-lib.sh is sourced, and bin/fm-teardown.sh is the heaviest root that sources it. Measured locally with `shellcheck --norc --external-sources bin/fm-teardown.sh`: - current main (fd325b1b): 7.29 GB - main merg…
umeranjum17 added a commit to umeranjum17/firstmate that referenced this pull request Oct 9, 2026
…upervision fixes (#44)

* fix(pi): hide duplicate assistant finals from hidden processing retries (#5863)

* fix(pi): silence unacknowledged processing retry replies

Suppress autonomous processing prose before persistence and during streaming while retaining tool calls, signed reasoning, and retryable outcomes. Restore ordinary output after acknowledgement or a user message.

Fixes #4954

* no-mistakes(review): Silence only processing retries, keep first presentation visible

* fix(pi): preserve differing processing retry replies

* test(pi): accept Pi 1.0.1's renamed HTML export renderer lookup (#6530)

Pi 1.0.1's createToolHtmlRenderer reads getToolRenderers and ignores
getToolDefinition. Calm /export still includes stock grep HTML; the
fixture has to pass the lookup key the installed Pi actually reads.

* fix(bin): tolerate transient quota read failures (#6490)

* fix(procevent-quota): tolerate consecutive slow quota-axi reads

The quota allowance poll treated any quota_json failure as terminal, so one
slow quota-axi --json (measured max ~29s under a 48s derived bound) shut the
watch down until someone re-armed it, and the detail always said
"missing/incompatible". Tolerate three consecutive failed or timed-out reads
before going terminal, reset the streak on any good read, and report a
timeout distinctly from a missing or incompatible tool. Each timed poll runs
exactly one bounded --version and one bounded --json: validate the captured
version text through fm_quota_axi_version_compatible rather than launching a
second probe, and describe a mixed failure streak by count plus last cause.

* no-mistakes(document): Document quota polling failure tolerance

* no-mistakes(ci): Fixed ci-2 and ci-3. Permanent quota read failures (rc 2 missing, rc 3 incompatible) now report on the first poll, while transient rc 1/4 failures retain the existing three-failure retry behavior. The missing-binary test now uses an isolated PATH without quota-axi and asserts both permanent failures stop at condition_polls: 1. Verification passed: tests/fm-procevent-quota.test.sh, canonical fast lint for both changed files, bash syntax checks, and git diff --check

* no-mistakes(review): Classify untimed quota version failures as transient

* no-mistakes(document): Clarify quota polling failure budget

* fix(bin): clarify scratch guidance and dirty teardown refusals (#6505)

* fix(teardown): clarify scratch guidance and dirty worktree refusals

Keep ship proof material outside the task worktree and distinguish untracked-only leftovers from tracked edits without changing cleanup guards.

Fixes #6319

* fix(ci): Fixed ci-3 only. Both promotion outputs now replace the scout restriction and require external scratch storage and a clean worktree before done. Verification: 27 delivery tests passed, five mutations caught, restored test passed, pinned ShellCheck and syntax/whitespace checks passed. ci-1, ci-2, and ci-4 remain untouched

* fix(bin): escalate inbox instructions blocked by busy workers (#6518)

* Escalate inbox instructions stuck behind a busy worker

Count consecutive busy-deferred due doorbells durably and escalate at the configured bound without typing into the worker pane.

Fixes #6445

* fix(review): Fix inbox escalation deduplication and busy streak resets

* fix(review): Preserve busy inbox escalations through daemon supervision

* fix(document): Correct busy-inbox escalation documentation

* fix(ci): Fixed SC2034 in tests/fm-task-inbox.test.sh by including the loop counter in the failure diagnostic. Source-aware lint, all 34 inbox tests, and git diff --check pass. Behavior portable serial 6 reproduces identically on base 1f3e7696 and target 78156b86 with Pi 1.0.1: an unrelated renderer API change breaks the unchanged Calm test. No Calm changes made; that failure is addressed separately by https://github.com/kunchenguid/firstmate/pull/6516. Logs retained in scratchpad-ci/

* fix(ci): Fixed ci-2, ci-3, and ci-4: successor failures surface, reset alerts deduplicate, and oversized busy limits fall back to two. Passed 47 inbox tests, 7 focused daemon checks, all 13 mutation checks, lint, documentation checks, and diff checks. Evidence: scratchpad-ci-selected/summary.json. ci-1 remains unchanged and unwaived. Fresh live Herdr proof remains with the outer driver

* fix: prevent unnecessary remote worker turnover (#6431)

* fix: prevent healthy remote job worker turnover

* no-mistakes(review): Serialize full LaunchAgent repair and verify launchd-tracked owners

* no-mistakes(review): Let launchd-tracked unpublished spawns start before reloading

* no-mistakes(document): Document remote worker heartbeat and serialized LaunchAgent recovery

* no-mistakes(ci): Full CI log showed the idle-worker regression exceeded its outdated command budget (82 versus 80) after independent heartbeat ownership checks were added. Raised the budget to 120 while retaining the separate busy-poll sleep limit. Remote-job and LaunchAgent executable tests passed, as did bash syntax validation and git diff --check. No production behavior changed

* no-mistakes(ci): Fixed missing readiness recovery under verified live ownership, preserving the serving PID and private file mode. Added executable regressions for deletion during a blocked sweep and stale readiness diagnostics without LaunchAgent reload. Deletion regression failed before the fix. Both remote-job test suites, bash syntax validation, and git diff --check passed

* no-mistakes(ci): Full CI log identified a flaky ownership-loss test racing an already-authorized heartbeat refresh. Replaced backdating and a fixed sleep with bounded observation of readiness expiry through the public probe. Production behavior unchanged. Remote-job and LaunchAgent executable suites passed; bash syntax validation and git diff --check passed

* fix: wait for launchd bootout cleanup

* no-mistakes(document): Document remote worker recovery and read-only turnover verification

* no-mistakes(review): Publish worker identity before lock owner records

* no-mistakes(document): Document worker identity publication safety invariant

* no-mistakes(ci): Fixed ci-1: replacement workers discard predecessor readiness before publishing identity and roll back identity if lock-owner recording fails. Added executable regressions reproducing both defects. LaunchAgent and orphan-reap suites, shell syntax checks, and git diff --check passed. No live service state was modified

* no-mistakes(ci): Restored lock-owner-before-identity publication and removed the identity rollback and reordering-only tests. Retained an executable regression proving predecessor readiness is rejected until replacement startup completes. LaunchAgent and orphan-reap suites, shell syntax checks, and git diff --check passed. Other changes remain intact; the retained-identity interrupted-repair edge remains out of scope. No live service state was modified

* fix: speed up remote-job sequence claim cleanup (#6575)

* fix(remote-job): reap expired seq claims with one directory walk

The hourly claim sweep forked uname+stat per .seq-claims entry and blocked
serving for ~85s at ~17k dirs. Delete expired empty claim dirs with a single
find -exec rmdir batch and cache the host uname for remaining mtime reads.

* no-mistakes(review): Restore original path mtime helper and drop uname cache

* no-mistakes(test): Restore claim retention eligibility; focused retention and serving tests pass

* no-mistakes(document): Document single-walk claim cleanup and regression entrypoints

* no-mistakes(ci): Fixed Lint 2’s reproduced SC1091 by adding the tests/lib.sh ShellCheck source directive to the retention test. Runtime behavior is unchanged. ShellCheck passed for both new claim tests; Bash syntax, the retention behavior test, and git diff --check passed

* fix(tests): keep fixture registries out of git worktree roots

A TMPDIR pointed at a repository root placed live .fm-test-* registries
beside tracked files, and a concurrent git add during the claim-walk CI
fix round committed three of them. Route registries and fixture roots
through a TMPDIR that refuses git worktree roots, remove the stray files,
and pin the escape with a behavioral cleanup test.

* no-mistakes(review): Preserve whole-second claim expiry in single-walk sweep

* no-mistakes(document): Correct temporary-directory resolution documentation

* no-mistakes(ci): Fixed ci-3 by changing only the stale, fresh, and read-only orphan fixture paths in tests/fm-test-fixture-cleanup.test.sh to use $FM_TEST_TMPDIR. All seven tests passed both normally and with TMPDIR set to the worktree root. Shell syntax and git diff --check passed

* feat(bin): record no-mistakes pipeline spend per task at teardown, opt-in (#5354)

* feat(bin): record each task's no-mistakes pipeline spend at cleanup

no-mistakes keeps every pipeline agent invocation's token usage only in its
local agent_invocations records, and cold pipeline agents (review, test,
document) leave no session log, so a task's review-loop cost never reached
Firstmate's records and could not be attributed after cleanup.

bin/fm-pipeline-spend.sh attributes a task's runs by the repository
no-mistakes resolves for the task copy, the task branch, and the branch's
creation (a relaunch mints a new spawn_gen while the branch keeps
validating), then sums every invocation, failed and cancelled included.
Token fields use no-mistakes' per-round deltas so resumed review rounds are
not counted twice, and unrecorded values stay unknown rather than zero.
show prints the record read-only; record appends it once per task
incarnation to data/pipeline-spend.jsonl. Teardown records it for every
ship task it cleans up, before deleting the branch and the task record.

fm_nm_state_db becomes the one owner of where no-mistakes' state database
lives, shared with the capped run-inventory reader.

* no-mistakes(review): Drop spend show command and timeout override

* style(bin): rewrap fm-pipeline-spend header comment

* Make pipeline spend recording opt-in

* feat(bin): record each task's no-mistakes pipeline spend at cleanup

no-mistakes keeps every pipeline agent invocation's token usage only in its
local agent_invocations records, and cold pipeline agents (review, test,
document) leave no session log, so a task's review-loop cost never reached
Firstmate's records and could not be attributed after cleanup.

bin/fm-pipeline-spend.sh attributes a task's runs by the repository
no-mistakes resolves for the task copy, the task branch, and the branch's
creation (a relaunch mints a new spawn_gen while the branch keeps
validating), then sums every invocation, failed and cancelled included.
Token fields use no-mistakes' per-round deltas so resumed review rounds are
not counted twice, and unrecorded values stay unknown rather than zero.
show prints the record read-only; record appends it once per task
incarnation to data/pipeline-spend.jsonl. Teardown records it for every
ship task it cleans up, before deleting the branch and the task record.

fm_nm_state_db becomes the one owner of where no-mistakes' state database
lives, shared with the capped run-inventory reader.

* no-mistakes(review): Drop spend show command and timeout override

* style(bin): rewrap fm-pipeline-spend header comment

* Make pipeline spend recording opt-in

* no-mistakes(document): Note pipeline-spend opt-in gate in teardown comments

* no-mistakes(ci): I fixed all three Greptile findings following your decision. Both test suites pass: tests/fm-teardown.test.sh (105 ok, exit 0) and tests/fm-pipeline-spend.test.sh (8 ok). I didn't run either new test against the old code, so the claim that they fail before the fix is from reading the code, not a run. Nothing was committed or pushed. **ci-1 / ci-2 (bin/fm-teardown.sh):** The rule is that every owned ship task in a home that has opted in gets a durable spend record before its task record is deleted. The `[ -d "$WT" ]` check skipped the recorder when the worktree folder was already gone, so no record was written. I removed that check and kept the other conditions: the task must be a ship task, teardown must own its worktree, and `config/pipeline-spend` must exist. `bin/fm-pipeline-spend.sh` already writes an unavailable-source line when the worktree is missing. The new test `test_teardown_records_unavailable_spend_for_a_gone_worktree` in tests/fm-teardown.test.sh sets up an owned ship task whose worktree is missing, with recording turned on. It checks that teardown succeeds, writes a line with `source=="unavailable"`, `total==null` and a reason saying the task copy is gone, and still removes the task record. Under the old check no ledger would have been written. **ci-3 (bin/fm-nm-run-lib.sh):** When `NM_HOME` and `HOME` were both unset, `fm_nm_state_db` fell back to `${HOME:-}/.no-mistakes`, which resolves to `/.no-mistakes`. It now uses `root=~/.no-mistakes`. Bash expands `~` to the account's home directory even when `HOME` is unset, which matches the previous Python `Path.home()` lookup. The new test `test_state_db_without_nm_home_or_home_uses_the_account_home` in tests/fm-pipeline-spend.test.sh runs the function with both variables unset and compares the result to the home directory from the system's password database. The old code would have returned `/.no-mistakes/state.sqlite`. shellcheck reports nothing new on the changed files. I added one `disable=SC2016` comment in the new test, where the single-quoted `$1` is meant to expand in the child shell

* feat(bin): start tasks from a named base branch (#6442)

* feat(bin): start tasks from a named base branch

Spawns always reset a task's pooled copy to origin's default branch,
so work that belongs on a feature, integration, or release branch
started from the wrong code and opened its PR against the default.

fm-brief.sh --base-branch records the base in the brief, fm-spawn.sh
resets the copy to origin/<base> and records base_branch= in task
meta, and the worker targets its PR at that branch. Review diffs,
the cleanup content check, and scout promotion read the recorded
base. local-only and Gerrit deliveries refuse a named base.

* no-mistakes(review): Anchor brief base-branch parse on scaffold Setup line

* no-mistakes(review): Require explicit --base-branch on spawn, matching brief lines

* no-mistakes(document): Note base-branch reset in fm-spawn freshness header

* no-mistakes(ci): I fixed all four Greptile findings the way you specified. The affected suites pass locally: fm-brief, fm-dod-lib, fm-spawn-pool-base-freshen, fm-review-diff, fm-teardown and fm-task-delivery. shellcheck on the changed files reports only info-level notes, no warnings or errors. I didn't revert the code to watch each new test fail before the fix. - **ci-1 (brief text blocked or redirected spawns):** In bin/fm-dod-lib.sh, `fm_brief_base_branches` now only accepts a `Base branch: <value>` line that comes directly after the base-variant Setup sentence fm-brief.sh writes. Any other `Base branch:` line is ignored. `--base-branch` is still the only authority, and the existing checks in fm-spawn.sh are unchanged. With the flag, the spawn is refused unless a matching pair exists and no pair names a different branch. Without the flag, it is refused only if such a pair exists. The header comment in fm-spawn.sh is updated to match. - The `brief_with_base` test helper now writes the real two-line pair. - The decoy refusal case now uses a full decoy pair in the captain's intent. - New test `test_prose_base_branch_line_is_ignored`: a brief with no base whose intent quotes `Base branch: release/1.2` spawns normally with no `base_branch=` recorded. A spawn with `--base-branch feature/hub` whose intent names release/1.2 starts from `origin/feature/hub` and records it. - **ci-2 (scout base not checked against the forge):** fm-spawn.sh now looks up the project's registered forge with `bin/fm-project-mode.sh --forge`, like the ship path does, before `fm_base_branch_valid` runs. Previously it passed `none`. New test `test_scout_base_branch_refused_on_gerrit_forge`: a scout with a base on a forge=gerrit project is refused at spawn and no meta file is written. - **ci-3 (base not shell-quoted in worker commands):** `fm_dod_block` now quotes the base with `printf %q` in the `--base` and `--base-branch` arguments, and the branch name in the prose stays readable. The test in fm-brief.test.sh uses the base `release/$HOTFIX` and checks that both direct-PR and no-mistakes briefs render `release/\$HOTFIX` in the commands. - **ci-4 (ambiguous PR sentence):** The direct-PR sentence now reads "open a PR with `gh-axi` that is ready for review, not a draft, against the base branch `X` (`--base X`), not the repository default." A brief with no base renders exactly as before. The existing assertion in fm-brief.test.sh is updated

* no-mistakes(ci): Both real failures on PR 6442 (run 37064756768) come from CI tooling and a Pi version bump, not from the base-branch feature. Fixes for both already exist on upstream main, so I applied those two commits to the working tree with `git cherry-pick --no-commit`. Nothing is committed or pushed, and no feature file changed. - **Lint 2:** ShellCheck ran out of memory on `bin/fm-teardown.sh` (rss about 8 GB, exit 251, "shellcheck: out of memory"). The rule that broke is that the lint gate must finish on any valid shell root without exceeding the runner's memory ceiling. Upstream commit d719ef3 (#6443) fixes this in `bin/fm-lint.sh`: a root that hits the memory ceiling is retried without `--external-sources`. It also updates `tests/fm-lint.test.sh` and `docs/fm-test-portable-shards.md`. - **Behavior portable serial 6:** `tests/fm-calm-pi-extension.test.sh` failed with "grep disappeared from /export calm.html HTML while calm mode was on". Pi 1.0.1 renamed its HTML export renderer lookup, and upstream commit fede619 (#6530) updates the test to accept the new name. This change is test-only. Both commits applied cleanly. The four touched files are `bin/fm-lint.sh`, `docs/fm-test-portable-shards.md`, `tests/fm-lint.test.sh` and `tests/fm-calm-pi-extension.test.sh`. None of them is in the feature diff, so the feature contract is unchanged. **Local checks:** - `tests/fm-lint.test.sh` passes. - `bin/fm-lint.sh` on `bin/fm-teardown.sh`, `bin/fm-spawn.sh` and `bin/fm-dod-lib.sh` exits 0 with full ShellCheck analysis. - `tests/fm-calm-pi-extension.test.sh` exits 0 with 7 ok and 0 not ok. One case skipped because the Pi package isn't installed locally, so the Pi 1.0.1 export path that failed in CI couldn't be reproduced here. CI needs to confirm it. - I didn't reproduce the 8 GB OOM locally. The PR has to be updated through the pipeline with a normal push: no force push, no replacement PR, no merge

* fix(bin): name the Firstmate skill file as a fallback in the worker role (#6647)

* fix(bin): point project workers at the Firstmate skill file

The Skill tool cannot resolve a Firstmate skill from another project's worktree, so the launch role names the readable skill file instead.

* no-mistakes(review): fix(bin): name Firstmate skill file as fallback only

* feat(bin): retire a contribution whose forge object is permanently gone (#6655)

* feat(bin): retire a contribution whose forge object is permanently gone

Add fm-contributions.sh retire <task> <url> <captain|fleet> <reason>,
which records actor, reason and time on the saved record and removes
that task/url pair from known, poll rotation and coverage even while a
backlog link remains. Repeating a retire keeps the first provenance;
an unrecorded pair, unknown actor, empty reason or unacknowledged
pending signal is refused.

* no-mistakes(review): Keep retirement per task when settling final owners

* no-mistakes(ci): Made the three Greptile fixes the captain chose. ci-1, retirement needs the captain's word: `retire` now accepts only the actor `captain`. Any other actor is refused, including `fleet`. I removed `fleet` from the usage line, the script header (`bin/fm-contributions.sh`), the check in `bin/fm-contributions.sh` and the retired-record validation in `bin/fm-contributions.jq`, which now requires `.actor == "captain"`. The script header now says a retirement always records the captain's word, because the script cannot verify who runs it and the authority to retire is the captain's. The Bearings skill line in `.agents/skills/bearings/SKILL.md` now says `retire` records the captain's word. In the tests, a `fleet` retire is now a refusal case and must leave the record unretired. Every other retire call in the tests uses `captain`. The PR description has not been changed yet: the ruling asks for the same captain's-word statement there, and that belongs to the PR phase. ci-2, blank reasons: the reason check is now `[ -n "${5//[[:space:]]/}" ]`, so a reason made only of spaces or tabs is refused. I added a refusal test that passes a space-tab-space reason. The header also lists a blank reason among the refusals. ci-3, the gone-object test: the test forge wrapper has a new `not-found` fault that returns HTTP 404 for every `api repos/o/r/...` read. The happy-path test `test_retire_ends_observation_of_a_gone_contribution` now uses it in place of the 502 `down` fault, so it models a permanently gone repository rather than an outage. Verification: `bash tests/fm-contributions.test.sh` passed with exit 0. The last lines of its output include the three retire tests passing and no failures. `shellcheck` flagged only SC2034 (`FM_WAKE_QUEUE` unused) and SC1091 (an unfollowed source of `bin/fm-path-lib.sh`), neither on a line this round changed

* test: extend remote-reply whole-log recapture waits (#6639)

* test: give the remote-reply whole-log recapture a longer wait

* no-mistakes(review): Extend both recapture waits and simplify retry handling

* fix(bin): reopen a pending-reply escalation after its resolve (#6654)

* fix(bin): reopen a pending-reply escalation after its resolve

A retry of an open decision still appends nothing. A later same-kind escalation after a resolved line for that key appends again, so the next loss is visible.

* no-mistakes(ci): Both Greptile findings are fixed in the working tree; the four directly affected test suites pass, and a wider run of related suites had not finished when I returned this result. Invariant: a line already in the status file is recorded, and only a caller with its own evidence of a new episode may append it again. It must hold for every caller of status_event_recorded: the continuity break and the other call in bin/fm-procevent-remote-reply.sh, fm_parent_channel_append_once, and the pending-reply escalation. Changes: - bin/fm-classify-lib.sh: status_event_recorded is back to the idempotent retry check. It returns at the first matching line, and a later resolved line no longer makes that line look new. This fixes ci-1 for every caller and restores the early exit for ci-2, with no separate scan optimization. - bin/fm-pending-reply-lib.sh: _fm_pending_reply_maybe_escalate_locked now owns the reopen. It appends a new blocked line when the record already has an escalated_epoch (it escalated before and was reset) and the keyed decision is no longer open. Otherwise it uses status_event_recorded, so a retry or a resend that leaves the decision open appends nothing. - Comments on status_event_recorded and the pending-reply header describe this scope. Tests: - tests/fm-pending-reply.test.sh: the regression for escalate, operator resolve, second escalate is unchanged and passes. - tests/fm-remote-reply.test.sh: new regression drives the real reader. After an operator resolve of remote-reply-continuity-ios, a repeated handle and ingest of the same break appends nothing and leaves the decision closed. I placed it beside the existing continuity-break test, not in the pending-reply file, because only that file has the reader harness. - tests/fm-classify-corr-token.test.sh and tests/fm-classify-decision-key.test.sh: the cases that expected a blocked line to append again after a resolve now assert it stays recorded. Verification: - fm-classify-decision-key, fm-classify-corr-token, fm-pending-reply and fm-remote-reply test suites all exit 0 with the fix. - With the two bin files reverted to the commit under review, the new continuity regression and the updated parent-publisher case both fail, so they reproduce ci-1. - shellcheck -x on the changed files reports nothing. - Not finished: a background run of every other suite that mentions the parent channel or pending replies had produced no output, pass or fail, when I returned. One known gap: the escalation writes escalated_epoch before phase. If the process dies between those two writes and an operator resolves the key before the next tick, that tick appends a blocked line for the same episode. I left it, because closing it needs new state. Issue 4755 and other escalation behavior are untouched. Nothing is committed

* fix: refuse a confirming Enter on the Claude background-task exit picker (#6666)

* fix: refuse a confirming Enter on the Claude exit picker

The background-task picker still classifies as pending, so a retried Enter confirms "Exit and stop tasks".
Stop after the Enter that opened it, and raise the existing stale wake with the dialog name.

A non-paused secondmate still skips that wake, so a secondmate parked on the picker still looks idle.
Model-downgrade confirmation, MCP approval, and any other Claude exit confirmation are not covered, because there is no recorded screen for them.

* no-mistakes(review): fix: anchor exit picker match and wake second mates

* fix: refuse a typed submit while the Claude exit picker is open

A pane that already shows the picker must not receive the message or a confirming Enter.
The match requires the recorded heading and footer lines, and it is skipped when no dialog name is being recorded.

* no-mistakes(review): restore watcher to main behaviour, drop dialog wake

* no-mistakes(test): test: align Herdr picker fixtures with the preflight read

* no-mistakes(document): document exit refusal on a recognised dialog

* fix: remove the dialog file when exit runs in a subshell

do_exit is invoked from a command substitution, so the parent cleanup never saw the path it is supposed to delete.

* no-mistakes(review): fix: set the dialog file path after the control lock

* no-mistakes(document): document why the dialog file path follows the lock

* no-mistakes(ci): The exit cleanup in bin/fm-control.sh now deletes the dialog file first and releases the control lock second. The changes are in the worktree and are not committed. Invariant: only the process that holds the control lock for a task may write or delete that task's dialog file ($STATE/<id>.composer-dialog, the file that carries the picker name from the composer read to the Enter checks). The old cleanup released the lock and then deleted the file, so the delete ran outside the lock. Places where this invariant must hold: control_cleanup is the only place that deletes the file, and the single assignment after CONTROL_LOCK_HELD=1 is the only place that names it. do_exit only truncates the file, and it runs while the parent holds the lock. A process that loses the lock never sets the path, so it deletes nothing (earlier fix, unchanged). No sibling site needed a change. Fix: in control_cleanup, the rm of the dialog file moved above the fm_lock_release call, with a two-line comment that gives the reason. No dialog recognition or supervision code changed. Test: added test_exit_removes_the_dialog_file_before_releasing_the_lock in tests/fm-control-relaunch.test.sh. It runs one `fm-control exit` with a recording rm on PATH. The lock release removes paths at or under the control lock with rm, so the recording rm writes whether the dialog file exists at that moment. The test asserts the last record is "absent". It starts no second command. Verification, all run locally: - Before the fix, the new test failed with "the dialog file must be gone when the control lock is released, got: present". - After the fix, tests/fm-control-relaunch.test.sh exits 0 with 75 ok lines and no "not ok" line, including the new test and the existing test that the file is gone after exit and after relaunch. - tests/fm-control.test.sh exits 0 with 44 ok lines. - shellcheck -x on both changed files exits 0 with no output. One thing I noticed and did not change: tests/fm-control-relaunch.test.sh prints 615 "rm: cannot remove ... Permission denied" lines during its temp-directory teardown. The count is the same with my change reverted, so this change does not cause it, and the suite still exits 0. I could not read the Greptile check log (the check run was not found), so I worked from the finding text and the user instructions

* fix: restore timeout watchdog compatibility with macOS Bash 3.2 (#6028)

* fix: support stock macOS Bash in timeout watchdog

* no-mistakes(ci): Fixed ci-1 in tests/fm-timeout-lib.test.sh: both startup-owner failure paths now kill the bounded command group before killing the watchdog. Fault injection reproduced both leaks before the fix and confirmed cleanup afterward. Bash 3.2 suite, syntax check, ShellCheck, and diff checks passed; GNU timeout coverage skipped because the binary is unavailable. Production code unchanged

* fix(project-management): use subshell form for Initialize command (#6699)

Wrap the cd command in a subshell to comply with the cd-guard policy
that blocks persistent top-level directory changes in the primary
firstmate checkout. The subshell form (cd projects/<name> && ...) is
accepted by the policy as documented in issue #6502.

Fixes #6502

* feat(bin): add armable daily startup growth check (#6725)

* Add daily startup growth check

* no-mistakes(review): watch printed startup memory, retain growth baselines, drop knobs

* no-mistakes(review): delegate budget verdict, report before publish, pin shim home

* no-mistakes(review): baseline first sightings silently, drop mtime, spare secondmates

* no-mistakes(review): cap the wake line, validate budget verdict fields

* no-mistakes(review): guard record schema, check appends, tighten assertions

* no-mistakes(review): drop arbitrary tracked library, make re-arm idempotent

* no-mistakes(review): correct tracked-set wording, restore shim guards, register artifacts

* no-mistakes(review): exit on signal instead of publishing partial record

* no-mistakes(document): correct startup-growth record removal cost in state registry

* no-mistakes(test): sweep orphaned empty startup-growth temp records on due evaluation

* no-mistakes(document): note watcher need for armed startup growth check

* no-mistakes(ci): Diagnosed "Behavior portable serial 5": the shard failed on tests/fm-contributions.test.sh in test_arm_plumbs_a_configured_budget_into_the_check_shim (inherited mode) with "generated check did not attempt a read" and a missing forge/calls file. Root cause: bin/fm-contributions.sh:379-382 computes DEADLINE=$(date +%s)+BUDGET and then gates each read on DEADLINE-$(date +%s) >= OBSERVATION_RESERVE (= min(BUDGET,15)). At the one-second budget this test uses, that gate demands zero elapsed whole seconds, so a wall-clock second boundary crossing between the two date calls makes the poll loop break before any forge read. The test calls wrap_forge (which installs a fake date reading $FORGE/clock only when that file exists) but never seeded forge/clock, so it ran against the real clock. The repo already documents this exact hazard at tests/fm-contributions.test.sh:675-677 and freezes the clock in every other budget-constrained test: test_budget_exhaustion_keeps_prior_record (both modes), test_genuine_failure_near_deadline_is_unavailable, test_reservation_defers_later_url_when_fifteen_seconds_do_not_remain, test_slow_read_deadline_kill_is_budget_refusal, test_budget_is_cut_down_to_the_watcher_check_bound. test_arm_plumbs was the only omission; its single loop body covers both the configured and inherited modes. The remaining wrap_forge users run the default 20-second budget (5 seconds of slack above the reserve) and are not reachable by this race. Fix (smallest, matching the sibling idiom): added the clock freeze `/bin/date +%s > "$home/forge/clock"` plus a two-line comment inside that test's loop, before the poll. Three added lines in tests/fm-contributions.test.sh; no production code and no other file touched. Verification: reproduced the exact CI failure message by running a scratch copy whose fake date advances one second per read (worst case of the unfrozen clock) - missing forge/calls, "not ok - generated check did not attempt a read". With the fix, the single test passes 10/10 standalone and the full tests/fm-contributions.test.sh passes all 46 assertions with exit 0. tests/fm-startup-growth-check.test.sh still exits 0. bash -n and shellcheck -S warning are clean on the edited file; scratch repro files were removed, leaving only the intended three-line change. Note for the author: this flake is not caused by this branch - tests/fm-contributions.test.sh is untouched by the PR and the race is pre-existing, surfacing only on a loaded runner (the shard's neighbouring assertions show 80+ second gaps). It was fixed because it is genuine nondeterminism with an established in-repo remedy and the check cannot otherwise go green. No work was done on the unselected greptile findings (ci-2..ci-5) or the cancelled "Behavior portable serial 2" check (ci-6)

* no-mistakes(ci): Diagnosed "Behavior portable serial 6": the omitted log section (fetched with gh api) shows the shard failed on tests/fm-calm-pi-extension.test.sh in test_hidden_block_geometry_e2e with "Pi Calm hidden-block geometry E2E did not complete the /reload viewport transition" (exit=1, duration_ms=23437; the family line pure-contract-unit count=3 duration_ms=51193 failed=1 matches devin-harness 3203 + calm-pi 23437 + task-delivery 24553). Root cause: tests/fm-calm-pi-extension.test.sh:2831 wait_for_geometry_transition detected the /reload by SAMPLING a single intermediate frame - it polled tmux capture-pane for Pi's transient "Reloading keybindings, extensions, skills, prompts, themes, and context files..." box and only accepted the rebuilt transcript afterwards (elif gated on saw_transient). Pi shows that box only while the reload runs and replaces it via dismissReloadBox when done, so on a loaded runner the box can live entirely between two polls; saw_transient then stays 0 forever and the 600-attempt wait times out although the reload fully succeeded. Reproduced locally under CPU oversubscription with CI's Pi 1.0.4: 1 failure in 8 runs, diagnostics proving the mechanism ("DIAG: transition timeout saw_transient=0 attempt=600") while the captured viewport already showed the completed reload status row and the intact collapsed transcript. Not a Pi-version regression: 1.0.4's handleReloadCommand and both banner strings are identical to the installed 0.85.1, and shard 6 of the previous pipeline run (job 112490351105) passed this same script on Pi 1.0.4. Invariant violated: a TUI E2E assertion must key off a durable state the UI retains, never off one intermediate frame polling may miss. Sibling enumeration: the helper was defined once and called once; grep for transient/saw_/two-stage waits across tests/, bin/ and .pi/ found no other one-shot-frame wait - every other wait in the file keys off durable text presence/absence or off continuously repeating animation frames (the working-ship motion checks) where sampling eventually connects; no doc or other test referenced the helper. Fix (smallest, one helper): renamed it wait_for_geometry_reload and made it wait for the durable status row Pi appends once the reload completed and the chat was rebuilt ("Reloaded keybindings, extensions, skills, prompts, themes, and context files", a persistent chat child via showStatus, present in both 0.85.1 and 1.0.4) together with CALM_GEOMETRY_FINAL. That marker is absent before the reload and never appears if the reload throws, so the check is strictly stronger than before (a failed reload now fails the test instead of passing on transient-then-final). One file changed: tests/fm-calm-pi-extension.test.sh, 2 hunks, no production code. Verification: deterministic before/after - with polls spaced 0.4s apart (the loaded-runner worst case) the pre-fix helper fails with exactly the CI message and the post-fix helper passes; 12/12 consecutive passes of the isolated test under load with Pi 1.0.4; full file 15/15 ok exit 0 under LANG=C.utf8 (matching the runner locale); bash -n, shellcheck -S warning and bin/fm-lint.sh clean; scratch repro files and the temp Pi 1.0.4 install removed, git status shows only the one intended modification. Notes: (1) the flake is not caused by this branch - that test file is untouched by the PR and the race is pre-existing; fixed because it is genuine nondeterminism and the check cannot otherwise go green. (2) Under far harsher load than CI applies a separate unrelated flake appeared in the same test: Pi's "Warning: tmux extended-keys is off..." row occasionally lands between the collapsed [skill] ahoy row and the final response, making assert_geometry_gap see 5 rows instead of 2. Different cause, never seen in CI, and its plausible remedies would change key delivery for the other TUI tests in the file, so it was left alone rather than growing this fix

---------

Co-authored-by: Quartermaster <quartermaster@users.noreply.github.com>

* fix(bin): reopen the remote-reply continuity decision on a later break (#6708)

* fix: reopen a remote-reply continuity break after repair

A later break for the same route and reason was swallowed after the
operator resolved the first one, because the status line matched for
the life of the log. The continuity ingest now appends again when the
cursor has moved or retirement has reset that episode, and an unchanged
re-read still appends nothing.

status_event_recorded is unchanged. Its other callers are the
pending-reply escalation, which already decides its own episode, the
parent-channel note append, and the remote document transfer note.

* no-mistakes(review): seed continuity episode for already recorded break line

* no-mistakes(document): document when a remote-reply continuity break reopens

* no-mistakes(ci): The adapter now stores the continuity episode record before it appends the `blocked` line, so a failed store appends nothing. The changes are uncommitted in the worktree, in `bin/fm-procevent-remote-reply.sh` and `tests/fm-remote-reply.test.sh`. **Invariant:** a continuity `blocked` line is on the parent status log only when the episode record for that break is already stored. Only the `continuity-broken` branch of `cmd_ingest` appends that line, so the fix is at that one place. **What changed in `cmd_ingest`:** - It decides first whether the break needs a line, then stores the episode record, then appends. - If storing the record fails, it stops with "cannot record continuity episode" and appends nothing. - If the line is already on the log and no episode record exists, it stores the record and does not append. **One addition you did not ask for:** if the append fails after the record is stored, the adapter puts the earlier record back (or deletes the new one when none existed). Without that, the retry would read as the unchanged repeat and append nothing, which is the issue 6701 failure again. **Deviation from your test instruction:** the test does not use `chmod` on the cursor directory. `write_continuity_episode` runs `chmod 700` on that directory before every write, so a read-only directory is made writable again. The test instead makes `mktemp` fail for the episode's temporary file in that directory, the same way the existing receipt-failure test does. **Tests added to `tests/fm-remote-reply.test.sh`:** - Store failure: the second break exits 1, appends no `blocked` line and opens no decision. One retry after storage recovers exits 3 and appends a single line. - Append failure: with the status log read-only, the third break exits 1 and appends nothing. One retry after the log is writable exits 3 and appends a single line. **Verification:** `bash tests/fm-remote-reply.test.sh` ends with "ALL TESTS PASSED" with the fix. Against the script at commit b750c828 the same test file fails at "a continuity break appended its line before its episode was stored". `shellcheck -S warning` reports only an unused loop variable at line 667 of the test file, which this change does not touch. I ran no other test files. The Greptile Review check log could not be retrieved, so I worked from the finding text alone

* fix: record a continuity break's reader position on its status line

A later break at another cursor is then a different line, so the existing
duplicate check appends it and reopens the decision. An unchanged re-read
builds the same line and appends nothing.

* fix: reopen a continuity break after an identical restore

A retirement that puts the same bytes back used to rebuild the recorded line, so the later break stayed closed. The retirement count on that line makes the later break distinct.

* no-mistakes(review): remove continuity match for full-prefix line without retirement count

* no-mistakes(document): clarify what a continuity break status line records

* fix: remove the reply cursor before recording retirement

A stop between those steps must leave the count unchanged, so an unchanged continuity break still builds the same line.

* fix(control): drop busy_gen from the task record when an incarnation is retired (#6733)

* fix(control): drop busy_gen when an incarnation is retired

A deliberate exit removed the busy sidecar and left busy_gen in the task record, so the two records disagreed about whether that incarnation was still observable.

* no-mistakes(review): drop GNU-only chmod and unreached sidecar-absent branch

* no-mistakes(review): correct lock comment to name the deadlock

* no-mistakes(ci): The test `test_exit_drops_meta_busy_gen_with_the_sidecar` in tests/fm-control.test.sh now compares the whole task record (the `state/<id>.meta` file), so the Greptile finding is fixed. Invariant: after `exit` retires an incarnation, the task record must equal the record from before `exit` with only the `busy_gen` line removed. This test is the only place in the change that asserts the record survives the rewrite, so it is the only site to fix. The other `busy_gen` tests assert that the line stays, and they do not go through the rewrite. What changed: before `exit`, the test writes the record without its `busy_gen` line to `expected.meta`. After `exit`, the test runs `diff` between that expected copy and the real record, and fails with the diff output if they differ. This one comparison replaces the two earlier checks (no `busy_gen` line left, and the `window` line present), because it covers both. I did not change bin/fm-control.sh or any other file. How I know it works: - I ran `bash tests/fm-control.test.sh`: exit code 0, 45 lines starting with `ok`, no other lines. - I temporarily changed the rewrite in bin/fm-control.sh to also drop the `harness` line. The test then failed with `not ok - exit should drop only busy_gen from the task record:` and the diff `< harness=codex`. The earlier `window`-only check would have passed that rewrite. I restored bin/fm-control.sh afterwards; `git status` shows only tests/fm-control.test.sh modified. - `bash -n` and `shellcheck` on the test file report no new warnings from the edit. The change is not committed; the working tree holds it

* test: cover PID collisions in harness ancestry detection (#6484)

* test(secondmate-harness): scope fake ps -codex label for pid 5252 to the liveness probe

Closes #6456

* no-mistakes(ci): Updated the collision test to log and assert that PID 5252 was queried before selecting 4242. Full fm-secondmate harness suite passes

---------

Co-authored-by: YifuGu <ironerumi@users.noreply.github.com>

* fix: share Pi Calm's working-ship widget slot (#1854)

* fix(calm): share the standalone Pi Calm working-ship widget slot

Firstmate Calm and the user-global standalone Pi Calm both install an
animated working-ship widget during agent runs. Each claimed its own Pi
widget key, so a session loading both (the main Firstmate home) rendered
two boats. Pi replaces widgets under one key, so claiming the shared
"calm-working-ship" slot keeps dual-install sessions to a single boat
while a Firstmate-only session is unchanged.

Pins the shared slot contract in the working-ship module test so the key
cannot silently diverge again.

* test(calm): pin the shared working-ship widget key in CI, document dual-install

The key-parity assertion inside the Pi fixture only runs where the
@earendil-works/pi-coding-agent package is installed, so CI never
exercised it. Add a source-level twin that needs nothing but the
tracked file, and note in docs/calm.md that the boat shares the
standalone Pi Calm working-row widget slot.

* no-mistakes(review): Add executable dual-install widget replacement coverage

* no-mistakes(review): Guard shared widget cleanup with disposal ownership

* no-mistakes(document): Document shared Calm working-ship slot behavior

* test(calm): read the standalone Calm slot from its own module

The dual-install check registered both boats itself under the shared
slot, so it could only prove that Pi replaces a widget under one key: it
would still pass if the standalone Pi Calm extension installed its boat
under a different key, which is the two-boat regression the check exists
to prevent.

Read the standalone extension's own working-ship module when it is
installed - FM_STANDALONE_CALM_SHIP, else ~/.pi/agent/extensions/calm -
and drive the check with the key that module exports, so a rename on
either side registers two widgets and fails naming both keys. A pinned
shared-slot contract still covers a machine without the extension, and
the run reports which side it used instead of passing silently over an
absent extension.

Verified: the touched Pi Calm suite passes and reads the installed
standalone extension; with a copy of it whose key is renamed to
calm-working-ship-v2 the suite fails naming the drift.

* no-mistakes(review): Gate stock-row restoration by shared-widget ownership

* no-mistakes(review): Removed redundant widget-key source assertions

* no-mistakes(document): Document shared Calm working-ship widget ownership

* feat(bin): add opt-in --herdr-resume-lock-wait to fm-spawn (#6649)

* fix(herdr): make exact-resume presentation-lock wait instead of a bounded timeout

The exact-resume path in bin/fm-spawn.sh used the same 50-attempt-then-
give-up lock acquire as the new-task-create path, but the two paths are
not equivalent on contention: a create has no prior state to strand and
can safely fall back to a flat layout, while a resume is recovering a
specific existing identity that a concurrent recovery may legitimately
be holding the lock for. Giving up there does not degrade gracefully,
it hard-fails the resume outright. The suite's own concurrent
cross-home recoveries test already asserts both concurrent recoveries
succeed with a genuine reclaim, and the file's header comment already
(inaccurately) claimed lock contention falls back to the ordinary flat
layout for both paths alike, so the intended contract was always that
recoveries serialize and both succeed, not that either one refuses
under a short bound.

Give spawn_herdr_presentation_order_lock_acquire a wait mode that uses
this file's own established fm_lock_acquire_wait idiom (already used
for its other fleet-shared locks) instead of the bounded loop, and use
it only at the exact-resume call site. The new-task-create call site
is unchanged and keeps its bounded-then-flat-fallback behavior, which
is already covered by its own passing test. Dead-owner PID-liveness
reclaim inside fm_lock_try_acquire still bounds the wait against a
holder that crashed mid-hold.

Adds a deterministic regression test that holds the shared session
lock from an unrelated process for well past the old bound, then
asserts the resume succeeds with a genuine reclaim and took close to
the full hold duration, so a fix that merely widens the bound rather
than genuinely waiting is still caught. The existing concurrent
cross-home recovery test exercises this under real timing but does not
reliably outlast a fixed bound on its own.

Corrects the header comment's claim that create and resume share one
bounded-then-flat-fallback behavior on lock contention; they no longer
do.

* no-mistakes(document): Document Herdr recovery waiting for presentation lock

* no-mistakes(document): Update stale hard-refusal claim in verification log

* no-mistakes(ci): Fixed the Greptile finding on tests/fm-backend-herdr-presentation-e2e.test.sh:1389 by bounding the resume lock-wait regression's spawn_task call. Added an optional 4th `deadline_seconds` arg to the `spawn_task` helper (defaults to empty, so all ~20 other existing call sites are unaffected and unwrapped by `timeout`). The lock-wait test now passes `LOCK_WAIT_HOLD_SECONDS + 60` (90s) as the deadline, and a dedicated check for exit code 124 emits a clear "hung for over Xs instead of waiting out a Ys lock hold" diagnostic before falling through to the existing pass/fail assertions, which are unchanged. No product code was touched. Verified with `bash -n`, `shellcheck -x` (no warnings), a standalone reproduction of the timeout/no-timeout/success paths, the project's `bin/fm-lint.sh --fast` on the file (clean), and the full `tests/fm-lint.test.sh` suite (all 46 assertions pass)

* no-mistakes(ci): Replaced the direct `timeout "$deadline_seconds"` call in `spawn_task()` (tests/fm-backend-herdr-presentation-e2e.test.sh) with the repo's portable bounded-execution helper: sourced `bin/fm-timeout-lib.sh` at the top of the file and changed `deadline_cmd=(timeout "$deadline_seconds")` to `deadline_cmd=(fm_run_timed "$deadline_seconds")`. This removes the GNU/BSD `timeout` dependency that would fail with exit 127 on a stock macOS host without coreutils, while preserving identical semantics (exit 124 on bound-hit, command's own exit otherwise), which the existing `[ "$LOCK_WAIT_STATUS" -eq 124 ]` diagnostic check already relies on. Verified: `bash -n` syntax check, `bin/fm-lint.sh --fast` clean, full `tests/fm-lint.test.sh` suite (46/46 pass), and a standalone repro confirming `fm_run_timed` returns 124 on timeout and 0 on success identically to the prior `timeout` call. No other direct `timeout` calls exist in this file or elsewhere in the PR's diff, so no sibling sites remain

* fix(herdr): gate exact-resume lock wait behind --herdr-resume-lock-wait

Keep refuse-by-default on presentation-order lock contention for Herdr
exact resume. Callers that need concurrent recoveries to serialize must
pass --herdr-resume-lock-wait; unbounded blocking on a third-party session
lock is never the default.

Update docs and the real-Herdr e2e suite so the default path asserts the
refusal and the opt-in path asserts the wait.

* no-mistakes(test): Fix e2e test's lost exit status after if/fi with no else branch

* docs(herdr): stop advertising --herdr-resume-lock-wait on --relaunch

The relaunch path reuses the recorded endpoint and never takes the
presentation-order lock, so the flag is inert there. Drop it from the
--relaunch usage line and state where the flag applies.

* no-mistakes(review): Clarify lock-wait docs; simplify bash-3.2-safe spawn_task helper

* no-mistakes(ci): Fixed ci-1 (Greptile P2). In tests/fm-backend-herdr-presentation-e2e.test.sh, the failure cleanup `cleanup_all` stopped only `LOCK_CONTENTION_OWNER_PID`. It now also stops `LOCK_REFUSE_HOLDER_PID` and `LOCK_WAIT_HOLDER_PID`, the holders of the two new contention cases, so a `fail` before their explicit `wait` no longer leaves them running. Both new PIDs are initialised empty next to the existing one, and each is cleared right after its successful `wait` so cleanup never touches a finished PID. I changed nothing else. `bash -n` passes. The real Herdr e2e run passed both new cases ("default resumed identity refuses session lock contention" and "--herdr-resume-lock-wait waits out session lock contention instead of refusing"). The full run hit my 550s timeout in a later, unrelated case, after the new cases passed

* fix(bin): let TERM stop a watcher blocked in a pane capture on bash 3.2 (#6762)

* fix(bin): let TERM stop a watcher blocked in a pane capture on bash 3.2

Stock macOS bash 3.2 holds a HUP or TERM until a running command
substitution's child exits, and the watcher read every pane through
$(fm_backend_capture ...). A blocked backend read therefore held the
watcher's stop for as long as the read lasted, and a stopped watcher left
the hung read orphaned. tests/fm-watch-triage.test.sh
test_term_stops_a_watcher_blocked_inside_a_poll failed on /bin/bash 3.2
for this reason while passing on bash 5.

Pane captures now go through watcher_capture, which runs the read as a
waited background process group recorded like a check's, so the stop is
honored at once and watcher_cleanup stops a read still in flight along
with its per-call output file.

* no-mistakes(review): Run drain-ring idle capture in watcher shell, add regression test

* no-mistakes(document): Document watcher TERM handling for blocked checks and captures

* no-mistakes(test): Silence bash 3.2 setpgid race noise from watcher captures

* fix(bin): verify the capture group and scope the stop claim to pane reads

watcher_capture now confirms its background read leads its own process
group, as run_check_capture already does, so watcher_cleanup never relies
on a group that set -m failed to create. The comment and continuity doc
now say only fm_backend_capture pane reads go through watcher_capture;
agent-state and composer-state reads still run inside command
substitutions.

* feat: add per-home worker tool exclusions (#6750)

* feat(spawn): add per-home worker tool exclusions

Add an optional per-home config/crew-exclude-tools file listing tool names to hide from workers, one per line, with blank lines and # comments allowed.
It applies to every ship and scout launch and relaunch in that home, is never inherited by another home, and does not affect secondmate agents.
Pi and pi-signed apply it through --exclude-tools, which also covers MCP tool names.
Any other runtime, and a raw launch command, refuses the launch when the list is non-empty rather than ignoring it.
Malformed entries are refused before provisioning, and before a relaunch stops a running worker.
Exclusions that match no tool in the worker's loaded registry are reported as unverified warnings in its status record instead of refusing the worker.

Closes #6744

* no-mistakes(review): Preserve UTF-8 exclusion paths and verify Pi lifecycle behavior

* no-mistakes(document): Clarify worker tool exclusion documentation

* no-mistakes(ci): Fixed ci-1 in bin/fm-exclude-tools-lib.sh: a failed read now returns an error before printing names, so all shared launch and relaunch callers refuse rather than silently dropping exclusions. Added deterministic regression coverage for a file disappearing after readability checks across Pi/pi-signed ship and scout launches. Reproduced the original failure; verified 83 spawn checks, 77 relaunch checks, direct parser/runtime failure cases, full targeted lint, Bash syntax, and git diff --check. Relaunch tests passed with existing fixture-cleanup permission warnings. ci-2 remains unchanged per the user's decision; the outer executor owns the fresh CI run

* feat(bin): defer spawns beyond a declared per-project capacity, opt-in (#5343)

* refactor(bin): share the local Firstmate home walk from the wake library

Teardown's walk over the root home and its registered local secondmate homes
moves into bin/fm-wake-lib.sh as fm_local_firstmate_state_dirs, next to
fm_firstmate_root_home, so a second consumer can count task records across
this machine's homes without a copy. Teardown keeps its exact refusal wording
through a thin wrapper.

* feat(bin): defer spawns beyond a project's declared machine capacity

A project whose machine-local resource only serves a few workers at once had
no way to tell Firstmate so: every queued item was launched, and the surplus
workers spent full-context turns retrying the resource.

config/project-capacity in the root home now declares how many workers each
named project admits at once on this machine. bin/fm-spawn.sh counts the ship
and scout records on the same project origin across the root and its local
secondmate homes, skipping ones whose ready PR is recorded, while holding the
shared project lock through publication. A spawn with every place held exits 75
before any brief render, endpoint, worktree, record, or backlog move, so the
item stays queued; batches report it as deferred. Undeclared projects keep
today's uncapped dispatch, and an unreadable declaration refuses rather than
guessing the limit.

Refs #4237

* no-mistakes(review): Document that capacity matches the clone directory name

* no-mistakes(document): Rewrap stale fm-wake-lib root-home doc comment

* no-mistakes(review): Dedupe local state dirs by identity to avoid double-counting

* no-mistakes(document): Rewrap fm_local_firstmate_state_dirs error doc comment

* no-mistakes(ci): I fixed all four Greptile findings. All 14 tests in tests/fm-project-capacity.test.sh pass, and shellcheck at warning level is clean on the changed files. Each new test failed against the old code and passes now. - **ci-1 (spaced names):** a declaration line must give a name its capacity whenever the name is a valid clone directory name. `fm_project_capacity_lookup` now trims each line, skips blank lines and lines whose first non-blank character is `#`, and takes the last field as the capacity. Everything before that field is the name, so it may contain spaces. The old error cases still refuse: a single field is rejected, and trailing text leaves a last field that is not an integer. The library header and docs/configuration.md now say a name starting with `#` cannot be declared. New test `test_spaced_project_name_is_declared` declares `my heavy project 1` next to an indented comment line and gets a deferral. - **ci-2 (unreadable records):** the holder count must never silently leave out a holder. `fm_project_capacity_occupants` now refuses when a local home's state directory exists but cannot be read or listed, or when a `.meta` file cannot be read. The error names the path, and `fm-spawn.sh` shows it in its existing refusal message. New test `test_unreadable_holders_refuse_admission` covers an unreadable record in the root home and an unreadable state directory in a registered local secondmate home, then checks that the spawn is admitted once both are readable. The test is skipped when run as root. - **ci-3 (Orca lock):** any spawn that can become a holder for a capped project must take that project's lock. The lookup now also reports whether the declaration caps any project at all, and an Orca spawn takes the per-origin lock whenever it does. This covers every capped same-origin clone. It also covers some cases where no same-origin clone is capped, because a spawn cannot find clones under other directory names without searching for them. With no declaration file, Orca still skips the lock. The comments in the library and in the `fm-spawn.sh` header are updated. The Orca test now clones the origin as `project-2`, which has no declaration, and checks that its Orca spawn refuses while the lock is held and publishes no record. - **ci-4 (worktrees):** `assert_nothing_created` now also compares the project's `git worktree list` from before and after a deferred spawn. Both tests that call it take that snapshot first. Files changed: bin/fm-project-capacity-lib.sh, bin/fm-spawn.sh, docs/configuration.md, tests/fm-project-capacity.test.sh

* fix(bin): declare capacity for a project name that begins with #

A clone directory whose name begins with # was skipped as a comment, so that project stayed uncapped. A line is a declaration when the # is written against the rest of the name and the line ends with a capacity; a # followed by whitespace stays a comment.

* no-mistakes(document): Rewrap project-capacity library header comment

* no-mistakes(ci): Lint 2 fails because this PR's code pushes ShellCheck past its memory cap. ShellCheck ran out of memory analyzing bin/fm-teardown.sh in CI (reason=memory, rc=251, peak about 8.39 GB). On current main the same file passes at about 7.29 GB. **Cause:** the new `fm_local_firstmate_state_dirs` function in bin/fm-wake-lib.sh had a conditional `. fm-secondmate-registry-lib.sh` with a `# shellcheck source=` directive inside the function. ShellCheck followed that source again, inside a function scope, wherever fm-wake-lib.sh is sourced, and bin/fm-teardown.sh is the heaviest root that sources it. Measured locally with `shellcheck --norc --external-sources bin/fm-teardown.sh`: - current main (fd325b1b): 7.29 GB - main merged with this PR: 7.86 GB - the same merge without the in-function source: 7.27 GB **Rule this restores:** this change must not make any lint root heavier than it is on main. That function holds the only new nested source in the change. **Fix:** I removed the in-function source, which no caller needs. Both callers already load the registry library at top level before calling the function: - bin/fm-teardown.sh sources it directly. - bin/fm-spawn.sh, the only user of bin/fm-project-capacity-lib.sh, gets it through bin/fm-ff-lib.sh. I also documented the requirement in the function's comment and in the "Requires" note in bin/fm-project-capacity-lib.sh. No behaviour changes. **Verification:** - ShellCheck on head: bin/fm-teardown.sh peaks at 7.12 GB and bin/fm-spawn.sh at 6.68 GB, both with rc=0. bin/fm-wake-lib.sh and bin/fm-project-capacity-lib.sh lint clean. - tests/fm-project-capacity.test.sh, tests/fm-teardown.test.sh (102 ok) and tests/fm-teardown-endpoint-safety.test.sh all pass. Files changed: bin/fm-wake-lib.sh, bin/fm-project-capacity-lib.sh

* fix(bin): release the Herdr session lock when reclaim finishes

A concurrent resume in another home waits five seconds for that lock.
Reclaim is the last presentation change on the recovery path, so holding
the lock through the launch tail made the waiter time out. The contributions
arm check also freezes its one-second clock, the same way the budget tests
do, because an unfrozen clock can tick past before the first forge read.

* no-mistakes(review): Keep Herdr session lock through launch handoff after reclaim

* no-mistakes(review): Skip the spawning task's own record in capacity count

* no-mistakes(review): Restore release test comment above its test

* docs: scope PR-ready re-evaluation to a declared project capacity

A ready pull request frees a place only when that project declares capacity, so the always-loaded backlog contract should re-evaluate on that handoff only in that case.

* fix(bin): name the Lavish read message count by the same label as its section (#6785)

fm-procevent-lavish.sh read labels tag=message rows SESSION-ENDING MESSAGE
only when session_ended is true and CAPTAIN MESSAGE otherwise, but the
count line always said session_ending_message_count. Several composer
messages on a still-open board were therefore counted as session-ending.

The count line now follows the same session_ended switch:
session_ending_message_count once the session ended, captain_message_count
otherwise. Message rows stay out of the annotation count, per triage.

Fixes #6743

* fix(herdr): make the presentation lock namespace per OS account (#6780)

The Herdr presentation lock namespace was the fixed machine-global
/tmp/firstmate-herdr-presentation, so on a host where two OS users run
Firstmate on Herdr the first account to create it owned it and every
teardown from the other account was refused with no way to clear it.

Suffix the namespace with the account uid. The owner-uid and mode-700
checks are unchanged, so a foreign-owned or wrong-mode name at this
account's path is still refused and never adopted, chowned, or removed.

Fixes #4716.

* Fix OpenCode arm plugin to decide with the shared supervision predicate (#6809)

The OpenCode session plugin's shouldArm kept its own copy of the need
test that only looked for in-flight task records, while the turn-end
guard decides with fm_supervision_needed in bin/fm-supervision-lib.sh,
which also counts registered process-event sources and trusted custom
checks. With an empty fleet but any registered source or check, the
guard blocked every turn end while the plugin declined to arm - a loop
the guard's own repair line could not resolve because it names the
plugin as the fix.

The plugin now delegates the decision to the shared predicate through
bash, keeping the local away-record decline and the x-mode.env arm
override. OpenCode plugin test fixtures now carry the real predicate
their arming path sources, and the arm suite gains six cases asserting
the plugin's decision against the shared verdict over the same
synthetic state directories.

Co-authored-by: Mia Sun <mia@Bigs-Mac-mini.localdomain>

* fix(bin): resolve a pending reply only from its own task's status line (#6792)

* fix(bin): resolve a pending reply only from its own task's status line

…
nathanjgaul-agile added a commit to nathanjgaul-agile/firstmate that referenced this pull request Oct 9, 2026
* feat(bin): defer spawns beyond a declared per-project capacity, opt-in (kunchenguid#5343)

* refactor(bin): share the local Firstmate home walk from the wake library

Teardown's walk over the root home and its registered local secondmate homes
moves into bin/fm-wake-lib.sh as fm_local_firstmate_state_dirs, next to
fm_firstmate_root_home, so a second consumer can count task records across
this machine's homes without a copy. Teardown keeps its exact refusal wording
through a thin wrapper.

* feat(bin): defer spawns beyond a project's declared machine capacity

A project whose machine-local resource only serves a few workers at once had
no way to tell Firstmate so: every queued item was launched, and the surplus
workers spent full-context turns retrying the resource.

config/project-capacity in the root home now declares how many workers each
named project admits at once on this machine. bin/fm-spawn.sh counts the ship
and scout records on the same project origin across the root and its local
secondmate homes, skipping ones whose ready PR is recorded, while holding the
shared project lock through publication. A spawn with every place held exits 75
before any brief render, endpoint, worktree, record, or backlog move, so the
item stays queued; batches report it as deferred. Undeclared projects keep
today's uncapped dispatch, and an unreadable declaration refuses rather than
guessing the limit.

Refs kunchenguid#4237

* no-mistakes(review): Document that capacity matches the clone directory name

* no-mistakes(document): Rewrap stale fm-wake-lib root-home doc comment

* no-mistakes(review): Dedupe local state dirs by identity to avoid double-counting

* no-mistakes(document): Rewrap fm_local_firstmate_state_dirs error doc comment

* no-mistakes(ci): I fixed all four Greptile findings. All 14 tests in tests/fm-project-capacity.test.sh pass, and shellcheck at warning level is clean on the changed files. Each new test failed against the old code and passes now. - **ci-1 (spaced names):** a declaration line must give a name its capacity whenever the name is a valid clone directory name. `fm_project_capacity_lookup` now trims each line, skips blank lines and lines whose first non-blank character is `#`, and takes the last field as the capacity. Everything before that field is the name, so it may contain spaces. The old error cases still refuse: a single field is rejected, and trailing text leaves a last field that is not an integer. The library header and docs/configuration.md now say a name starting with `#` cannot be declared. New test `test_spaced_project_name_is_declared` declares `my heavy project 1` next to an indented comment line and gets a deferral. - **ci-2 (unreadable records):** the holder count must never silently leave out a holder. `fm_project_capacity_occupants` now refuses when a local home's state directory exists but cannot be read or listed, or when a `.meta` file cannot be read. The error names the path, and `fm-spawn.sh` shows it in its existing refusal message. New test `test_unreadable_holders_refuse_admission` covers an unreadable record in the root home and an unreadable state directory in a registered local secondmate home, then checks that the spawn is admitted once both are readable. The test is skipped when run as root. - **ci-3 (Orca lock):** any spawn that can become a holder for a capped project must take that project's lock. The lookup now also reports whether the declaration caps any project at all, and an Orca spawn takes the per-origin lock whenever it does. This covers every capped same-origin clone. It also covers some cases where no same-origin clone is capped, because a spawn cannot find clones under other directory names without searching for them. With no declaration file, Orca still skips the lock. The comments in the library and in the `fm-spawn.sh` header are updated. The Orca test now clones the origin as `project-2`, which has no declaration, and checks that its Orca spawn refuses while the lock is held and publishes no record. - **ci-4 (worktrees):** `assert_nothing_created` now also compares the project's `git worktree list` from before and after a deferred spawn. Both tests that call it take that snapshot first. Files changed: bin/fm-project-capacity-lib.sh, bin/fm-spawn.sh, docs/configuration.md, tests/fm-project-capacity.test.sh

* fix(bin): declare capacity for a project name that begins with #

A clone directory whose name begins with # was skipped as a comment, so that project stayed uncapped. A line is a declaration when the # is written against the rest of the name and the line ends with a capacity; a # followed by whitespace stays a comment.

* no-mistakes(document): Rewrap project-capacity library header comment

* no-mistakes(ci): Lint 2 fails because this PR's code pushes ShellCheck past its memory cap. ShellCheck ran out of memory analyzing bin/fm-teardown.sh in CI (reason=memory, rc=251, peak about 8.39 GB). On current main the same file passes at about 7.29 GB. **Cause:** the new `fm_local_firstmate_state_dirs` function in bin/fm-wake-lib.sh had a conditional `. fm-secondmate-registry-lib.sh` with a `# shellcheck source=` directive inside the function. ShellCheck followed that source again, inside a function scope, wherever fm-wake-lib.sh is sourced, and bin/fm-teardown.sh is the heaviest root that sources it. Measured locally with `shellcheck --norc --external-sources bin/fm-teardown.sh`: - current main (fd325b1): 7.29 GB - main merged with this PR: 7.86 GB - the same merge without the in-function source: 7.27 GB **Rule this restores:** this change must not make any lint root heavier than it is on main. That function holds the only new nested source in the change. **Fix:** I removed the in-function source, which no caller needs. Both callers already load the registry library at top level before calling the function: - bin/fm-teardown.sh sources it directly. - bin/fm-spawn.sh, the only user of bin/fm-project-capacity-lib.sh, gets it through bin/fm-ff-lib.sh. I also documented the requirement in the function's comment and in the "Requires" note in bin/fm-project-capacity-lib.sh. No behaviour changes. **Verification:** - ShellCheck on head: bin/fm-teardown.sh peaks at 7.12 GB and bin/fm-spawn.sh at 6.68 GB, both with rc=0. bin/fm-wake-lib.sh and bin/fm-project-capacity-lib.sh lint clean. - tests/fm-project-capacity.test.sh, tests/fm-teardown.test.sh (102 ok) and tests/fm-teardown-endpoint-safety.test.sh all pass. Files changed: bin/fm-wake-lib.sh, bin/fm-project-capacity-lib.sh

* fix(bin): release the Herdr session lock when reclaim finishes

A concurrent resume in another home waits five seconds for that lock.
Reclaim is the last presentation change on the recovery path, so holding
the lock through the launch tail made the waiter time out. The contributions
arm check also freezes its one-second clock, the same way the budget tests
do, because an unfrozen clock can tick past before the first forge read.

* no-mistakes(review): Keep Herdr session lock through launch handoff after reclaim

* no-mistakes(review): Skip the spawning task's own record in capacity count

* no-mistakes(review): Restore release test comment above its test

* docs: scope PR-ready re-evaluation to a declared project capacity

A ready pull request frees a place only when that project declares capacity, so the always-loaded backlog contract should re-evaluate on that handoff only in that case.

* fix(bin): name the Lavish read message count by the same label as its section (kunchenguid#6785)

fm-procevent-lavish.sh read labels tag=message rows SESSION-ENDING MESSAGE
only when session_ended is true and CAPTAIN MESSAGE otherwise, but the
count line always said session_ending_message_count. Several composer
messages on a still-open board were therefore counted as session-ending.

The count line now follows the same session_ended switch:
session_ending_message_count once the session ended, captain_message_count
otherwise. Message rows stay out of the annotation count, per triage.

Fixes kunchenguid#6743

* fix(herdr): make the presentation lock namespace per OS account (kunchenguid#6780)

The Herdr presentation lock namespace was the fixed machine-global
/tmp/firstmate-herdr-presentation, so on a host where two OS users run
Firstmate on Herdr the first account to create it owned it and every
teardown from the other account was refused with no way to clear it.

Suffix the namespace with the account uid. The owner-uid and mode-700
checks are unchanged, so a foreign-owned or wrong-mode name at this
account's path is still refused and never adopted, chowned, or removed.

Fixes kunchenguid#4716.

* Fix OpenCode arm plugin to decide with the shared supervision predicate (kunchenguid#6809)

The OpenCode session plugin's shouldArm kept its own copy of the need
test that only looked for in-flight task records, while the turn-end
guard decides with fm_supervision_needed in bin/fm-supervision-lib.sh,
which also counts registered process-event sources and trusted custom
checks. With an empty fleet but any registered source or check, the
guard blocked every turn end while the plugin declined to arm - a loop
the guard's own repair line could not resolve because it names the
plugin as the fix.

The plugin now delegates the decision to the shared predicate through
bash, keeping the local away-record decline and the x-mode.env arm
override. OpenCode plugin test fixtures now carry the real predicate
their arming path sources, and the arm suite gains six cases asserting
the plugin's decision against the shared verdict over the same
synthetic state directories.

Co-authored-by: Mia Sun <mia@Bigs-Mac-mini.localdomain>

* fix(bin): resolve a pending reply only from its own task's status line (kunchenguid#6792)

* fix(bin): resolve a pending reply only from its own task's status line

Remote reply ingestion handed every corr= token in a mate's payload to
fm_pending_reply_try_resolve together with that mate's own status log,
so one mate echoing another mate's token resolved the other request.
Honor a status-file override only when it is the record's own
parent_status, and match the corr= token as a whole word.

Fixes kunchenguid#6538

* no-mistakes(document): docs: scope remote reply settlement to the asked mate

* docs(skills): index the six missing agent-only skill triggers (kunchenguid#6784)

agent-skill-trigger-index claims to be the complete agent-only trigger
index but omitted operational-home-layout, session-start-recovery,
validation-supervision, ship-landing, scout-completion, and
away-quiet-supervision. Add each with its own description's trigger,
placed beside the related entries.

The decision-hold-lifecycle redirect stub stays out, per triage.

Fixes kunchenguid#6503

* test: make agent process fixtures compatible with multicall sleep (kunchenguid#6814)

* test: share a rename-safe agent stand-in across liveness suites

On Ubuntu 26.04, `sleep` is the uutils multicall binary, which refuses to
run when invoked through a symlink named after another utility. The Herdr
descendant process-walk tests built their agent-named process as a `pi`
symlink to the host `sleep`, so the process exited at once, its parent shell
was gone before the walk ran, and both cases read `unknown unreadable` and
failed on that host. The suite stops at its first failure, so every later
case went unrun. The Herdr control smoke test's `claude` symlink has the
same construction.

The tmux liveness suite already solved this with a host-compiled spinner and
a survival-checked `sleep` fallback. That builder moves into tests/lib.sh as
fm_agent_standin, and the tmux suite, both Herdr descendant cases, and the
Herdr control smoke test now use it. When no stand-in can survive a foreign
name, a case skips with the reason instead of failing.

tests/fm-test-fixtures.test.sh gains a portable regression with a fake
multicall `sleep`, so it bites on hosts whose own `sleep` is single-purpose.

* no-mistakes(document): Correct Herdr verification fixture reference

* ci: retrigger cancelled shard

* fix(bin): keep the supervision host's successor watcher alive after the Stop hook's group is torn down (kunchenguid#6787)

* fix(bin): keep the supervision host's pass-through successor out of the hook's process group

The successor a main-only pass-through leaves for main shared the Stop hook's
process group, so the harness tearing that group down after the exit-2 rewake
stopped it. The stop published downtime and the next park's first cycle
announced an empty check: rearm-resurface, which woke main again in a loop.
Start that successor in a process group of its own, as the hook's own
handling successor already is.

* no-mistakes(review): Give the at-turn successor left for main its own group

* no-mistakes(document): Document own-group successor for turn-start hand-back too

* no-mistakes(ci): I made the change you asked for: both new teardown tests in tests/fm-supervision-host.test.sh now call the existing `stop_home_processes "$home"` just before `pass`. The tests are `test_successor_left_at_the_turn_survives_the_hook_process_group_teardown` and `test_pass_through_successor_survives_the_hook_process_group_teardown`. No production code and no other tests changed. The rule broken was that a test must not leave a home's watcher or arm processes running after it passes. These two were the only cases in the changed area that broke it. The other host+hook tests already stop their home, and `test_successor_close_during_main_turn_is_delivered_at_the_next_turn_end` leaves its watcher behind too, but it is an older test you said not to touch. The only reason anything was left over is that the successor's arm now sits in its own process group, outside the hook's teardown. `stop_home_processes` kills the watcher by the pid in its lock file, which stops it no matter which group it is in. **Checks run:** - I ran just these two tests from a scratch copy of the suite (since deleted). Both pass in about 13 seconds. - After each test, a process listing filtered to that test's home directory came back empty once the processes had about a second to exit after TERM. - `bash -n` on the test file passes. - `shellcheck` is not installed here, so I did not lint the file. - I did not run the full serial-2 suite locally. Whether it now finishes under its 30-minute limit will only show on the next CI run

* test: make tmux Claude readiness checks independent of permission footers (kunchenguid#6823)

* test: use idle composer readiness for Claude tmux guards

* no-mistakes(test): Fix attended supervision test expectations and isolate worker state

* no-mistakes(document): Correct live guard coverage and readiness documentation

* no-mistakes(ci): Captain, fixed SC2100 by quoting the cursor-agent assignment in tests/fm-host-mirror-live-e2e.test.sh. Reproduced the failure before editing; pinned ShellCheck lint on both PR test files, bash syntax checks, and git diff --check now pass

* test: preserve attended successor close assertions

* no-mistakes(test): Fix attended live test watcher takeover expectations

* no-mistakes(document): Correct stale attended guard documentation

* Revert "no-mistakes(document): Correct stale attended guard documentation"

This reverts commit 8e59d89.

* Revert "no-mistakes(test): Fix attended live test watcher takeover expectations"

This reverts commit c0b8510.

* test(herdr): accept safe exit refusals and clean lab once (kunchenguid#6818)

* test: stabilize watcher lock race fixtures (kunchenguid#6887)

* test: order stale watcher-lock fixture races explicitly

* no-mistakes(review): Removed duplicate stale-steal reap call

* Rerun CI after the merge

---------

Co-authored-by: Tiago <tiagop@hey.com>
Co-authored-by: yairtech <39274208+falkoro@users.noreply.github.com>
Co-authored-by: Asser AboElkhair <asser.aboelkhair@gmail.com>
Co-authored-by: dubiousenvelope <greg.ecklin@gmail.com>
Co-authored-by: Mia Sun <mia@Bigs-Mac-mini.localdomain>
Co-authored-by: Pedro Guimarães <21346846+0x7067@users.noreply.github.com>
Co-authored-by: menidi <menidi@users.noreply.github.com>
Co-authored-by: YifuGu <39033099+ironerumi@users.noreply.github.com>
witt3rd added a commit to witt3rd/firstmate that referenced this pull request Oct 10, 2026
* fix: prevent unnecessary remote worker turnover (#6431)

* fix: prevent healthy remote job worker turnover

* no-mistakes(review): Serialize full LaunchAgent repair and verify launchd-tracked owners

* no-mistakes(review): Let launchd-tracked unpublished spawns start before reloading

* no-mistakes(document): Document remote worker heartbeat and serialized LaunchAgent recovery

* no-mistakes(ci): Full CI log showed the idle-worker regression exceeded its outdated command budget (82 versus 80) after independent heartbeat ownership checks were added. Raised the budget to 120 while retaining the separate busy-poll sleep limit. Remote-job and LaunchAgent executable tests passed, as did bash syntax validation and git diff --check. No production behavior changed

* no-mistakes(ci): Fixed missing readiness recovery under verified live ownership, preserving the serving PID and private file mode. Added executable regressions for deletion during a blocked sweep and stale readiness diagnostics without LaunchAgent reload. Deletion regression failed before the fix. Both remote-job test suites, bash syntax validation, and git diff --check passed

* no-mistakes(ci): Full CI log identified a flaky ownership-loss test racing an already-authorized heartbeat refresh. Replaced backdating and a fixed sleep with bounded observation of readiness expiry through the public probe. Production behavior unchanged. Remote-job and LaunchAgent executable suites passed; bash syntax validation and git diff --check passed

* fix: wait for launchd bootout cleanup

* no-mistakes(document): Document remote worker recovery and read-only turnover verification

* no-mistakes(review): Publish worker identity before lock owner records

* no-mistakes(document): Document worker identity publication safety invariant

* no-mistakes(ci): Fixed ci-1: replacement workers discard predecessor readiness before publishing identity and roll back identity if lock-owner recording fails. Added executable regressions reproducing both defects. LaunchAgent and orphan-reap suites, shell syntax checks, and git diff --check passed. No live service state was modified

* no-mistakes(ci): Restored lock-owner-before-identity publication and removed the identity rollback and reordering-only tests. Retained an executable regression proving predecessor readiness is rejected until replacement startup completes. LaunchAgent and orphan-reap suites, shell syntax checks, and git diff --check passed. Other changes remain intact; the retained-identity interrupted-repair edge remains out of scope. No live service state was modified

* fix: speed up remote-job sequence claim cleanup (#6575)

* fix(remote-job): reap expired seq claims with one directory walk

The hourly claim sweep forked uname+stat per .seq-claims entry and blocked
serving for ~85s at ~17k dirs. Delete expired empty claim dirs with a single
find -exec rmdir batch and cache the host uname for remaining mtime reads.

* no-mistakes(review): Restore original path mtime helper and drop uname cache

* no-mistakes(test): Restore claim retention eligibility; focused retention and serving tests pass

* no-mistakes(document): Document single-walk claim cleanup and regression entrypoints

* no-mistakes(ci): Fixed Lint 2’s reproduced SC1091 by adding the tests/lib.sh ShellCheck source directive to the retention test. Runtime behavior is unchanged. ShellCheck passed for both new claim tests; Bash syntax, the retention behavior test, and git diff --check passed

* fix(tests): keep fixture registries out of git worktree roots

A TMPDIR pointed at a repository root placed live .fm-test-* registries
beside tracked files, and a concurrent git add during the claim-walk CI
fix round committed three of them. Route registries and fixture roots
through a TMPDIR that refuses git worktree roots, remove the stray files,
and pin the escape with a behavioral cleanup test.

* no-mistakes(review): Preserve whole-second claim expiry in single-walk sweep

* no-mistakes(document): Correct temporary-directory resolution documentation

* no-mistakes(ci): Fixed ci-3 by changing only the stale, fresh, and read-only orphan fixture paths in tests/fm-test-fixture-cleanup.test.sh to use $FM_TEST_TMPDIR. All seven tests passed both normally and with TMPDIR set to the worktree root. Shell syntax and git diff --check passed

* feat(bin): record no-mistakes pipeline spend per task at teardown, opt-in (#5354)

* feat(bin): record each task's no-mistakes pipeline spend at cleanup

no-mistakes keeps every pipeline agent invocation's token usage only in its
local agent_invocations records, and cold pipeline agents (review, test,
document) leave no session log, so a task's review-loop cost never reached
Firstmate's records and could not be attributed after cleanup.

bin/fm-pipeline-spend.sh attributes a task's runs by the repository
no-mistakes resolves for the task copy, the task branch, and the branch's
creation (a relaunch mints a new spawn_gen while the branch keeps
validating), then sums every invocation, failed and cancelled included.
Token fields use no-mistakes' per-round deltas so resumed review rounds are
not counted twice, and unrecorded values stay unknown rather than zero.
show prints the record read-only; record appends it once per task
incarnation to data/pipeline-spend.jsonl. Teardown records it for every
ship task it cleans up, before deleting the branch and the task record.

fm_nm_state_db becomes the one owner of where no-mistakes' state database
lives, shared with the capped run-inventory reader.

* no-mistakes(review): Drop spend show command and timeout override

* style(bin): rewrap fm-pipeline-spend header comment

* Make pipeline spend recording opt-in

* feat(bin): record each task's no-mistakes pipeline spend at cleanup

no-mistakes keeps every pipeline agent invocation's token usage only in its
local agent_invocations records, and cold pipeline agents (review, test,
document) leave no session log, so a task's review-loop cost never reached
Firstmate's records and could not be attributed after cleanup.

bin/fm-pipeline-spend.sh attributes a task's runs by the repository
no-mistakes resolves for the task copy, the task branch, and the branch's
creation (a relaunch mints a new spawn_gen while the branch keeps
validating), then sums every invocation, failed and cancelled included.
Token fields use no-mistakes' per-round deltas so resumed review rounds are
not counted twice, and unrecorded values stay unknown rather than zero.
show prints the record read-only; record appends it once per task
incarnation to data/pipeline-spend.jsonl. Teardown records it for every
ship task it cleans up, before deleting the branch and the task record.

fm_nm_state_db becomes the one owner of where no-mistakes' state database
lives, shared with the capped run-inventory reader.

* no-mistakes(review): Drop spend show command and timeout override

* style(bin): rewrap fm-pipeline-spend header comment

* Make pipeline spend recording opt-in

* no-mistakes(document): Note pipeline-spend opt-in gate in teardown comments

* no-mistakes(ci): I fixed all three Greptile findings following your decision. Both test suites pass: tests/fm-teardown.test.sh (105 ok, exit 0) and tests/fm-pipeline-spend.test.sh (8 ok). I didn't run either new test against the old code, so the claim that they fail before the fix is from reading the code, not a run. Nothing was committed or pushed. **ci-1 / ci-2 (bin/fm-teardown.sh):** The rule is that every owned ship task in a home that has opted in gets a durable spend record before its task record is deleted. The `[ -d "$WT" ]` check skipped the recorder when the worktree folder was already gone, so no record was written. I removed that check and kept the other conditions: the task must be a ship task, teardown must own its worktree, and `config/pipeline-spend` must exist. `bin/fm-pipeline-spend.sh` already writes an unavailable-source line when the worktree is missing. The new test `test_teardown_records_unavailable_spend_for_a_gone_worktree` in tests/fm-teardown.test.sh sets up an owned ship task whose worktree is missing, with recording turned on. It checks that teardown succeeds, writes a line with `source=="unavailable"`, `total==null` and a reason saying the task copy is gone, and still removes the task record. Under the old check no ledger would have been written. **ci-3 (bin/fm-nm-run-lib.sh):** When `NM_HOME` and `HOME` were both unset, `fm_nm_state_db` fell back to `${HOME:-}/.no-mistakes`, which resolves to `/.no-mistakes`. It now uses `root=~/.no-mistakes`. Bash expands `~` to the account's home directory even when `HOME` is unset, which matches the previous Python `Path.home()` lookup. The new test `test_state_db_without_nm_home_or_home_uses_the_account_home` in tests/fm-pipeline-spend.test.sh runs the function with both variables unset and compares the result to the home directory from the system's password database. The old code would have returned `/.no-mistakes/state.sqlite`. shellcheck reports nothing new on the changed files. I added one `disable=SC2016` comment in the new test, where the single-quoted `$1` is meant to expand in the child shell

* feat(bin): start tasks from a named base branch (#6442)

* feat(bin): start tasks from a named base branch

Spawns always reset a task's pooled copy to origin's default branch,
so work that belongs on a feature, integration, or release branch
started from the wrong code and opened its PR against the default.

fm-brief.sh --base-branch records the base in the brief, fm-spawn.sh
resets the copy to origin/<base> and records base_branch= in task
meta, and the worker targets its PR at that branch. Review diffs,
the cleanup content check, and scout promotion read the recorded
base. local-only and Gerrit deliveries refuse a named base.

* no-mistakes(review): Anchor brief base-branch parse on scaffold Setup line

* no-mistakes(review): Require explicit --base-branch on spawn, matching brief lines

* no-mistakes(document): Note base-branch reset in fm-spawn freshness header

* no-mistakes(ci): I fixed all four Greptile findings the way you specified. The affected suites pass locally: fm-brief, fm-dod-lib, fm-spawn-pool-base-freshen, fm-review-diff, fm-teardown and fm-task-delivery. shellcheck on the changed files reports only info-level notes, no warnings or errors. I didn't revert the code to watch each new test fail before the fix. - **ci-1 (brief text blocked or redirected spawns):** In bin/fm-dod-lib.sh, `fm_brief_base_branches` now only accepts a `Base branch: <value>` line that comes directly after the base-variant Setup sentence fm-brief.sh writes. Any other `Base branch:` line is ignored. `--base-branch` is still the only authority, and the existing checks in fm-spawn.sh are unchanged. With the flag, the spawn is refused unless a matching pair exists and no pair names a different branch. Without the flag, it is refused only if such a pair exists. The header comment in fm-spawn.sh is updated to match. - The `brief_with_base` test helper now writes the real two-line pair. - The decoy refusal case now uses a full decoy pair in the captain's intent. - New test `test_prose_base_branch_line_is_ignored`: a brief with no base whose intent quotes `Base branch: release/1.2` spawns normally with no `base_branch=` recorded. A spawn with `--base-branch feature/hub` whose intent names release/1.2 starts from `origin/feature/hub` and records it. - **ci-2 (scout base not checked against the forge):** fm-spawn.sh now looks up the project's registered forge with `bin/fm-project-mode.sh --forge`, like the ship path does, before `fm_base_branch_valid` runs. Previously it passed `none`. New test `test_scout_base_branch_refused_on_gerrit_forge`: a scout with a base on a forge=gerrit project is refused at spawn and no meta file is written. - **ci-3 (base not shell-quoted in worker commands):** `fm_dod_block` now quotes the base with `printf %q` in the `--base` and `--base-branch` arguments, and the branch name in the prose stays readable. The test in fm-brief.test.sh uses the base `release/$HOTFIX` and checks that both direct-PR and no-mistakes briefs render `release/\$HOTFIX` in the commands. - **ci-4 (ambiguous PR sentence):** The direct-PR sentence now reads "open a PR with `gh-axi` that is ready for review, not a draft, against the base branch `X` (`--base X`), not the repository default." A brief with no base renders exactly as before. The existing assertion in fm-brief.test.sh is updated

* no-mistakes(ci): Both real failures on PR 6442 (run 37064756768) come from CI tooling and a Pi version bump, not from the base-branch feature. Fixes for both already exist on upstream main, so I applied those two commits to the working tree with `git cherry-pick --no-commit`. Nothing is committed or pushed, and no feature file changed. - **Lint 2:** ShellCheck ran out of memory on `bin/fm-teardown.sh` (rss about 8 GB, exit 251, "shellcheck: out of memory"). The rule that broke is that the lint gate must finish on any valid shell root without exceeding the runner's memory ceiling. Upstream commit d719ef3 (#6443) fixes this in `bin/fm-lint.sh`: a root that hits the memory ceiling is retried without `--external-sources`. It also updates `tests/fm-lint.test.sh` and `docs/fm-test-portable-shards.md`. - **Behavior portable serial 6:** `tests/fm-calm-pi-extension.test.sh` failed with "grep disappeared from /export calm.html HTML while calm mode was on". Pi 1.0.1 renamed its HTML export renderer lookup, and upstream commit fede619 (#6530) updates the test to accept the new name. This change is test-only. Both commits applied cleanly. The four touched files are `bin/fm-lint.sh`, `docs/fm-test-portable-shards.md`, `tests/fm-lint.test.sh` and `tests/fm-calm-pi-extension.test.sh`. None of them is in the feature diff, so the feature contract is unchanged. **Local checks:** - `tests/fm-lint.test.sh` passes. - `bin/fm-lint.sh` on `bin/fm-teardown.sh`, `bin/fm-spawn.sh` and `bin/fm-dod-lib.sh` exits 0 with full ShellCheck analysis. - `tests/fm-calm-pi-extension.test.sh` exits 0 with 7 ok and 0 not ok. One case skipped because the Pi package isn't installed locally, so the Pi 1.0.1 export path that failed in CI couldn't be reproduced here. CI needs to confirm it. - I didn't reproduce the 8 GB OOM locally. The PR has to be updated through the pipeline with a normal push: no force push, no replacement PR, no merge

* fix(bin): name the Firstmate skill file as a fallback in the worker role (#6647)

* fix(bin): point project workers at the Firstmate skill file

The Skill tool cannot resolve a Firstmate skill from another project's worktree, so the launch role names the readable skill file instead.

* no-mistakes(review): fix(bin): name Firstmate skill file as fallback only

* feat(bin): retire a contribution whose forge object is permanently gone (#6655)

* feat(bin): retire a contribution whose forge object is permanently gone

Add fm-contributions.sh retire <task> <url> <captain|fleet> <reason>,
which records actor, reason and time on the saved record and removes
that task/url pair from known, poll rotation and coverage even while a
backlog link remains. Repeating a retire keeps the first provenance;
an unrecorded pair, unknown actor, empty reason or unacknowledged
pending signal is refused.

* no-mistakes(review): Keep retirement per task when settling final owners

* no-mistakes(ci): Made the three Greptile fixes the captain chose. ci-1, retirement needs the captain's word: `retire` now accepts only the actor `captain`. Any other actor is refused, including `fleet`. I removed `fleet` from the usage line, the script header (`bin/fm-contributions.sh`), the check in `bin/fm-contributions.sh` and the retired-record validation in `bin/fm-contributions.jq`, which now requires `.actor == "captain"`. The script header now says a retirement always records the captain's word, because the script cannot verify who runs it and the authority to retire is the captain's. The Bearings skill line in `.agents/skills/bearings/SKILL.md` now says `retire` records the captain's word. In the tests, a `fleet` retire is now a refusal case and must leave the record unretired. Every other retire call in the tests uses `captain`. The PR description has not been changed yet: the ruling asks for the same captain's-word statement there, and that belongs to the PR phase. ci-2, blank reasons: the reason check is now `[ -n "${5//[[:space:]]/}" ]`, so a reason made only of spaces or tabs is refused. I added a refusal test that passes a space-tab-space reason. The header also lists a blank reason among the refusals. ci-3, the gone-object test: the test forge wrapper has a new `not-found` fault that returns HTTP 404 for every `api repos/o/r/...` read. The happy-path test `test_retire_ends_observation_of_a_gone_contribution` now uses it in place of the 502 `down` fault, so it models a permanently gone repository rather than an outage. Verification: `bash tests/fm-contributions.test.sh` passed with exit 0. The last lines of its output include the three retire tests passing and no failures. `shellcheck` flagged only SC2034 (`FM_WAKE_QUEUE` unused) and SC1091 (an unfollowed source of `bin/fm-path-lib.sh`), neither on a line this round changed

* test: extend remote-reply whole-log recapture waits (#6639)

* test: give the remote-reply whole-log recapture a longer wait

* no-mistakes(review): Extend both recapture waits and simplify retry handling

* fix(bin): reopen a pending-reply escalation after its resolve (#6654)

* fix(bin): reopen a pending-reply escalation after its resolve

A retry of an open decision still appends nothing. A later same-kind escalation after a resolved line for that key appends again, so the next loss is visible.

* no-mistakes(ci): Both Greptile findings are fixed in the working tree; the four directly affected test suites pass, and a wider run of related suites had not finished when I returned this result. Invariant: a line already in the status file is recorded, and only a caller with its own evidence of a new episode may append it again. It must hold for every caller of status_event_recorded: the continuity break and the other call in bin/fm-procevent-remote-reply.sh, fm_parent_channel_append_once, and the pending-reply escalation. Changes: - bin/fm-classify-lib.sh: status_event_recorded is back to the idempotent retry check. It returns at the first matching line, and a later resolved line no longer makes that line look new. This fixes ci-1 for every caller and restores the early exit for ci-2, with no separate scan optimization. - bin/fm-pending-reply-lib.sh: _fm_pending_reply_maybe_escalate_locked now owns the reopen. It appends a new blocked line when the record already has an escalated_epoch (it escalated before and was reset) and the keyed decision is no longer open. Otherwise it uses status_event_recorded, so a retry or a resend that leaves the decision open appends nothing. - Comments on status_event_recorded and the pending-reply header describe this scope. Tests: - tests/fm-pending-reply.test.sh: the regression for escalate, operator resolve, second escalate is unchanged and passes. - tests/fm-remote-reply.test.sh: new regression drives the real reader. After an operator resolve of remote-reply-continuity-ios, a repeated handle and ingest of the same break appends nothing and leaves the decision closed. I placed it beside the existing continuity-break test, not in the pending-reply file, because only that file has the reader harness. - tests/fm-classify-corr-token.test.sh and tests/fm-classify-decision-key.test.sh: the cases that expected a blocked line to append again after a resolve now assert it stays recorded. Verification: - fm-classify-decision-key, fm-classify-corr-token, fm-pending-reply and fm-remote-reply test suites all exit 0 with the fix. - With the two bin files reverted to the commit under review, the new continuity regression and the updated parent-publisher case both fail, so they reproduce ci-1. - shellcheck -x on the changed files reports nothing. - Not finished: a background run of every other suite that mentions the parent channel or pending replies had produced no output, pass or fail, when I returned. One known gap: the escalation writes escalated_epoch before phase. If the process dies between those two writes and an operator resolves the key before the next tick, that tick appends a blocked line for the same episode. I left it, because closing it needs new state. Issue 4755 and other escalation behavior are untouched. Nothing is committed

* fix: refuse a confirming Enter on the Claude background-task exit picker (#6666)

* fix: refuse a confirming Enter on the Claude exit picker

The background-task picker still classifies as pending, so a retried Enter confirms "Exit and stop tasks".
Stop after the Enter that opened it, and raise the existing stale wake with the dialog name.

A non-paused secondmate still skips that wake, so a secondmate parked on the picker still looks idle.
Model-downgrade confirmation, MCP approval, and any other Claude exit confirmation are not covered, because there is no recorded screen for them.

* no-mistakes(review): fix: anchor exit picker match and wake second mates

* fix: refuse a typed submit while the Claude exit picker is open

A pane that already shows the picker must not receive the message or a confirming Enter.
The match requires the recorded heading and footer lines, and it is skipped when no dialog name is being recorded.

* no-mistakes(review): restore watcher to main behaviour, drop dialog wake

* no-mistakes(test): test: align Herdr picker fixtures with the preflight read

* no-mistakes(document): document exit refusal on a recognised dialog

* fix: remove the dialog file when exit runs in a subshell

do_exit is invoked from a command substitution, so the parent cleanup never saw the path it is supposed to delete.

* no-mistakes(review): fix: set the dialog file path after the control lock

* no-mistakes(document): document why the dialog file path follows the lock

* no-mistakes(ci): The exit cleanup in bin/fm-control.sh now deletes the dialog file first and releases the control lock second. The changes are in the worktree and are not committed. Invariant: only the process that holds the control lock for a task may write or delete that task's dialog file ($STATE/<id>.composer-dialog, the file that carries the picker name from the composer read to the Enter checks). The old cleanup released the lock and then deleted the file, so the delete ran outside the lock. Places where this invariant must hold: control_cleanup is the only place that deletes the file, and the single assignment after CONTROL_LOCK_HELD=1 is the only place that names it. do_exit only truncates the file, and it runs while the parent holds the lock. A process that loses the lock never sets the path, so it deletes nothing (earlier fix, unchanged). No sibling site needed a change. Fix: in control_cleanup, the rm of the dialog file moved above the fm_lock_release call, with a two-line comment that gives the reason. No dialog recognition or supervision code changed. Test: added test_exit_removes_the_dialog_file_before_releasing_the_lock in tests/fm-control-relaunch.test.sh. It runs one `fm-control exit` with a recording rm on PATH. The lock release removes paths at or under the control lock with rm, so the recording rm writes whether the dialog file exists at that moment. The test asserts the last record is "absent". It starts no second command. Verification, all run locally: - Before the fix, the new test failed with "the dialog file must be gone when the control lock is released, got: present". - After the fix, tests/fm-control-relaunch.test.sh exits 0 with 75 ok lines and no "not ok" line, including the new test and the existing test that the file is gone after exit and after relaunch. - tests/fm-control.test.sh exits 0 with 44 ok lines. - shellcheck -x on both changed files exits 0 with no output. One thing I noticed and did not change: tests/fm-control-relaunch.test.sh prints 615 "rm: cannot remove ... Permission denied" lines during its temp-directory teardown. The count is the same with my change reverted, so this change does not cause it, and the suite still exits 0. I could not read the Greptile check log (the check run was not found), so I worked from the finding text and the user instructions

* fix: restore timeout watchdog compatibility with macOS Bash 3.2 (#6028)

* fix: support stock macOS Bash in timeout watchdog

* no-mistakes(ci): Fixed ci-1 in tests/fm-timeout-lib.test.sh: both startup-owner failure paths now kill the bounded command group before killing the watchdog. Fault injection reproduced both leaks before the fix and confirmed cleanup afterward. Bash 3.2 suite, syntax check, ShellCheck, and diff checks passed; GNU timeout coverage skipped because the binary is unavailable. Production code unchanged

* fix(project-management): use subshell form for Initialize command (#6699)

Wrap the cd command in a subshell to comply with the cd-guard policy
that blocks persistent top-level directory changes in the primary
firstmate checkout. The subshell form (cd projects/<name> && ...) is
accepted by the policy as documented in issue #6502.

Fixes #6502

* feat(bin): add armable daily startup growth check (#6725)

* Add daily startup growth check

* no-mistakes(review): watch printed startup memory, retain growth baselines, drop knobs

* no-mistakes(review): delegate budget verdict, report before publish, pin shim home

* no-mistakes(review): baseline first sightings silently, drop mtime, spare secondmates

* no-mistakes(review): cap the wake line, validate budget verdict fields

* no-mistakes(review): guard record schema, check appends, tighten assertions

* no-mistakes(review): drop arbitrary tracked library, make re-arm idempotent

* no-mistakes(review): correct tracked-set wording, restore shim guards, register artifacts

* no-mistakes(review): exit on signal instead of publishing partial record

* no-mistakes(document): correct startup-growth record removal cost in state registry

* no-mistakes(test): sweep orphaned empty startup-growth temp records on due evaluation

* no-mistakes(document): note watcher need for armed startup growth check

* no-mistakes(ci): Diagnosed "Behavior portable serial 5": the shard failed on tests/fm-contributions.test.sh in test_arm_plumbs_a_configured_budget_into_the_check_shim (inherited mode) with "generated check did not attempt a read" and a missing forge/calls file. Root cause: bin/fm-contributions.sh:379-382 computes DEADLINE=$(date +%s)+BUDGET and then gates each read on DEADLINE-$(date +%s) >= OBSERVATION_RESERVE (= min(BUDGET,15)). At the one-second budget this test uses, that gate demands zero elapsed whole seconds, so a wall-clock second boundary crossing between the two date calls makes the poll loop break before any forge read. The test calls wrap_forge (which installs a fake date reading $FORGE/clock only when that file exists) but never seeded forge/clock, so it ran against the real clock. The repo already documents this exact hazard at tests/fm-contributions.test.sh:675-677 and freezes the clock in every other budget-constrained test: test_budget_exhaustion_keeps_prior_record (both modes), test_genuine_failure_near_deadline_is_unavailable, test_reservation_defers_later_url_when_fifteen_seconds_do_not_remain, test_slow_read_deadline_kill_is_budget_refusal, test_budget_is_cut_down_to_the_watcher_check_bound. test_arm_plumbs was the only omission; its single loop body covers both the configured and inherited modes. The remaining wrap_forge users run the default 20-second budget (5 seconds of slack above the reserve) and are not reachable by this race. Fix (smallest, matching the sibling idiom): added the clock freeze `/bin/date +%s > "$home/forge/clock"` plus a two-line comment inside that test's loop, before the poll. Three added lines in tests/fm-contributions.test.sh; no production code and no other file touched. Verification: reproduced the exact CI failure message by running a scratch copy whose fake date advances one second per read (worst case of the unfrozen clock) - missing forge/calls, "not ok - generated check did not attempt a read". With the fix, the single test passes 10/10 standalone and the full tests/fm-contributions.test.sh passes all 46 assertions with exit 0. tests/fm-startup-growth-check.test.sh still exits 0. bash -n and shellcheck -S warning are clean on the edited file; scratch repro files were removed, leaving only the intended three-line change. Note for the author: this flake is not caused by this branch - tests/fm-contributions.test.sh is untouched by the PR and the race is pre-existing, surfacing only on a loaded runner (the shard's neighbouring assertions show 80+ second gaps). It was fixed because it is genuine nondeterminism with an established in-repo remedy and the check cannot otherwise go green. No work was done on the unselected greptile findings (ci-2..ci-5) or the cancelled "Behavior portable serial 2" check (ci-6)

* no-mistakes(ci): Diagnosed "Behavior portable serial 6": the omitted log section (fetched with gh api) shows the shard failed on tests/fm-calm-pi-extension.test.sh in test_hidden_block_geometry_e2e with "Pi Calm hidden-block geometry E2E did not complete the /reload viewport transition" (exit=1, duration_ms=23437; the family line pure-contract-unit count=3 duration_ms=51193 failed=1 matches devin-harness 3203 + calm-pi 23437 + task-delivery 24553). Root cause: tests/fm-calm-pi-extension.test.sh:2831 wait_for_geometry_transition detected the /reload by SAMPLING a single intermediate frame - it polled tmux capture-pane for Pi's transient "Reloading keybindings, extensions, skills, prompts, themes, and context files..." box and only accepted the rebuilt transcript afterwards (elif gated on saw_transient). Pi shows that box only while the reload runs and replaces it via dismissReloadBox when done, so on a loaded runner the box can live entirely between two polls; saw_transient then stays 0 forever and the 600-attempt wait times out although the reload fully succeeded. Reproduced locally under CPU oversubscription with CI's Pi 1.0.4: 1 failure in 8 runs, diagnostics proving the mechanism ("DIAG: transition timeout saw_transient=0 attempt=600") while the captured viewport already showed the completed reload status row and the intact collapsed transcript. Not a Pi-version regression: 1.0.4's handleReloadCommand and both banner strings are identical to the installed 0.85.1, and shard 6 of the previous pipeline run (job 112490351105) passed this same script on Pi 1.0.4. Invariant violated: a TUI E2E assertion must key off a durable state the UI retains, never off one intermediate frame polling may miss. Sibling enumeration: the helper was defined once and called once; grep for transient/saw_/two-stage waits across tests/, bin/ and .pi/ found no other one-shot-frame wait - every other wait in the file keys off durable text presence/absence or off continuously repeating animation frames (the working-ship motion checks) where sampling eventually connects; no doc or other test referenced the helper. Fix (smallest, one helper): renamed it wait_for_geometry_reload and made it wait for the durable status row Pi appends once the reload completed and the chat was rebuilt ("Reloaded keybindings, extensions, skills, prompts, themes, and context files", a persistent chat child via showStatus, present in both 0.85.1 and 1.0.4) together with CALM_GEOMETRY_FINAL. That marker is absent before the reload and never appears if the reload throws, so the check is strictly stronger than before (a failed reload now fails the test instead of passing on transient-then-final). One file changed: tests/fm-calm-pi-extension.test.sh, 2 hunks, no production code. Verification: deterministic before/after - with polls spaced 0.4s apart (the loaded-runner worst case) the pre-fix helper fails with exactly the CI message and the post-fix helper passes; 12/12 consecutive passes of the isolated test under load with Pi 1.0.4; full file 15/15 ok exit 0 under LANG=C.utf8 (matching the runner locale); bash -n, shellcheck -S warning and bin/fm-lint.sh clean; scratch repro files and the temp Pi 1.0.4 install removed, git status shows only the one intended modification. Notes: (1) the flake is not caused by this branch - that test file is untouched by the PR and the race is pre-existing; fixed because it is genuine nondeterminism and the check cannot otherwise go green. (2) Under far harsher load than CI applies a separate unrelated flake appeared in the same test: Pi's "Warning: tmux extended-keys is off..." row occasionally lands between the collapsed [skill] ahoy row and the final response, making assert_geometry_gap see 5 rows instead of 2. Different cause, never seen in CI, and its plausible remedies would change key delivery for the other TUI tests in the file, so it was left alone rather than growing this fix

---------

Co-authored-by: Quartermaster <quartermaster@users.noreply.github.com>

* fix(bin): reopen the remote-reply continuity decision on a later break (#6708)

* fix: reopen a remote-reply continuity break after repair

A later break for the same route and reason was swallowed after the
operator resolved the first one, because the status line matched for
the life of the log. The continuity ingest now appends again when the
cursor has moved or retirement has reset that episode, and an unchanged
re-read still appends nothing.

status_event_recorded is unchanged. Its other callers are the
pending-reply escalation, which already decides its own episode, the
parent-channel note append, and the remote document transfer note.

* no-mistakes(review): seed continuity episode for already recorded break line

* no-mistakes(document): document when a remote-reply continuity break reopens

* no-mistakes(ci): The adapter now stores the continuity episode record before it appends the `blocked` line, so a failed store appends nothing. The changes are uncommitted in the worktree, in `bin/fm-procevent-remote-reply.sh` and `tests/fm-remote-reply.test.sh`. **Invariant:** a continuity `blocked` line is on the parent status log only when the episode record for that break is already stored. Only the `continuity-broken` branch of `cmd_ingest` appends that line, so the fix is at that one place. **What changed in `cmd_ingest`:** - It decides first whether the break needs a line, then stores the episode record, then appends. - If storing the record fails, it stops with "cannot record continuity episode" and appends nothing. - If the line is already on the log and no episode record exists, it stores the record and does not append. **One addition you did not ask for:** if the append fails after the record is stored, the adapter puts the earlier record back (or deletes the new one when none existed). Without that, the retry would read as the unchanged repeat and append nothing, which is the issue 6701 failure again. **Deviation from your test instruction:** the test does not use `chmod` on the cursor directory. `write_continuity_episode` runs `chmod 700` on that directory before every write, so a read-only directory is made writable again. The test instead makes `mktemp` fail for the episode's temporary file in that directory, the same way the existing receipt-failure test does. **Tests added to `tests/fm-remote-reply.test.sh`:** - Store failure: the second break exits 1, appends no `blocked` line and opens no decision. One retry after storage recovers exits 3 and appends a single line. - Append failure: with the status log read-only, the third break exits 1 and appends nothing. One retry after the log is writable exits 3 and appends a single line. **Verification:** `bash tests/fm-remote-reply.test.sh` ends with "ALL TESTS PASSED" with the fix. Against the script at commit b750c828 the same test file fails at "a continuity break appended its line before its episode was stored". `shellcheck -S warning` reports only an unused loop variable at line 667 of the test file, which this change does not touch. I ran no other test files. The Greptile Review check log could not be retrieved, so I worked from the finding text alone

* fix: record a continuity break's reader position on its status line

A later break at another cursor is then a different line, so the existing
duplicate check appends it and reopens the decision. An unchanged re-read
builds the same line and appends nothing.

* fix: reopen a continuity break after an identical restore

A retirement that puts the same bytes back used to rebuild the recorded line, so the later break stayed closed. The retirement count on that line makes the later break distinct.

* no-mistakes(review): remove continuity match for full-prefix line without retirement count

* no-mistakes(document): clarify what a continuity break status line records

* fix: remove the reply cursor before recording retirement

A stop between those steps must leave the count unchanged, so an unchanged continuity break still builds the same line.

* fix(control): drop busy_gen from the task record when an incarnation is retired (#6733)

* fix(control): drop busy_gen when an incarnation is retired

A deliberate exit removed the busy sidecar and left busy_gen in the task record, so the two records disagreed about whether that incarnation was still observable.

* no-mistakes(review): drop GNU-only chmod and unreached sidecar-absent branch

* no-mistakes(review): correct lock comment to name the deadlock

* no-mistakes(ci): The test `test_exit_drops_meta_busy_gen_with_the_sidecar` in tests/fm-control.test.sh now compares the whole task record (the `state/<id>.meta` file), so the Greptile finding is fixed. Invariant: after `exit` retires an incarnation, the task record must equal the record from before `exit` with only the `busy_gen` line removed. This test is the only place in the change that asserts the record survives the rewrite, so it is the only site to fix. The other `busy_gen` tests assert that the line stays, and they do not go through the rewrite. What changed: before `exit`, the test writes the record without its `busy_gen` line to `expected.meta`. After `exit`, the test runs `diff` between that expected copy and the real record, and fails with the diff output if they differ. This one comparison replaces the two earlier checks (no `busy_gen` line left, and the `window` line present), because it covers both. I did not change bin/fm-control.sh or any other file. How I know it works: - I ran `bash tests/fm-control.test.sh`: exit code 0, 45 lines starting with `ok`, no other lines. - I temporarily changed the rewrite in bin/fm-control.sh to also drop the `harness` line. The test then failed with `not ok - exit should drop only busy_gen from the task record:` and the diff `< harness=codex`. The earlier `window`-only check would have passed that rewrite. I restored bin/fm-control.sh afterwards; `git status` shows only tests/fm-control.test.sh modified. - `bash -n` and `shellcheck` on the test file report no new warnings from the edit. The change is not committed; the working tree holds it

* test: cover PID collisions in harness ancestry detection (#6484)

* test(secondmate-harness): scope fake ps -codex label for pid 5252 to the liveness probe

Closes #6456

* no-mistakes(ci): Updated the collision test to log and assert that PID 5252 was queried before selecting 4242. Full fm-secondmate harness suite passes

---------

Co-authored-by: YifuGu <ironerumi@users.noreply.github.com>

* fix: share Pi Calm's working-ship widget slot (#1854)

* fix(calm): share the standalone Pi Calm working-ship widget slot

Firstmate Calm and the user-global standalone Pi Calm both install an
animated working-ship widget during agent runs. Each claimed its own Pi
widget key, so a session loading both (the main Firstmate home) rendered
two boats. Pi replaces widgets under one key, so claiming the shared
"calm-working-ship" slot keeps dual-install sessions to a single boat
while a Firstmate-only session is unchanged.

Pins the shared slot contract in the working-ship module test so the key
cannot silently diverge again.

* test(calm): pin the shared working-ship widget key in CI, document dual-install

The key-parity assertion inside the Pi fixture only runs where the
@earendil-works/pi-coding-agent package is installed, so CI never
exercised it. Add a source-level twin that needs nothing but the
tracked file, and note in docs/calm.md that the boat shares the
standalone Pi Calm working-row widget slot.

* no-mistakes(review): Add executable dual-install widget replacement coverage

* no-mistakes(review): Guard shared widget cleanup with disposal ownership

* no-mistakes(document): Document shared Calm working-ship slot behavior

* test(calm): read the standalone Calm slot from its own module

The dual-install check registered both boats itself under the shared
slot, so it could only prove that Pi replaces a widget under one key: it
would still pass if the standalone Pi Calm extension installed its boat
under a different key, which is the two-boat regression the check exists
to prevent.

Read the standalone extension's own working-ship module when it is
installed - FM_STANDALONE_CALM_SHIP, else ~/.pi/agent/extensions/calm -
and drive the check with the key that module exports, so a rename on
either side registers two widgets and fails naming both keys. A pinned
shared-slot contract still covers a machine without the extension, and
the run reports which side it used instead of passing silently over an
absent extension.

Verified: the touched Pi Calm suite passes and reads the installed
standalone extension; with a copy of it whose key is renamed to
calm-working-ship-v2 the suite fails naming the drift.

* no-mistakes(review): Gate stock-row restoration by shared-widget ownership

* no-mistakes(review): Removed redundant widget-key source assertions

* no-mistakes(document): Document shared Calm working-ship widget ownership

* feat(bin): add opt-in --herdr-resume-lock-wait to fm-spawn (#6649)

* fix(herdr): make exact-resume presentation-lock wait instead of a bounded timeout

The exact-resume path in bin/fm-spawn.sh used the same 50-attempt-then-
give-up lock acquire as the new-task-create path, but the two paths are
not equivalent on contention: a create has no prior state to strand and
can safely fall back to a flat layout, while a resume is recovering a
specific existing identity that a concurrent recovery may legitimately
be holding the lock for. Giving up there does not degrade gracefully,
it hard-fails the resume outright. The suite's own concurrent
cross-home recoveries test already asserts both concurrent recoveries
succeed with a genuine reclaim, and the file's header comment already
(inaccurately) claimed lock contention falls back to the ordinary flat
layout for both paths alike, so the intended contract was always that
recoveries serialize and both succeed, not that either one refuses
under a short bound.

Give spawn_herdr_presentation_order_lock_acquire a wait mode that uses
this file's own established fm_lock_acquire_wait idiom (already used
for its other fleet-shared locks) instead of the bounded loop, and use
it only at the exact-resume call site. The new-task-create call site
is unchanged and keeps its bounded-then-flat-fallback behavior, which
is already covered by its own passing test. Dead-owner PID-liveness
reclaim inside fm_lock_try_acquire still bounds the wait against a
holder that crashed mid-hold.

Adds a deterministic regression test that holds the shared session
lock from an unrelated process for well past the old bound, then
asserts the resume succeeds with a genuine reclaim and took close to
the full hold duration, so a fix that merely widens the bound rather
than genuinely waiting is still caught. The existing concurrent
cross-home recovery test exercises this under real timing but does not
reliably outlast a fixed bound on its own.

Corrects the header comment's claim that create and resume share one
bounded-then-flat-fallback behavior on lock contention; they no longer
do.

* no-mistakes(document): Document Herdr recovery waiting for presentation lock

* no-mistakes(document): Update stale hard-refusal claim in verification log

* no-mistakes(ci): Fixed the Greptile finding on tests/fm-backend-herdr-presentation-e2e.test.sh:1389 by bounding the resume lock-wait regression's spawn_task call. Added an optional 4th `deadline_seconds` arg to the `spawn_task` helper (defaults to empty, so all ~20 other existing call sites are unaffected and unwrapped by `timeout`). The lock-wait test now passes `LOCK_WAIT_HOLD_SECONDS + 60` (90s) as the deadline, and a dedicated check for exit code 124 emits a clear "hung for over Xs instead of waiting out a Ys lock hold" diagnostic before falling through to the existing pass/fail assertions, which are unchanged. No product code was touched. Verified with `bash -n`, `shellcheck -x` (no warnings), a standalone reproduction of the timeout/no-timeout/success paths, the project's `bin/fm-lint.sh --fast` on the file (clean), and the full `tests/fm-lint.test.sh` suite (all 46 assertions pass)

* no-mistakes(ci): Replaced the direct `timeout "$deadline_seconds"` call in `spawn_task()` (tests/fm-backend-herdr-presentation-e2e.test.sh) with the repo's portable bounded-execution helper: sourced `bin/fm-timeout-lib.sh` at the top of the file and changed `deadline_cmd=(timeout "$deadline_seconds")` to `deadline_cmd=(fm_run_timed "$deadline_seconds")`. This removes the GNU/BSD `timeout` dependency that would fail with exit 127 on a stock macOS host without coreutils, while preserving identical semantics (exit 124 on bound-hit, command's own exit otherwise), which the existing `[ "$LOCK_WAIT_STATUS" -eq 124 ]` diagnostic check already relies on. Verified: `bash -n` syntax check, `bin/fm-lint.sh --fast` clean, full `tests/fm-lint.test.sh` suite (46/46 pass), and a standalone repro confirming `fm_run_timed` returns 124 on timeout and 0 on success identically to the prior `timeout` call. No other direct `timeout` calls exist in this file or elsewhere in the PR's diff, so no sibling sites remain

* fix(herdr): gate exact-resume lock wait behind --herdr-resume-lock-wait

Keep refuse-by-default on presentation-order lock contention for Herdr
exact resume. Callers that need concurrent recoveries to serialize must
pass --herdr-resume-lock-wait; unbounded blocking on a third-party session
lock is never the default.

Update docs and the real-Herdr e2e suite so the default path asserts the
refusal and the opt-in path asserts the wait.

* no-mistakes(test): Fix e2e test's lost exit status after if/fi with no else branch

* docs(herdr): stop advertising --herdr-resume-lock-wait on --relaunch

The relaunch path reuses the recorded endpoint and never takes the
presentation-order lock, so the flag is inert there. Drop it from the
--relaunch usage line and state where the flag applies.

* no-mistakes(review): Clarify lock-wait docs; simplify bash-3.2-safe spawn_task helper

* no-mistakes(ci): Fixed ci-1 (Greptile P2). In tests/fm-backend-herdr-presentation-e2e.test.sh, the failure cleanup `cleanup_all` stopped only `LOCK_CONTENTION_OWNER_PID`. It now also stops `LOCK_REFUSE_HOLDER_PID` and `LOCK_WAIT_HOLDER_PID`, the holders of the two new contention cases, so a `fail` before their explicit `wait` no longer leaves them running. Both new PIDs are initialised empty next to the existing one, and each is cleared right after its successful `wait` so cleanup never touches a finished PID. I changed nothing else. `bash -n` passes. The real Herdr e2e run passed both new cases ("default resumed identity refuses session lock contention" and "--herdr-resume-lock-wait waits out session lock contention instead of refusing"). The full run hit my 550s timeout in a later, unrelated case, after the new cases passed

* fix(bin): let TERM stop a watcher blocked in a pane capture on bash 3.2 (#6762)

* fix(bin): let TERM stop a watcher blocked in a pane capture on bash 3.2

Stock macOS bash 3.2 holds a HUP or TERM until a running command
substitution's child exits, and the watcher read every pane through
$(fm_backend_capture ...). A blocked backend read therefore held the
watcher's stop for as long as the read lasted, and a stopped watcher left
the hung read orphaned. tests/fm-watch-triage.test.sh
test_term_stops_a_watcher_blocked_inside_a_poll failed on /bin/bash 3.2
for this reason while passing on bash 5.

Pane captures now go through watcher_capture, which runs the read as a
waited background process group recorded like a check's, so the stop is
honored at once and watcher_cleanup stops a read still in flight along
with its per-call output file.

* no-mistakes(review): Run drain-ring idle capture in watcher shell, add regression test

* no-mistakes(document): Document watcher TERM handling for blocked checks and captures

* no-mistakes(test): Silence bash 3.2 setpgid race noise from watcher captures

* fix(bin): verify the capture group and scope the stop claim to pane reads

watcher_capture now confirms its background read leads its own process
group, as run_check_capture already does, so watcher_cleanup never relies
on a group that set -m failed to create. The comment and continuity doc
now say only fm_backend_capture pane reads go through watcher_capture;
agent-state and composer-state reads still run inside command
substitutions.

* feat: add per-home worker tool exclusions (#6750)

* feat(spawn): add per-home worker tool exclusions

Add an optional per-home config/crew-exclude-tools file listing tool names to hide from workers, one per line, with blank lines and # comments allowed.
It applies to every ship and scout launch and relaunch in that home, is never inherited by another home, and does not affect secondmate agents.
Pi and pi-signed apply it through --exclude-tools, which also covers MCP tool names.
Any other runtime, and a raw launch command, refuses the launch when the list is non-empty rather than ignoring it.
Malformed entries are refused before provisioning, and before a relaunch stops a running worker.
Exclusions that match no tool in the worker's loaded registry are reported as unverified warnings in its status record instead of refusing the worker.

Closes #6744

* no-mistakes(review): Preserve UTF-8 exclusion paths and verify Pi lifecycle behavior

* no-mistakes(document): Clarify worker tool exclusion documentation

* no-mistakes(ci): Fixed ci-1 in bin/fm-exclude-tools-lib.sh: a failed read now returns an error before printing names, so all shared launch and relaunch callers refuse rather than silently dropping exclusions. Added deterministic regression coverage for a file disappearing after readability checks across Pi/pi-signed ship and scout launches. Reproduced the original failure; verified 83 spawn checks, 77 relaunch checks, direct parser/runtime failure cases, full targeted lint, Bash syntax, and git diff --check. Relaunch tests passed with existing fixture-cleanup permission warnings. ci-2 remains unchanged per the user's decision; the outer executor owns the fresh CI run

* feat(bin): defer spawns beyond a declared per-project capacity, opt-in (#5343)

* refactor(bin): share the local Firstmate home walk from the wake library

Teardown's walk over the root home and its registered local secondmate homes
moves into bin/fm-wake-lib.sh as fm_local_firstmate_state_dirs, next to
fm_firstmate_root_home, so a second consumer can count task records across
this machine's homes without a copy. Teardown keeps its exact refusal wording
through a thin wrapper.

* feat(bin): defer spawns beyond a project's declared machine capacity

A project whose machine-local resource only serves a few workers at once had
no way to tell Firstmate so: every queued item was launched, and the surplus
workers spent full-context turns retrying the resource.

config/project-capacity in the root home now declares how many workers each
named project admits at once on this machine. bin/fm-spawn.sh counts the ship
and scout records on the same project origin across the root and its local
secondmate homes, skipping ones whose ready PR is recorded, while holding the
shared project lock through publication. A spawn with every place held exits 75
before any brief render, endpoint, worktree, record, or backlog move, so the
item stays queued; batches report it as deferred. Undeclared projects keep
today's uncapped dispatch, and an unreadable declaration refuses rather than
guessing the limit.

Refs #4237

* no-mistakes(review): Document that capacity matches the clone directory name

* no-mistakes(document): Rewrap stale fm-wake-lib root-home doc comment

* no-mistakes(review): Dedupe local state dirs by identity to avoid double-counting

* no-mistakes(document): Rewrap fm_local_firstmate_state_dirs error doc comment

* no-mistakes(ci): I fixed all four Greptile findings. All 14 tests in tests/fm-project-capacity.test.sh pass, and shellcheck at warning level is clean on the changed files. Each new test failed against the old code and passes now. - **ci-1 (spaced names):** a declaration line must give a name its capacity whenever the name is a valid clone directory name. `fm_project_capacity_lookup` now trims each line, skips blank lines and lines whose first non-blank character is `#`, and takes the last field as the capacity. Everything before that field is the name, so it may contain spaces. The old error cases still refuse: a single field is rejected, and trailing text leaves a last field that is not an integer. The library header and docs/configuration.md now say a name starting with `#` cannot be declared. New test `test_spaced_project_name_is_declared` declares `my heavy project 1` next to an indented comment line and gets a deferral. - **ci-2 (unreadable records):** the holder count must never silently leave out a holder. `fm_project_capacity_occupants` now refuses when a local home's state directory exists but cannot be read or listed, or when a `.meta` file cannot be read. The error names the path, and `fm-spawn.sh` shows it in its existing refusal message. New test `test_unreadable_holders_refuse_admission` covers an unreadable record in the root home and an unreadable state directory in a registered local secondmate home, then checks that the spawn is admitted once both are readable. The test is skipped when run as root. - **ci-3 (Orca lock):** any spawn that can become a holder for a capped project must take that project's lock. The lookup now also reports whether the declaration caps any project at all, and an Orca spawn takes the per-origin lock whenever it does. This covers every capped same-origin clone. It also covers some cases where no same-origin clone is capped, because a spawn cannot find clones under other directory names without searching for them. With no declaration file, Orca still skips the lock. The comments in the library and in the `fm-spawn.sh` header are updated. The Orca test now clones the origin as `project-2`, which has no declaration, and checks that its Orca spawn refuses while the lock is held and publishes no record. - **ci-4 (worktrees):** `assert_nothing_created` now also compares the project's `git worktree list` from before and after a deferred spawn. Both tests that call it take that snapshot first. Files changed: bin/fm-project-capacity-lib.sh, bin/fm-spawn.sh, docs/configuration.md, tests/fm-project-capacity.test.sh

* fix(bin): declare capacity for a project name that begins with #

A clone directory whose name begins with # was skipped as a comment, so that project stayed uncapped. A line is a declaration when the # is written against the rest of the name and the line ends with a capacity; a # followed by whitespace stays a comment.

* no-mistakes(document): Rewrap project-capacity library header comment

* no-mistakes(ci): Lint 2 fails because this PR's code pushes ShellCheck past its memory cap. ShellCheck ran out of memory analyzing bin/fm-teardown.sh in CI (reason=memory, rc=251, peak about 8.39 GB). On current main the same file passes at about 7.29 GB. **Cause:** the new `fm_local_firstmate_state_dirs` function in bin/fm-wake-lib.sh had a conditional `. fm-secondmate-registry-lib.sh` with a `# shellcheck source=` directive inside the function. ShellCheck followed that source again, inside a function scope, wherever fm-wake-lib.sh is sourced, and bin/fm-teardown.sh is the heaviest root that sources it. Measured locally with `shellcheck --norc --external-sources bin/fm-teardown.sh`: - current main (fd325b1b): 7.29 GB - main merged with this PR: 7.86 GB - the same merge without the in-function source: 7.27 GB **Rule this restores:** this change must not make any lint root heavier than it is on main. That function holds the only new nested source in the change. **Fix:** I removed the in-function source, which no caller needs. Both callers already load the registry library at top level before calling the function: - bin/fm-teardown.sh sources it directly. - bin/fm-spawn.sh, the only user of bin/fm-project-capacity-lib.sh, gets it through bin/fm-ff-lib.sh. I also documented the requirement in the function's comment and in the "Requires" note in bin/fm-project-capacity-lib.sh. No behaviour changes. **Verification:** - ShellCheck on head: bin/fm-teardown.sh peaks at 7.12 GB and bin/fm-spawn.sh at 6.68 GB, both with rc=0. bin/fm-wake-lib.sh and bin/fm-project-capacity-lib.sh lint clean. - tests/fm-project-capacity.test.sh, tests/fm-teardown.test.sh (102 ok) and tests/fm-teardown-endpoint-safety.test.sh all pass. Files changed: bin/fm-wake-lib.sh, bin/fm-project-capacity-lib.sh

* fix(bin): release the Herdr session lock when reclaim finishes

A concurrent resume in another home waits five seconds for that lock.
Reclaim is the last presentation change on the recovery path, so holding
the lock through the launch tail made the waiter time out. The contributions
arm check also freezes its one-second clock, the same way the budget tests
do, because an unfrozen clock can tick past before the first forge read.

* no-mistakes(review): Keep Herdr session lock through launch handoff after reclaim

* no-mistakes(review): Skip the spawning task's own record in capacity count

* no-mistakes(review): Restore release test comment above its test

* docs: scope PR-ready re-evaluation to a declared project capacity

A ready pull request frees a place only when that project declares capacity, so the always-loaded backlog contract should re-evaluate on that handoff only in that case.

* fix(bin): name the Lavish read message count by the same label as its section (#6785)

fm-procevent-lavish.sh read labels tag=message rows SESSION-ENDING MESSAGE
only when session_ended is true and CAPTAIN MESSAGE otherwise, but the
count line always said session_ending_message_count. Several composer
messages on a still-open board were therefore counted as session-ending.

The count line now follows the same session_ended switch:
session_ending_message_count once the session ended, captain_message_count
otherwise. Message rows stay out of the annotation count, per triage.

Fixes #6743

* fix(herdr): make the presentation lock namespace per OS account (#6780)

The Herdr presentation lock namespace was the fixed machine-global
/tmp/firstmate-herdr-presentation, so on a host where two OS users run
Firstmate on Herdr the first account to create it owned it and every
teardown from the other account was refused with no way to clear it.

Suffix the namespace with the account uid. The owner-uid and mode-700
checks are unchanged, so a foreign-owned or wrong-mode name at this
account's path is still refused and never adopted, chowned, or removed.

Fixes #4716.

* Fix OpenCode arm plugin to decide with the shared supervision predicate (#6809)

The OpenCode session plugin's shouldArm kept its own copy of the need
test that only looked for in-flight task records, while the turn-end
guard decides with fm_supervision_needed in bin/fm-supervision-lib.sh,
which also counts registered process-event sources and trusted custom
checks. With an empty fleet but any registered source or check, the
guard blocked every turn end while the plugin declined to arm - a loop
the guard's own repair line could not resolve because it names the
plugin as the fix.

The plugin now delegates the decision to the shared predicate through
bash, keeping the local away-record decline and the x-mode.env arm
override. OpenCode plugin test fixtures now carry the real predicate
their arming path sources, and the arm suite gains six cases asserting
the plugin's decision against the shared verdict over the same
synthetic state directories.

Co-authored-by: Mia Sun <mia@Bigs-Mac-mini.localdomain>

* fix(bin): resolve a pending reply only from its own task's status line (#6792)

* fix(bin): resolve a pending reply only from its own task's status line

Remote reply ingestion handed every corr= token in a mate's payload to
fm_pending_reply_try_resolve together with that mate's own status log,
so one mate echoing another mate's token resolved the other request.
Honor a status-file override only when it is the record's own
parent_status, and match the corr= token as a whole word.

Fixes #6538

* no-mistakes(document): docs: scope remote reply settlement to the asked mate

* docs(skills): index the six missing agent-only skill triggers (#6784)

agent-skill-trigger-index claims to be the complete agent-only trigger
index but omitted operational-home-layout, session-start-recovery,
validation-supervision, ship-landing, scout-completion, and
away-quiet-supervision. Add each with its own description's trigger,
placed beside the related entries.

The decision-hold-lifecycle redirect stub stays out, per triage.

Fixes #6503

* test: make agent process fixtures compatible with multicall sleep (#6814)

* test: share a rename-safe agent stand-in across liveness suites

On Ubuntu 26.04, `sleep` is the uutils multicall binary, which refuses to
run when invoked through a symlink named after another utility. The Herdr
descendant process-walk tests built their agent-named process as a `pi`
symlink to the host `sleep`, so the process exited at once, its parent shell
was gone before the walk ran, and both cases read `unknown unreadable` and
failed on that host. The suite stops at its first failure, so every later
case went unrun. The Herdr control smoke test's `claude` symlink has the
same construction.

The tmux liveness suite already solved this with a host-compiled spinner and
a survival-checked `sleep` fallback. That builder moves into tests/lib.sh as
fm_agent_standin, and the tmux suite, both Herdr descendant cases, and the
Herdr control smoke test now use it. When no stand-in can survive a foreign
name, a case skips with the reason instead of failing.

tests/fm-test-fixtures.test.sh gains a portable regression with a fake
multicall `sleep`, so it bites on hosts whose own `sleep` is single-purpose.

* no-mistakes(document): Correct Herdr verification fixture reference

* ci: retrigger cancelled shard

* fix(bin): keep the supervision host's successor watcher alive after the Stop hook's group is torn down (#6787)

* fix(bin): keep the supervision host's pass-through successor out of the hook's process group

The successor a main-only pass-through leaves for main shared the Stop hook's
process group, so the harness tearing that group down after the exit-2 rewake
stopped it. The stop published downtime and the next park's first cycle
announced an empty check: rearm-resurface, which woke main again in a loop.
Start that successor in a process group of its own, as the hook's own
handling successor already is.

* no-mistakes(review): Give the at-turn successor left for main its own group

* no-mistakes(document): Document own-group successor for turn-start hand-back too

* no-mistakes(ci): I made the change you asked for: both new teardown tests in tests/fm-supervision-host.test.sh now call the existing `stop_home_processes "$home"` just before `pass`. The tests are `test_successor_left_at_the_turn_survives_the_hook_process_group_teardown` and `test_pass_through_successor_survives_the_hook_process_group_teardown`. No production code and no other tests changed. The rule broken was that a test must not leave a home's watcher or arm processes running after it passes. These two were the only cases in the changed area that broke it. The other host+hook tests already stop their home, and `test_successor_close_during_main_turn_is_delivered_at_the_next_turn_end` leaves its watcher behind too, but it is an older test you said not to touch. The only reason anything was left over is that the successor's arm now sits in its own process group, outside the hook's teardown. `stop_home_processes` kills the watcher by the pid in its lock file, which stops it no matter which group it is in. **Checks run:** - I ran just these two tests from a scratch copy of the suite (since deleted). Both pass in about 13 seconds. - After each test, a process listing filtered to that test's home directory came back empty once the processes had about a second to exit after TERM. - `bash -n` on the test file passes. - `shellcheck` is not installed…
witt3rd added a commit to witt3rd/firstmate that referenced this pull request Oct 10, 2026
* fix: prevent unnecessary remote worker turnover (#6431)

* fix: prevent healthy remote job worker turnover

* no-mistakes(review): Serialize full LaunchAgent repair and verify launchd-tracked owners

* no-mistakes(review): Let launchd-tracked unpublished spawns start before reloading

* no-mistakes(document): Document remote worker heartbeat and serialized LaunchAgent recovery

* no-mistakes(ci): Full CI log showed the idle-worker regression exceeded its outdated command budget (82 versus 80) after independent heartbeat ownership checks were added. Raised the budget to 120 while retaining the separate busy-poll sleep limit. Remote-job and LaunchAgent executable tests passed, as did bash syntax validation and git diff --check. No production behavior changed

* no-mistakes(ci): Fixed missing readiness recovery under verified live ownership, preserving the serving PID and private file mode. Added executable regressions for deletion during a blocked sweep and stale readiness diagnostics without LaunchAgent reload. Deletion regression failed before the fix. Both remote-job test suites, bash syntax validation, and git diff --check passed

* no-mistakes(ci): Full CI log identified a flaky ownership-loss test racing an already-authorized heartbeat refresh. Replaced backdating and a fixed sleep with bounded observation of readiness expiry through the public probe. Production behavior unchanged. Remote-job and LaunchAgent executable suites passed; bash syntax validation and git diff --check passed

* fix: wait for launchd bootout cleanup

* no-mistakes(document): Document remote worker recovery and read-only turnover verification

* no-mistakes(review): Publish worker identity before lock owner records

* no-mistakes(document): Document worker identity publication safety invariant

* no-mistakes(ci): Fixed ci-1: replacement workers discard predecessor readiness before publishing identity and roll back identity if lock-owner recording fails. Added executable regressions reproducing both defects. LaunchAgent and orphan-reap suites, shell syntax checks, and git diff --check passed. No live service state was modified

* no-mistakes(ci): Restored lock-owner-before-identity publication and removed the identity rollback and reordering-only tests. Retained an executable regression proving predecessor readiness is rejected until replacement startup completes. LaunchAgent and orphan-reap suites, shell syntax checks, and git diff --check passed. Other changes remain intact; the retained-identity interrupted-repair edge remains out of scope. No live service state was modified

* fix: speed up remote-job sequence claim cleanup (#6575)

* fix(remote-job): reap expired seq claims with one directory walk

The hourly claim sweep forked uname+stat per .seq-claims entry and blocked
serving for ~85s at ~17k dirs. Delete expired empty claim dirs with a single
find -exec rmdir batch and cache the host uname for remaining mtime reads.

* no-mistakes(review): Restore original path mtime helper and drop uname cache

* no-mistakes(test): Restore claim retention eligibility; focused retention and serving tests pass

* no-mistakes(document): Document single-walk claim cleanup and regression entrypoints

* no-mistakes(ci): Fixed Lint 2’s reproduced SC1091 by adding the tests/lib.sh ShellCheck source directive to the retention test. Runtime behavior is unchanged. ShellCheck passed for both new claim tests; Bash syntax, the retention behavior test, and git diff --check passed

* fix(tests): keep fixture registries out of git worktree roots

A TMPDIR pointed at a repository root placed live .fm-test-* registries
beside tracked files, and a concurrent git add during the claim-walk CI
fix round committed three of them. Route registries and fixture roots
through a TMPDIR that refuses git worktree roots, remove the stray files,
and pin the escape with a behavioral cleanup test.

* no-mistakes(review): Preserve whole-second claim expiry in single-walk sweep

* no-mistakes(document): Correct temporary-directory resolution documentation

* no-mistakes(ci): Fixed ci-3 by changing only the stale, fresh, and read-only orphan fixture paths in tests/fm-test-fixture-cleanup.test.sh to use $FM_TEST_TMPDIR. All seven tests passed both normally and with TMPDIR set to the worktree root. Shell syntax and git diff --check passed

* feat(bin): record no-mistakes pipeline spend per task at teardown, opt-in (#5354)

* feat(bin): record each task's no-mistakes pipeline spend at cleanup

no-mistakes keeps every pipeline agent invocation's token usage only in its
local agent_invocations records, and cold pipeline agents (review, test,
document) leave no session log, so a task's review-loop cost never reached
Firstmate's records and could not be attributed after cleanup.

bin/fm-pipeline-spend.sh attributes a task's runs by the repository
no-mistakes resolves for the task copy, the task branch, and the branch's
creation (a relaunch mints a new spawn_gen while the branch keeps
validating), then sums every invocation, failed and cancelled included.
Token fields use no-mistakes' per-round deltas so resumed review rounds are
not counted twice, and unrecorded values stay unknown rather than zero.
show prints the record read-only; record appends it once per task
incarnation to data/pipeline-spend.jsonl. Teardown records it for every
ship task it cleans up, before deleting the branch and the task record.

fm_nm_state_db becomes the one owner of where no-mistakes' state database
lives, shared with the capped run-inventory reader.

* no-mistakes(review): Drop spend show command and timeout override

* style(bin): rewrap fm-pipeline-spend header comment

* Make pipeline spend recording opt-in

* feat(bin): record each task's no-mistakes pipeline spend at cleanup

no-mistakes keeps every pipeline agent invocation's token usage only in its
local agent_invocations records, and cold pipeline agents (review, test,
document) leave no session log, so a task's review-loop cost never reached
Firstmate's records and could not be attributed after cleanup.

bin/fm-pipeline-spend.sh attributes a task's runs by the repository
no-mistakes resolves for the task copy, the task branch, and the branch's
creation (a relaunch mints a new spawn_gen while the branch keeps
validating), then sums every invocation, failed and cancelled included.
Token fields use no-mistakes' per-round deltas so resumed review rounds are
not counted twice, and unrecorded values stay unknown rather than zero.
show prints the record read-only; record appends it once per task
incarnation to data/pipeline-spend.jsonl. Teardown records it for every
ship task it cleans up, before deleting the branch and the task record.

fm_nm_state_db becomes the one owner of where no-mistakes' state database
lives, shared with the capped run-inventory reader.

* no-mistakes(review): Drop spend show command and timeout override

* style(bin): rewrap fm-pipeline-spend header comment

* Make pipeline spend recording opt-in

* no-mistakes(document): Note pipeline-spend opt-in gate in teardown comments

* no-mistakes(ci): I fixed all three Greptile findings following your decision. Both test suites pass: tests/fm-teardown.test.sh (105 ok, exit 0) and tests/fm-pipeline-spend.test.sh (8 ok). I didn't run either new test against the old code, so the claim that they fail before the fix is from reading the code, not a run. Nothing was committed or pushed. **ci-1 / ci-2 (bin/fm-teardown.sh):** The rule is that every owned ship task in a home that has opted in gets a durable spend record before its task record is deleted. The `[ -d "$WT" ]` check skipped the recorder when the worktree folder was already gone, so no record was written. I removed that check and kept the other conditions: the task must be a ship task, teardown must own its worktree, and `config/pipeline-spend` must exist. `bin/fm-pipeline-spend.sh` already writes an unavailable-source line when the worktree is missing. The new test `test_teardown_records_unavailable_spend_for_a_gone_worktree` in tests/fm-teardown.test.sh sets up an owned ship task whose worktree is missing, with recording turned on. It checks that teardown succeeds, writes a line with `source=="unavailable"`, `total==null` and a reason saying the task copy is gone, and still removes the task record. Under the old check no ledger would have been written. **ci-3 (bin/fm-nm-run-lib.sh):** When `NM_HOME` and `HOME` were both unset, `fm_nm_state_db` fell back to `${HOME:-}/.no-mistakes`, which resolves to `/.no-mistakes`. It now uses `root=~/.no-mistakes`. Bash expands `~` to the account's home directory even when `HOME` is unset, which matches the previous Python `Path.home()` lookup. The new test `test_state_db_without_nm_home_or_home_uses_the_account_home` in tests/fm-pipeline-spend.test.sh runs the function with both variables unset and compares the result to the home directory from the system's password database. The old code would have returned `/.no-mistakes/state.sqlite`. shellcheck reports nothing new on the changed files. I added one `disable=SC2016` comment in the new test, where the single-quoted `$1` is meant to expand in the child shell

* feat(bin): start tasks from a named base branch (#6442)

* feat(bin): start tasks from a named base branch

Spawns always reset a task's pooled copy to origin's default branch,
so work that belongs on a feature, integration, or release branch
started from the wrong code and opened its PR against the default.

fm-brief.sh --base-branch records the base in the brief, fm-spawn.sh
resets the copy to origin/<base> and records base_branch= in task
meta, and the worker targets its PR at that branch. Review diffs,
the cleanup content check, and scout promotion read the recorded
base. local-only and Gerrit deliveries refuse a named base.

* no-mistakes(review): Anchor brief base-branch parse on scaffold Setup line

* no-mistakes(review): Require explicit --base-branch on spawn, matching brief lines

* no-mistakes(document): Note base-branch reset in fm-spawn freshness header

* no-mistakes(ci): I fixed all four Greptile findings the way you specified. The affected suites pass locally: fm-brief, fm-dod-lib, fm-spawn-pool-base-freshen, fm-review-diff, fm-teardown and fm-task-delivery. shellcheck on the changed files reports only info-level notes, no warnings or errors. I didn't revert the code to watch each new test fail before the fix. - **ci-1 (brief text blocked or redirected spawns):** In bin/fm-dod-lib.sh, `fm_brief_base_branches` now only accepts a `Base branch: <value>` line that comes directly after the base-variant Setup sentence fm-brief.sh writes. Any other `Base branch:` line is ignored. `--base-branch` is still the only authority, and the existing checks in fm-spawn.sh are unchanged. With the flag, the spawn is refused unless a matching pair exists and no pair names a different branch. Without the flag, it is refused only if such a pair exists. The header comment in fm-spawn.sh is updated to match. - The `brief_with_base` test helper now writes the real two-line pair. - The decoy refusal case now uses a full decoy pair in the captain's intent. - New test `test_prose_base_branch_line_is_ignored`: a brief with no base whose intent quotes `Base branch: release/1.2` spawns normally with no `base_branch=` recorded. A spawn with `--base-branch feature/hub` whose intent names release/1.2 starts from `origin/feature/hub` and records it. - **ci-2 (scout base not checked against the forge):** fm-spawn.sh now looks up the project's registered forge with `bin/fm-project-mode.sh --forge`, like the ship path does, before `fm_base_branch_valid` runs. Previously it passed `none`. New test `test_scout_base_branch_refused_on_gerrit_forge`: a scout with a base on a forge=gerrit project is refused at spawn and no meta file is written. - **ci-3 (base not shell-quoted in worker commands):** `fm_dod_block` now quotes the base with `printf %q` in the `--base` and `--base-branch` arguments, and the branch name in the prose stays readable. The test in fm-brief.test.sh uses the base `release/$HOTFIX` and checks that both direct-PR and no-mistakes briefs render `release/\$HOTFIX` in the commands. - **ci-4 (ambiguous PR sentence):** The direct-PR sentence now reads "open a PR with `gh-axi` that is ready for review, not a draft, against the base branch `X` (`--base X`), not the repository default." A brief with no base renders exactly as before. The existing assertion in fm-brief.test.sh is updated

* no-mistakes(ci): Both real failures on PR 6442 (run 37064756768) come from CI tooling and a Pi version bump, not from the base-branch feature. Fixes for both already exist on upstream main, so I applied those two commits to the working tree with `git cherry-pick --no-commit`. Nothing is committed or pushed, and no feature file changed. - **Lint 2:** ShellCheck ran out of memory on `bin/fm-teardown.sh` (rss about 8 GB, exit 251, "shellcheck: out of memory"). The rule that broke is that the lint gate must finish on any valid shell root without exceeding the runner's memory ceiling. Upstream commit d719ef3 (#6443) fixes this in `bin/fm-lint.sh`: a root that hits the memory ceiling is retried without `--external-sources`. It also updates `tests/fm-lint.test.sh` and `docs/fm-test-portable-shards.md`. - **Behavior portable serial 6:** `tests/fm-calm-pi-extension.test.sh` failed with "grep disappeared from /export calm.html HTML while calm mode was on". Pi 1.0.1 renamed its HTML export renderer lookup, and upstream commit fede619 (#6530) updates the test to accept the new name. This change is test-only. Both commits applied cleanly. The four touched files are `bin/fm-lint.sh`, `docs/fm-test-portable-shards.md`, `tests/fm-lint.test.sh` and `tests/fm-calm-pi-extension.test.sh`. None of them is in the feature diff, so the feature contract is unchanged. **Local checks:** - `tests/fm-lint.test.sh` passes. - `bin/fm-lint.sh` on `bin/fm-teardown.sh`, `bin/fm-spawn.sh` and `bin/fm-dod-lib.sh` exits 0 with full ShellCheck analysis. - `tests/fm-calm-pi-extension.test.sh` exits 0 with 7 ok and 0 not ok. One case skipped because the Pi package isn't installed locally, so the Pi 1.0.1 export path that failed in CI couldn't be reproduced here. CI needs to confirm it. - I didn't reproduce the 8 GB OOM locally. The PR has to be updated through the pipeline with a normal push: no force push, no replacement PR, no merge

* fix(bin): name the Firstmate skill file as a fallback in the worker role (#6647)

* fix(bin): point project workers at the Firstmate skill file

The Skill tool cannot resolve a Firstmate skill from another project's worktree, so the launch role names the readable skill file instead.

* no-mistakes(review): fix(bin): name Firstmate skill file as fallback only

* feat(bin): retire a contribution whose forge object is permanently gone (#6655)

* feat(bin): retire a contribution whose forge object is permanently gone

Add fm-contributions.sh retire <task> <url> <captain|fleet> <reason>,
which records actor, reason and time on the saved record and removes
that task/url pair from known, poll rotation and coverage even while a
backlog link remains. Repeating a retire keeps the first provenance;
an unrecorded pair, unknown actor, empty reason or unacknowledged
pending signal is refused.

* no-mistakes(review): Keep retirement per task when settling final owners

* no-mistakes(ci): Made the three Greptile fixes the captain chose. ci-1, retirement needs the captain's word: `retire` now accepts only the actor `captain`. Any other actor is refused, including `fleet`. I removed `fleet` from the usage line, the script header (`bin/fm-contributions.sh`), the check in `bin/fm-contributions.sh` and the retired-record validation in `bin/fm-contributions.jq`, which now requires `.actor == "captain"`. The script header now says a retirement always records the captain's word, because the script cannot verify who runs it and the authority to retire is the captain's. The Bearings skill line in `.agents/skills/bearings/SKILL.md` now says `retire` records the captain's word. In the tests, a `fleet` retire is now a refusal case and must leave the record unretired. Every other retire call in the tests uses `captain`. The PR description has not been changed yet: the ruling asks for the same captain's-word statement there, and that belongs to the PR phase. ci-2, blank reasons: the reason check is now `[ -n "${5//[[:space:]]/}" ]`, so a reason made only of spaces or tabs is refused. I added a refusal test that passes a space-tab-space reason. The header also lists a blank reason among the refusals. ci-3, the gone-object test: the test forge wrapper has a new `not-found` fault that returns HTTP 404 for every `api repos/o/r/...` read. The happy-path test `test_retire_ends_observation_of_a_gone_contribution` now uses it in place of the 502 `down` fault, so it models a permanently gone repository rather than an outage. Verification: `bash tests/fm-contributions.test.sh` passed with exit 0. The last lines of its output include the three retire tests passing and no failures. `shellcheck` flagged only SC2034 (`FM_WAKE_QUEUE` unused) and SC1091 (an unfollowed source of `bin/fm-path-lib.sh`), neither on a line this round changed

* test: extend remote-reply whole-log recapture waits (#6639)

* test: give the remote-reply whole-log recapture a longer wait

* no-mistakes(review): Extend both recapture waits and simplify retry handling

* fix(bin): reopen a pending-reply escalation after its resolve (#6654)

* fix(bin): reopen a pending-reply escalation after its resolve

A retry of an open decision still appends nothing. A later same-kind escalation after a resolved line for that key appends again, so the next loss is visible.

* no-mistakes(ci): Both Greptile findings are fixed in the working tree; the four directly affected test suites pass, and a wider run of related suites had not finished when I returned this result. Invariant: a line already in the status file is recorded, and only a caller with its own evidence of a new episode may append it again. It must hold for every caller of status_event_recorded: the continuity break and the other call in bin/fm-procevent-remote-reply.sh, fm_parent_channel_append_once, and the pending-reply escalation. Changes: - bin/fm-classify-lib.sh: status_event_recorded is back to the idempotent retry check. It returns at the first matching line, and a later resolved line no longer makes that line look new. This fixes ci-1 for every caller and restores the early exit for ci-2, with no separate scan optimization. - bin/fm-pending-reply-lib.sh: _fm_pending_reply_maybe_escalate_locked now owns the reopen. It appends a new blocked line when the record already has an escalated_epoch (it escalated before and was reset) and the keyed decision is no longer open. Otherwise it uses status_event_recorded, so a retry or a resend that leaves the decision open appends nothing. - Comments on status_event_recorded and the pending-reply header describe this scope. Tests: - tests/fm-pending-reply.test.sh: the regression for escalate, operator resolve, second escalate is unchanged and passes. - tests/fm-remote-reply.test.sh: new regression drives the real reader. After an operator resolve of remote-reply-continuity-ios, a repeated handle and ingest of the same break appends nothing and leaves the decision closed. I placed it beside the existing continuity-break test, not in the pending-reply file, because only that file has the reader harness. - tests/fm-classify-corr-token.test.sh and tests/fm-classify-decision-key.test.sh: the cases that expected a blocked line to append again after a resolve now assert it stays recorded. Verification: - fm-classify-decision-key, fm-classify-corr-token, fm-pending-reply and fm-remote-reply test suites all exit 0 with the fix. - With the two bin files reverted to the commit under review, the new continuity regression and the updated parent-publisher case both fail, so they reproduce ci-1. - shellcheck -x on the changed files reports nothing. - Not finished: a background run of every other suite that mentions the parent channel or pending replies had produced no output, pass or fail, when I returned. One known gap: the escalation writes escalated_epoch before phase. If the process dies between those two writes and an operator resolves the key before the next tick, that tick appends a blocked line for the same episode. I left it, because closing it needs new state. Issue 4755 and other escalation behavior are untouched. Nothing is committed

* fix: refuse a confirming Enter on the Claude background-task exit picker (#6666)

* fix: refuse a confirming Enter on the Claude exit picker

The background-task picker still classifies as pending, so a retried Enter confirms "Exit and stop tasks".
Stop after the Enter that opened it, and raise the existing stale wake with the dialog name.

A non-paused secondmate still skips that wake, so a secondmate parked on the picker still looks idle.
Model-downgrade confirmation, MCP approval, and any other Claude exit confirmation are not covered, because there is no recorded screen for them.

* no-mistakes(review): fix: anchor exit picker match and wake second mates

* fix: refuse a typed submit while the Claude exit picker is open

A pane that already shows the picker must not receive the message or a confirming Enter.
The match requires the recorded heading and footer lines, and it is skipped when no dialog name is being recorded.

* no-mistakes(review): restore watcher to main behaviour, drop dialog wake

* no-mistakes(test): test: align Herdr picker fixtures with the preflight read

* no-mistakes(document): document exit refusal on a recognised dialog

* fix: remove the dialog file when exit runs in a subshell

do_exit is invoked from a command substitution, so the parent cleanup never saw the path it is supposed to delete.

* no-mistakes(review): fix: set the dialog file path after the control lock

* no-mistakes(document): document why the dialog file path follows the lock

* no-mistakes(ci): The exit cleanup in bin/fm-control.sh now deletes the dialog file first and releases the control lock second. The changes are in the worktree and are not committed. Invariant: only the process that holds the control lock for a task may write or delete that task's dialog file ($STATE/<id>.composer-dialog, the file that carries the picker name from the composer read to the Enter checks). The old cleanup released the lock and then deleted the file, so the delete ran outside the lock. Places where this invariant must hold: control_cleanup is the only place that deletes the file, and the single assignment after CONTROL_LOCK_HELD=1 is the only place that names it. do_exit only truncates the file, and it runs while the parent holds the lock. A process that loses the lock never sets the path, so it deletes nothing (earlier fix, unchanged). No sibling site needed a change. Fix: in control_cleanup, the rm of the dialog file moved above the fm_lock_release call, with a two-line comment that gives the reason. No dialog recognition or supervision code changed. Test: added test_exit_removes_the_dialog_file_before_releasing_the_lock in tests/fm-control-relaunch.test.sh. It runs one `fm-control exit` with a recording rm on PATH. The lock release removes paths at or under the control lock with rm, so the recording rm writes whether the dialog file exists at that moment. The test asserts the last record is "absent". It starts no second command. Verification, all run locally: - Before the fix, the new test failed with "the dialog file must be gone when the control lock is released, got: present". - After the fix, tests/fm-control-relaunch.test.sh exits 0 with 75 ok lines and no "not ok" line, including the new test and the existing test that the file is gone after exit and after relaunch. - tests/fm-control.test.sh exits 0 with 44 ok lines. - shellcheck -x on both changed files exits 0 with no output. One thing I noticed and did not change: tests/fm-control-relaunch.test.sh prints 615 "rm: cannot remove ... Permission denied" lines during its temp-directory teardown. The count is the same with my change reverted, so this change does not cause it, and the suite still exits 0. I could not read the Greptile check log (the check run was not found), so I worked from the finding text and the user instructions

* fix: restore timeout watchdog compatibility with macOS Bash 3.2 (#6028)

* fix: support stock macOS Bash in timeout watchdog

* no-mistakes(ci): Fixed ci-1 in tests/fm-timeout-lib.test.sh: both startup-owner failure paths now kill the bounded command group before killing the watchdog. Fault injection reproduced both leaks before the fix and confirmed cleanup afterward. Bash 3.2 suite, syntax check, ShellCheck, and diff checks passed; GNU timeout coverage skipped because the binary is unavailable. Production code unchanged

* fix(project-management): use subshell form for Initialize command (#6699)

Wrap the cd command in a subshell to comply with the cd-guard policy
that blocks persistent top-level directory changes in the primary
firstmate checkout. The subshell form (cd projects/<name> && ...) is
accepted by the policy as documented in issue #6502.

Fixes #6502

* feat(bin): add armable daily startup growth check (#6725)

* Add daily startup growth check

* no-mistakes(review): watch printed startup memory, retain growth baselines, drop knobs

* no-mistakes(review): delegate budget verdict, report before publish, pin shim home

* no-mistakes(review): baseline first sightings silently, drop mtime, spare secondmates

* no-mistakes(review): cap the wake line, validate budget verdict fields

* no-mistakes(review): guard record schema, check appends, tighten assertions

* no-mistakes(review): drop arbitrary tracked library, make re-arm idempotent

* no-mistakes(review): correct tracked-set wording, restore shim guards, register artifacts

* no-mistakes(review): exit on signal instead of publishing partial record

* no-mistakes(document): correct startup-growth record removal cost in state registry

* no-mistakes(test): sweep orphaned empty startup-growth temp records on due evaluation

* no-mistakes(document): note watcher need for armed startup growth check

* no-mistakes(ci): Diagnosed "Behavior portable serial 5": the shard failed on tests/fm-contributions.test.sh in test_arm_plumbs_a_configured_budget_into_the_check_shim (inherited mode) with "generated check did not attempt a read" and a missing forge/calls file. Root cause: bin/fm-contributions.sh:379-382 computes DEADLINE=$(date +%s)+BUDGET and then gates each read on DEADLINE-$(date +%s) >= OBSERVATION_RESERVE (= min(BUDGET,15)). At the one-second budget this test uses, that gate demands zero elapsed whole seconds, so a wall-clock second boundary crossing between the two date calls makes the poll loop break before any forge read. The test calls wrap_forge (which installs a fake date reading $FORGE/clock only when that file exists) but never seeded forge/clock, so it ran against the real clock. The repo already documents this exact hazard at tests/fm-contributions.test.sh:675-677 and freezes the clock in every other budget-constrained test: test_budget_exhaustion_keeps_prior_record (both modes), test_genuine_failure_near_deadline_is_unavailable, test_reservation_defers_later_url_when_fifteen_seconds_do_not_remain, test_slow_read_deadline_kill_is_budget_refusal, test_budget_is_cut_down_to_the_watcher_check_bound. test_arm_plumbs was the only omission; its single loop body covers both the configured and inherited modes. The remaining wrap_forge users run the default 20-second budget (5 seconds of slack above the reserve) and are not reachable by this race. Fix (smallest, matching the sibling idiom): added the clock freeze `/bin/date +%s > "$home/forge/clock"` plus a two-line comment inside that test's loop, before the poll. Three added lines in tests/fm-contributions.test.sh; no production code and no other file touched. Verification: reproduced the exact CI failure message by running a scratch copy whose fake date advances one second per read (worst case of the unfrozen clock) - missing forge/calls, "not ok - generated check did not attempt a read". With the fix, the single test passes 10/10 standalone and the full tests/fm-contributions.test.sh passes all 46 assertions with exit 0. tests/fm-startup-growth-check.test.sh still exits 0. bash -n and shellcheck -S warning are clean on the edited file; scratch repro files were removed, leaving only the intended three-line change. Note for the author: this flake is not caused by this branch - tests/fm-contributions.test.sh is untouched by the PR and the race is pre-existing, surfacing only on a loaded runner (the shard's neighbouring assertions show 80+ second gaps). It was fixed because it is genuine nondeterminism with an established in-repo remedy and the check cannot otherwise go green. No work was done on the unselected greptile findings (ci-2..ci-5) or the cancelled "Behavior portable serial 2" check (ci-6)

* no-mistakes(ci): Diagnosed "Behavior portable serial 6": the omitted log section (fetched with gh api) shows the shard failed on tests/fm-calm-pi-extension.test.sh in test_hidden_block_geometry_e2e with "Pi Calm hidden-block geometry E2E did not complete the /reload viewport transition" (exit=1, duration_ms=23437; the family line pure-contract-unit count=3 duration_ms=51193 failed=1 matches devin-harness 3203 + calm-pi 23437 + task-delivery 24553). Root cause: tests/fm-calm-pi-extension.test.sh:2831 wait_for_geometry_transition detected the /reload by SAMPLING a single intermediate frame - it polled tmux capture-pane for Pi's transient "Reloading keybindings, extensions, skills, prompts, themes, and context files..." box and only accepted the rebuilt transcript afterwards (elif gated on saw_transient). Pi shows that box only while the reload runs and replaces it via dismissReloadBox when done, so on a loaded runner the box can live entirely between two polls; saw_transient then stays 0 forever and the 600-attempt wait times out although the reload fully succeeded. Reproduced locally under CPU oversubscription with CI's Pi 1.0.4: 1 failure in 8 runs, diagnostics proving the mechanism ("DIAG: transition timeout saw_transient=0 attempt=600") while the captured viewport already showed the completed reload status row and the intact collapsed transcript. Not a Pi-version regression: 1.0.4's handleReloadCommand and both banner strings are identical to the installed 0.85.1, and shard 6 of the previous pipeline run (job 112490351105) passed this same script on Pi 1.0.4. Invariant violated: a TUI E2E assertion must key off a durable state the UI retains, never off one intermediate frame polling may miss. Sibling enumeration: the helper was defined once and called once; grep for transient/saw_/two-stage waits across tests/, bin/ and .pi/ found no other one-shot-frame wait - every other wait in the file keys off durable text presence/absence or off continuously repeating animation frames (the working-ship motion checks) where sampling eventually connects; no doc or other test referenced the helper. Fix (smallest, one helper): renamed it wait_for_geometry_reload and made it wait for the durable status row Pi appends once the reload completed and the chat was rebuilt ("Reloaded keybindings, extensions, skills, prompts, themes, and context files", a persistent chat child via showStatus, present in both 0.85.1 and 1.0.4) together with CALM_GEOMETRY_FINAL. That marker is absent before the reload and never appears if the reload throws, so the check is strictly stronger than before (a failed reload now fails the test instead of passing on transient-then-final). One file changed: tests/fm-calm-pi-extension.test.sh, 2 hunks, no production code. Verification: deterministic before/after - with polls spaced 0.4s apart (the loaded-runner worst case) the pre-fix helper fails with exactly the CI message and the post-fix helper passes; 12/12 consecutive passes of the isolated test under load with Pi 1.0.4; full file 15/15 ok exit 0 under LANG=C.utf8 (matching the runner locale); bash -n, shellcheck -S warning and bin/fm-lint.sh clean; scratch repro files and the temp Pi 1.0.4 install removed, git status shows only the one intended modification. Notes: (1) the flake is not caused by this branch - that test file is untouched by the PR and the race is pre-existing; fixed because it is genuine nondeterminism and the check cannot otherwise go green. (2) Under far harsher load than CI applies a separate unrelated flake appeared in the same test: Pi's "Warning: tmux extended-keys is off..." row occasionally lands between the collapsed [skill] ahoy row and the final response, making assert_geometry_gap see 5 rows instead of 2. Different cause, never seen in CI, and its plausible remedies would change key delivery for the other TUI tests in the file, so it was left alone rather than growing this fix

---------

Co-authored-by: Quartermaster <quartermaster@users.noreply.github.com>

* fix(bin): reopen the remote-reply continuity decision on a later break (#6708)

* fix: reopen a remote-reply continuity break after repair

A later break for the same route and reason was swallowed after the
operator resolved the first one, because the status line matched for
the life of the log. The continuity ingest now appends again when the
cursor has moved or retirement has reset that episode, and an unchanged
re-read still appends nothing.

status_event_recorded is unchanged. Its other callers are the
pending-reply escalation, which already decides its own episode, the
parent-channel note append, and the remote document transfer note.

* no-mistakes(review): seed continuity episode for already recorded break line

* no-mistakes(document): document when a remote-reply continuity break reopens

* no-mistakes(ci): The adapter now stores the continuity episode record before it appends the `blocked` line, so a failed store appends nothing. The changes are uncommitted in the worktree, in `bin/fm-procevent-remote-reply.sh` and `tests/fm-remote-reply.test.sh`. **Invariant:** a continuity `blocked` line is on the parent status log only when the episode record for that break is already stored. Only the `continuity-broken` branch of `cmd_ingest` appends that line, so the fix is at that one place. **What changed in `cmd_ingest`:** - It decides first whether the break needs a line, then stores the episode record, then appends. - If storing the record fails, it stops with "cannot record continuity episode" and appends nothing. - If the line is already on the log and no episode record exists, it stores the record and does not append. **One addition you did not ask for:** if the append fails after the record is stored, the adapter puts the earlier record back (or deletes the new one when none existed). Without that, the retry would read as the unchanged repeat and append nothing, which is the issue 6701 failure again. **Deviation from your test instruction:** the test does not use `chmod` on the cursor directory. `write_continuity_episode` runs `chmod 700` on that directory before every write, so a read-only directory is made writable again. The test instead makes `mktemp` fail for the episode's temporary file in that directory, the same way the existing receipt-failure test does. **Tests added to `tests/fm-remote-reply.test.sh`:** - Store failure: the second break exits 1, appends no `blocked` line and opens no decision. One retry after storage recovers exits 3 and appends a single line. - Append failure: with the status log read-only, the third break exits 1 and appends nothing. One retry after the log is writable exits 3 and appends a single line. **Verification:** `bash tests/fm-remote-reply.test.sh` ends with "ALL TESTS PASSED" with the fix. Against the script at commit b750c828 the same test file fails at "a continuity break appended its line before its episode was stored". `shellcheck -S warning` reports only an unused loop variable at line 667 of the test file, which this change does not touch. I ran no other test files. The Greptile Review check log could not be retrieved, so I worked from the finding text alone

* fix: record a continuity break's reader position on its status line

A later break at another cursor is then a different line, so the existing
duplicate check appends it and reopens the decision. An unchanged re-read
builds the same line and appends nothing.

* fix: reopen a continuity break after an identical restore

A retirement that puts the same bytes back used to rebuild the recorded line, so the later break stayed closed. The retirement count on that line makes the later break distinct.

* no-mistakes(review): remove continuity match for full-prefix line without retirement count

* no-mistakes(document): clarify what a continuity break status line records

* fix: remove the reply cursor before recording retirement

A stop between those steps must leave the count unchanged, so an unchanged continuity break still builds the same line.

* fix(control): drop busy_gen from the task record when an incarnation is retired (#6733)

* fix(control): drop busy_gen when an incarnation is retired

A deliberate exit removed the busy sidecar and left busy_gen in the task record, so the two records disagreed about whether that incarnation was still observable.

* no-mistakes(review): drop GNU-only chmod and unreached sidecar-absent branch

* no-mistakes(review): correct lock comment to name the deadlock

* no-mistakes(ci): The test `test_exit_drops_meta_busy_gen_with_the_sidecar` in tests/fm-control.test.sh now compares the whole task record (the `state/<id>.meta` file), so the Greptile finding is fixed. Invariant: after `exit` retires an incarnation, the task record must equal the record from before `exit` with only the `busy_gen` line removed. This test is the only place in the change that asserts the record survives the rewrite, so it is the only site to fix. The other `busy_gen` tests assert that the line stays, and they do not go through the rewrite. What changed: before `exit`, the test writes the record without its `busy_gen` line to `expected.meta`. After `exit`, the test runs `diff` between that expected copy and the real record, and fails with the diff output if they differ. This one comparison replaces the two earlier checks (no `busy_gen` line left, and the `window` line present), because it covers both. I did not change bin/fm-control.sh or any other file. How I know it works: - I ran `bash tests/fm-control.test.sh`: exit code 0, 45 lines starting with `ok`, no other lines. - I temporarily changed the rewrite in bin/fm-control.sh to also drop the `harness` line. The test then failed with `not ok - exit should drop only busy_gen from the task record:` and the diff `< harness=codex`. The earlier `window`-only check would have passed that rewrite. I restored bin/fm-control.sh afterwards; `git status` shows only tests/fm-control.test.sh modified. - `bash -n` and `shellcheck` on the test file report no new warnings from the edit. The change is not committed; the working tree holds it

* test: cover PID collisions in harness ancestry detection (#6484)

* test(secondmate-harness): scope fake ps -codex label for pid 5252 to the liveness probe

Closes #6456

* no-mistakes(ci): Updated the collision test to log and assert that PID 5252 was queried before selecting 4242. Full fm-secondmate harness suite passes

---------

Co-authored-by: YifuGu <ironerumi@users.noreply.github.com>

* fix: share Pi Calm's working-ship widget slot (#1854)

* fix(calm): share the standalone Pi Calm working-ship widget slot

Firstmate Calm and the user-global standalone Pi Calm both install an
animated working-ship widget during agent runs. Each claimed its own Pi
widget key, so a session loading both (the main Firstmate home) rendered
two boats. Pi replaces widgets under one key, so claiming the shared
"calm-working-ship" slot keeps dual-install sessions to a single boat
while a Firstmate-only session is unchanged.

Pins the shared slot contract in the working-ship module test so the key
cannot silently diverge again.

* test(calm): pin the shared working-ship widget key in CI, document dual-install

The key-parity assertion inside the Pi fixture only runs where the
@earendil-works/pi-coding-agent package is installed, so CI never
exercised it. Add a source-level twin that needs nothing but the
tracked file, and note in docs/calm.md that the boat shares the
standalone Pi Calm working-row widget slot.

* no-mistakes(review): Add executable dual-install widget replacement coverage

* no-mistakes(review): Guard shared widget cleanup with disposal ownership

* no-mistakes(document): Document shared Calm working-ship slot behavior

* test(calm): read the standalone Calm slot from its own module

The dual-install check registered both boats itself under the shared
slot, so it could only prove that Pi replaces a widget under one key: it
would still pass if the standalone Pi Calm extension installed its boat
under a different key, which is the two-boat regression the check exists
to prevent.

Read the standalone extension's own working-ship module when it is
installed - FM_STANDALONE_CALM_SHIP, else ~/.pi/agent/extensions/calm -
and drive the check with the key that module exports, so a rename on
either side registers two widgets and fails naming both keys. A pinned
shared-slot contract still covers a machine without the extension, and
the run reports which side it used instead of passing silently over an
absent extension.

Verified: the touched Pi Calm suite passes and reads the installed
standalone extension; with a copy of it whose key is renamed to
calm-working-ship-v2 the suite fails naming the drift.

* no-mistakes(review): Gate stock-row restoration by shared-widget ownership

* no-mistakes(review): Removed redundant widget-key source assertions

* no-mistakes(document): Document shared Calm working-ship widget ownership

* feat(bin): add opt-in --herdr-resume-lock-wait to fm-spawn (#6649)

* fix(herdr): make exact-resume presentation-lock wait instead of a bounded timeout

The exact-resume path in bin/fm-spawn.sh used the same 50-attempt-then-
give-up lock acquire as the new-task-create path, but the two paths are
not equivalent on contention: a create has no prior state to strand and
can safely fall back to a flat layout, while a resume is recovering a
specific existing identity that a concurrent recovery may legitimately
be holding the lock for. Giving up there does not degrade gracefully,
it hard-fails the resume outright. The suite's own concurrent
cross-home recoveries test already asserts both concurrent recoveries
succeed with a genuine reclaim, and the file's header comment already
(inaccurately) claimed lock contention falls back to the ordinary flat
layout for both paths alike, so the intended contract was always that
recoveries serialize and both succeed, not that either one refuses
under a short bound.

Give spawn_herdr_presentation_order_lock_acquire a wait mode that uses
this file's own established fm_lock_acquire_wait idiom (already used
for its other fleet-shared locks) instead of the bounded loop, and use
it only at the exact-resume call site. The new-task-create call site
is unchanged and keeps its bounded-then-flat-fallback behavior, which
is already covered by its own passing test. Dead-owner PID-liveness
reclaim inside fm_lock_try_acquire still bounds the wait against a
holder that crashed mid-hold.

Adds a deterministic regression test that holds the shared session
lock from an unrelated process for well past the old bound, then
asserts the resume succeeds with a genuine reclaim and took close to
the full hold duration, so a fix that merely widens the bound rather
than genuinely waiting is still caught. The existing concurrent
cross-home recovery test exercises this under real timing but does not
reliably outlast a fixed bound on its own.

Corrects the header comment's claim that create and resume share one
bounded-then-flat-fallback behavior on lock contention; they no longer
do.

* no-mistakes(document): Document Herdr recovery waiting for presentation lock

* no-mistakes(document): Update stale hard-refusal claim in verification log

* no-mistakes(ci): Fixed the Greptile finding on tests/fm-backend-herdr-presentation-e2e.test.sh:1389 by bounding the resume lock-wait regression's spawn_task call. Added an optional 4th `deadline_seconds` arg to the `spawn_task` helper (defaults to empty, so all ~20 other existing call sites are unaffected and unwrapped by `timeout`). The lock-wait test now passes `LOCK_WAIT_HOLD_SECONDS + 60` (90s) as the deadline, and a dedicated check for exit code 124 emits a clear "hung for over Xs instead of waiting out a Ys lock hold" diagnostic before falling through to the existing pass/fail assertions, which are unchanged. No product code was touched. Verified with `bash -n`, `shellcheck -x` (no warnings), a standalone reproduction of the timeout/no-timeout/success paths, the project's `bin/fm-lint.sh --fast` on the file (clean), and the full `tests/fm-lint.test.sh` suite (all 46 assertions pass)

* no-mistakes(ci): Replaced the direct `timeout "$deadline_seconds"` call in `spawn_task()` (tests/fm-backend-herdr-presentation-e2e.test.sh) with the repo's portable bounded-execution helper: sourced `bin/fm-timeout-lib.sh` at the top of the file and changed `deadline_cmd=(timeout "$deadline_seconds")` to `deadline_cmd=(fm_run_timed "$deadline_seconds")`. This removes the GNU/BSD `timeout` dependency that would fail with exit 127 on a stock macOS host without coreutils, while preserving identical semantics (exit 124 on bound-hit, command's own exit otherwise), which the existing `[ "$LOCK_WAIT_STATUS" -eq 124 ]` diagnostic check already relies on. Verified: `bash -n` syntax check, `bin/fm-lint.sh --fast` clean, full `tests/fm-lint.test.sh` suite (46/46 pass), and a standalone repro confirming `fm_run_timed` returns 124 on timeout and 0 on success identically to the prior `timeout` call. No other direct `timeout` calls exist in this file or elsewhere in the PR's diff, so no sibling sites remain

* fix(herdr): gate exact-resume lock wait behind --herdr-resume-lock-wait

Keep refuse-by-default on presentation-order lock contention for Herdr
exact resume. Callers that need concurrent recoveries to serialize must
pass --herdr-resume-lock-wait; unbounded blocking on a third-party session
lock is never the default.

Update docs and the real-Herdr e2e suite so the default path asserts the
refusal and the opt-in path asserts the wait.

* no-mistakes(test): Fix e2e test's lost exit status after if/fi with no else branch

* docs(herdr): stop advertising --herdr-resume-lock-wait on --relaunch

The relaunch path reuses the recorded endpoint and never takes the
presentation-order lock, so the flag is inert there. Drop it from the
--relaunch usage line and state where the flag applies.

* no-mistakes(review): Clarify lock-wait docs; simplify bash-3.2-safe spawn_task helper

* no-mistakes(ci): Fixed ci-1 (Greptile P2). In tests/fm-backend-herdr-presentation-e2e.test.sh, the failure cleanup `cleanup_all` stopped only `LOCK_CONTENTION_OWNER_PID`. It now also stops `LOCK_REFUSE_HOLDER_PID` and `LOCK_WAIT_HOLDER_PID`, the holders of the two new contention cases, so a `fail` before their explicit `wait` no longer leaves them running. Both new PIDs are initialised empty next to the existing one, and each is cleared right after its successful `wait` so cleanup never touches a finished PID. I changed nothing else. `bash -n` passes. The real Herdr e2e run passed both new cases ("default resumed identity refuses session lock contention" and "--herdr-resume-lock-wait waits out session lock contention instead of refusing"). The full run hit my 550s timeout in a later, unrelated case, after the new cases passed

* fix(bin): let TERM stop a watcher blocked in a pane capture on bash 3.2 (#6762)

* fix(bin): let TERM stop a watcher blocked in a pane capture on bash 3.2

Stock macOS bash 3.2 holds a HUP or TERM until a running command
substitution's child exits, and the watcher read every pane through
$(fm_backend_capture ...). A blocked backend read therefore held the
watcher's stop for as long as the read lasted, and a stopped watcher left
the hung read orphaned. tests/fm-watch-triage.test.sh
test_term_stops_a_watcher_blocked_inside_a_poll failed on /bin/bash 3.2
for this reason while passing on bash 5.

Pane captures now go through watcher_capture, which runs the read as a
waited background process group recorded like a check's, so the stop is
honored at once and watcher_cleanup stops a read still in flight along
with its per-call output file.

* no-mistakes(review): Run drain-ring idle capture in watcher shell, add regression test

* no-mistakes(document): Document watcher TERM handling for blocked checks and captures

* no-mistakes(test): Silence bash 3.2 setpgid race noise from watcher captures

* fix(bin): verify the capture group and scope the stop claim to pane reads

watcher_capture now confirms its background read leads its own process
group, as run_check_capture already does, so watcher_cleanup never relies
on a group that set -m failed to create. The comment and continuity doc
now say only fm_backend_capture pane reads go through watcher_capture;
agent-state and composer-state reads still run inside command
substitutions.

* feat: add per-home worker tool exclusions (#6750)

* feat(spawn): add per-home worker tool exclusions

Add an optional per-home config/crew-exclude-tools file listing tool names to hide from workers, one per line, with blank lines and # comments allowed.
It applies to every ship and scout launch and relaunch in that home, is never inherited by another home, and does not affect secondmate agents.
Pi and pi-signed apply it through --exclude-tools, which also covers MCP tool names.
Any other runtime, and a raw launch command, refuses the launch when the list is non-empty rather than ignoring it.
Malformed entries are refused before provisioning, and before a relaunch stops a running worker.
Exclusions that match no tool in the worker's loaded registry are reported as unverified warnings in its status record instead of refusing the worker.

Closes #6744

* no-mistakes(review): Preserve UTF-8 exclusion paths and verify Pi lifecycle behavior

* no-mistakes(document): Clarify worker tool exclusion documentation

* no-mistakes(ci): Fixed ci-1 in bin/fm-exclude-tools-lib.sh: a failed read now returns an error before printing names, so all shared launch and relaunch callers refuse rather than silently dropping exclusions. Added deterministic regression coverage for a file disappearing after readability checks across Pi/pi-signed ship and scout launches. Reproduced the original failure; verified 83 spawn checks, 77 relaunch checks, direct parser/runtime failure cases, full targeted lint, Bash syntax, and git diff --check. Relaunch tests passed with existing fixture-cleanup permission warnings. ci-2 remains unchanged per the user's decision; the outer executor owns the fresh CI run

* feat(bin): defer spawns beyond a declared per-project capacity, opt-in (#5343)

* refactor(bin): share the local Firstmate home walk from the wake library

Teardown's walk over the root home and its registered local secondmate homes
moves into bin/fm-wake-lib.sh as fm_local_firstmate_state_dirs, next to
fm_firstmate_root_home, so a second consumer can count task records across
this machine's homes without a copy. Teardown keeps its exact refusal wording
through a thin wrapper.

* feat(bin): defer spawns beyond a project's declared machine capacity

A project whose machine-local resource only serves a few workers at once had
no way to tell Firstmate so: every queued item was launched, and the surplus
workers spent full-context turns retrying the resource.

config/project-capacity in the root home now declares how many workers each
named project admits at once on this machine. bin/fm-spawn.sh counts the ship
and scout records on the same project origin across the root and its local
secondmate homes, skipping ones whose ready PR is recorded, while holding the
shared project lock through publication. A spawn with every place held exits 75
before any brief render, endpoint, worktree, record, or backlog move, so the
item stays queued; batches report it as deferred. Undeclared projects keep
today's uncapped dispatch, and an unreadable declaration refuses rather than
guessing the limit.

Refs #4237

* no-mistakes(review): Document that capacity matches the clone directory name

* no-mistakes(document): Rewrap stale fm-wake-lib root-home doc comment

* no-mistakes(review): Dedupe local state dirs by identity to avoid double-counting

* no-mistakes(document): Rewrap fm_local_firstmate_state_dirs error doc comment

* no-mistakes(ci): I fixed all four Greptile findings. All 14 tests in tests/fm-project-capacity.test.sh pass, and shellcheck at warning level is clean on the changed files. Each new test failed against the old code and passes now. - **ci-1 (spaced names):** a declaration line must give a name its capacity whenever the name is a valid clone directory name. `fm_project_capacity_lookup` now trims each line, skips blank lines and lines whose first non-blank character is `#`, and takes the last field as the capacity. Everything before that field is the name, so it may contain spaces. The old error cases still refuse: a single field is rejected, and trailing text leaves a last field that is not an integer. The library header and docs/configuration.md now say a name starting with `#` cannot be declared. New test `test_spaced_project_name_is_declared` declares `my heavy project 1` next to an indented comment line and gets a deferral. - **ci-2 (unreadable records):** the holder count must never silently leave out a holder. `fm_project_capacity_occupants` now refuses when a local home's state directory exists but cannot be read or listed, or when a `.meta` file cannot be read. The error names the path, and `fm-spawn.sh` shows it in its existing refusal message. New test `test_unreadable_holders_refuse_admission` covers an unreadable record in the root home and an unreadable state directory in a registered local secondmate home, then checks that the spawn is admitted once both are readable. The test is skipped when run as root. - **ci-3 (Orca lock):** any spawn that can become a holder for a capped project must take that project's lock. The lookup now also reports whether the declaration caps any project at all, and an Orca spawn takes the per-origin lock whenever it does. This covers every capped same-origin clone. It also covers some cases where no same-origin clone is capped, because a spawn cannot find clones under other directory names without searching for them. With no declaration file, Orca still skips the lock. The comments in the library and in the `fm-spawn.sh` header are updated. The Orca test now clones the origin as `project-2`, which has no declaration, and checks that its Orca spawn refuses while the lock is held and publishes no record. - **ci-4 (worktrees):** `assert_nothing_created` now also compares the project's `git worktree list` from before and after a deferred spawn. Both tests that call it take that snapshot first. Files changed: bin/fm-project-capacity-lib.sh, bin/fm-spawn.sh, docs/configuration.md, tests/fm-project-capacity.test.sh

* fix(bin): declare capacity for a project name that begins with #

A clone directory whose name begins with # was skipped as a comment, so that project stayed uncapped. A line is a declaration when the # is written against the rest of the name and the line ends with a capacity; a # followed by whitespace stays a comment.

* no-mistakes(document): Rewrap project-capacity library header comment

* no-mistakes(ci): Lint 2 fails because this PR's code pushes ShellCheck past its memory cap. ShellCheck ran out of memory analyzing bin/fm-teardown.sh in CI (reason=memory, rc=251, peak about 8.39 GB). On current main the same file passes at about 7.29 GB. **Cause:** the new `fm_local_firstmate_state_dirs` function in bin/fm-wake-lib.sh had a conditional `. fm-secondmate-registry-lib.sh` with a `# shellcheck source=` directive inside the function. ShellCheck followed that source again, inside a function scope, wherever fm-wake-lib.sh is sourced, and bin/fm-teardown.sh is the heaviest root that sources it. Measured locally with `shellcheck --norc --external-sources bin/fm-teardown.sh`: - current main (fd325b1b): 7.29 GB - main merged with this PR: 7.86 GB - the same merge without the in-function source: 7.27 GB **Rule this restores:** this change must not make any lint root heavier than it is on main. That function holds the only new nested source in the change. **Fix:** I removed the in-function source, which no caller needs. Both callers already load the registry library at top level before calling the function: - bin/fm-teardown.sh sources it directly. - bin/fm-spawn.sh, the only user of bin/fm-project-capacity-lib.sh, gets it through bin/fm-ff-lib.sh. I also documented the requirement in the function's comment and in the "Requires" note in bin/fm-project-capacity-lib.sh. No behaviour changes. **Verification:** - ShellCheck on head: bin/fm-teardown.sh peaks at 7.12 GB and bin/fm-spawn.sh at 6.68 GB, both with rc=0. bin/fm-wake-lib.sh and bin/fm-project-capacity-lib.sh lint clean. - tests/fm-project-capacity.test.sh, tests/fm-teardown.test.sh (102 ok) and tests/fm-teardown-endpoint-safety.test.sh all pass. Files changed: bin/fm-wake-lib.sh, bin/fm-project-capacity-lib.sh

* fix(bin): release the Herdr session lock when reclaim finishes

A concurrent resume in another home waits five seconds for that lock.
Reclaim is the last presentation change on the recovery path, so holding
the lock through the launch tail made the waiter time out. The contributions
arm check also freezes its one-second clock, the same way the budget tests
do, because an unfrozen clock can tick past before the first forge read.

* no-mistakes(review): Keep Herdr session lock through launch handoff after reclaim

* no-mistakes(review): Skip the spawning task's own record in capacity count

* no-mistakes(review): Restore release test comment above its test

* docs: scope PR-ready re-evaluation to a declared project capacity

A ready pull request frees a place only when that project declares capacity, so the always-loaded backlog contract should re-evaluate on that handoff only in that case.

* fix(bin): name the Lavish read message count by the same label as its section (#6785)

fm-procevent-lavish.sh read labels tag=message rows SESSION-ENDING MESSAGE
only when session_ended is true and CAPTAIN MESSAGE otherwise, but the
count line always said session_ending_message_count. Several composer
messages on a still-open board were therefore counted as session-ending.

The count line now follows the same session_ended switch:
session_ending_message_count once the session ended, captain_message_count
otherwise. Message rows stay out of the annotation count, per triage.

Fixes #6743

* fix(herdr): make the presentation lock namespace per OS account (#6780)

The Herdr presentation lock namespace was the fixed machine-global
/tmp/firstmate-herdr-presentation, so on a host where two OS users run
Firstmate on Herdr the first account to create it owned it and every
teardown from the other account was refused with no way to clear it.

Suffix the namespace with the account uid. The owner-uid and mode-700
checks are unchanged, so a foreign-owned or wrong-mode name at this
account's path is still refused and never adopted, chowned, or removed.

Fixes #4716.

* Fix OpenCode arm plugin to decide with the shared supervision predicate (#6809)

The OpenCode session plugin's shouldArm kept its own copy of the need
test that only looked for in-flight task records, while the turn-end
guard decides with fm_supervision_needed in bin/fm-supervision-lib.sh,
which also counts registered process-event sources and trusted custom
checks. With an empty fleet but any registered source or check, the
guard blocked every turn end while the plugin declined to arm - a loop
the guard's own repair line could not resolve because it names the
plugin as the fix.

The plugin now delegates the decision to the shared predicate through
bash, keeping the local away-record decline and the x-mode.env arm
override. OpenCode plugin test fixtures now carry the real predicate
their arming path sources, and the arm suite gains six cases asserting
the plugin's decision against the shared verdict over the same
synthetic state directories.

Co-authored-by: Mia Sun <mia@Bigs-Mac-mini.localdomain>

* fix(bin): resolve a pending reply only from its own task's status line (#6792)

* fix(bin): resolve a pending reply only from its own task's status line

Remote reply ingestion handed every corr= token in a mate's payload to
fm_pending_reply_try_resolve together with that mate's own status log,
so one mate echoing another mate's token resolved the other request.
Honor a status-file override only when it is the record's own
parent_status, and match the corr= token as a whole word.

Fixes #6538

* no-mistakes(document): docs: scope remote reply settlement to the asked mate

* docs(skills): index the six missing agent-only skill triggers (#6784)

agent-skill-trigger-index claims to be the complete agent-only trigger
index but omitted operational-home-layout, session-start-recovery,
validation-supervision, ship-landing, scout-completion, and
away-quiet-supervision. Add each with its own description's trigger,
placed beside the related entries.

The decision-hold-lifecycle redirect stub stays out, per triage.

Fixes #6503

* test: make agent process fixtures compatible with multicall sleep (#6814)

* test: share a rename-safe agent stand-in across liveness suites

On Ubuntu 26.04, `sleep` is the uutils multicall binary, which refuses to
run when invoked through a symlink named after another utility. The Herdr
descendant process-walk tests built their agent-named process as a `pi`
symlink to the host `sleep`, so the process exited at once, its parent shell
was gone before the walk ran, and both cases read `unknown unreadable` and
failed on that host. The suite stops at its first failure, so every later
case went unrun. The Herdr control smoke test's `claude` symlink has the
same construction.

The tmux liveness suite already solved this with a host-compiled spinner and
a survival-checked `sleep` fallback. That builder moves into tests/lib.sh as
fm_agent_standin, and the tmux suite, both Herdr descendant cases, and the
Herdr control smoke test now use it. When no stand-in can survive a foreign
name, a case skips with the reason instead of failing.

tests/fm-test-fixtures.test.sh gains a portable regression with a fake
multicall `sleep`, so it bites on hosts whose own `sleep` is single-purpose.

* no-mistakes(document): Correct Herdr verification fixture reference

* ci: retrigger cancelled shard

* fix(bin): keep the supervision host's successor watcher alive after the Stop hook's group is torn down (#6787)

* fix(bin): keep the supervision host's pass-through successor out of the hook's process group

The successor a main-only pass-through leaves for main shared the Stop hook's
process group, so the harness tearing that group down after the exit-2 rewake
stopped it. The stop published downtime and the next park's first cycle
announced an empty check: rearm-resurface, which woke main again in a loop.
Start that successor in a process group of its own, as the hook's own
handling successor already is.

* no-mistakes(review): Give the at-turn successor left for main its own group

* no-mistakes(document): Document own-group successor for turn-start hand-back too

* no-mistakes(ci): I made the change you asked for: both new teardown tests in tests/fm-supervision-host.test.sh now call the existing `stop_home_processes "$home"` just before `pass`. The tests are `test_successor_left_at_the_turn_survives_the_hook_process_group_teardown` and `test_pass_through_successor_survives_the_hook_process_group_teardown`. No production code and no other tests changed. The rule broken was that a test must not leave a home's watcher or arm processes running after it passes. These two were the only cases in the changed area that broke it. The other host+hook tests already stop their home, and `test_successor_close_during_main_turn_is_delivered_at_the_next_turn_end` leaves its watcher behind too, but it is an older test you said not to touch. The only reason anything was left over is that the successor's arm now sits in its own process group, outside the hook's teardown. `stop_home_processes` kills the watcher by the pid in its lock file, which stops it no matter which group it is in. **Checks run:** - I ran just these two tests from a scratch copy of the suite (since deleted). Both pass in about 13 seconds. - After each test, a process listing filtered to that test's home directory came back empty once the processes had about a second to exit after TERM. - `bash -n` on the test file passes. - `shellcheck` is not insta…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Workers beyond a project resource's capacity are run anyway and spend turns retrying admission

2 participants