fix(bin): republish parent metadata after a remote secondmate relaunch - #5583
Conversation
A host-local relaunch rewrote only the far endpoint, so this home kept the old harness, model, and effort, and appending those keys after pr= broke pull-request poll authentication.
|
Speaking as Kun's firstmate: Verdict: Diff reviewed ( VISION: (1) One interface — aligns (record honesty). (2) Authority explicit — aligns (no new grant). (3) Scripts/agents — aligns (deterministic metadata rewrite). (4) Restart non-event — aligns (parent record matches host). (5) Delegation spine — aligns. (6) Fleet outlives vendor — aligns (Herdr/SSH fixtures unchanged). (7) Scope — aligns (bin/docs/tests only). Aligns overall. Attestation: MATCH |
|
Friendly nudge: head |
|
Speaking as Kun's firstmate: Verdict: Re-reviewed after author nudge (2026-09-26). Diff is the parent-side wrapper + route-block model/effort + restart caller. Author body Fixes #5332 verified against ready-for-pr #5332 (documented VISION (per rule):
Aligns overall. Attestation MATCH. CI green / CLEAN. Auto-merging squash. |
|
Speaking as Kun's firstmate: this is merged. Thank you @tiago-peixoto — really appreciate you taking the time on this. |
…used two of them, and both are now fixed. 1. **fm-live-gate.test.sh (shard 7), caused by this PR.** Rule it broke: every guard in the live-harness family has to start with the shared `fm_live_gate`, so that `FM_LIVE=0` switches them all off together. The new `tests/fm-calm-branch-outcomes-live-e2e.test.sh` used its own on/off check and its own pi/tmux checks instead. I replaced those with `fm_live_gate opt-in FM_CALM_BRANCH_OUTCOMES_LIVE_E2E pi tmux`, which keeps the "refuse to pass having checked nothing" behaviour when the guard is asked to run. It was the only file in this diff with its own gate. `tests/fm-live-gate.test.sh` now passes: all 40 live guards skip together under `FM_LIVE=0`. 2. **fm-pi-branch-responsiveness-live-e2e.test.sh (shard 6), caused by this PR.** Rule it broke: any test that copies `fm-branch-supervision.ts` into a fixture must also copy every `./lib` file it imports. This PR added an import of `./lib/fm-calm-branch-outcomes.ts`, but the fixture didn't copy that file, so Pi couldn't load the extension and its TUI never drew. I checked every test that copies the extension. The same gap was in `tests/fm-pi-branch-live-e2e.test.sh`, which CI didn't catch because it only runs when asked, so I fixed both. `fm-pi-codex-native` loads the extension straight from the repo, so it needed no change. The responsiveness guard now passes locally against the real Pi 0.87.1: "ok - supervision outcome delivery keeps the real Pi 0.87.1 TUI echoing keystrokes at its unloaded floor". 3. **fm-remote-secondmate-relaunch.test.sh (shard 6), not caused by this PR.** It fails with "could not arm the PR poll fixture for the relaunch-ordering test". That part of the test was added upstream in e789e52 (kunchenguid#5583). It fails the same way when run on a clean export of the base commit ea7c7f7, with none of this branch's changes, and this PR touches nothing it uses. I made no change for it. While debugging number 3, a temporary `sed` edit I made emptied `bin/fm-pr-check.sh` in the worktree. I restored it with `git checkout`, and it is not part of the changes. Also checked: `shellcheck -x` is clean on the three changed test files, and `tests/fm-calm-branch-outcomes.test.sh` still passes
* fix(bin): report a Lavish source armed only after its listener is running (#5566)
* fix(bin): report a Lavish source armed only after its listener is running
Registration alone was treated as ready, so arm could succeed before anything was collecting from the board.
* no-mistakes(review): Guard Lavish arm launches, keep retire refusals, report live prior listener
* no-mistakes(review): Keep polling through window before reporting a still-live prior listener
* test: wait for a capture's claim to drop before the next arm
The result is stored before the runner exits, so a re-arm in that gap was meeting a live claim.
* no-mistakes(document): Record Lavish arm readiness evidence in verification doc
* no-mistakes(ci): Both failures were caused by this PR, and both are fixed with test-only edits. Lint 2 (ShellCheck SC2034): this branch removed the only use of `reply_id` (a `start "$reply_id"` call) from tests/fm-procevent.test.sh, which left the assignment at line 1450 unused. I deleted that assignment. It was the only `reply_id` in the file. ShellCheck is now clean on both test files. Behavior portable serial 4: the failing test was tests/fm-bearings-board.test.sh, in the check "registration consumed its answer before the any-origin binding existed". I reproduced it locally: the hold was still `state: queued` when the test checked it. - What must hold: the test's check that the hold is closed must run after the listener has captured the answer. - Why it broke: the test used a stand-in adapter that ran `fm-procevent.sh start` in the foreground after `arm`, so capture finished before build returned. On this branch, `arm` starts the listener itself in the background, so the real listener captures the answer and closes the hold a moment after build returns. - Fix: removed the now-redundant stand-in adapter, the copied runtime directory, and its extra environment variables. The test now runs the real build through the existing `run_board` helper and waits up to about 10s for the hold to reach `state: done`. The checks that follow are unchanged: `Resolution mode: answered` and the any-origin binding. - Other tests: this was the only test in the file that stood in for the adapter this way. The shard's other pure-contract-unit test (tests/fm-trace-context-lib.test.sh) passed unchanged. Verification: - tests/fm-bearings-board.test.sh passed 3 times in a row via bin/fm-test-run.sh, all 18 checks, about 53s per run. - tests/fm-procevent.test.sh was not rerun, because the lint fix only removed an unused assignment
* fix(bin): stop repeating unknown-wake escalations that were already delivered (#5599)
* fix(bin): acknowledge a delivered unknown-wake escalation
The same unrecognized wake was escalated again after it had already been handled, because delivery never recorded that identity.
* no-mistakes(review): Scope unknown-wake acknowledgements to one away session
* no-mistakes(review): Clear delivered digest when unknown-wake ack write fails
* no-mistakes(review): Limit unknown-wake suppression to acknowledged lines
* no-mistakes(document): List unknown-wake ack file among away-session artifacts
* fix(bin): keep a stated default-key retraction from cancelling a keyless wait (#5587)
* fix(bin): keep a stated default retraction from cancelling a keyless wait
A resolved line that names the shared default decision bucket was closing the keyless live wait that only prints as that same key. Keyless self-retraction still closes the keyless wait.
* no-mistakes(review): Keep declared waits standing past foreign-key resolved lines
* no-mistakes(review): Bound declared-wait read and share one decision-key parser
* no-mistakes(document): Document supervisors' key-aware declared-wait read
* fix(bin): terminate a remote job worker that lost ownership on TERM (#5544)
* fix(bin): terminate a remote job worker that lost ownership when it receives TERM
A serving worker whose lock directory is gone can no longer quarantine shutdown, and resuming service publishes a false ready heartbeat. Exit after stopping only that worker's own command tree, without removing a replacement owner's lock.
* no-mistakes(review): Check worker lock ownership before publishing shutdown quarantine
* no-mistakes(document): Correct worker shutdown comment on replacement-owned lock
* fix(bin): keep an ousted remote job worker off the replacement quarantine
Shutdown can lose the lock after the first ownership check and before it
writes or clears quarantine. Bind both operations to the directory object
this process still owns so a replacement's quarantine stays untouched.
* no-mistakes(review): Make ousted-worker shutdown test reliably reach quarantine clear
* no-mistakes(document): Reattach worker_shutdown doc comment to its function
* no-mistakes(ci): Fixed the failing check (Behavior portable serial 7) with a test-only change to the stall test in tests/fm-remote-job.test.sh. Product code is unchanged; no other test changed. Cause: after the decoy dies, both workers run the same check-exists, read, delete sequence on the job records. On the CI runner the replacement deleted a record between the ousted worker's check and its read. The ousted worker exited 125, and because the file runs under set -e the unguarded `wait` ended the test with 125. The exit trap then killed the replacement, which produced the "Killed" line. Reproduction: a temporary 0.3 s delay between the check and the read, applied to the ousted worker only, made the committed test fail exactly as in CI (exit 125 and the "Killed" line). The new test passed with the same delay. The delay is reverted, along with a similar debug hook that the timed-out attempt had left in bin/fm-remote-job-worker.sh. Test changes: - The replacement is frozen (and confirmed stopped) before the decoy is killed and resumed only after the ousted worker exits, so only one worker touches the job records at a time. - The ousted worker is stopped only once its quarantine exists and its lane is reaped, which places it inside its stop loop. - Every fixed poll loop is now a wait on a named condition with a 30 s deadline and an explicit failure message. Exit detection also handles zombies. - The exit trap kills and waits for the decoy and both workers on every path. - A non-zero exit from the ousted worker now fails with its exit code and stderr instead of silently ending the file. The test still proves that the resumed ousted worker exits 0 and leaves the replacement's lock, quarantine contents and quarantine inode unchanged. Verification: the full test file passed four times on its own and three times under nice -n 10 with four busy-loop CPU hogs; bin/fm-lint.sh passes. Changes are not committed
* no-mistakes(ci): I fixed the failing check (Behavior portable serial 7) by changing only the stall test in tests/fm-remote-job.test.sh. Product code is unchanged. **What failed:** "an ousted worker in shutdown leaves the replacement quarantine untouched" failed on CI with the ousted worker exiting 125 ("could not stop the active command tree"). **Why:** during shutdown, the worker retries the still-running decoy command group a fixed 100 times, 0.01 s apart, then gives up and exits 125. The test tried to freeze the worker partway through those retries by sending SIGSTOP from outside. On a slow runner the retries ran out before the stop arrived, so the worker had already given up. The invariant is that the test must hold the ousted worker inside that retry loop until the replacement owns the lock. That was the only place the test depended on timing. The other waits already watch for a named state change with a 30 s deadline. **Fix:** - The ousted worker now starts with a small `sleep` wrapper at the front of its PATH, and the SIGSTOP race is gone. - The wrapper only holds a `sleep` called directly by that worker's own process (it checks its parent pid against a hold file) while its quarantine file exists. - The only such `sleep` is the first retry in the shutdown stop loop, so the worker waits there as long as needed. - The wrapper writes a marker when it starts holding. The test waits for that marker, then hands the lock to the replacement, freezes the replacement, and kills the decoy. - The test releases the worker by deleting the hold file. Deleting the whole temp directory also releases it, so a failed run cannot leave the wrapper looping. - A process leak: the test overwrites the job's command-group record with the decoy, so no worker ever stopped the job's real command. `fm-hold-job.sh` and its `sleep 30` stayed running for up to 30 s after the test. The test now records that group before overwriting it and kills it at the end of the test and in the exit cleanup. - The test still asserts the same things: the ousted worker exits 0, and the replacement's lock, quarantine contents and quarantine inode are unchanged. **Verification:** - The full file passed twice on its own, twice under `nice -n 10` with six busy-loop CPU hogs, and twice more after the leak fix. - `pgrep` found no leftover processes afterwards. - With the worker from just before the fix commit (cf45cb6^), the test still fails with "the ousted worker wrote or cleared the replacement quarantine during shutdown", so it still proves the fix. - `bin/fm-lint.sh` passes. - I did not reproduce the CI failure locally. The cause comes from the fixed retry limit and the CI error message. The changes are not committed
* no-mistakes(ci): I changed only the stall test ("an ousted worker in shutdown leaves the replacement quarantine untouched") in tests/fm-remote-job.test.sh. Product code is unchanged, and so is every other test. **Invariant:** the pid written to the job's group record must be a process-group leader whose group dies when that one process is killed. Otherwise the worker's bounded stop loop never sees the group die, gives up, and exits 125 ("could not stop the active command tree") before it reaches the lost-ownership exit. The decoy is the only place in this test that depends on this. **Fix:** - The decoy used to be `set -m; sleep 30 &`. It now starts as `perl -MPOSIX=setsid -e 'setsid() >= 0 or exit 1; exec @ARGV' sleep 30 &`, which gets its own session and group without shell job control. tests/fm-procevent.test.sh already uses the same idiom. - The test now waits, with the file's usual 30 s deadline and a named failure, until `ps -o pgid=` of the decoy equals its pid before writing it into the group record. This way the worker can never read the record before `setsid` has run. - The existing steps are unchanged: the test kills the decoy, reaps it with `wait` before releasing the hold file, and the exit trap still kills and reaps the decoy and both workers. - The assertions are unchanged: the ousted worker exits 0, and the replacement's lock pid, quarantine text and quarantine inode stay the same. **Cleanup:** I reverted a debug `printf` hook that the timed-out previous attempt had left in bin/fm-remote-job-worker.sh, and deleted its untracked `.tmp-repro/` directory. Neither was committed. **Verification:** - The full tests/fm-remote-job.test.sh passed twice normally and once under `setsid -w` with stdin from /dev/null (no controlling terminal). - `bin/fm-lint.sh` passes. - No leftover `sleep 30` processes afterwards. **Not reproduced:** I could not reproduce the CI failure locally. On this host `set -m` made the decoy its own group leader even without a controlling terminal, so the cause on the runner is not confirmed. The change removes the test's reliance on shell job control, as the user asked. Changes are not committed
* no-mistakes(ci): I changed only the stall test ("an ousted worker in shutdown leaves the replacement quarantine untouched") in tests/fm-remote-job.test.sh. Product code is unchanged, and so is every other test. **Invariant:** the group record the ousted worker checks in its stop loop must stay the job's own command group, and the test must stop that group before it releases the hold. Otherwise the bounded retry keeps seeing a live group, gives up, and exits 125 ("could not stop the active command tree") before it reaches the lost-ownership exit. The test overwrote this record in one place (the decoy) and stopped the group in one place (killing the decoy); both are changed. **Fix:** - I removed the setsid decoy and the overwrite of `.claim/group`. The record keeps the job's real command group, which the test still saves as `STALL_JOB_GROUP`. - The two-line `group_start` stays. It is still needed: without it the worker kills the real group on its first pass, before the replacement takes over, so the hold would never matter. - The `sleep` wrapper that holds the worker at its first stop-loop retry is unchanged. - After the replacement owns the lock, its quarantine is planted and it is frozen, the test runs `kill -KILL -- -$STALL_JOB_GROUP`. It then waits, with the file's usual 30 s deadline and a named failure, until `kill -0` on the group fails. Only then does it remove the hold file. The worker therefore always sees its own command already stopped and never races its retry budget. - The exit trap still kills the saved command group if the test fails. It can't `wait` on that group because the group is not a child of the test shell. The decoy variable and its cleanup entry are gone. - The assertions are unchanged: the ousted worker exits 0, and the replacement's lock pid, quarantine text and quarantine inode stay the same. **Verification:** - The full tests/fm-remote-job.test.sh passed twice normally. - It passed once under `setsid -w` with stdin from /dev/null (no controlling terminal). - It passed once under `nice -n 10` with six busy-loop CPU hogs. - With the worker from before the fix (cf45cb6^), the test still fails with "the ousted worker wrote or cleared the replacement quarantine during shutdown", so it still proves the fix. - No `fm-hold-job` or `sleep 30` processes were left afterwards. - `bin/fm-lint.sh` passes. **Not reproduced:** I couldn't reproduce the CI failure locally; the decoy version also passed on this host. So I can't confirm why the decoy group stayed alive on the runner. The new wait turns any leftover live group into a clear named failure instead of an exit 125. The changes are not committed
* fix(bin): keep a dead command group dead on bash 5.2
A bare return inside the liveness check drops the failing kill status when the check runs in a conditional, so shutdown keeps treating a stopped group as alive and exits 125.
* no-mistakes(review): Use bash 3.2 fd syntax and fix trap return comments
* docs: make configuration settings easier to find and understand (#5589)
* docs: make configuration settings easier to find and understand
* no-mistakes(review): Restore dropped qualifiers and fix misplaced config doc labels
* no-mistakes(review): Restore three dropped qualifiers in configuration reference
* fix: limit project memory edits to factual corrections (#5636)
* fix: bound worker edits of project AGENTS.md/CLAUDE.md to factual corrections
These files are loaded into every agent session of a project, so additions
should be a deliberate human choice rather than automated task output. The
ship brief's project-memory section and AGENTS.md section 6 previously invited
workers to record durable knowledge, which let project AGENTS.md files accrete
detail the codebase or README already carries. Workers now edit only to fix
factually wrong content - including content their own change made wrong - and
fm-ensure-agents-md.sh runs only alongside such a correction. Stow no longer
routes project-memory additions through ship tasks, and the generated skeleton
no longer invites discovery-driven additions.
* no-mistakes(review): Stop running fm-ensure-agents-md.sh on memory-file corrections
* no-mistakes(document): Clarify manual project-memory initialization and remove duplicate guidance
* feat: permit gate lifecycle calls against disposable lab homes (#5635)
* fix(bin): let gate agents drive lifecycle against marked lab homes
Part 2 of the #5615 split. A no-mistakes gate agent runs inside a
checkout carrying the fleet-captain identity, so fm-gate-refuse-lib
refuses fleet mutation on the gate signal. That refusal was absolute,
which kept gate validation from ever exercising the real lifecycle.
Stamp a disposable lab FM_HOME with a .fm-lab-home marker file that
only bin/fm-lab-home.sh writes, and only onto a fresh empty dir, so
no call path can mark a populated real home. fm_refuse_if_gate_agent
then permits lifecycle only when FM_HOME carries the marker and is
driven through its stock layout - any FM_*_OVERRIDE relocation stays
refused so part of the "lab" cannot be split back onto the real fleet.
The threat model is a confused agent touching the real fleet, not
deliberate forgery, so the marker is a plain token file rather than a
bound record. FM_GATE_REFUSE_BYPASS is unchanged: it still serves the
test harness, which cannot mark hundreds of temp homes.
Teardown's slot-ownership scan compared state-dir paths textually
while fm_firstmate_root_home canonicalizes, so a lab home under a
symlinked TMPDIR scanned its own record twice and self-collided;
compare file identity (-ef) instead.
* no-mistakes(review): Refuse unlistable lab homes and hardlinked slot records
* no-mistakes(review): Mint lab markers only on verified-empty fresh dirs
* no-mistakes(document): Clarify lab-home gate documentation and comment contracts
* no-mistakes(document): Clarify lab-home gate documentation and remove stale claims
* no-mistakes(document): Clarify gate lab-home documentation and boundary wording
* fix(bin): evict a watcher whose beacon stalls past a hard bound instead of refusing every re-arm (#5594)
* fix(bin): replace a watcher whose beacon stalls past a hard bound instead of refusing every re-arm
A fleet watcher that is alive but whose liveness beacon has gone stale could
never be replaced: every re-arm was refused because the lock holder was a live
pid, and the holder was never evicted because it was not dead. Add
FM_WATCHER_STALL_BOUND (default 3x the stale grace): below it the refusal is
unchanged; at or past it the arm re-verifies the holder against the lock's
recorded identity, sends TERM, waits boundedly, and takes the lock the normal
way, ledgering a stalled-holder-replaced row. A holder that survives TERM keeps
the old refusal.
Fixes #4400
* no-mistakes(test): poll for replacement message to fix watcher-lock test flake
* no-mistakes(document): document FM_WATCHER_STALL_BOUND in config inventory
* fix(pi): hide queued Firstmate inputs under Calm only when the session can keep them (#5563)
* fix(pi): hide queued Firstmate notifications under Calm only when the session can keep them
Calm now keeps authenticated Firstmate operational inputs out of Pi's queued-message
listing, but only after proving the live session exposes every member needed to keep
them across Escape. A session missing any of them keeps stock rows and Escape and shows
one generic warning. Escape and the dequeue key return only captain-authored messages to
the editor and re-queue hidden notifications in order; after an abort that kept any in
Pi's agent queue, the adapter starts the delivery turn itself because Pi 0.87.1 does not
continue an aborted run. Compaction-held notifications stay with Pi's compaction flush and
never start or announce a turn.
Fixes #1588
* docs(calm): record Pi 0.87.1 queued-row retention verification
* no-mistakes(review): Deliver kept Calm notifications after tree-navigation aborts too
* no-mistakes(review): Defer Calm notification turn until tree navigation finishes
* no-mistakes(lint): Silence SC2016 for literal JavaScript in queue-retention e2e test
* fix(bin): refuse teardown when a required source disappears (#5548)
* fix(bin): refuse teardown when a required source disappears
A missing sibling was sourced after cleanup had started, so Bash 3.2
exited 0 from the EXIT trap and Bash 5 continued and reported success.
* no-mistakes(review): Remove unused FM_TEST_ONLY hook from teardown tests
* no-mistakes(review): Check task backend sources before any teardown cleanup
* test(gotmp): give teardown fixtures every tmux adapter sibling
Teardown now refuses when a sibling the recorded backend's adapter sources
is missing, so the fake bin must carry fm-session-lock-lib.sh,
fm-agent-process-lib.sh and fm-gemini-lib.sh.
* docs: make supervision-host easier to read (#5605)
Restructure the supervision host doc's prose into shorter sections, lists,
and tables without changing documented behavior. Every original heading,
anchor, identifier, number, quoted string, and link target is preserved.
* docs: make the Herdr backend guide easier to read (#5606)
* docs: make herdr-backend easier to read
Restructure the Herdr backend doc's prose into shorter sections, lists, numbered procedures, and tables without changing documented behavior. Every original heading, anchor, fenced code block, link target, and documented fact is kept.
* no-mistakes(document): Restore composer-proof reason and complete Herdr topic table
* docs: make pi-supervision-branch easier to read (#5607)
Restructure the prose into sections, lists, and tables without changing
documented behavior. Every original heading and anchor, inline-code span,
link target, number, and quoted string is kept, and each sentence sits on
its own line. Adds a topic navigation table and short subsections under
the existing headings.
* docs: make watcher-continuity easier to read (#5608)
* docs: make watcher-continuity easier to read
Restructure the prose into sections, lists, and tables without changing documented behavior. Every original heading, anchor, identifier, link target, and number is kept.
* no-mistakes(review): Fix actor and supervision-host scope in watcher-continuity doc
* no-mistakes(review): Make readiness TERM and retry conditional on unready successor
* docs: make sessionstart-nudge easier to read (#5609)
* docs: make sessionstart-nudge easier to read
Restructure the prose into sections, lists, and tables without changing documented behavior. Every original heading, inline-code span, link target, number, and fact is preserved, and a harness-to-tier table now sits near the top.
* no-mistakes(review): Drop helm glossary line and dedupe exit-code lead-in
* docs: make captain-hold-lifecycle easier to read (#5610)
* docs: make captain-hold-lifecycle easier to read
Restructure the captain-hold lifecycle prose into sections, lists, and tables without changing documented behavior. Every original heading, anchor, identifier, number, quoted string, and link target is kept.
* no-mistakes(review): Fix verification record subjects and grouping headings
* no-mistakes(review): Clarify task-body read-back cases belong to the suite
* docs: make remote-secondmates easier to read (#5612)
* docs: make remote-secondmates easier to read
Restructure the remote second mates prose into sections, lists, numbered procedures, and tables without changing documented behavior. Every original heading, anchor, fenced code block, identifier, link target, and qualifier is preserved.
* no-mistakes(review): Merge remote-home table cell into one sentence
* no-mistakes(review): Tighten readiness lead-in, restore causal link, fix dangling reference
* fix(bin): bound the away digest and log why a delivery failed (#5554)
* fix(bin): bound the away digest and log why a delivery failed
The away daemon joined every buffered escalation into one unbounded
digest. A start-up catch-all span can exceed what one transport argument
carries (tmux rejects the send-keys command; Linux refuses to exec any
argument above 131,071 bytes, which is how herdr receives it), so the
initial send failed on every housekeeping pass and was logged as an
unconfirmed Enter with text possibly in the composer.
escalate_flush now builds the injected digest under a fixed byte budget:
each event is cut at a UTF-8 boundary with an omitted-bytes marker, the
joined events stop with a "+K more event(s)" tail, and a bounded digest
names a state/.subsuper-digests/ file that keeps every buffered event
verbatim. The buffer itself is untouched, so the return catch-up stays
complete.
The tmux submit core and the herdr literal send now replay the
transport's stderr on failure, and inject_msg logs the failing stage
(initial send versus Enter confirmation) with the byte count and that
stderr. The wedge alarm line and marker carry the last failure reason.
Fixes #4382
* no-mistakes(review): Drop digest pruning; label send-failed as send-or-Enter stage
* no-mistakes(review): Keep digest full text once submit ran; reuse on retry
* no-mistakes(lint): Count digest files with find instead of ls
---------
Co-authored-by: firstmate-oss <firstmate@kunchenguid.local>
* test: add a gated harness seam and stabilize lifecycle fixtures (#5638)
* feat(tests): add FM_TEST_SEAM launch seam and gate lab-primary recipe
Part 1 of the #5615 split: the pieces that let the no-mistakes pipeline
live-validate firstmate changes, without the gate-refusal rescoping.
- bin/fm-afk-launch.sh: FM_TEST_HARNESS pins the detected harness only
alongside the FM_TEST_SEAM=1 marker test suites set, so a leaked variable
in a real primary's environment stays inert and unknown tokens fall
through to real detection.
- tests/lib.sh: export FM_TEST_SEAM=1 for every suite.
- .no-mistakes.yaml: per-harness recipe for running a real fixture primary
from a gate run - a plain mktemp lab FM_HOME on a private tmux socket,
with FM_GATE_REFUSE_BYPASS=1 scoped to it and NO_MISTAKES_GATE scrubbed.
- tests/fm-wake-queue.test.sh: stop the owned watcher fixture with KILL and
clear its lifecycle state so the next leg starts clean; TERM could leave
bash waiting in a child on some runners.
- tests/fm-remote-secondmate-lifecycle-e2e.test.sh: wait for the liveness
lock holder's post-acquire marker instead of the lock dir, which is
published before the claim finishes.
* no-mistakes(review): Scrub lab home overrides and require FM_TEST_SEAM separately
* no-mistakes(document): Clarify test seam and disposable lab bypass documentation
* no-mistakes(document): Clarify lab isolation and test-seam documentation
* no-mistakes(ci): Fixed the CI failure: test cleanup killed the remote worker child but left its supervisor able to restart it during fixture removal. Cleanup now stops the worker tree. The lifecycle test passed locally; ShellCheck and diff checks passed
* fix(bin): refuse watchers from disposable checkouts and exit when the home is gone (#5552)
* fix(bin): refuse watchers from disposable checkouts and exit when the home is gone
Fixes #321
Fixes #4760
A watcher armed from a disposable no-mistakes validation checkout under
.no-mistakes/worktrees/ outlived the validation step and kept writing the
real home's state, and a running watcher never noticed when its home,
state directory, or code root disappeared. The arm now refuses from such
a checkout with the typed failure line, the watcher checks once per poll
that its home, state directory (or its own lock holder record), and bin
directory still exist and exits with a logged reason scoped to itself,
and the shared test helpers reap every watcher a suite armed for a
temporary home through the home-scoped stop.
* no-mistakes(lint): fix SC1007 by assigning empty string in watch-arm test
* no-mistakes(ci): Found and fixed a genuine, reproducible hang introduced by this branch's test-watcher reaper, which is what killed both CI checks (serial-2 cancelled at the 30-min cap; Lint 2 exit 143 = the suite's own TERM-trap code). Root cause: test_drain_asserts_watcher_liveness (tests/fm-wake-queue.test.sh) fabricates a .watch.lock whose pid is the test runner's own $$ with the runner's real identity, to make the drain believe a live watcher exists. The new make_case tracking registers that state dir for reaping, so at fm_test_cleanup the new fm_test_reap_watchers drives fm-watch-arm.sh --stop; its identity check matches (the fixture recorded the runner's identity) and it kill -TERMs the test runner. tests/lib.sh:231 is `trap 'fm_test_cleanup; exit 143' TERM`, so the TERM re-enters cleanup -> reap -> kills $$ again -> infinite loop until the runner cap. I reproduced this locally: the suite ran all tests then looped forever in cleanup spawning fm-watch-arm.sh --stop against a lock naming its own PID. Fix (tests/lib.sh, +5 lines): in fm_test_reap_watchers, skip any tracked lock whose pid equals our own $$ before driving --stop. This is the single shared reap boundary; seven $$-self-lock fixtures across four test files are all covered by the one guard, and real armed watchers (pid != $$) are still reaped. Invariant: the test reaper must only signal real armed watcher processes, never the test runner itself. Verified locally: tests/fm-wake-queue.test.sh -> EXIT 0 (63 ok, no hang); tests/fm-watch-arm.test.sh -> EXIT 0 (21 ok, including test_reaper_stops_a_tracked_watcher, confirming the guard does not over-skip). Lint 2's exit 143 was the same shard/cap signature; a fresh CI run on this new commit will re-evaluate it
---------
Co-authored-by: firstmate-oss <firstmate@kunchenguid.local>
* fix(bin): surface unrecognized status prefixes instead of reading them as silence (#5588)
* fix(bin): surface an unrecognized status prefix instead of dropping it
A parked or holding declaration, and a verb whose correlation token did not parse, never became an event, so the supervisor still saw the earlier line.
* no-mistakes(review): Require verb-shaped unrecognized status prefixes, add continuation tests
* no-mistakes(document): Document unrecognized status prefix escalation in afk skill
* no-mistakes(ci): I reproduced the "Behavior portable serial 6" failure locally and fixed it by changing the test data in one test. No product code changed. **What failed:** `tests/fm-session-start.test.sh`, in `test_orphan_status_logs_are_printed`, with "matched status log was printed 2 times". **Why:** the test writes status lines with made-up prefixes, `matched: surfaced once` and `orphan: step N`. The test only uses them as placeholder text. It checks that the session-start digest prints each task's status tail exactly once. This PR (#4763) deliberately makes an unrecognized one-word lowercase prefix a status event. So those lines now surface as captain-relevant events, and the wake queue's STATUS OUTCOME BACKSTOP section prints them a second time. The code under review is behaving as the issue asks. Only the test's placeholder data had become meaningful. **Rule the test depends on:** its status lines must not be captain-relevant, so the digest is the only place they are printed. Both lines in this test broke that rule. The orphan line would have failed the same count check right after the matched line did. **Fix:** in that test only, I switched both lines to the recognized, non-captain verb `working:`: `working: surfaced once` and `working: orphan step 1..6`. I updated the matching assertions and counts to use the new text. What the test checks is unchanged: orphan logs are labelled, the tail is bounded, the log path is printed, and each tail appears once. **Verification:** before the fix, the test failed locally the same way as in CI. After it, `bash tests/fm-session-start.test.sh` reports "all assertions passed
* no-mistakes(review): Detect unrecognized prefixes on unstamped lines; share verb list
---------
Co-authored-by: Kun's firstmate <kunchenguid+firstmate@users.noreply.github.com>
* fix(bin): report the failing item when remote inheritance fails (#5658)
Fixes #5295
Session start now reports a remote inheritance failure using the
push's own error line instead of the first unchanged item that
happened to print before it, and the shared captain preferences
header check now names the first required phrase it did not find,
on both the local and remote inheritance paths.
* fix(bin): stand down the Claude Stop auto-arm on pi-code-delivered payloads (#5657)
* fix(bin): stand down the Claude Stop auto-arm on pi-code-delivered payloads
pi-code loads the tracked Claude settings but has no asyncRewake, so it
awaits every Stop hook; without a stand-down the auto-arm runs
synchronously inside Pi's turn end and holds it open for the declared
multi-hour timeout. Stand down when the payload's transcript_path
contains a /.pi/ path component, the same discriminator the closed-but-
unmerged fix in #3352 used, with an explicit string-type check on the
jq filter.
Fixes #3343
* no-mistakes(document): document pi-code stand-down in harness integrations reference
* fix(bin): match whole multi-word project names in the registry lookup (#5659)
* fix(bin): match whole multi-word project names in the registry lookup
bin/fm-project-mode.sh matched a registered project name against only the
first whitespace-delimited token of a registry row, so a name containing a
space never matched, silently defaulting the project to no-mistakes off
instead of its declared posture.
The lookup now matches the whole registered name against the raw line text,
so a name is compared literally (never as a regex) and a name that is a
leading prefix of another registered name still resolves to its own row.
* no-mistakes(document): docs already accurate for multiword registry name match
* chore: drop accidental empty err file
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
* fix(bin): classify shell stdin payloads after -s operands (#5546)
* fix(bin): classify the stdin program of `bash -s` with operands in the arm policy
With -s, sh/bash/zsh read the program from stdin even when operands follow;
the operands are only positional parameters. The arm policy treated the first
operand as a script path, so heredoc and here-string payloads were never
classified and a hidden bin/fm-watch.sh execution was allowed.
A protected path in the operand position still fails closed as before.
Fixes #1489
* no-mistakes(document): Clarify stdin shell operand documentation
* no-mistakes(ci): Captain, fixed `shellInvocation` so `bash -- -s` treats `-s` as a script name, and updated R21 to test the exact command. The targeted policy suite, lint, documentation check, and diff check pass. Both hosted workflows show `action_required` before any jobs ran; that external approval state remains unresolved
* fix(bin): keep main's handling of words after a leading `--`
Revert the pipeline CI-step change that made the first word after a leading
`--` always a script. It turned forms that main denies today into allow
(for example `bash -- -c 'bin/fm-watch.sh'`), which is outside #1489 and
loosens a fail-closed policy. `--` after `-s` still ends option parsing.
* fix(bin): strip AI co-author trailers from fleet-launched commits (#5695)
* fix(bin): strip AI co-author trailers from fleet-launched commits
Cursor and other non-Claude runtimes append the trailer after the typed
message. A per-task commit-msg hook removes it and leaves human co-authors
and the author identity untouched.
* no-mistakes(review): Export pane hooksPath override and drop generated-with stripping
* no-mistakes(ci): This PR caused all three CI failures, and the fix is test-only: 4 test files change, no product code. **Cause.** `fm-spawn.sh` now installs the AI-trailer strip hooks for every spawn, secondmates included. The installer refuses a worktree that is not a git repository, and the PR deliberately keeps that fail-closed rule because real secondmate homes are firstmate clones. Four test fixtures still gave secondmates a plain directory as their home, so each spawn failed with "not a git worktree ... could not install the AI-trailer strip hooks": - serial 5: `tests/fm-backlog-atomicity.test.sh` ("secondmate spawn failed"). - serial 8: `tests/fm-secondmate-harness.test.sh` ("split: no meta written"). - Herdr: `tests/fm-backend-herdr-launcher-workspace-e2e.test.sh` and `tests/fm-backend-herdr-workspace-per-home-e2e.test.sh`. **Rule that must hold.** Every home a test spawns as a secondmate must be a git worktree. I checked the other places in the changed area: the only secondmate spawns in these tests are the ones listed. The earlier rounds already fixed the other fixtures (`fm-secondmate-liveness`, `fm-secondmate-safety`) the same way. **Fix.** - Each of those four secondmate homes now gets the same `.gitignore` plus `git init -q -b main` that the liveness and safety tests already use. - The two Herdr tests clean up with their own plain `rm -rf "$TMP_ROOT"`, not the shared `tests/lib.sh` helper. Because the installer leaves each `state/<id>.git-hooks` directory read-only, that cleanup printed "Permission denied" and left the directories behind. Both cleanups now restore the owner's write bit on every directory before removing (`find ... -exec chmod u+rwx`), which is what `fm_test_remove_tree` in `tests/lib.sh` does. **Verification.** - `tests/fm-secondmate-harness.test.sh` passes. - `tests/fm-backlog-atomicity.test.sh` passes (99 ok, exit 0). - shellcheck is clean on all four files. - I could not run the two real-Herdr tests locally: the Herdr lab on this host refuses to start because it needs exactly one running default session, and I did not change the host's Herdr state to get around that. Instead I checked their two changed steps directly: the installer succeeds on a home set up the new way, and the new cleanup removes the read-only hooks directory completely. Those two tests will only be proven on CI
* fix(bin): treat Pi's dollar-first cost footer as furniture (#5683)
* fix(bin): treat Pi's dollar-first cost footer as furniture
An idle Pi status row opening with $0.000 was read as a dead-shell prompt, so exit and relaunch refused on an empty composer.
* test: wait for the draining holder to exec sleep before reading its identity
The procevent drain fixture read fm_pid_identity immediately after
backgrounding setsid sleep, racing the child's exec chain. Mid-exec the
cmdline can read empty, failing the fixture on a loaded CI runner.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(bin): refuse merges with unreported required checks (#5534)
* fix(bin): refuse a merge when a required check never reported
fm-pr-merge.sh built its GitHub refusals only from checks present in
statusCheckRollup, so a required check that never ran was simply absent and
the merge proceeded on the subset that reported, contradicting its own
"every required check green" claim.
The GitHub verify now reads the base branch's required contexts from the
forge itself - the classic branch protection summary on
GET repos/{o}/{r}/branches/{b} and the active ruleset rules on
GET repos/{o}/{r}/rules/branches/{b} - and refuses when a required context
has no entry in the same rollup, at the same head, that the merge is bound
to. Absence reads as unknown, never green. The required-set read joins the
existing refusal list, so a draft, a red check, and an unreported required
check are all reported together.
Could not read vs nothing required: both endpoints need only repository
read access. The admin-only GET .../branches/{b}/protection endpoint is
deliberately not used: it answers a non-admin token with the same 404 an
unprotected branch gets (observed live on kunchenguid/firstmate main with
this token), which would read a missing permission as "nothing required".
Any failed or malformed read of either source (auth, missing fine-grained
permission, rate limit, network, 404, unexpected shape) refuses the merge
with a line naming the unreadable source. The one exception is GitHub's
plan-gated 403 on the rules endpoint ("Upgrade to GitHub Pro or make this
repository public"), which already means "this repository has no branch
rules" for the merge-queue reader; that check moves into one shared helper
and the classic summary still decides for such a repository.
Attended waiver: --allow-missing <check-name> is the twin of --allow-red and
follows the same design and recording path: once, separate name argument,
waives only that exact unreported required check, still requires every
other required check reported and every check green, never waives an
unreadable required set, refused while the away-posture record exists, and
refused on GitLab. Merge-state BLOCKED policy is unchanged.
How this differs from the withdrawn #5353 (read from its diff):
- #5353 read the admin-only branches/{b}/protection endpoint and treated
its 404 as "no required checks", so for any non-admin token the required
set silently read as empty; this change reads the read-access branch
summary and treats every failure as unreadable.
- #5353 ignored rulesets; this change also reads required_status_checks
rules from the effective branch rules.
- #5353 made separate per-head REST reads of statuses and check-runs capped
at per_page=100 with no pagination; this change checks presence in the
same statusCheckRollup view the red-check gate already reads at the
verified head.
- #5353 stopped at the first unreadable read; this change reports it as one
refusal among all the others.
- #5353 also claimed #5345 (lock stealing) and changed 39 files, most
unrelated; this change is #5344 only.
Live proof, read-only (a gh wrapper refused every merge and mutating call):
- cli/cli#14474 (trunk requires 3 classic build contexts, none ran):
refused, naming build (macos-latest), build (ubuntu-latest),
build (windows-latest); with --allow-missing "build (macos-latest)" it
still refused, naming the other two.
- cli/cli#13665 (ran build (ubuntu-24.04-firewall) instead): refused,
naming build (ubuntu-latest).
- hashicorp/terraform#39262 (ruleset-required checks absent): refused,
naming Code Consistency Checks, End-to-end Tests, Race Tests, Unit Tests.
- cli/cli#14485 (all required reported and green): verified; the wrapper
blocked the merge call and the pull request read back open.
Fixes #5344
* fix(review): Preserve required-check producers and aggregate independent read failures
* fix(document): Clarify required-check verification and waiver documentation
* fix(bin): match an app-bound required commit status by name
The producer-identity check resolved an app-bound required context only
against check runs, so a required context that the required app reports as
a commit status could never match and always read as "has not reported".
A commit status carries no app id to compare, so an app-bound requirement
that arrives as a status now matches by name, as before producer binding;
check runs keep requiring the configured producer app.
Live, read-only: hashicorp/terraform#39262 requires license/cla from
integration 865473, reported green as a commit status by the CLA app. The
previous head refused it as unreported; this head no longer does, while
still naming the four required check runs that never ran there.
Refs #5344
* fix(document): Clarify accepted commit-status producer verification limitation
---------
Co-authored-by: firstmate-oss <firstmate-oss@kunchenguid.local>
* fix: keep persistent secondmates out of landed-work cleanup (#5696)
* fix(bin): never offer a persistent secondmate for teardown
The return brief's "Landed, cleanup due" scan listed every state/*.meta
record carrying a pr= and a merge-notified marker without regard to kind, so
a secondmate record holding a relayed child's merged PR put the mate itself
up for "bin/fm-teardown.sh <mate>" cleanup. A secondmate is a persistent
worker, never landed work.
- bin/fm-afk-return.sh: skip kind=secondmate in the landed-cleanup scan.
- bin/fm-pr-check.sh: refuse to record pr= or arm a merge watch on a
kind=secondmate record before any side effect; a PR reported on its routed
status channel belongs to a task in the mate's own home, which arms its
own watch.
- bin/fm-watch.sh: a merged result from a poll already armed on a secondmate
retires the poll silently - no merge outcome, marker, or wake.
* no-mistakes(document): Document secondmate merge-watch and return-brief exclusions
* no-mistakes(ci): CI failed in an unchanged watcher-shutdown test whose three-second wait was sensitive to runner load. Increased the wait for both state- and home-deletion cases without changing watcher behavior. The full fm-watch-arm suite passed locally; syntax and diff checks passed
* fix: reject invalid X reply and follow-up arguments before posting (#5702)
* fix(bin): refuse unknown dash-leading args in public-posting fm-x scripts
fm-x-reply.sh collected any unrecognized argument into the positional
pool and took the first one as the reply text, so an invocation like
"fm-x-reply.sh <id> --followup --final <text>" posted the literal
string "--final" to X and silently dropped the real text.
Make argument parsing strict in every script that can post publicly:
an unknown dash-leading argument, a dash-leading request_id/task id, a
dash-leading option value, or a surplus positional now exits 2 with a
usage error before any config load, outbox write, or network call. Reply
text starting with '-' is still accepted via --text-file or stdin, and
--help is honored wherever it appears instead of becoming text (a --help
forwarded through fm-x-followup.sh would have counted as a posted
follow-up and mutated the link).
fm-x-link.sh and the fm-public-followup scripts already refuse unknown
arguments; fm-x-poll.sh takes none.
* no-mistakes(review): Refuse surplus follow-up text sources; drop post-ID help branches
* no-mistakes(document): Clarify reply and follow-up argument usage
* no-mistakes(review): Refuse dash-leading --text-file operands in fm-x-reply
* no-mistakes(document): Correct follow-up argument parsing comment
* no-mistakes(document): Document dismiss argument rejection in script header
* fix: pause broken supervision-host sessions between engine probes (#5701)
* feat(bin): latch the supervision host after repeated engine errors
Rung 3c-1 of the PR 5631 re-cut: the host copies the Pi branch's
broken-session policy. Two consecutive engine errors latch the session;
every away wake then reaches main with one supervision-host line for a
five-minute cooldown, after which one wake probes the engine, and each
failed probe doubles the cooldown up to one hour. A reported turn without
an engine error clears it. The latch is kept per main session, engine, and
model in state/.supervision-host-health, and the engine conversation now
uses the same main-session key, which includes the lock holder's process
identity so a recycled pid never shares either.
Lifted from the validated 5631 tree and adapted to main's away-only host:
the attended recovery line and attended cooldown pass-through are left for
the attended core, so a recovery is only logged.
* no-mistakes(document): Consolidate supervision-host latch documentation
* fix(bin): absorb routine secondmate working and paused status appends (#5535)
* fix(bin): absorb routine second-mate progress while surfacing routed replies
Fixes #2959
A kind=secondmate task's status signal was never absorbable, so a healthy
mate's routine working: and paused: appends woke the primary every time.
signal_crew_provably_working now reads the mate's lines new since the
watcher's classified position: a decision, blocker, terminal outcome, note:,
correlation-marked line, or unknown verb still surfaces regardless of busy
evidence, while unmarked working:, paused:, and resolved: fall through to the
same provably-working absorb an ordinary crewmate gets.
* no-mistakes(review): narrow secondmate routine absorb to working and paused
* fix(bin): use gh-axi for the ship DoD draft check (#5519)
* fix(bin): use gh-axi for the ship DoD draft check
* no-mistakes(review): use PR number not URL in gh-axi draft check
* fix(bin): bind inactive-outcome receipt identity to structured fields only (#5520)
* fix(bin): drop status prose from the inactive-outcome dedupe identity
The inactive-outcome receipt fingerprint included the child's sanitized
last status line, so a persistent child appending routine prose after one
terminal outcome minted a fresh parent event per sentence. Bind the
identity to incarnation, task id, terminal state, and PR only, keeping
the last line in the record as status_head evidence.
Fixes #2960
* no-mistakes(document): note structured-only inactive receipt identity in regression coverage
* feat: record the Claude and Cursor dialog for the supervision host (#5707)
* feat(bin): record the supervision host's dialog mirror on Claude and Cursor
Add bin/fm-host-mirror.sh, the one owner of the supervision host's dialog
mirror file, cursor, lock, and feed, plus the main-session key it keys
entries to. The tracked Claude UserPromptSubmit and Stop hooks and the
Cursor beforeSubmitPrompt and afterAgentResponse hooks record the captain's
prompt and main's reply, only on a home with config/supervision-host, from
a genuine primary checkout, for the lock-owning session. The mirror lands
inert: writers record and nothing reads it yet; attended supervision on the
host is the later step that consumes the feed.
Codex, Grok, OpenCode, and omp have no writer here.
* no-mistakes(review): Scope mirror dedup to session, atomic appends, marker-inclusive caps
* no-mistakes(document): Clarify dialog mirror scope and remove duplicate contract details
* no-mistakes(document): Correct Cursor hook documentation for dialog mirror registration
* no-mistakes(review): Pass mirrored dialog text to jq via stdin
* no-mistakes(document): Clarify dialog mirror documentation and remove duplicate claims
* no-mistakes(review): Preserve internal dialog whitespace; drop mirror check and verified modes
* no-mistakes(review): Drop only identical mirror repeats; remove redundant chmod guard
* fix(bin): retire check-row receipts on branch acknowledgement so away escalations are not repeated (#5731)
* fix(bin): retire check-row receipts on branch acks and report an unchanged situation once
* fix(bin): scope a branch acknowledgement's check-row receipt retirement to
its granted sequences
The away posture lifts the attended partition's check/decision exclusions, so
a branch grant can name check-kind rows - but the branch-actor ack still
assumed check rows were main-only and skipped every receipt scan. The queue
row was consumed while its terminal-outcome .pending receipt stayed behind,
and each inactive-reconcile cadence scan re-queued the same fingerprint. In
the first real away window on the supervision host that re-escalated one
unchanged held-PR situation on every cycle (~1,734 of 4,149 outcomes).
A branch ack now scans inactive-outcome and inactive-reconcile receipts and
commits secondmate stall receipts against exactly the sequences in its
eligible-row snapshot - the same rows it consumes - instead of none. Attended
grants still name no check row, so the scans find nothing.
* fix(bin): store a repeated captain verdict as routine while the task's
durable situation is provably unchanged
fm-branch-outcome.sh append computes a mechanical situation key per captain
row - metadata bytes, captured status-log endpoint and identity, live
crew-state verb, worktree head - and anchors it in
state/.<task>.branch-captain-key. A later captain verdict whose recomputed key
matches is stored as routine with "unchanged since seq <N>:" prefixed to its
summary, so one situation escalates once until something provably changes. A
task with no readable status ledger is never demoted, an unreadable record
fails toward reporting, and teardown removes the sidecar with the task's
other branch records. The append-only store schema is unchanged.
This covers both hosts: the Pi supervision branch and the supervision host
both funnel reports through append.
* docs: check rows are main-owned only while attended; the away posture grants
them to the branch, whose ack retires their receipts exactly
* test: the away-flood reproduction as a regression test (branch ack retires
the receipt and later scans stay quiet), store-level dedupe coverage, and a
branch-ack secondmate stall receipt case
* fix(bin): restore the secondmate child devin-config cleanup path
The branch-captain-key sidecar addition mistyped the sibling entry as
.$child_id.devin-config.json, so a forced secondmate teardown would have
stopped removing each child's real <id>.devin-config.json. Restore the
original path and add a behavioral test that stops the child sweep mid-loop
on a refused close, proving the cleaned child's devin config and captain
anchor are both removed while the unconsumed child's records are retained.
* no-mistakes(review): Key captain dedupe on the covered wake rows' fingerprint
* no-mistakes(review): Drop captain-key demotion; prove one escalation on both surfaces
* no-mistakes(review): Drop unrelated teardown test; cite both receipt test files
* no-mistakes(document): Docs already match branch-ack check-receipt retirement
* fix(bin): deliver Claude-bound operational input as a record-backed doorbell (#5664)
* fix(calm): deliver Claude-bound operational input as a record-backed doorbell
Claude Code 2.1.280 removes U+2063 from every submitted prompt, so a typed
operational envelope reaches a Claude Code primary as plain text. The away
daemon now writes the envelope to a record under state/operational-inbox and
types only a plain doorbell naming it; the /afk return check and the Calm mod
recognize the doorbell only when that record holds a current envelope. Marker-
preserving harnesses keep the typed envelope. The live Calm guard accepts the
2.1.280 module-load log line, drives the doorbell, and asserts thinking stays
hidden.
* no-mistakes(review): Fix operational record retention at 7 days and document prune limit
* no-mistakes(document): Point Calm bounds at 2.1.280 evidence; fix afk-exit comment
* no-mistakes(lint): Pick newest Calm e2e transcript without parsing ls
* docs(calm): add a minimal turning-Calm-on step for Claude Code
* fix(spawn): deliver the Claude launch brief as a record-backed doorbell
Claude Code strips U+2063 from the launch-prompt argument too, so a
worker's launch brief arrived with its operational marker removed.
Publish the brief as a record in the receiving home's operational
inbox - a secondmate's own state, not the primary's - and pass only
the printable doorbell naming it, falling back to the typed envelope
when the record cannot be published so the brief body still delivers.
Unwrap doorbell-carried digests in the daemon digest tests that still
read the raw send log under the claude pin, and update the documented
bounds now that launch briefs hide like the other operational rows.
* test(spawn): cover a secondmate's launch-brief record landing in its own home
The record-backed doorbell resolves its state through the receiving
pane's home, so prove a claude secondmate launch publishes into the
seeded secondmate's operational inbox and never leaks a record into
the primary's.
* no-mistakes(review): Pass primary harness to daemon, tighten retention, refresh verdicts
* no-mistakes(review): Prune operational records by exact seven-day elapsed age
* no-mistakes(review): Batch record pruning so large inboxes still expire
* no-mistakes(review): Refuse Claude spawn when brief record cannot publish
* no-mistakes(review): Drop thinking probe from Claude Calm live test and docs
* no-mistakes(review): Record dated Claude Code 2.1.282 reproduction evidence
* no-mistakes(document): Clarify operational doorbell documentation and record expiry
* no-mistakes(document): Correct AFK escalation carrier guidance
* no-mistakes(review): Describe operational record retention as about seven days
* no-mistakes(document): Clarify Calm delivery and operational record retention
* no-mistakes(review): Remove out-of-scope Calm launch guide from Claude docs
* no-mistakes(document): Document Claude launch-brief delivery and refusal
* no-mistakes(document): Correct stale operational-input documentation
* no-mistakes(ci): Fixed the stale Claude trust test to verify that worker and secondmate launches deliver readable, record-backed briefs instead of expecting brief paths in their commands. Annotated the daemon’s output variable for ShellCheck without changing behavior. The affected tests, daemon tests, ShellCheck, and diff check pass locally
* no-mistakes(ci): parse rebased Claude launch after trailer hook prefix
* no-mistakes(review): Trust launch-brief record and restore thinking bound doc
* no-mistakes(review): Parse final Claude launch statement; drop Stop-hook docs
---------
Co-authored-by: Mike Sewell <maikunari@protonmail.com>
Co-authored-by: no-mistakes <no-mistakes@localhost>
* test: isolate lint fixture from tracked suite (#5727)
* fix(bin): republish parent metadata after a remote secondmate relaunch (#5583)
A host-local relaunch rewrote only the far endpoint, so this home kept the old harness, model, and effort, and appending those keys after pr= broke pull-request poll authentication.
* fix(bin): give slow watcher suites headroom under the changed-suite bound (#5516)
tests/fm-watch-triage.test.sh finishes in about 434s alone and about 698s
under CI load, so the 900s bound the changed-suite runner applies produced
a false timeout under ordinary concurrent validation. Raise the automatic
bound to 1500s, which keeps every measured script under it while staying
below the 30-minute normal CI tier so a genuinely hung script still fails
here with its output before the job cap cancels the lane.
Fixes #3869
Refs #3565
* fix(bin): stop nested steal-lock recursion and mid-steal watcher TERM (#5728)
* Fix nested watcher lock reclaim
* no-mistakes(review): Elect a single steal-mutex reaper and bound arm TERM wait
* no-mistakes(review): Reclaim self-held steal mutex and unify autoarm steal reaping
* no-mistakes(review): Resume own interrupted steal reap from its tombstone
* test: make supervision-host park-boundary tests deterministic (#5710)
* test: hold the back-to-back boundary close on the host's own clock
test_park_boundary_holds_under_back_to_back_closes assumed two engine
turns fit in the ~16s pre-refusal window and that the stub finished a
turn in 3s. Under load the stub's real drain, report, and
acknowledgement take ~13s, so the turn either died at its bound (which
hands the wake to main, no boundary line) or the second close landed
past the window and the fixture failed while the boundary held. 3
failures in 5 runs at a load average near 11.
Hold the first turn on a release file instead: once the engine is in
flight, a second close is appended mid-turn and the turn is released as
the refusal window opens (park bound minus turn bound and grace, read
off the host's own start record). The queued close can then only wait
for the boundary on any machine speed, which is what the test asserts:
the boundary line ends the output, the demo.status row stays queued for
main, and no second engine turn ever starts. A host too loaded to start
the turn at all hands the first close to the same boundary exit.
After: 12/12 at load ~15-42.
* no-mistakes(review): Print boundary test deadline as a decimal integer
* no-mistakes(review): Hold boundary test turn on a FIFO, require full sequence
* no-mistakes(review): Remove stray before/after supervision-host test copies
* test: hold the late close's render until the refusal window opens
The boundary recheck test's node shim slept a fixed 10s, which assumed
the first close was read before the host's refusal window opened. Under
load the close arrived after the refusal check, so the host correctly
refused it before the successor started and the render snapshot never
appeared. Block the wake-prompt render on a FIFO released at the
refusal-open instant read from the host's own start record, so the
pre-turn recheck must refuse on any machine speed.
* no-mistakes(review): Derive minimal park bounds and refresh supervision-host shard hint
* no-mistakes(review): Drive park-boundary tests from a seam-gated host test clock
* fix: stage remote home clones before publication (#5733)
* fix(bin): stage remote home clones before publishing them
A remote home provision cloned the code root directly into the public
FM_HOME path while rollback() claimed rm -rf of that same path on any
failure. Bash defers trapped signals past a foreground child, but any
other cleanup or lifecycle path that removes the home directory races
the live clone's object copy, producing the CI flake "fatal: failed to
copy file to .../.git/objects/...: No such file or directory".
Clone into a private staging directory beside the home and publish with
an atomic rename once complete, so no cleanup can remove a directory a
live clone is still writing; a home that appears mid-provision now dies
cleanly instead of inheriting torn state. The regression coverage holds
a real clone mid-copy, removes the public path, and requires the
provision to finish and publish intact.
* no-mistakes(review): Prove home ownership by sentinel and hold only a live clone
* no-mistakes(review): Assert raced provision publishes a complete, intact clone
* no-mistakes(document): Document remote home staging and publication safety
* no-mistakes(lint): Fix ShellCheck warning in clone integrity assertion
* no-mistakes(document): Clarify remote home publication and rollback guarantees
* fix: preserve Herdr status on Pi relaunch (#5161)
* fix(control): keep a relaunched Pi worker's herdr pane status authority alive
Defect: after `bin/fm-control.sh <id> relaunch` (observed live on a herdr
Pi crewmate whose pane read idle while it ran its validation pipeline),
the pane froze at whatever its previous agent had last reported.
Cause, measured on herdr 0.9.1 against a real Pi: a pane has one status
authority, and for Pi with its integration installed that authority is
the lifecycle hooks, so herdr also skips screen detection for the pane.
In the crew shape the registration outlives its agent process (upstream
issue #4115; docs/herdr-backend.md "Restart and liveness behavior"), and
herdr applies only reports carrying the session identity it bound. A
replacement started fresh in that pane reports a NEW session, so its
state reports are ignored and the pane stays frozen. Nothing from
outside repairs it: `pane report-agent-session` and `pane report-agent`
for `herdr:pi` are accepted (rc=0) without being applied unless the
reporter is the registered pane agent, and `pane release-agent` on the
stale record changes nothing.
Fix: a relaunch preserves the binding instead of fighting it. The launch
owner reads the session reference the endpoint's own runtime recorded
(`fm_backend_herdr_pane_agent_session_ref`) and passes it back as Pi's
own `--session <path-or-id>` (`relaunch_resume_args`;
`fm_control_relaunch_resume_flag` owns which adapters and which
registered-agent labels qualify). That is the same reference herdr
itself resumes Pi panes with after a server restart, and the resumed
session's reports land again, which the live check confirmed: the pane
returned to working while the replacement worked and idle when it
settled, on the same session identity.
Safety: relaunch-only (a fresh spawn binds nothing), herdr-only (the one
adapter that records a per-pane session), Pi-family only, and only when
the registration's own agent label matches - so no other adapter's
conversation can be handed to a Pi launch. An unreadable, missing, or
malformed reference degrades to exactly the fresh-session launch that
existed before. No lifecycle, liveness, isolation, or merge guard is
touched, and an empty result leaves every non-Pi launch byte-identical.
`resume` remains a refused verb; docs/agent-control.md and the
harness-adapters references are corrected where they claimed Pi had no
verified resume form at all.
* no-mistakes(document): docs: correct relaunch session-authority ownership and skill paths
* no-mistakes(document): docs: correct stale control-plane ownership claim
* no-mistakes(document): docs: drop unverified Herdr restart resume claim
* no-mistakes(test): Added offline Herdr Pi session-authority relaunch coverage
* no-mistakes(document): Document Herdr Pi relaunch session continuity
* no-mistakes(ci): The failing remote relaunch test tried to arm a PR poll for a secondmate, which `fm-pr-check.sh` correctly refuses. Removed that invalid test scenario; the remaining remote relaunch tests pass, and `git diff --check` is clean
* Seed the relaunch-ordering PR poll fixture without fm-pr-check.sh (#5758)
Main has been red since fm-pr-check.sh began refusing to arm a merge poll
on a kind=secondmate record (#5696): the relaunch-ordering case in
tests/fm-remote-secondmate-relaunch.test.sh armed its fixture through that
entry point and could no longer be set up.
The ordering guarantee still matters: a secondmate record armed before the
refusal can legitimately carry a trailing pr=/pr_head= identity block until
the watcher retires it, and fm-remote-secondmate-relaunch.sh must still keep
that block last when republishing harness/model/effort. Seed the fixture the
way such a record was really written - pr= appended last to the meta, then
the poll artifacts published through…
* docs: make configuration settings easier to find and understand (#5589)
* docs: make configuration settings easier to find and understand
* no-mistakes(review): Restore dropped qualifiers and fix misplaced config doc labels
* no-mistakes(review): Restore three dropped qualifiers in configuration reference
* fix: limit project memory edits to factual corrections (#5636)
* fix: bound worker edits of project AGENTS.md/CLAUDE.md to factual corrections
These files are loaded into every agent session of a project, so additions
should be a deliberate human choice rather than automated task output. The
ship brief's project-memory section and AGENTS.md section 6 previously invited
workers to record durable knowledge, which let project AGENTS.md files accrete
detail the codebase or README already carries. Workers now edit only to fix
factually wrong content - including content their own change made wrong - and
fm-ensure-agents-md.sh runs only alongside such a correction. Stow no longer
routes project-memory additions through ship tasks, and the generated skeleton
no longer invites discovery-driven additions.
* no-mistakes(review): Stop running fm-ensure-agents-md.sh on memory-file corrections
* no-mistakes(document): Clarify manual project-memory initialization and remove duplicate guidance
* feat: permit gate lifecycle calls against disposable lab homes (#5635)
* fix(bin): let gate agents drive lifecycle against marked lab homes
Part 2 of the #5615 split. A no-mistakes gate agent runs inside a
checkout carrying the fleet-captain identity, so fm-gate-refuse-lib
refuses fleet mutation on the gate signal. That refusal was absolute,
which kept gate validation from ever exercising the real lifecycle.
Stamp a disposable lab FM_HOME with a .fm-lab-home marker file that
only bin/fm-lab-home.sh writes, and only onto a fresh empty dir, so
no call path can mark a populated real home. fm_refuse_if_gate_agent
then permits lifecycle only when FM_HOME carries the marker and is
driven through its stock layout - any FM_*_OVERRIDE relocation stays
refused so part of the "lab" cannot be split back onto the real fleet.
The threat model is a confused agent touching the real fleet, not
deliberate forgery, so the marker is a plain token file rather than a
bound record. FM_GATE_REFUSE_BYPASS is unchanged: it still serves the
test harness, which cannot mark hundreds of temp homes.
Teardown's slot-ownership scan compared state-dir paths textually
while fm_firstmate_root_home canonicalizes, so a lab home under a
symlinked TMPDIR scanned its own record twice and self-collided;
compare file identity (-ef) instead.
* no-mistakes(review): Refuse unlistable lab homes and hardlinked slot records
* no-mistakes(review): Mint lab markers only on verified-empty fresh dirs
* no-mistakes(document): Clarify lab-home gate documentation and comment contracts
* no-mistakes(document): Clarify lab-home gate documentation and remove stale claims
* no-mistakes(document): Clarify gate lab-home documentation and boundary wording
* fix(bin): evict a watcher whose beacon stalls past a hard bound instead of refusing every re-arm (#5594)
* fix(bin): replace a watcher whose beacon stalls past a hard bound instead of refusing every re-arm
A fleet watcher that is alive but whose liveness beacon has gone stale could
never be replaced: every re-arm was refused because the lock holder was a live
pid, and the holder was never evicted because it was not dead. Add
FM_WATCHER_STALL_BOUND (default 3x the stale grace): below it the refusal is
unchanged; at or past it the arm re-verifies the holder against the lock's
recorded identity, sends TERM, waits boundedly, and takes the lock the normal
way, ledgering a stalled-holder-replaced row. A holder that survives TERM keeps
the old refusal.
Fixes #4400
* no-mistakes(test): poll for replacement message to fix watcher-lock test flake
* no-mistakes(document): document FM_WATCHER_STALL_BOUND in config inventory
* fix(pi): hide queued Firstmate inputs under Calm only when the session can keep them (#5563)
* fix(pi): hide queued Firstmate notifications under Calm only when the session can keep them
Calm now keeps authenticated Firstmate operational inputs out of Pi's queued-message
listing, but only after proving the live session exposes every member needed to keep
them across Escape. A session missing any of them keeps stock rows and Escape and shows
one generic warning. Escape and the dequeue key return only captain-authored messages to
the editor and re-queue hidden notifications in order; after an abort that kept any in
Pi's agent queue, the adapter starts the delivery turn itself because Pi 0.87.1 does not
continue an aborted run. Compaction-held notifications stay with Pi's compaction flush and
never start or announce a turn.
Fixes #1588
* docs(calm): record Pi 0.87.1 queued-row retention verification
* no-mistakes(review): Deliver kept Calm notifications after tree-navigation aborts too
* no-mistakes(review): Defer Calm notification turn until tree navigation finishes
* no-mistakes(lint): Silence SC2016 for literal JavaScript in queue-retention e2e test
* fix(bin): refuse teardown when a required source disappears (#5548)
* fix(bin): refuse teardown when a required source disappears
A missing sibling was sourced after cleanup had started, so Bash 3.2
exited 0 from the EXIT trap and Bash 5 continued and reported success.
* no-mistakes(review): Remove unused FM_TEST_ONLY hook from teardown tests
* no-mistakes(review): Check task backend sources before any teardown cleanup
* test(gotmp): give teardown fixtures every tmux adapter sibling
Teardown now refuses when a sibling the recorded backend's adapter sources
is missing, so the fake bin must carry fm-session-lock-lib.sh,
fm-agent-process-lib.sh and fm-gemini-lib.sh.
* docs: make supervision-host easier to read (#5605)
Restructure the supervision host doc's prose into shorter sections, lists,
and tables without changing documented behavior. Every original heading,
anchor, identifier, number, quoted string, and link target is preserved.
* docs: make the Herdr backend guide easier to read (#5606)
* docs: make herdr-backend easier to read
Restructure the Herdr backend doc's prose into shorter sections, lists, numbered procedures, and tables without changing documented behavior. Every original heading, anchor, fenced code block, link target, and documented fact is kept.
* no-mistakes(document): Restore composer-proof reason and complete Herdr topic table
* docs: make pi-supervision-branch easier to read (#5607)
Restructure the prose into sections, lists, and tables without changing
documented behavior. Every original heading and anchor, inline-code span,
link target, number, and quoted string is kept, and each sentence sits on
its own line. Adds a topic navigation table and short subsections under
the existing headings.
* docs: make watcher-continuity easier to read (#5608)
* docs: make watcher-continuity easier to read
Restructure the prose into sections, lists, and tables without changing documented behavior. Every original heading, anchor, identifier, link target, and number is kept.
* no-mistakes(review): Fix actor and supervision-host scope in watcher-continuity doc
* no-mistakes(review): Make readiness TERM and retry conditional on unready successor
* docs: make sessionstart-nudge easier to read (#5609)
* docs: make sessionstart-nudge easier to read
Restructure the prose into sections, lists, and tables without changing documented behavior. Every original heading, inline-code span, link target, number, and fact is preserved, and a harness-to-tier table now sits near the top.
* no-mistakes(review): Drop helm glossary line and dedupe exit-code lead-in
* docs: make captain-hold-lifecycle easier to read (#5610)
* docs: make captain-hold-lifecycle easier to read
Restructure the captain-hold lifecycle prose into sections, lists, and tables without changing documented behavior. Every original heading, anchor, identifier, number, quoted string, and link target is kept.
* no-mistakes(review): Fix verification record subjects and grouping headings
* no-mistakes(review): Clarify task-body read-back cases belong to the suite
* docs: make remote-secondmates easier to read (#5612)
* docs: make remote-secondmates easier to read
Restructure the remote second mates prose into sections, lists, numbered procedures, and tables without changing documented behavior. Every original heading, anchor, fenced code block, identifier, link target, and qualifier is preserved.
* no-mistakes(review): Merge remote-home table cell into one sentence
* no-mistakes(review): Tighten readiness lead-in, restore causal link, fix dangling reference
* fix(bin): bound the away digest and log why a delivery failed (#5554)
* fix(bin): bound the away digest and log why a delivery failed
The away daemon joined every buffered escalation into one unbounded
digest. A start-up catch-all span can exceed what one transport argument
carries (tmux rejects the send-keys command; Linux refuses to exec any
argument above 131,071 bytes, which is how herdr receives it), so the
initial send failed on every housekeeping pass and was logged as an
unconfirmed Enter with text possibly in the composer.
escalate_flush now builds the injected digest under a fixed byte budget:
each event is cut at a UTF-8 boundary with an omitted-bytes marker, the
joined events stop with a "+K more event(s)" tail, and a bounded digest
names a state/.subsuper-digests/ file that keeps every buffered event
verbatim. The buffer itself is untouched, so the return catch-up stays
complete.
The tmux submit core and the herdr literal send now replay the
transport's stderr on failure, and inject_msg logs the failing stage
(initial send versus Enter confirmation) with the byte count and that
stderr. The wedge alarm line and marker carry the last failure reason.
Fixes #4382
* no-mistakes(review): Drop digest pruning; label send-failed as send-or-Enter stage
* no-mistakes(review): Keep digest full text once submit ran; reuse on retry
* no-mistakes(lint): Count digest files with find instead of ls
---------
Co-authored-by: firstmate-oss <firstmate@kunchenguid.local>
* test: add a gated harness seam and stabilize lifecycle fixtures (#5638)
* feat(tests): add FM_TEST_SEAM launch seam and gate lab-primary recipe
Part 1 of the #5615 split: the pieces that let the no-mistakes pipeline
live-validate firstmate changes, without the gate-refusal rescoping.
- bin/fm-afk-launch.sh: FM_TEST_HARNESS pins the detected harness only
alongside the FM_TEST_SEAM=1 marker test suites set, so a leaked variable
in a real primary's environment stays inert and unknown tokens fall
through to real detection.
- tests/lib.sh: export FM_TEST_SEAM=1 for every suite.
- .no-mistakes.yaml: per-harness recipe for running a real fixture primary
from a gate run - a plain mktemp lab FM_HOME on a private tmux socket,
with FM_GATE_REFUSE_BYPASS=1 scoped to it and NO_MISTAKES_GATE scrubbed.
- tests/fm-wake-queue.test.sh: stop the owned watcher fixture with KILL and
clear its lifecycle state so the next leg starts clean; TERM could leave
bash waiting in a child on some runners.
- tests/fm-remote-secondmate-lifecycle-e2e.test.sh: wait for the liveness
lock holder's post-acquire marker instead of the lock dir, which is
published before the claim finishes.
* no-mistakes(review): Scrub lab home overrides and require FM_TEST_SEAM separately
* no-mistakes(document): Clarify test seam and disposable lab bypass documentation
* no-mistakes(document): Clarify lab isolation and test-seam documentation
* no-mistakes(ci): Fixed the CI failure: test cleanup killed the remote worker child but left its supervisor able to restart it during fixture removal. Cleanup now stops the worker tree. The lifecycle test passed locally; ShellCheck and diff checks passed
* fix(bin): refuse watchers from disposable checkouts and exit when the home is gone (#5552)
* fix(bin): refuse watchers from disposable checkouts and exit when the home is gone
Fixes #321
Fixes #4760
A watcher armed from a disposable no-mistakes validation checkout under
.no-mistakes/worktrees/ outlived the validation step and kept writing the
real home's state, and a running watcher never noticed when its home,
state directory, or code root disappeared. The arm now refuses from such
a checkout with the typed failure line, the watcher checks once per poll
that its home, state directory (or its own lock holder record), and bin
directory still exist and exits with a logged reason scoped to itself,
and the shared test helpers reap every watcher a suite armed for a
temporary home through the home-scoped stop.
* no-mistakes(lint): fix SC1007 by assigning empty string in watch-arm test
* no-mistakes(ci): Found and fixed a genuine, reproducible hang introduced by this branch's test-watcher reaper, which is what killed both CI checks (serial-2 cancelled at the 30-min cap; Lint 2 exit 143 = the suite's own TERM-trap code). Root cause: test_drain_asserts_watcher_liveness (tests/fm-wake-queue.test.sh) fabricates a .watch.lock whose pid is the test runner's own $$ with the runner's real identity, to make the drain believe a live watcher exists. The new make_case tracking registers that state dir for reaping, so at fm_test_cleanup the new fm_test_reap_watchers drives fm-watch-arm.sh --stop; its identity check matches (the fixture recorded the runner's identity) and it kill -TERMs the test runner. tests/lib.sh:231 is `trap 'fm_test_cleanup; exit 143' TERM`, so the TERM re-enters cleanup -> reap -> kills $$ again -> infinite loop until the runner cap. I reproduced this locally: the suite ran all tests then looped forever in cleanup spawning fm-watch-arm.sh --stop against a lock naming its own PID. Fix (tests/lib.sh, +5 lines): in fm_test_reap_watchers, skip any tracked lock whose pid equals our own $$ before driving --stop. This is the single shared reap boundary; seven $$-self-lock fixtures across four test files are all covered by the one guard, and real armed watchers (pid != $$) are still reaped. Invariant: the test reaper must only signal real armed watcher processes, never the test runner itself. Verified locally: tests/fm-wake-queue.test.sh -> EXIT 0 (63 ok, no hang); tests/fm-watch-arm.test.sh -> EXIT 0 (21 ok, including test_reaper_stops_a_tracked_watcher, confirming the guard does not over-skip). Lint 2's exit 143 was the same shard/cap signature; a fresh CI run on this new commit will re-evaluate it
---------
Co-authored-by: firstmate-oss <firstmate@kunchenguid.local>
* fix(bin): surface unrecognized status prefixes instead of reading them as silence (#5588)
* fix(bin): surface an unrecognized status prefix instead of dropping it
A parked or holding declaration, and a verb whose correlation token did not parse, never became an event, so the supervisor still saw the earlier line.
* no-mistakes(review): Require verb-shaped unrecognized status prefixes, add continuation tests
* no-mistakes(document): Document unrecognized status prefix escalation in afk skill
* no-mistakes(ci): I reproduced the "Behavior portable serial 6" failure locally and fixed it by changing the test data in one test. No product code changed. **What failed:** `tests/fm-session-start.test.sh`, in `test_orphan_status_logs_are_printed`, with "matched status log was printed 2 times". **Why:** the test writes status lines with made-up prefixes, `matched: surfaced once` and `orphan: step N`. The test only uses them as placeholder text. It checks that the session-start digest prints each task's status tail exactly once. This PR (#4763) deliberately makes an unrecognized one-word lowercase prefix a status event. So those lines now surface as captain-relevant events, and the wake queue's STATUS OUTCOME BACKSTOP section prints them a second time. The code under review is behaving as the issue asks. Only the test's placeholder data had become meaningful. **Rule the test depends on:** its status lines must not be captain-relevant, so the digest is the only place they are printed. Both lines in this test broke that rule. The orphan line would have failed the same count check right after the matched line did. **Fix:** in that test only, I switched both lines to the recognized, non-captain verb `working:`: `working: surfaced once` and `working: orphan step 1..6`. I updated the matching assertions and counts to use the new text. What the test checks is unchanged: orphan logs are labelled, the tail is bounded, the log path is printed, and each tail appears once. **Verification:** before the fix, the test failed locally the same way as in CI. After it, `bash tests/fm-session-start.test.sh` reports "all assertions passed
* no-mistakes(review): Detect unrecognized prefixes on unstamped lines; share verb list
---------
Co-authored-by: Kun's firstmate <kunchenguid+firstmate@users.noreply.github.com>
* fix(bin): report the failing item when remote inheritance fails (#5658)
Fixes #5295
Session start now reports a remote inheritance failure using the
push's own error line instead of the first unchanged item that
happened to print before it, and the shared captain preferences
header check now names the first required phrase it did not find,
on both the local and remote inheritance paths.
* fix(bin): stand down the Claude Stop auto-arm on pi-code-delivered payloads (#5657)
* fix(bin): stand down the Claude Stop auto-arm on pi-code-delivered payloads
pi-code loads the tracked Claude settings but has no asyncRewake, so it
awaits every Stop hook; without a stand-down the auto-arm runs
synchronously inside Pi's turn end and holds it open for the declared
multi-hour timeout. Stand down when the payload's transcript_path
contains a /.pi/ path component, the same discriminator the closed-but-
unmerged fix in #3352 used, with an explicit string-type check on the
jq filter.
Fixes #3343
* no-mistakes(document): document pi-code stand-down in harness integrations reference
* fix(bin): match whole multi-word project names in the registry lookup (#5659)
* fix(bin): match whole multi-word project names in the registry lookup
bin/fm-project-mode.sh matched a registered project name against only the
first whitespace-delimited token of a registry row, so a name containing a
space never matched, silently defaulting the project to no-mistakes off
instead of its declared posture.
The lookup now matches the whole registered name against the raw line text,
so a name is compared literally (never as a regex) and a name that is a
leading prefix of another registered name still resolves to its own row.
* no-mistakes(document): docs already accurate for multiword registry name match
* chore: drop accidental empty err file
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
* fix(bin): classify shell stdin payloads after -s operands (#5546)
* fix(bin): classify the stdin program of `bash -s` with operands in the arm policy
With -s, sh/bash/zsh read the program from stdin even when operands follow;
the operands are only positional parameters. The arm policy treated the first
operand as a script path, so heredoc and here-string payloads were never
classified and a hidden bin/fm-watch.sh execution was allowed.
A protected path in the operand position still fails closed as before.
Fixes #1489
* no-mistakes(document): Clarify stdin shell operand documentation
* no-mistakes(ci): Captain, fixed `shellInvocation` so `bash -- -s` treats `-s` as a script name, and updated R21 to test the exact command. The targeted policy suite, lint, documentation check, and diff check pass. Both hosted workflows show `action_required` before any jobs ran; that external approval state remains unresolved
* fix(bin): keep main's handling of words after a leading `--`
Revert the pipeline CI-step change that made the first word after a leading
`--` always a script. It turned forms that main denies today into allow
(for example `bash -- -c 'bin/fm-watch.sh'`), which is outside #1489 and
loosens a fail-closed policy. `--` after `-s` still ends option parsing.
* fix(bin): strip AI co-author trailers from fleet-launched commits (#5695)
* fix(bin): strip AI co-author trailers from fleet-launched commits
Cursor and other non-Claude runtimes append the trailer after the typed
message. A per-task commit-msg hook removes it and leaves human co-authors
and the author identity untouched.
* no-mistakes(review): Export pane hooksPath override and drop generated-with stripping
* no-mistakes(ci): This PR caused all three CI failures, and the fix is test-only: 4 test files change, no product code. **Cause.** `fm-spawn.sh` now installs the AI-trailer strip hooks for every spawn, secondmates included. The installer refuses a worktree that is not a git repository, and the PR deliberately keeps that fail-closed rule because real secondmate homes are firstmate clones. Four test fixtures still gave secondmates a plain directory as their home, so each spawn failed with "not a git worktree ... could not install the AI-trailer strip hooks": - serial 5: `tests/fm-backlog-atomicity.test.sh` ("secondmate spawn failed"). - serial 8: `tests/fm-secondmate-harness.test.sh` ("split: no meta written"). - Herdr: `tests/fm-backend-herdr-launcher-workspace-e2e.test.sh` and `tests/fm-backend-herdr-workspace-per-home-e2e.test.sh`. **Rule that must hold.** Every home a test spawns as a secondmate must be a git worktree. I checked the other places in the changed area: the only secondmate spawns in these tests are the ones listed. The earlier rounds already fixed the other fixtures (`fm-secondmate-liveness`, `fm-secondmate-safety`) the same way. **Fix.** - Each of those four secondmate homes now gets the same `.gitignore` plus `git init -q -b main` that the liveness and safety tests already use. - The two Herdr tests clean up with their own plain `rm -rf "$TMP_ROOT"`, not the shared `tests/lib.sh` helper. Because the installer leaves each `state/<id>.git-hooks` directory read-only, that cleanup printed "Permission denied" and left the directories behind. Both cleanups now restore the owner's write bit on every directory before removing (`find ... -exec chmod u+rwx`), which is what `fm_test_remove_tree` in `tests/lib.sh` does. **Verification.** - `tests/fm-secondmate-harness.test.sh` passes. - `tests/fm-backlog-atomicity.test.sh` passes (99 ok, exit 0). - shellcheck is clean on all four files. - I could not run the two real-Herdr tests locally: the Herdr lab on this host refuses to start because it needs exactly one running default session, and I did not change the host's Herdr state to get around that. Instead I checked their two changed steps directly: the installer succeeds on a home set up the new way, and the new cleanup removes the read-only hooks directory completely. Those two tests will only be proven on CI
* fix(bin): treat Pi's dollar-first cost footer as furniture (#5683)
* fix(bin): treat Pi's dollar-first cost footer as furniture
An idle Pi status row opening with $0.000 was read as a dead-shell prompt, so exit and relaunch refused on an empty composer.
* test: wait for the draining holder to exec sleep before reading its identity
The procevent drain fixture read fm_pid_identity immediately after
backgrounding setsid sleep, racing the child's exec chain. Mid-exec the
cmdline can read empty, failing the fixture on a loaded CI runner.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(bin): refuse merges with unreported required checks (#5534)
* fix(bin): refuse a merge when a required check never reported
fm-pr-merge.sh built its GitHub refusals only from checks present in
statusCheckRollup, so a required check that never ran was simply absent and
the merge proceeded on the subset that reported, contradicting its own
"every required check green" claim.
The GitHub verify now reads the base branch's required contexts from the
forge itself - the classic branch protection summary on
GET repos/{o}/{r}/branches/{b} and the active ruleset rules on
GET repos/{o}/{r}/rules/branches/{b} - and refuses when a required context
has no entry in the same rollup, at the same head, that the merge is bound
to. Absence reads as unknown, never green. The required-set read joins the
existing refusal list, so a draft, a red check, and an unreported required
check are all reported together.
Could not read vs nothing required: both endpoints need only repository
read access. The admin-only GET .../branches/{b}/protection endpoint is
deliberately not used: it answers a non-admin token with the same 404 an
unprotected branch gets (observed live on kunchenguid/firstmate main with
this token), which would read a missing permission as "nothing required".
Any failed or malformed read of either source (auth, missing fine-grained
permission, rate limit, network, 404, unexpected shape) refuses the merge
with a line naming the unreadable source. The one exception is GitHub's
plan-gated 403 on the rules endpoint ("Upgrade to GitHub Pro or make this
repository public"), which already means "this repository has no branch
rules" for the merge-queue reader; that check moves into one shared helper
and the classic summary still decides for such a repository.
Attended waiver: --allow-missing <check-name> is the twin of --allow-red and
follows the same design and recording path: once, separate name argument,
waives only that exact unreported required check, still requires every
other required check reported and every check green, never waives an
unreadable required set, refused while the away-posture record exists, and
refused on GitLab. Merge-state BLOCKED policy is unchanged.
How this differs from the withdrawn #5353 (read from its diff):
- #5353 read the admin-only branches/{b}/protection endpoint and treated
its 404 as "no required checks", so for any non-admin token the required
set silently read as empty; this change reads the read-access branch
summary and treats every failure as unreadable.
- #5353 ignored rulesets; this change also reads required_status_checks
rules from the effective branch rules.
- #5353 made separate per-head REST reads of statuses and check-runs capped
at per_page=100 with no pagination; this change checks presence in the
same statusCheckRollup view the red-check gate already reads at the
verified head.
- #5353 stopped at the first unreadable read; this change reports it as one
refusal among all the others.
- #5353 also claimed #5345 (lock stealing) and changed 39 files, most
unrelated; this change is #5344 only.
Live proof, read-only (a gh wrapper refused every merge and mutating call):
- cli/cli#14474 (trunk requires 3 classic build contexts, none ran):
refused, naming build (macos-latest), build (ubuntu-latest),
build (windows-latest); with --allow-missing "build (macos-latest)" it
still refused, naming the other two.
- cli/cli#13665 (ran build (ubuntu-24.04-firewall) instead): refused,
naming build (ubuntu-latest).
- hashicorp/terraform#39262 (ruleset-required checks absent): refused,
naming Code Consistency Checks, End-to-end Tests, Race Tests, Unit Tests.
- cli/cli#14485 (all required reported and green): verified; the wrapper
blocked the merge call and the pull request read back open.
Fixes #5344
* fix(review): Preserve required-check producers and aggregate independent read failures
* fix(document): Clarify required-check verification and waiver documentation
* fix(bin): match an app-bound required commit status by name
The producer-identity check resolved an app-bound required context only
against check runs, so a required context that the required app reports as
a commit status could never match and always read as "has not reported".
A commit status carries no app id to compare, so an app-bound requirement
that arrives as a status now matches by name, as before producer binding;
check runs keep requiring the configured producer app.
Live, read-only: hashicorp/terraform#39262 requires license/cla from
integration 865473, reported green as a commit status by the CLA app. The
previous head refused it as unreported; this head no longer does, while
still naming the four required check runs that never ran there.
Refs #5344
* fix(document): Clarify accepted commit-status producer verification limitation
---------
Co-authored-by: firstmate-oss <firstmate-oss@kunchenguid.local>
* fix: keep persistent secondmates out of landed-work cleanup (#5696)
* fix(bin): never offer a persistent secondmate for teardown
The return brief's "Landed, cleanup due" scan listed every state/*.meta
record carrying a pr= and a merge-notified marker without regard to kind, so
a secondmate record holding a relayed child's merged PR put the mate itself
up for "bin/fm-teardown.sh <mate>" cleanup. A secondmate is a persistent
worker, never landed work.
- bin/fm-afk-return.sh: skip kind=secondmate in the landed-cleanup scan.
- bin/fm-pr-check.sh: refuse to record pr= or arm a merge watch on a
kind=secondmate record before any side effect; a PR reported on its routed
status channel belongs to a task in the mate's own home, which arms its
own watch.
- bin/fm-watch.sh: a merged result from a poll already armed on a secondmate
retires the poll silently - no merge outcome, marker, or wake.
* no-mistakes(document): Document secondmate merge-watch and return-brief exclusions
* no-mistakes(ci): CI failed in an unchanged watcher-shutdown test whose three-second wait was sensitive to runner load. Increased the wait for both state- and home-deletion cases without changing watcher behavior. The full fm-watch-arm suite passed locally; syntax and diff checks passed
* fix: reject invalid X reply and follow-up arguments before posting (#5702)
* fix(bin): refuse unknown dash-leading args in public-posting fm-x scripts
fm-x-reply.sh collected any unrecognized argument into the positional
pool and took the first one as the reply text, so an invocation like
"fm-x-reply.sh <id> --followup --final <text>" posted the literal
string "--final" to X and silently dropped the real text.
Make argument parsing strict in every script that can post publicly:
an unknown dash-leading argument, a dash-leading request_id/task id, a
dash-leading option value, or a surplus positional now exits 2 with a
usage error before any config load, outbox write, or network call. Reply
text starting with '-' is still accepted via --text-file or stdin, and
--help is honored wherever it appears instead of becoming text (a --help
forwarded through fm-x-followup.sh would have counted as a posted
follow-up and mutated the link).
fm-x-link.sh and the fm-public-followup scripts already refuse unknown
arguments; fm-x-poll.sh takes none.
* no-mistakes(review): Refuse surplus follow-up text sources; drop post-ID help branches
* no-mistakes(document): Clarify reply and follow-up argument usage
* no-mistakes(review): Refuse dash-leading --text-file operands in fm-x-reply
* no-mistakes(document): Correct follow-up argument parsing comment
* no-mistakes(document): Document dismiss argument rejection in script header
* fix: pause broken supervision-host sessions between engine probes (#5701)
* feat(bin): latch the supervision host after repeated engine errors
Rung 3c-1 of the PR 5631 re-cut: the host copies the Pi branch's
broken-session policy. Two consecutive engine errors latch the session;
every away wake then reaches main with one supervision-host line for a
five-minute cooldown, after which one wake probes the engine, and each
failed probe doubles the cooldown up to one hour. A reported turn without
an engine error clears it. The latch is kept per main session, engine, and
model in state/.supervision-host-health, and the engine conversation now
uses the same main-session key, which includes the lock holder's process
identity so a recycled pid never shares either.
Lifted from the validated 5631 tree and adapted to main's away-only host:
the attended recovery line and attended cooldown pass-through are left for
the attended core, so a recovery is only logged.
* no-mistakes(document): Consolidate supervision-host latch documentation
* fix(bin): absorb routine secondmate working and paused status appends (#5535)
* fix(bin): absorb routine second-mate progress while surfacing routed replies
Fixes #2959
A kind=secondmate task's status signal was never absorbable, so a healthy
mate's routine working: and paused: appends woke the primary every time.
signal_crew_provably_working now reads the mate's lines new since the
watcher's classified position: a decision, blocker, terminal outcome, note:,
correlation-marked line, or unknown verb still surfaces regardless of busy
evidence, while unmarked working:, paused:, and resolved: fall through to the
same provably-working absorb an ordinary crewmate gets.
* no-mistakes(review): narrow secondmate routine absorb to working and paused
* fix(bin): use gh-axi for the ship DoD draft check (#5519)
* fix(bin): use gh-axi for the ship DoD draft check
* no-mistakes(review): use PR number not URL in gh-axi draft check
* fix(bin): bind inactive-outcome receipt identity to structured fields only (#5520)
* fix(bin): drop status prose from the inactive-outcome dedupe identity
The inactive-outcome receipt fingerprint included the child's sanitized
last status line, so a persistent child appending routine prose after one
terminal outcome minted a fresh parent event per sentence. Bind the
identity to incarnation, task id, terminal state, and PR only, keeping
the last line in the record as status_head evidence.
Fixes #2960
* no-mistakes(document): note structured-only inactive receipt identity in regression coverage
* feat: record the Claude and Cursor dialog for the supervision host (#5707)
* feat(bin): record the supervision host's dialog mirror on Claude and Cursor
Add bin/fm-host-mirror.sh, the one owner of the supervision host's dialog
mirror file, cursor, lock, and feed, plus the main-session key it keys
entries to. The tracked Claude UserPromptSubmit and Stop hooks and the
Cursor beforeSubmitPrompt and afterAgentResponse hooks record the captain's
prompt and main's reply, only on a home with config/supervision-host, from
a genuine primary checkout, for the lock-owning session. The mirror lands
inert: writers record and nothing reads it yet; attended supervision on the
host is the later step that consumes the feed.
Codex, Grok, OpenCode, and omp have no writer here.
* no-mistakes(review): Scope mirror dedup to session, atomic appends, marker-inclusive caps
* no-mistakes(document): Clarify dialog mirror scope and remove duplicate contract details
* no-mistakes(document): Correct Cursor hook documentation for dialog mirror registration
* no-mistakes(review): Pass mirrored dialog text to jq via stdin
* no-mistakes(document): Clarify dialog mirror documentation and remove duplicate claims
* no-mistakes(review): Preserve internal dialog whitespace; drop mirror check and verified modes
* no-mistakes(review): Drop only identical mirror repeats; remove redundant chmod guard
* fix(bin): retire check-row receipts on branch acknowledgement so away escalations are not repeated (#5731)
* fix(bin): retire check-row receipts on branch acks and report an unchanged situation once
* fix(bin): scope a branch acknowledgement's check-row receipt retirement to
its granted sequences
The away posture lifts the attended partition's check/decision exclusions, so
a branch grant can name check-kind rows - but the branch-actor ack still
assumed check rows were main-only and skipped every receipt scan. The queue
row was consumed while its terminal-outcome .pending receipt stayed behind,
and each inactive-reconcile cadence scan re-queued the same fingerprint. In
the first real away window on the supervision host that re-escalated one
unchanged held-PR situation on every cycle (~1,734 of 4,149 outcomes).
A branch ack now scans inactive-outcome and inactive-reconcile receipts and
commits secondmate stall receipts against exactly the sequences in its
eligible-row snapshot - the same rows it consumes - instead of none. Attended
grants still name no check row, so the scans find nothing.
* fix(bin): store a repeated captain verdict as routine while the task's
durable situation is provably unchanged
fm-branch-outcome.sh append computes a mechanical situation key per captain
row - metadata bytes, captured status-log endpoint and identity, live
crew-state verb, worktree head - and anchors it in
state/.<task>.branch-captain-key. A later captain verdict whose recomputed key
matches is stored as routine with "unchanged since seq <N>:" prefixed to its
summary, so one situation escalates once until something provably changes. A
task with no readable status ledger is never demoted, an unreadable record
fails toward reporting, and teardown removes the sidecar with the task's
other branch records. The append-only store schema is unchanged.
This covers both hosts: the Pi supervision branch and the supervision host
both funnel reports through append.
* docs: check rows are main-owned only while attended; the away posture grants
them to the branch, whose ack retires their receipts exactly
* test: the away-flood reproduction as a regression test (branch ack retires
the receipt and later scans stay quiet), store-level dedupe coverage, and a
branch-ack secondmate stall receipt case
* fix(bin): restore the secondmate child devin-config cleanup path
The branch-captain-key sidecar addition mistyped the sibling entry as
.$child_id.devin-config.json, so a forced secondmate teardown would have
stopped removing each child's real <id>.devin-config.json. Restore the
original path and add a behavioral test that stops the child sweep mid-loop
on a refused close, proving the cleaned child's devin config and captain
anchor are both removed while the unconsumed child's records are retained.
* no-mistakes(review): Key captain dedupe on the covered wake rows' fingerprint
* no-mistakes(review): Drop captain-key demotion; prove one escalation on both surfaces
* no-mistakes(review): Drop unrelated teardown test; cite both receipt test files
* no-mistakes(document): Docs already match branch-ack check-receipt retirement
* fix(bin): deliver Claude-bound operational input as a record-backed doorbell (#5664)
* fix(calm): deliver Claude-bound operational input as a record-backed doorbell
Claude Code 2.1.280 removes U+2063 from every submitted prompt, so a typed
operational envelope reaches a Claude Code primary as plain text. The away
daemon now writes the envelope to a record under state/operational-inbox and
types only a plain doorbell naming it; the /afk return check and the Calm mod
recognize the doorbell only when that record holds a current envelope. Marker-
preserving harnesses keep the typed envelope. The live Calm guard accepts the
2.1.280 module-load log line, drives the doorbell, and asserts thinking stays
hidden.
* no-mistakes(review): Fix operational record retention at 7 days and document prune limit
* no-mistakes(document): Point Calm bounds at 2.1.280 evidence; fix afk-exit comment
* no-mistakes(lint): Pick newest Calm e2e transcript without parsing ls
* docs(calm): add a minimal turning-Calm-on step for Claude Code
* fix(spawn): deliver the Claude launch brief as a record-backed doorbell
Claude Code strips U+2063 from the launch-prompt argument too, so a
worker's launch brief arrived with its operational marker removed.
Publish the brief as a record in the receiving home's operational
inbox - a secondmate's own state, not the primary's - and pass only
the printable doorbell naming it, falling back to the typed envelope
when the record cannot be published so the brief body still delivers.
Unwrap doorbell-carried digests in the daemon digest tests that still
read the raw send log under the claude pin, and update the documented
bounds now that launch briefs hide like the other operational rows.
* test(spawn): cover a secondmate's launch-brief record landing in its own home
The record-backed doorbell resolves its state through the receiving
pane's home, so prove a claude secondmate launch publishes into the
seeded secondmate's operational inbox and never leaks a record into
the primary's.
* no-mistakes(review): Pass primary harness to daemon, tighten retention, refresh verdicts
* no-mistakes(review): Prune operational records by exact seven-day elapsed age
* no-mistakes(review): Batch record pruning so large inboxes still expire
* no-mistakes(review): Refuse Claude spawn when brief record cannot publish
* no-mistakes(review): Drop thinking probe from Claude Calm live test and docs
* no-mistakes(review): Record dated Claude Code 2.1.282 reproduction evidence
* no-mistakes(document): Clarify operational doorbell documentation and record expiry
* no-mistakes(document): Correct AFK escalation carrier guidance
* no-mistakes(review): Describe operational record retention as about seven days
* no-mistakes(document): Clarify Calm delivery and operational record retention
* no-mistakes(review): Remove out-of-scope Calm launch guide from Claude docs
* no-mistakes(document): Document Claude launch-brief delivery and refusal
* no-mistakes(document): Correct stale operational-input documentation
* no-mistakes(ci): Fixed the stale Claude trust test to verify that worker and secondmate launches deliver readable, record-backed briefs instead of expecting brief paths in their commands. Annotated the daemon’s output variable for ShellCheck without changing behavior. The affected tests, daemon tests, ShellCheck, and diff check pass locally
* no-mistakes(ci): parse rebased Claude launch after trailer hook prefix
* no-mistakes(review): Trust launch-brief record and restore thinking bound doc
* no-mistakes(review): Parse final Claude launch statement; drop Stop-hook docs
---------
Co-authored-by: Mike Sewell <maikunari@protonmail.com>
Co-authored-by: no-mistakes <no-mistakes@localhost>
* test: isolate lint fixture from tracked suite (#5727)
* fix(bin): republish parent metadata after a remote secondmate relaunch (#5583)
A host-local relaunch rewrote only the far endpoint, so this home kept the old harness, model, and effort, and appending those keys after pr= broke pull-request poll authentication.
* fix(bin): give slow watcher suites headroom under the changed-suite bound (#5516)
tests/fm-watch-triage.test.sh finishes in about 434s alone and about 698s
under CI load, so the 900s bound the changed-suite runner applies produced
a false timeout under ordinary concurrent validation. Raise the automatic
bound to 1500s, which keeps every measured script under it while staying
below the 30-minute normal CI tier so a genuinely hung script still fails
here with its output before the job cap cancels the lane.
Fixes #3869
Refs #3565
* fix(bin): stop nested steal-lock recursion and mid-steal watcher TERM (#5728)
* Fix nested watcher lock reclaim
* no-mistakes(review): Elect a single steal-mutex reaper and bound arm TERM wait
* no-mistakes(review): Reclaim self-held steal mutex and unify autoarm steal reaping
* no-mistakes(review): Resume own interrupted steal reap from its tombstone
* test: make supervision-host park-boundary tests deterministic (#5710)
* test: hold the back-to-back boundary close on the host's own clock
test_park_boundary_holds_under_back_to_back_closes assumed two engine
turns fit in the ~16s pre-refusal window and that the stub finished a
turn in 3s. Under load the stub's real drain, report, and
acknowledgement take ~13s, so the turn either died at its bound (which
hands the wake to main, no boundary line) or the second close landed
past the window and the fixture failed while the boundary held. 3
failures in 5 runs at a load average near 11.
Hold the first turn on a release file instead: once the engine is in
flight, a second close is appended mid-turn and the turn is released as
the refusal window opens (park bound minus turn bound and grace, read
off the host's own start record). The queued close can then only wait
for the boundary on any machine speed, which is what the test asserts:
the boundary line ends the output, the demo.status row stays queued for
main, and no second engine turn ever starts. A host too loaded to start
the turn at all hands the first close to the same boundary exit.
After: 12/12 at load ~15-42.
* no-mistakes(review): Print boundary test deadline as a decimal integer
* no-mistakes(review): Hold boundary test turn on a FIFO, require full sequence
* no-mistakes(review): Remove stray before/after supervision-host test copies
* test: hold the late close's render until the refusal window opens
The boundary recheck test's node shim slept a fixed 10s, which assumed
the first close was read before the host's refusal window opened. Under
load the close arrived after the refusal check, so the host correctly
refused it before the successor started and the render snapshot never
appeared. Block the wake-prompt render on a FIFO released at the
refusal-open instant read from the host's own start record, so the
pre-turn recheck must refuse on any machine speed.
* no-mistakes(review): Derive minimal park bounds and refresh supervision-host shard hint
* no-mistakes(review): Drive park-boundary tests from a seam-gated host test clock
* fix: stage remote home clones before publication (#5733)
* fix(bin): stage remote home clones before publishing them
A remote home provision cloned the code root directly into the public
FM_HOME path while rollback() claimed rm -rf of that same path on any
failure. Bash defers trapped signals past a foreground child, but any
other cleanup or lifecycle path that removes the home directory races
the live clone's object copy, producing the CI flake "fatal: failed to
copy file to .../.git/objects/...: No such file or directory".
Clone into a private staging directory beside the home and publish with
an atomic rename once complete, so no cleanup can remove a directory a
live clone is still writing; a home that appears mid-provision now dies
cleanly instead of inheriting torn state. The regression coverage holds
a real clone mid-copy, removes the public path, and requires the
provision to finish and publish intact.
* no-mistakes(review): Prove home ownership by sentinel and hold only a live clone
* no-mistakes(review): Assert raced provision publishes a complete, intact clone
* no-mistakes(document): Document remote home staging and publication safety
* no-mistakes(lint): Fix ShellCheck warning in clone integrity assertion
* no-mistakes(document): Clarify remote home publication and rollback guarantees
* fix: preserve Herdr status on Pi relaunch (#5161)
* fix(control): keep a relaunched Pi worker's herdr pane status authority alive
Defect: after `bin/fm-control.sh <id> relaunch` (observed live on a herdr
Pi crewmate whose pane read idle while it ran its validation pipeline),
the pane froze at whatever its previous agent had last reported.
Cause, measured on herdr 0.9.1 against a real Pi: a pane has one status
authority, and for Pi with its integration installed that authority is
the lifecycle hooks, so herdr also skips screen detection for the pane.
In the crew shape the registration outlives its agent process (upstream
issue #4115; docs/herdr-backend.md "Restart and liveness behavior"), and
herdr applies only reports carrying the session identity it bound. A
replacement started fresh in that pane reports a NEW session, so its
state reports are ignored and the pane stays frozen. Nothing from
outside repairs it: `pane report-agent-session` and `pane report-agent`
for `herdr:pi` are accepted (rc=0) without being applied unless the
reporter is the registered pane agent, and `pane release-agent` on the
stale record changes nothing.
Fix: a relaunch preserves the binding instead of fighting it. The launch
owner reads the session reference the endpoint's own runtime recorded
(`fm_backend_herdr_pane_agent_session_ref`) and passes it back as Pi's
own `--session <path-or-id>` (`relaunch_resume_args`;
`fm_control_relaunch_resume_flag` owns which adapters and which
registered-agent labels qualify). That is the same reference herdr
itself resumes Pi panes with after a server restart, and the resumed
session's reports land again, which the live check confirmed: the pane
returned to working while the replacement worked and idle when it
settled, on the same session identity.
Safety: relaunch-only (a fresh spawn binds nothing), herdr-only (the one
adapter that records a per-pane session), Pi-family only, and only when
the registration's own agent label matches - so no other adapter's
conversation can be handed to a Pi launch. An unreadable, missing, or
malformed reference degrades to exactly the fresh-session launch that
existed before. No lifecycle, liveness, isolation, or merge guard is
touched, and an empty result leaves every non-Pi launch byte-identical.
`resume` remains a refused verb; docs/agent-control.md and the
harness-adapters references are corrected where they claimed Pi had no
verified resume form at all.
* no-mistakes(document): docs: correct relaunch session-authority ownership and skill paths
* no-mistakes(document): docs: correct stale control-plane ownership claim
* no-mistakes(document): docs: drop unverified Herdr restart resume claim
* no-mistakes(test): Added offline Herdr Pi session-authority relaunch coverage
* no-mistakes(document): Document Herdr Pi relaunch session continuity
* no-mistakes(ci): The failing remote relaunch test tried to arm a PR poll for a secondmate, which `fm-pr-check.sh` correctly refuses. Removed that invalid test scenario; the remaining remote relaunch tests pass, and `git diff --check` is clean
* Seed the relaunch-ordering PR poll fixture without fm-pr-check.sh (#5758)
Main has been red since fm-pr-check.sh began refusing to arm a merge poll
on a kind=secondmate record (#5696): the relaunch-ordering case in
tests/fm-remote-secondmate-relaunch.test.sh armed its fixture through that
entry point and could no longer be set up.
The ordering guarantee still matters: a secondmate record armed before the
refusal can legitimately carry a trailing pr=/pr_head= identity block until
the watcher retires it, and fm-remote-secondmate-relaunch.sh must still keep
that block last when republishing harness/model/effort. Seed the fixture the
way such a record was really written - pr= appended last to the meta, then
the poll artifacts published through the same
fm_pr_poll_prepare/fm_pr_poll_publish_prepared pair fm-pr-check.sh uses, a
pattern tests/fm-pr-check-security.test.sh already follows - and drop the
now-unused fake gh fixture. The #5696 refusal itself stays pinned by the
security suite's secondmate-record case.
* feat: add attended supervision for Claude and Cursor hosts (#5748)
* feat: run attended supervision on the host for Claude and Cursor
On a home opted into config/supervision-host with a Claude or Cursor
primary, the supervision host now takes the attended wakes the Pi branch
would take: routine outcomes stay off main, and a captain outcome wakes
main once with a branch-outcome line and waits in the drain's new
BRANCH OUTCOMES section until main acknowledges it with mark-processed.
- The offer rule moves into branchOfferForWake, shared by the Pi watcher
and the host through bin/fm-branch-dispatch.mjs offer.
- The host feeds the dialog mirror at the head of each attended wake and
passes a close through unchanged when it is main-only, the engine or a
tool is missing, the primary has no verified mirror, the main session
cannot be identified, or the session is cooling down.
- The drain presents captain outcomes first, one line per task, never
behind older routine outcomes, and collapses routine overflow into a
count that is marked read.
- The return advances the store's read cursor through the away window
once the brief has rendered, so the first drain does not replay it.
- The branch prompt's mirror wording is host-neutral, and the rule to
report what main must act on as captain, once per unchanged situation,
applies only to the attended posture on the host.
* docs: record the attended supervision host live check
* no-mistakes(review): Present pre-window unread outcomes and contiguous captain prefix
* no-mistakes(review): Return brief presents every row it marks read
* no-mistakes(review): Return brief lists every unread outcome in one list
* no-mistakes(review): Keep return list in store order and gate cursor failures
* no-mistakes(review): Make the drain the only branch-outcome presenter after return
* no-mistakes(review): Gate return on drain outcome failures; byte-count outcome budgets
* no-mistakes(review): Gate drain on projection failures; UTF-8-safe byte cuts
* no-mistakes(review): Fail drain without jq; hand unreadable prompt mirror to main
* no-mistakes(document): Correct supervision-host return and drain documentation
* no-mistakes(review): Recheck attended offer at turn start; honest failed-drain brief
* no-mistakes(document): Correct supervision-host posture and drain documentation
* no-mistakes(document): Documentation remains accurate for attended supervision
* fix(bin): dedup directed source expansions in fm-pending-reply-lib (#5753)
Each '# shellcheck source=' directive makes ShellCheck's external-source
traversal expand that library's whole transitive graph again at the site.
fm-pending-reply-lib carried three directed lazy sources of fm-wake-lib and
two of fm-parent-channel-lib on identical per-call re-source sites, so one
file analysis peaked above 4 GiB and every caller (fm-watch, fm-teardown)
inherited the multiplier - the root cause of the PR #5732 Lint 1 OOM kill.
Keep the runtime '.' commands byte-identical: the lazy re-source under
'local STATE FM_WAKE_QUEUE FM_WAKE_QUEUE_LOCK' is real behavior. Drop the
duplicate directives so each library expands once per unit, and drop the
tmux/classify directives since classify already arrives through the kept
fm-wake-lib expansion and no tmux symbol is referenced here. The directive
above the lib-dir assignment is kept - it binds the bin/ prefix so the
undirected sites still resolve without SC1091.
Measured peak RSS, ShellCheck 0.11.0 -x on Linux arm64:
bin/fm-pending-reply-lib.sh 4.06 GiB -> 1.96 GiB, zero findings
* fix(bin): avoid bash 5.2 sibling $() in recovery mint and delivery log (#5773)
* fix: split bash 5.2 sibling $() in recovery mint and delivery log
Sibling command substitutions on one line can empty a recovery generation
under bash 5.2 when a CHLD trap is set. Mint pid/epoch sequentially, refuse
empty tokens before write, and clean delivery fields before printf.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: split bash 5.2 sibling $() in recovery mint and delivery log
Sibling command substitutions on one line can empty a recovery generation
under bash 5.2 when a CHLD trap is set. Mint pid/epoch sequentially, refuse
empty tokens before write, and clean delivery fields before printf.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: keep recovery mint failure semantics after sibling $() split
Remove the new pid/date refusal and grammar guard so a mint miss still
yields a grammar-valid token and a durable wake row, matching accepted
review intent. Drop the fake-failing-date case that locked in the refuse.
Co-authored-by: Cursor <cursoragent@cursor.com>
* no-mistakes(document): Point recovery-mint hazard comment at its regression test
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(bin): name the recovery for a declined Claude imports dialog and stop calling Escape safe there (#5791)
Fixes #4520
* fix(bin): report the newest status event in the voice status reader (#5790)
Fixes #4756
The voice status reader in bin/fm_voice_records.py reports each
worker's state from the last non-blank line of its status log. When a
worker appends a status line and then a line of plain prose, the
reader reported "note" with the prose line instead of the declared
state, diverging from bin/fm-classify-lib.sh's shell scan.
Scan back through the tail for the newest line whose prefix is a
single lowercase verb-shaped word (letters and hyphens), and report
that event's verb instead of always taking the last line. An
unrecognised verb-shaped prefix still reports "note" rather than
letting an earlier recognised line answer for it, and free text with
no colon is skipped as prose. When the tail holds no such event, the
last line is reported exactly as before.
* fix(bin): gate a self-announcing tool's update-available report on a newer version (#5786)
* fix(bin): stop reporting an already-installed version as an available update
An update announcement named its version first ("current -> new"), so
reading the first dotted number as the announced version compared the
current version against itself and always looked newer. Read the last
dotted number instead, and only report an available update when that
announced version is newer than the newest installed copy found; when
that version is already installed, report only PATH skew.
Fixes #5151
* no-mistakes(document): docs: gate announce update-available report on newer-than-installed
* fix(bin): pass the dispatch profile effort to OpenCode workers through their launch config (#5799)
* fix(bin): pass the profile effort to OpenCode workers through their launch config
The dispatch profile's effort axis was recorded in task metadata but never
reached an OpenCode worker: the launch wrote only a permission grant into
the config it constructs.
OpenCode 1.18.32's config schema carries per-model reasoning effort as
agent.<name>.variant, so the chosen effort is now merged into the same
OPENCODE_CONFIG_CONTENT JSON as the default build agent's variant, keyed to
the resolved model. With no effort chosen the launch stays byte-identical.
Fixes #1373
* no-mistakes(review): gate OpenCode effort variant by model provider family
* no-mistakes(document): docs(opencode): note provider-family gating for effort variant
* fix: stop cancelled validation runs from reporting false failures (#5815)
* fix: preserve cancellation as no verdict in crew state
Reuse the green-delivery safeguard for cancelled CI monitors and permit a skipped rebase. Other cancelled outcomes and coarse ledger records use the existing unknown state.
Four delivered-PR regressions failed before the fix and pass afterward. The isolated public resolver and fleet-summary tests prove that undelivered cancellation no longer creates a failure contradiction, while preserving historical records and the terminal_in_flight invariant. Evidence uses fixture no-mistakes responses, not a live daemon cancellation.
Update the existing coarse cancellation assertion from failed to unknown because it encoded this defect; retain its newest-run precedence check. Full fm-crew-state suite and pinned lint pass.
* fix(review): Verify PR disposition before reclassifying terminal validation runs
* fix(test): Add captured cancellation replay coverage for resolver and fleet
* fix(document): Clarify cancellation and terminal delivery documentation
* fix: require declared waits for workers awaiting their own work (#5812)
* fix: declare worker background and pipeline waits
Require ship and scout workers to declare owned-work waits with the existing
paused verb before ending a turn or waiting on a pipeline or long command.
Keep the first-sight alert and existing liveness classification unchanged;
subsequent inspection follows the existing long pause cadence.
Validation: emitted brief regression failed before the instruction change
and passes afterward. Public watcher/drain regressions cover the first
alert, repeated wedge suppression, bounded rechecks, and undeclared idle
alarms using isolated backend fixtures. Brief suite, pinned lint, Bash
syntax, documentation inventory, and whitespace checks pass.
No real worker harness was exercised for wait behavior.
* fix(document): Clarify declared worker waits and documentation ownership
* fix(ci): Captain, fixed the cadence test to age both the declaration and first-alert throttle while preserving declaration identity. Reproduced the CI failure using stable identity; corrected tests pass with both stable identity and native macOS behavior. Focused ShellCheck and diff checks pass. Production behavior is unchanged; Linux CI was not rerun locally
* fix: bound ShellCheck to one canonical root per process (#5770)
* fix(bin): bound each lint root in its own ShellCheck process
CI job "Lint 1" died twice at about ten minutes because the two shard
workers each packed about 110 canonical roots into one unbounded ShellCheck
process, and a byte-weight rebalance moved the analysis-heavy fm-watch.sh
into a partition with other heavy roots, so the pair outgrew the 16 GiB
runner before anything could name a culprit.
Run one canonical root per ShellCheck process under an enforced envelope:
a wall deadline plus terminate-then-kill grace via the shared
fm-timeout-lib.sh watchdog, and a per-root rlimit spec applied inside the
child before exec (default a 4 GiB address-space cap, so two workers stay
inside a 16 GiB job with headroom). A root that exceeds the envelope fails
by name with a recorded reason - timeout, memory, signal, or
limit-unavailable - instead of taking the runner down. The per-root
watchdog runs in its own process group so the owner's group sweep cannot
orphan the bounded subtree, and fm_exec_timed now starts the same
escalation when its parent dies before it can be signalled.
FM_LINT_REQUIRE_BOUNDS=1, set in CI, refuses the run outright when a
configured bound cannot be enforced on the host rather than lint uncapped.
Each root's begin/end, reason, duration, and peak RSS stream to stderr in
partition mode and append to a retained <telemetry>.roots.tsv sidecar
uploaded beside the partition telemetry.
Coverage is unchanged: pinned ShellCheck 0.11.0, --norc, --external-sources
full analysis, complete and disjoint partition inventory, workflow lint,
and the backend-purity check, with byte-identical diagnostics across
jobs=1/2 proven by tests/fm-lint.test.sh.
* fix(bin): fail closed on unenforceable lint bounds and size the cap
Required-bounds mode (FM_LINT_REQUIRE_BOUNDS=1, set by CI) now refuses the
run with named errors before any root starts: a missing fm-timeout-lib.sh,
a watchdog that cannot actually bound a probe command, or a host that
rejects the address-space limit all stop the run rather than lint uncapped.
The generalized FM_LINT_ROOT_RLIMITS flag:value interface is replaced by a
single FM_LINT_ROOT_MEMORY_KIB, and the roots sidecar and telemetry record
the run's final exit status after backend-purity and workflow checks
instead of the pre-check lint status.
The default cap is 6 GiB of address space per root, not 4 GiB: ulimit -v
bounds virtual address space rather than resident memory, and ShellCheck's
GHC runtime keeps roughly a third of that space as reservation, so 6 GiB
yields about a 4 GiB working heap budget. A Linux measurement during this
change showed eleven real canonical roots running out of memory under the
earlier 4 GiB cap while the largest passing root peaked near 2.8 GiB
resident; two 6 GiB roots plus runner overhead still fit the 16 GiB job.
Roots that still exceed the cap keep failing by name, and the sidecar's
per-root peak RSS keeps roots approaching the budget visible.
tests/fm-lint.test.sh now proves the memory primitive where it can be
proven: on hosts that accept ulimit -v a perl allocator is refused under a
256 MiB limit and reported by name as a memory death, the pinned ShellCheck
lints a small file under the configured cap and is named when a far smaller
cap binds it, and a watchdog-less copy refuses under REQUIRE_BOUNDS; the
bounded cases skip on macOS, which cannot enforce the address-space limit.
* no-mistakes(review): Prove memory cap binds, pass watchdog owner, drop unused modes
* no-mistakes(review): Capture watchdog owner before startup for every fm_exec_timed caller
* no-mistakes(document): Clarify bounded lint documentation and telemetry
* no-mistakes(document): Correct bounded lint documentation and sidecar path
* docs(bin): restore the per-root memory cap sizing rationale
The pipeline's document step rewrote the ROOT_MEMORY_KIB comment and
dropped the sizing reasoning the change is required to record: address
space vs resident memory, the GHC reservation share, the measured 4 GiB
failures and ~2.8 GiB peak, and the two-roots-plus-runner capacity
arithmetic. Restore it beside the default while keeping the corrected
"not a resident-memory ceiling" framing.
* no-mistakes(review): Document memory cap RSS reduction threshold and first candidate
* no-mistakes(review): Scope owner-death escalation docs to the perl watchdog
* no-mistakes(document): Clarify bounded lint and timeout documentation
* no-mistakes(review): Install perl watchdog signal handlers before forking the command
* no-mistakes(document): Correct bounded l…
kunchenguid#5583) A host-local relaunch rewrote only the far endpoint, so this home kept the old harness, model, and effort, and appending those keys after pr= broke pull-request poll authentication.
kunchenguid#5583) A host-local relaunch rewrote only the far endpoint, so this home kept the old harness, model, and effort, and appending those keys after pr= broke pull-request poll authentication.
kunchenguid#5583) A host-local relaunch rewrote only the far endpoint, so this home kept the old harness, model, and effort, and appending those keys after pr= broke pull-request poll authentication.
* fix(bin): report a Lavish source armed only after its listener is running (#5566)
* fix(bin): report a Lavish source armed only after its listener is running
Registration alone was treated as ready, so arm could succeed before anything was collecting from the board.
* no-mistakes(review): Guard Lavish arm launches, keep retire refusals, report live prior listener
* no-mistakes(review): Keep polling through window before reporting a still-live prior listener
* test: wait for a capture's claim to drop before the next arm
The result is stored before the runner exits, so a re-arm in that gap was meeting a live claim.
* no-mistakes(document): Record Lavish arm readiness evidence in verification doc
* no-mistakes(ci): Both failures were caused by this PR, and both are fixed with test-only edits. Lint 2 (ShellCheck SC2034): this branch removed the only use of `reply_id` (a `start "$reply_id"` call) from tests/fm-procevent.test.sh, which left the assignment at line 1450 unused. I deleted that assignment. It was the only `reply_id` in the file. ShellCheck is now clean on both test files. Behavior portable serial 4: the failing test was tests/fm-bearings-board.test.sh, in the check "registration consumed its answer before the any-origin binding existed". I reproduced it locally: the hold was still `state: queued` when the test checked it. - What must hold: the test's check that the hold is closed must run after the listener has captured the answer. - Why it broke: the test used a stand-in adapter that ran `fm-procevent.sh start` in the foreground after `arm`, so capture finished before build returned. On this branch, `arm` starts the listener itself in the background, so the real listener captures the answer and closes the hold a moment after build returns. - Fix: removed the now-redundant stand-in adapter, the copied runtime directory, and its extra environment variables. The test now runs the real build through the existing `run_board` helper and waits up to about 10s for the hold to reach `state: done`. The checks that follow are unchanged: `Resolution mode: answered` and the any-origin binding. - Other tests: this was the only test in the file that stood in for the adapter this way. The shard's other pure-contract-unit test (tests/fm-trace-context-lib.test.sh) passed unchanged. Verification: - tests/fm-bearings-board.test.sh passed 3 times in a row via bin/fm-test-run.sh, all 18 checks, about 53s per run. - tests/fm-procevent.test.sh was not rerun, because the lint fix only removed an unused assignment
* fix(bin): stop repeating unknown-wake escalations that were already delivered (#5599)
* fix(bin): acknowledge a delivered unknown-wake escalation
The same unrecognized wake was escalated again after it had already been handled, because delivery never recorded that identity.
* no-mistakes(review): Scope unknown-wake acknowledgements to one away session
* no-mistakes(review): Clear delivered digest when unknown-wake ack write fails
* no-mistakes(review): Limit unknown-wake suppression to acknowledged lines
* no-mistakes(document): List unknown-wake ack file among away-session artifacts
* fix(bin): keep a stated default-key retraction from cancelling a keyless wait (#5587)
* fix(bin): keep a stated default retraction from cancelling a keyless wait
A resolved line that names the shared default decision bucket was closing the keyless live wait that only prints as that same key. Keyless self-retraction still closes the keyless wait.
* no-mistakes(review): Keep declared waits standing past foreign-key resolved lines
* no-mistakes(review): Bound declared-wait read and share one decision-key parser
* no-mistakes(document): Document supervisors' key-aware declared-wait read
* fix(bin): terminate a remote job worker that lost ownership on TERM (#5544)
* fix(bin): terminate a remote job worker that lost ownership when it receives TERM
A serving worker whose lock directory is gone can no longer quarantine shutdown, and resuming service publishes a false ready heartbeat. Exit after stopping only that worker's own command tree, without removing a replacement owner's lock.
* no-mistakes(review): Check worker lock ownership before publishing shutdown quarantine
* no-mistakes(document): Correct worker shutdown comment on replacement-owned lock
* fix(bin): keep an ousted remote job worker off the replacement quarantine
Shutdown can lose the lock after the first ownership check and before it
writes or clears quarantine. Bind both operations to the directory object
this process still owns so a replacement's quarantine stays untouched.
* no-mistakes(review): Make ousted-worker shutdown test reliably reach quarantine clear
* no-mistakes(document): Reattach worker_shutdown doc comment to its function
* no-mistakes(ci): Fixed the failing check (Behavior portable serial 7) with a test-only change to the stall test in tests/fm-remote-job.test.sh. Product code is unchanged; no other test changed. Cause: after the decoy dies, both workers run the same check-exists, read, delete sequence on the job records. On the CI runner the replacement deleted a record between the ousted worker's check and its read. The ousted worker exited 125, and because the file runs under set -e the unguarded `wait` ended the test with 125. The exit trap then killed the replacement, which produced the "Killed" line. Reproduction: a temporary 0.3 s delay between the check and the read, applied to the ousted worker only, made the committed test fail exactly as in CI (exit 125 and the "Killed" line). The new test passed with the same delay. The delay is reverted, along with a similar debug hook that the timed-out attempt had left in bin/fm-remote-job-worker.sh. Test changes: - The replacement is frozen (and confirmed stopped) before the decoy is killed and resumed only after the ousted worker exits, so only one worker touches the job records at a time. - The ousted worker is stopped only once its quarantine exists and its lane is reaped, which places it inside its stop loop. - Every fixed poll loop is now a wait on a named condition with a 30 s deadline and an explicit failure message. Exit detection also handles zombies. - The exit trap kills and waits for the decoy and both workers on every path. - A non-zero exit from the ousted worker now fails with its exit code and stderr instead of silently ending the file. The test still proves that the resumed ousted worker exits 0 and leaves the replacement's lock, quarantine contents and quarantine inode unchanged. Verification: the full test file passed four times on its own and three times under nice -n 10 with four busy-loop CPU hogs; bin/fm-lint.sh passes. Changes are not committed
* no-mistakes(ci): I fixed the failing check (Behavior portable serial 7) by changing only the stall test in tests/fm-remote-job.test.sh. Product code is unchanged. **What failed:** "an ousted worker in shutdown leaves the replacement quarantine untouched" failed on CI with the ousted worker exiting 125 ("could not stop the active command tree"). **Why:** during shutdown, the worker retries the still-running decoy command group a fixed 100 times, 0.01 s apart, then gives up and exits 125. The test tried to freeze the worker partway through those retries by sending SIGSTOP from outside. On a slow runner the retries ran out before the stop arrived, so the worker had already given up. The invariant is that the test must hold the ousted worker inside that retry loop until the replacement owns the lock. That was the only place the test depended on timing. The other waits already watch for a named state change with a 30 s deadline. **Fix:** - The ousted worker now starts with a small `sleep` wrapper at the front of its PATH, and the SIGSTOP race is gone. - The wrapper only holds a `sleep` called directly by that worker's own process (it checks its parent pid against a hold file) while its quarantine file exists. - The only such `sleep` is the first retry in the shutdown stop loop, so the worker waits there as long as needed. - The wrapper writes a marker when it starts holding. The test waits for that marker, then hands the lock to the replacement, freezes the replacement, and kills the decoy. - The test releases the worker by deleting the hold file. Deleting the whole temp directory also releases it, so a failed run cannot leave the wrapper looping. - A process leak: the test overwrites the job's command-group record with the decoy, so no worker ever stopped the job's real command. `fm-hold-job.sh` and its `sleep 30` stayed running for up to 30 s after the test. The test now records that group before overwriting it and kills it at the end of the test and in the exit cleanup. - The test still asserts the same things: the ousted worker exits 0, and the replacement's lock, quarantine contents and quarantine inode are unchanged. **Verification:** - The full file passed twice on its own, twice under `nice -n 10` with six busy-loop CPU hogs, and twice more after the leak fix. - `pgrep` found no leftover processes afterwards. - With the worker from just before the fix commit (cf45cb6^), the test still fails with "the ousted worker wrote or cleared the replacement quarantine during shutdown", so it still proves the fix. - `bin/fm-lint.sh` passes. - I did not reproduce the CI failure locally. The cause comes from the fixed retry limit and the CI error message. The changes are not committed
* no-mistakes(ci): I changed only the stall test ("an ousted worker in shutdown leaves the replacement quarantine untouched") in tests/fm-remote-job.test.sh. Product code is unchanged, and so is every other test. **Invariant:** the pid written to the job's group record must be a process-group leader whose group dies when that one process is killed. Otherwise the worker's bounded stop loop never sees the group die, gives up, and exits 125 ("could not stop the active command tree") before it reaches the lost-ownership exit. The decoy is the only place in this test that depends on this. **Fix:** - The decoy used to be `set -m; sleep 30 &`. It now starts as `perl -MPOSIX=setsid -e 'setsid() >= 0 or exit 1; exec @ARGV' sleep 30 &`, which gets its own session and group without shell job control. tests/fm-procevent.test.sh already uses the same idiom. - The test now waits, with the file's usual 30 s deadline and a named failure, until `ps -o pgid=` of the decoy equals its pid before writing it into the group record. This way the worker can never read the record before `setsid` has run. - The existing steps are unchanged: the test kills the decoy, reaps it with `wait` before releasing the hold file, and the exit trap still kills and reaps the decoy and both workers. - The assertions are unchanged: the ousted worker exits 0, and the replacement's lock pid, quarantine text and quarantine inode stay the same. **Cleanup:** I reverted a debug `printf` hook that the timed-out previous attempt had left in bin/fm-remote-job-worker.sh, and deleted its untracked `.tmp-repro/` directory. Neither was committed. **Verification:** - The full tests/fm-remote-job.test.sh passed twice normally and once under `setsid -w` with stdin from /dev/null (no controlling terminal). - `bin/fm-lint.sh` passes. - No leftover `sleep 30` processes afterwards. **Not reproduced:** I could not reproduce the CI failure locally. On this host `set -m` made the decoy its own group leader even without a controlling terminal, so the cause on the runner is not confirmed. The change removes the test's reliance on shell job control, as the user asked. Changes are not committed
* no-mistakes(ci): I changed only the stall test ("an ousted worker in shutdown leaves the replacement quarantine untouched") in tests/fm-remote-job.test.sh. Product code is unchanged, and so is every other test. **Invariant:** the group record the ousted worker checks in its stop loop must stay the job's own command group, and the test must stop that group before it releases the hold. Otherwise the bounded retry keeps seeing a live group, gives up, and exits 125 ("could not stop the active command tree") before it reaches the lost-ownership exit. The test overwrote this record in one place (the decoy) and stopped the group in one place (killing the decoy); both are changed. **Fix:** - I removed the setsid decoy and the overwrite of `.claim/group`. The record keeps the job's real command group, which the test still saves as `STALL_JOB_GROUP`. - The two-line `group_start` stays. It is still needed: without it the worker kills the real group on its first pass, before the replacement takes over, so the hold would never matter. - The `sleep` wrapper that holds the worker at its first stop-loop retry is unchanged. - After the replacement owns the lock, its quarantine is planted and it is frozen, the test runs `kill -KILL -- -$STALL_JOB_GROUP`. It then waits, with the file's usual 30 s deadline and a named failure, until `kill -0` on the group fails. Only then does it remove the hold file. The worker therefore always sees its own command already stopped and never races its retry budget. - The exit trap still kills the saved command group if the test fails. It can't `wait` on that group because the group is not a child of the test shell. The decoy variable and its cleanup entry are gone. - The assertions are unchanged: the ousted worker exits 0, and the replacement's lock pid, quarantine text and quarantine inode stay the same. **Verification:** - The full tests/fm-remote-job.test.sh passed twice normally. - It passed once under `setsid -w` with stdin from /dev/null (no controlling terminal). - It passed once under `nice -n 10` with six busy-loop CPU hogs. - With the worker from before the fix (cf45cb6^), the test still fails with "the ousted worker wrote or cleared the replacement quarantine during shutdown", so it still proves the fix. - No `fm-hold-job` or `sleep 30` processes were left afterwards. - `bin/fm-lint.sh` passes. **Not reproduced:** I couldn't reproduce the CI failure locally; the decoy version also passed on this host. So I can't confirm why the decoy group stayed alive on the runner. The new wait turns any leftover live group into a clear named failure instead of an exit 125. The changes are not committed
* fix(bin): keep a dead command group dead on bash 5.2
A bare return inside the liveness check drops the failing kill status when the check runs in a conditional, so shutdown keeps treating a stopped group as alive and exits 125.
* no-mistakes(review): Use bash 3.2 fd syntax and fix trap return comments
* docs: make configuration settings easier to find and understand (#5589)
* docs: make configuration settings easier to find and understand
* no-mistakes(review): Restore dropped qualifiers and fix misplaced config doc labels
* no-mistakes(review): Restore three dropped qualifiers in configuration reference
* fix: limit project memory edits to factual corrections (#5636)
* fix: bound worker edits of project AGENTS.md/CLAUDE.md to factual corrections
These files are loaded into every agent session of a project, so additions
should be a deliberate human choice rather than automated task output. The
ship brief's project-memory section and AGENTS.md section 6 previously invited
workers to record durable knowledge, which let project AGENTS.md files accrete
detail the codebase or README already carries. Workers now edit only to fix
factually wrong content - including content their own change made wrong - and
fm-ensure-agents-md.sh runs only alongside such a correction. Stow no longer
routes project-memory additions through ship tasks, and the generated skeleton
no longer invites discovery-driven additions.
* no-mistakes(review): Stop running fm-ensure-agents-md.sh on memory-file corrections
* no-mistakes(document): Clarify manual project-memory initialization and remove duplicate guidance
* feat: permit gate lifecycle calls against disposable lab homes (#5635)
* fix(bin): let gate agents drive lifecycle against marked lab homes
Part 2 of the #5615 split. A no-mistakes gate agent runs inside a
checkout carrying the fleet-captain identity, so fm-gate-refuse-lib
refuses fleet mutation on the gate signal. That refusal was absolute,
which kept gate validation from ever exercising the real lifecycle.
Stamp a disposable lab FM_HOME with a .fm-lab-home marker file that
only bin/fm-lab-home.sh writes, and only onto a fresh empty dir, so
no call path can mark a populated real home. fm_refuse_if_gate_agent
then permits lifecycle only when FM_HOME carries the marker and is
driven through its stock layout - any FM_*_OVERRIDE relocation stays
refused so part of the "lab" cannot be split back onto the real fleet.
The threat model is a confused agent touching the real fleet, not
deliberate forgery, so the marker is a plain token file rather than a
bound record. FM_GATE_REFUSE_BYPASS is unchanged: it still serves the
test harness, which cannot mark hundreds of temp homes.
Teardown's slot-ownership scan compared state-dir paths textually
while fm_firstmate_root_home canonicalizes, so a lab home under a
symlinked TMPDIR scanned its own record twice and self-collided;
compare file identity (-ef) instead.
* no-mistakes(review): Refuse unlistable lab homes and hardlinked slot records
* no-mistakes(review): Mint lab markers only on verified-empty fresh dirs
* no-mistakes(document): Clarify lab-home gate documentation and comment contracts
* no-mistakes(document): Clarify lab-home gate documentation and remove stale claims
* no-mistakes(document): Clarify gate lab-home documentation and boundary wording
* fix(bin): evict a watcher whose beacon stalls past a hard bound instead of refusing every re-arm (#5594)
* fix(bin): replace a watcher whose beacon stalls past a hard bound instead of refusing every re-arm
A fleet watcher that is alive but whose liveness beacon has gone stale could
never be replaced: every re-arm was refused because the lock holder was a live
pid, and the holder was never evicted because it was not dead. Add
FM_WATCHER_STALL_BOUND (default 3x the stale grace): below it the refusal is
unchanged; at or past it the arm re-verifies the holder against the lock's
recorded identity, sends TERM, waits boundedly, and takes the lock the normal
way, ledgering a stalled-holder-replaced row. A holder that survives TERM keeps
the old refusal.
Fixes #4400
* no-mistakes(test): poll for replacement message to fix watcher-lock test flake
* no-mistakes(document): document FM_WATCHER_STALL_BOUND in config inventory
* fix(pi): hide queued Firstmate inputs under Calm only when the session can keep them (#5563)
* fix(pi): hide queued Firstmate notifications under Calm only when the session can keep them
Calm now keeps authenticated Firstmate operational inputs out of Pi's queued-message
listing, but only after proving the live session exposes every member needed to keep
them across Escape. A session missing any of them keeps stock rows and Escape and shows
one generic warning. Escape and the dequeue key return only captain-authored messages to
the editor and re-queue hidden notifications in order; after an abort that kept any in
Pi's agent queue, the adapter starts the delivery turn itself because Pi 0.87.1 does not
continue an aborted run. Compaction-held notifications stay with Pi's compaction flush and
never start or announce a turn.
Fixes #1588
* docs(calm): record Pi 0.87.1 queued-row retention verification
* no-mistakes(review): Deliver kept Calm notifications after tree-navigation aborts too
* no-mistakes(review): Defer Calm notification turn until tree navigation finishes
* no-mistakes(lint): Silence SC2016 for literal JavaScript in queue-retention e2e test
* fix(bin): refuse teardown when a required source disappears (#5548)
* fix(bin): refuse teardown when a required source disappears
A missing sibling was sourced after cleanup had started, so Bash 3.2
exited 0 from the EXIT trap and Bash 5 continued and reported success.
* no-mistakes(review): Remove unused FM_TEST_ONLY hook from teardown tests
* no-mistakes(review): Check task backend sources before any teardown cleanup
* test(gotmp): give teardown fixtures every tmux adapter sibling
Teardown now refuses when a sibling the recorded backend's adapter sources
is missing, so the fake bin must carry fm-session-lock-lib.sh,
fm-agent-process-lib.sh and fm-gemini-lib.sh.
* docs: make supervision-host easier to read (#5605)
Restructure the supervision host doc's prose into shorter sections, lists,
and tables without changing documented behavior. Every original heading,
anchor, identifier, number, quoted string, and link target is preserved.
* docs: make the Herdr backend guide easier to read (#5606)
* docs: make herdr-backend easier to read
Restructure the Herdr backend doc's prose into shorter sections, lists, numbered procedures, and tables without changing documented behavior. Every original heading, anchor, fenced code block, link target, and documented fact is kept.
* no-mistakes(document): Restore composer-proof reason and complete Herdr topic table
* docs: make pi-supervision-branch easier to read (#5607)
Restructure the prose into sections, lists, and tables without changing
documented behavior. Every original heading and anchor, inline-code span,
link target, number, and quoted string is kept, and each sentence sits on
its own line. Adds a topic navigation table and short subsections under
the existing headings.
* docs: make watcher-continuity easier to read (#5608)
* docs: make watcher-continuity easier to read
Restructure the prose into sections, lists, and tables without changing documented behavior. Every original heading, anchor, identifier, link target, and number is kept.
* no-mistakes(review): Fix actor and supervision-host scope in watcher-continuity doc
* no-mistakes(review): Make readiness TERM and retry conditional on unready successor
* docs: make sessionstart-nudge easier to read (#5609)
* docs: make sessionstart-nudge easier to read
Restructure the prose into sections, lists, and tables without changing documented behavior. Every original heading, inline-code span, link target, number, and fact is preserved, and a harness-to-tier table now sits near the top.
* no-mistakes(review): Drop helm glossary line and dedupe exit-code lead-in
* docs: make captain-hold-lifecycle easier to read (#5610)
* docs: make captain-hold-lifecycle easier to read
Restructure the captain-hold lifecycle prose into sections, lists, and tables without changing documented behavior. Every original heading, anchor, identifier, number, quoted string, and link target is kept.
* no-mistakes(review): Fix verification record subjects and grouping headings
* no-mistakes(review): Clarify task-body read-back cases belong to the suite
* docs: make remote-secondmates easier to read (#5612)
* docs: make remote-secondmates easier to read
Restructure the remote second mates prose into sections, lists, numbered procedures, and tables without changing documented behavior. Every original heading, anchor, fenced code block, identifier, link target, and qualifier is preserved.
* no-mistakes(review): Merge remote-home table cell into one sentence
* no-mistakes(review): Tighten readiness lead-in, restore causal link, fix dangling reference
* fix(bin): bound the away digest and log why a delivery failed (#5554)
* fix(bin): bound the away digest and log why a delivery failed
The away daemon joined every buffered escalation into one unbounded
digest. A start-up catch-all span can exceed what one transport argument
carries (tmux rejects the send-keys command; Linux refuses to exec any
argument above 131,071 bytes, which is how herdr receives it), so the
initial send failed on every housekeeping pass and was logged as an
unconfirmed Enter with text possibly in the composer.
escalate_flush now builds the injected digest under a fixed byte budget:
each event is cut at a UTF-8 boundary with an omitted-bytes marker, the
joined events stop with a "+K more event(s)" tail, and a bounded digest
names a state/.subsuper-digests/ file that keeps every buffered event
verbatim. The buffer itself is untouched, so the return catch-up stays
complete.
The tmux submit core and the herdr literal send now replay the
transport's stderr on failure, and inject_msg logs the failing stage
(initial send versus Enter confirmation) with the byte count and that
stderr. The wedge alarm line and marker carry the last failure reason.
Fixes #4382
* no-mistakes(review): Drop digest pruning; label send-failed as send-or-Enter stage
* no-mistakes(review): Keep digest full text once submit ran; reuse on retry
* no-mistakes(lint): Count digest files with find instead of ls
---------
Co-authored-by: firstmate-oss <firstmate@kunchenguid.local>
* test: add a gated harness seam and stabilize lifecycle fixtures (#5638)
* feat(tests): add FM_TEST_SEAM launch seam and gate lab-primary recipe
Part 1 of the #5615 split: the pieces that let the no-mistakes pipeline
live-validate firstmate changes, without the gate-refusal rescoping.
- bin/fm-afk-launch.sh: FM_TEST_HARNESS pins the detected harness only
alongside the FM_TEST_SEAM=1 marker test suites set, so a leaked variable
in a real primary's environment stays inert and unknown tokens fall
through to real detection.
- tests/lib.sh: export FM_TEST_SEAM=1 for every suite.
- .no-mistakes.yaml: per-harness recipe for running a real fixture primary
from a gate run - a plain mktemp lab FM_HOME on a private tmux socket,
with FM_GATE_REFUSE_BYPASS=1 scoped to it and NO_MISTAKES_GATE scrubbed.
- tests/fm-wake-queue.test.sh: stop the owned watcher fixture with KILL and
clear its lifecycle state so the next leg starts clean; TERM could leave
bash waiting in a child on some runners.
- tests/fm-remote-secondmate-lifecycle-e2e.test.sh: wait for the liveness
lock holder's post-acquire marker instead of the lock dir, which is
published before the claim finishes.
* no-mistakes(review): Scrub lab home overrides and require FM_TEST_SEAM separately
* no-mistakes(document): Clarify test seam and disposable lab bypass documentation
* no-mistakes(document): Clarify lab isolation and test-seam documentation
* no-mistakes(ci): Fixed the CI failure: test cleanup killed the remote worker child but left its supervisor able to restart it during fixture removal. Cleanup now stops the worker tree. The lifecycle test passed locally; ShellCheck and diff checks passed
* fix(bin): refuse watchers from disposable checkouts and exit when the home is gone (#5552)
* fix(bin): refuse watchers from disposable checkouts and exit when the home is gone
Fixes #321
Fixes #4760
A watcher armed from a disposable no-mistakes validation checkout under
.no-mistakes/worktrees/ outlived the validation step and kept writing the
real home's state, and a running watcher never noticed when its home,
state directory, or code root disappeared. The arm now refuses from such
a checkout with the typed failure line, the watcher checks once per poll
that its home, state directory (or its own lock holder record), and bin
directory still exist and exits with a logged reason scoped to itself,
and the shared test helpers reap every watcher a suite armed for a
temporary home through the home-scoped stop.
* no-mistakes(lint): fix SC1007 by assigning empty string in watch-arm test
* no-mistakes(ci): Found and fixed a genuine, reproducible hang introduced by this branch's test-watcher reaper, which is what killed both CI checks (serial-2 cancelled at the 30-min cap; Lint 2 exit 143 = the suite's own TERM-trap code). Root cause: test_drain_asserts_watcher_liveness (tests/fm-wake-queue.test.sh) fabricates a .watch.lock whose pid is the test runner's own $$ with the runner's real identity, to make the drain believe a live watcher exists. The new make_case tracking registers that state dir for reaping, so at fm_test_cleanup the new fm_test_reap_watchers drives fm-watch-arm.sh --stop; its identity check matches (the fixture recorded the runner's identity) and it kill -TERMs the test runner. tests/lib.sh:231 is `trap 'fm_test_cleanup; exit 143' TERM`, so the TERM re-enters cleanup -> reap -> kills $$ again -> infinite loop until the runner cap. I reproduced this locally: the suite ran all tests then looped forever in cleanup spawning fm-watch-arm.sh --stop against a lock naming its own PID. Fix (tests/lib.sh, +5 lines): in fm_test_reap_watchers, skip any tracked lock whose pid equals our own $$ before driving --stop. This is the single shared reap boundary; seven $$-self-lock fixtures across four test files are all covered by the one guard, and real armed watchers (pid != $$) are still reaped. Invariant: the test reaper must only signal real armed watcher processes, never the test runner itself. Verified locally: tests/fm-wake-queue.test.sh -> EXIT 0 (63 ok, no hang); tests/fm-watch-arm.test.sh -> EXIT 0 (21 ok, including test_reaper_stops_a_tracked_watcher, confirming the guard does not over-skip). Lint 2's exit 143 was the same shard/cap signature; a fresh CI run on this new commit will re-evaluate it
---------
Co-authored-by: firstmate-oss <firstmate@kunchenguid.local>
* fix(bin): surface unrecognized status prefixes instead of reading them as silence (#5588)
* fix(bin): surface an unrecognized status prefix instead of dropping it
A parked or holding declaration, and a verb whose correlation token did not parse, never became an event, so the supervisor still saw the earlier line.
* no-mistakes(review): Require verb-shaped unrecognized status prefixes, add continuation tests
* no-mistakes(document): Document unrecognized status prefix escalation in afk skill
* no-mistakes(ci): I reproduced the "Behavior portable serial 6" failure locally and fixed it by changing the test data in one test. No product code changed. **What failed:** `tests/fm-session-start.test.sh`, in `test_orphan_status_logs_are_printed`, with "matched status log was printed 2 times". **Why:** the test writes status lines with made-up prefixes, `matched: surfaced once` and `orphan: step N`. The test only uses them as placeholder text. It checks that the session-start digest prints each task's status tail exactly once. This PR (#4763) deliberately makes an unrecognized one-word lowercase prefix a status event. So those lines now surface as captain-relevant events, and the wake queue's STATUS OUTCOME BACKSTOP section prints them a second time. The code under review is behaving as the issue asks. Only the test's placeholder data had become meaningful. **Rule the test depends on:** its status lines must not be captain-relevant, so the digest is the only place they are printed. Both lines in this test broke that rule. The orphan line would have failed the same count check right after the matched line did. **Fix:** in that test only, I switched both lines to the recognized, non-captain verb `working:`: `working: surfaced once` and `working: orphan step 1..6`. I updated the matching assertions and counts to use the new text. What the test checks is unchanged: orphan logs are labelled, the tail is bounded, the log path is printed, and each tail appears once. **Verification:** before the fix, the test failed locally the same way as in CI. After it, `bash tests/fm-session-start.test.sh` reports "all assertions passed
* no-mistakes(review): Detect unrecognized prefixes on unstamped lines; share verb list
---------
Co-authored-by: Kun's firstmate <kunchenguid+firstmate@users.noreply.github.com>
* fix(bin): report the failing item when remote inheritance fails (#5658)
Fixes #5295
Session start now reports a remote inheritance failure using the
push's own error line instead of the first unchanged item that
happened to print before it, and the shared captain preferences
header check now names the first required phrase it did not find,
on both the local and remote inheritance paths.
* fix(bin): stand down the Claude Stop auto-arm on pi-code-delivered payloads (#5657)
* fix(bin): stand down the Claude Stop auto-arm on pi-code-delivered payloads
pi-code loads the tracked Claude settings but has no asyncRewake, so it
awaits every Stop hook; without a stand-down the auto-arm runs
synchronously inside Pi's turn end and holds it open for the declared
multi-hour timeout. Stand down when the payload's transcript_path
contains a /.pi/ path component, the same discriminator the closed-but-
unmerged fix in #3352 used, with an explicit string-type check on the
jq filter.
Fixes #3343
* no-mistakes(document): document pi-code stand-down in harness integrations reference
* fix(bin): match whole multi-word project names in the registry lookup (#5659)
* fix(bin): match whole multi-word project names in the registry lookup
bin/fm-project-mode.sh matched a registered project name against only the
first whitespace-delimited token of a registry row, so a name containing a
space never matched, silently defaulting the project to no-mistakes off
instead of its declared posture.
The lookup now matches the whole registered name against the raw line text,
so a name is compared literally (never as a regex) and a name that is a
leading prefix of another registered name still resolves to its own row.
* no-mistakes(document): docs already accurate for multiword registry name match
* chore: drop accidental empty err file
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
* fix(bin): classify shell stdin payloads after -s operands (#5546)
* fix(bin): classify the stdin program of `bash -s` with operands in the arm policy
With -s, sh/bash/zsh read the program from stdin even when operands follow;
the operands are only positional parameters. The arm policy treated the first
operand as a script path, so heredoc and here-string payloads were never
classified and a hidden bin/fm-watch.sh execution was allowed.
A protected path in the operand position still fails closed as before.
Fixes #1489
* no-mistakes(document): Clarify stdin shell operand documentation
* no-mistakes(ci): Captain, fixed `shellInvocation` so `bash -- -s` treats `-s` as a script name, and updated R21 to test the exact command. The targeted policy suite, lint, documentation check, and diff check pass. Both hosted workflows show `action_required` before any jobs ran; that external approval state remains unresolved
* fix(bin): keep main's handling of words after a leading `--`
Revert the pipeline CI-step change that made the first word after a leading
`--` always a script. It turned forms that main denies today into allow
(for example `bash -- -c 'bin/fm-watch.sh'`), which is outside #1489 and
loosens a fail-closed policy. `--` after `-s` still ends option parsing.
* fix(bin): strip AI co-author trailers from fleet-launched commits (#5695)
* fix(bin): strip AI co-author trailers from fleet-launched commits
Cursor and other non-Claude runtimes append the trailer after the typed
message. A per-task commit-msg hook removes it and leaves human co-authors
and the author identity untouched.
* no-mistakes(review): Export pane hooksPath override and drop generated-with stripping
* no-mistakes(ci): This PR caused all three CI failures, and the fix is test-only: 4 test files change, no product code. **Cause.** `fm-spawn.sh` now installs the AI-trailer strip hooks for every spawn, secondmates included. The installer refuses a worktree that is not a git repository, and the PR deliberately keeps that fail-closed rule because real secondmate homes are firstmate clones. Four test fixtures still gave secondmates a plain directory as their home, so each spawn failed with "not a git worktree ... could not install the AI-trailer strip hooks": - serial 5: `tests/fm-backlog-atomicity.test.sh` ("secondmate spawn failed"). - serial 8: `tests/fm-secondmate-harness.test.sh` ("split: no meta written"). - Herdr: `tests/fm-backend-herdr-launcher-workspace-e2e.test.sh` and `tests/fm-backend-herdr-workspace-per-home-e2e.test.sh`. **Rule that must hold.** Every home a test spawns as a secondmate must be a git worktree. I checked the other places in the changed area: the only secondmate spawns in these tests are the ones listed. The earlier rounds already fixed the other fixtures (`fm-secondmate-liveness`, `fm-secondmate-safety`) the same way. **Fix.** - Each of those four secondmate homes now gets the same `.gitignore` plus `git init -q -b main` that the liveness and safety tests already use. - The two Herdr tests clean up with their own plain `rm -rf "$TMP_ROOT"`, not the shared `tests/lib.sh` helper. Because the installer leaves each `state/<id>.git-hooks` directory read-only, that cleanup printed "Permission denied" and left the directories behind. Both cleanups now restore the owner's write bit on every directory before removing (`find ... -exec chmod u+rwx`), which is what `fm_test_remove_tree` in `tests/lib.sh` does. **Verification.** - `tests/fm-secondmate-harness.test.sh` passes. - `tests/fm-backlog-atomicity.test.sh` passes (99 ok, exit 0). - shellcheck is clean on all four files. - I could not run the two real-Herdr tests locally: the Herdr lab on this host refuses to start because it needs exactly one running default session, and I did not change the host's Herdr state to get around that. Instead I checked their two changed steps directly: the installer succeeds on a home set up the new way, and the new cleanup removes the read-only hooks directory completely. Those two tests will only be proven on CI
* fix(bin): treat Pi's dollar-first cost footer as furniture (#5683)
* fix(bin): treat Pi's dollar-first cost footer as furniture
An idle Pi status row opening with $0.000 was read as a dead-shell prompt, so exit and relaunch refused on an empty composer.
* test: wait for the draining holder to exec sleep before reading its identity
The procevent drain fixture read fm_pid_identity immediately after
backgrounding setsid sleep, racing the child's exec chain. Mid-exec the
cmdline can read empty, failing the fixture on a loaded CI runner.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(bin): refuse merges with unreported required checks (#5534)
* fix(bin): refuse a merge when a required check never reported
fm-pr-merge.sh built its GitHub refusals only from checks present in
statusCheckRollup, so a required check that never ran was simply absent and
the merge proceeded on the subset that reported, contradicting its own
"every required check green" claim.
The GitHub verify now reads the base branch's required contexts from the
forge itself - the classic branch protection summary on
GET repos/{o}/{r}/branches/{b} and the active ruleset rules on
GET repos/{o}/{r}/rules/branches/{b} - and refuses when a required context
has no entry in the same rollup, at the same head, that the merge is bound
to. Absence reads as unknown, never green. The required-set read joins the
existing refusal list, so a draft, a red check, and an unreported required
check are all reported together.
Could not read vs nothing required: both endpoints need only repository
read access. The admin-only GET .../branches/{b}/protection endpoint is
deliberately not used: it answers a non-admin token with the same 404 an
unprotected branch gets (observed live on kunchenguid/firstmate main with
this token), which would read a missing permission as "nothing required".
Any failed or malformed read of either source (auth, missing fine-grained
permission, rate limit, network, 404, unexpected shape) refuses the merge
with a line naming the unreadable source. The one exception is GitHub's
plan-gated 403 on the rules endpoint ("Upgrade to GitHub Pro or make this
repository public"), which already means "this repository has no branch
rules" for the merge-queue reader; that check moves into one shared helper
and the classic summary still decides for such a repository.
Attended waiver: --allow-missing <check-name> is the twin of --allow-red and
follows the same design and recording path: once, separate name argument,
waives only that exact unreported required check, still requires every
other required check reported and every check green, never waives an
unreadable required set, refused while the away-posture record exists, and
refused on GitLab. Merge-state BLOCKED policy is unchanged.
How this differs from the withdrawn #5353 (read from its diff):
- #5353 read the admin-only branches/{b}/protection endpoint and treated
its 404 as "no required checks", so for any non-admin token the required
set silently read as empty; this change reads the read-access branch
summary and treats every failure as unreadable.
- #5353 ignored rulesets; this change also reads required_status_checks
rules from the effective branch rules.
- #5353 made separate per-head REST reads of statuses and check-runs capped
at per_page=100 with no pagination; this change checks presence in the
same statusCheckRollup view the red-check gate already reads at the
verified head.
- #5353 stopped at the first unreadable read; this change reports it as one
refusal among all the others.
- #5353 also claimed #5345 (lock stealing) and changed 39 files, most
unrelated; this change is #5344 only.
Live proof, read-only (a gh wrapper refused every merge and mutating call):
- cli/cli#14474 (trunk requires 3 classic build contexts, none ran):
refused, naming build (macos-latest), build (ubuntu-latest),
build (windows-latest); with --allow-missing "build (macos-latest)" it
still refused, naming the other two.
- cli/cli#13665 (ran build (ubuntu-24.04-firewall) instead): refused,
naming build (ubuntu-latest).
- hashicorp/terraform#39262 (ruleset-required checks absent): refused,
naming Code Consistency Checks, End-to-end Tests, Race Tests, Unit Tests.
- cli/cli#14485 (all required reported and green): verified; the wrapper
blocked the merge call and the pull request read back open.
Fixes #5344
* fix(review): Preserve required-check producers and aggregate independent read failures
* fix(document): Clarify required-check verification and waiver documentation
* fix(bin): match an app-bound required commit status by name
The producer-identity check resolved an app-bound required context only
against check runs, so a required context that the required app reports as
a commit status could never match and always read as "has not reported".
A commit status carries no app id to compare, so an app-bound requirement
that arrives as a status now matches by name, as before producer binding;
check runs keep requiring the configured producer app.
Live, read-only: hashicorp/terraform#39262 requires license/cla from
integration 865473, reported green as a commit status by the CLA app. The
previous head refused it as unreported; this head no longer does, while
still naming the four required check runs that never ran there.
Refs #5344
* fix(document): Clarify accepted commit-status producer verification limitation
---------
Co-authored-by: firstmate-oss <firstmate-oss@kunchenguid.local>
* fix: keep persistent secondmates out of landed-work cleanup (#5696)
* fix(bin): never offer a persistent secondmate for teardown
The return brief's "Landed, cleanup due" scan listed every state/*.meta
record carrying a pr= and a merge-notified marker without regard to kind, so
a secondmate record holding a relayed child's merged PR put the mate itself
up for "bin/fm-teardown.sh <mate>" cleanup. A secondmate is a persistent
worker, never landed work.
- bin/fm-afk-return.sh: skip kind=secondmate in the landed-cleanup scan.
- bin/fm-pr-check.sh: refuse to record pr= or arm a merge watch on a
kind=secondmate record before any side effect; a PR reported on its routed
status channel belongs to a task in the mate's own home, which arms its
own watch.
- bin/fm-watch.sh: a merged result from a poll already armed on a secondmate
retires the poll silently - no merge outcome, marker, or wake.
* no-mistakes(document): Document secondmate merge-watch and return-brief exclusions
* no-mistakes(ci): CI failed in an unchanged watcher-shutdown test whose three-second wait was sensitive to runner load. Increased the wait for both state- and home-deletion cases without changing watcher behavior. The full fm-watch-arm suite passed locally; syntax and diff checks passed
* fix: reject invalid X reply and follow-up arguments before posting (#5702)
* fix(bin): refuse unknown dash-leading args in public-posting fm-x scripts
fm-x-reply.sh collected any unrecognized argument into the positional
pool and took the first one as the reply text, so an invocation like
"fm-x-reply.sh <id> --followup --final <text>" posted the literal
string "--final" to X and silently dropped the real text.
Make argument parsing strict in every script that can post publicly:
an unknown dash-leading argument, a dash-leading request_id/task id, a
dash-leading option value, or a surplus positional now exits 2 with a
usage error before any config load, outbox write, or network call. Reply
text starting with '-' is still accepted via --text-file or stdin, and
--help is honored wherever it appears instead of becoming text (a --help
forwarded through fm-x-followup.sh would have counted as a posted
follow-up and mutated the link).
fm-x-link.sh and the fm-public-followup scripts already refuse unknown
arguments; fm-x-poll.sh takes none.
* no-mistakes(review): Refuse surplus follow-up text sources; drop post-ID help branches
* no-mistakes(document): Clarify reply and follow-up argument usage
* no-mistakes(review): Refuse dash-leading --text-file operands in fm-x-reply
* no-mistakes(document): Correct follow-up argument parsing comment
* no-mistakes(document): Document dismiss argument rejection in script header
* fix: pause broken supervision-host sessions between engine probes (#5701)
* feat(bin): latch the supervision host after repeated engine errors
Rung 3c-1 of the PR 5631 re-cut: the host copies the Pi branch's
broken-session policy. Two consecutive engine errors latch the session;
every away wake then reaches main with one supervision-host line for a
five-minute cooldown, after which one wake probes the engine, and each
failed probe doubles the cooldown up to one hour. A reported turn without
an engine error clears it. The latch is kept per main session, engine, and
model in state/.supervision-host-health, and the engine conversation now
uses the same main-session key, which includes the lock holder's process
identity so a recycled pid never shares either.
Lifted from the validated 5631 tree and adapted to main's away-only host:
the attended recovery line and attended cooldown pass-through are left for
the attended core, so a recovery is only logged.
* no-mistakes(document): Consolidate supervision-host latch documentation
* fix(bin): absorb routine secondmate working and paused status appends (#5535)
* fix(bin): absorb routine second-mate progress while surfacing routed replies
Fixes #2959
A kind=secondmate task's status signal was never absorbable, so a healthy
mate's routine working: and paused: appends woke the primary every time.
signal_crew_provably_working now reads the mate's lines new since the
watcher's classified position: a decision, blocker, terminal outcome, note:,
correlation-marked line, or unknown verb still surfaces regardless of busy
evidence, while unmarked working:, paused:, and resolved: fall through to the
same provably-working absorb an ordinary crewmate gets.
* no-mistakes(review): narrow secondmate routine absorb to working and paused
* fix(bin): use gh-axi for the ship DoD draft check (#5519)
* fix(bin): use gh-axi for the ship DoD draft check
* no-mistakes(review): use PR number not URL in gh-axi draft check
* fix(bin): bind inactive-outcome receipt identity to structured fields only (#5520)
* fix(bin): drop status prose from the inactive-outcome dedupe identity
The inactive-outcome receipt fingerprint included the child's sanitized
last status line, so a persistent child appending routine prose after one
terminal outcome minted a fresh parent event per sentence. Bind the
identity to incarnation, task id, terminal state, and PR only, keeping
the last line in the record as status_head evidence.
Fixes #2960
* no-mistakes(document): note structured-only inactive receipt identity in regression coverage
* feat: record the Claude and Cursor dialog for the supervision host (#5707)
* feat(bin): record the supervision host's dialog mirror on Claude and Cursor
Add bin/fm-host-mirror.sh, the one owner of the supervision host's dialog
mirror file, cursor, lock, and feed, plus the main-session key it keys
entries to. The tracked Claude UserPromptSubmit and Stop hooks and the
Cursor beforeSubmitPrompt and afterAgentResponse hooks record the captain's
prompt and main's reply, only on a home with config/supervision-host, from
a genuine primary checkout, for the lock-owning session. The mirror lands
inert: writers record and nothing reads it yet; attended supervision on the
host is the later step that consumes the feed.
Codex, Grok, OpenCode, and omp have no writer here.
* no-mistakes(review): Scope mirror dedup to session, atomic appends, marker-inclusive caps
* no-mistakes(document): Clarify dialog mirror scope and remove duplicate contract details
* no-mistakes(document): Correct Cursor hook documentation for dialog mirror registration
* no-mistakes(review): Pass mirrored dialog text to jq via stdin
* no-mistakes(document): Clarify dialog mirror documentation and remove duplicate claims
* no-mistakes(review): Preserve internal dialog whitespace; drop mirror check and verified modes
* no-mistakes(review): Drop only identical mirror repeats; remove redundant chmod guard
* fix(bin): retire check-row receipts on branch acknowledgement so away escalations are not repeated (#5731)
* fix(bin): retire check-row receipts on branch acks and report an unchanged situation once
* fix(bin): scope a branch acknowledgement's check-row receipt retirement to
its granted sequences
The away posture lifts the attended partition's check/decision exclusions, so
a branch grant can name check-kind rows - but the branch-actor ack still
assumed check rows were main-only and skipped every receipt scan. The queue
row was consumed while its terminal-outcome .pending receipt stayed behind,
and each inactive-reconcile cadence scan re-queued the same fingerprint. In
the first real away window on the supervision host that re-escalated one
unchanged held-PR situation on every cycle (~1,734 of 4,149 outcomes).
A branch ack now scans inactive-outcome and inactive-reconcile receipts and
commits secondmate stall receipts against exactly the sequences in its
eligible-row snapshot - the same rows it consumes - instead of none. Attended
grants still name no check row, so the scans find nothing.
* fix(bin): store a repeated captain verdict as routine while the task's
durable situation is provably unchanged
fm-branch-outcome.sh append computes a mechanical situation key per captain
row - metadata bytes, captured status-log endpoint and identity, live
crew-state verb, worktree head - and anchors it in
state/.<task>.branch-captain-key. A later captain verdict whose recomputed key
matches is stored as routine with "unchanged since seq <N>:" prefixed to its
summary, so one situation escalates once until something provably changes. A
task with no readable status ledger is never demoted, an unreadable record
fails toward reporting, and teardown removes the sidecar with the task's
other branch records. The append-only store schema is unchanged.
This covers both hosts: the Pi supervision branch and the supervision host
both funnel reports through append.
* docs: check rows are main-owned only while attended; the away posture grants
them to the branch, whose ack retires their receipts exactly
* test: the away-flood reproduction as a regression test (branch ack retires
the receipt and later scans stay quiet), store-level dedupe coverage, and a
branch-ack secondmate stall receipt case
* fix(bin): restore the secondmate child devin-config cleanup path
The branch-captain-key sidecar addition mistyped the sibling entry as
.$child_id.devin-config.json, so a forced secondmate teardown would have
stopped removing each child's real <id>.devin-config.json. Restore the
original path and add a behavioral test that stops the child sweep mid-loop
on a refused close, proving the cleaned child's devin config and captain
anchor are both removed while the unconsumed child's records are retained.
* no-mistakes(review): Key captain dedupe on the covered wake rows' fingerprint
* no-mistakes(review): Drop captain-key demotion; prove one escalation on both surfaces
* no-mistakes(review): Drop unrelated teardown test; cite both receipt test files
* no-mistakes(document): Docs already match branch-ack check-receipt retirement
* fix(bin): deliver Claude-bound operational input as a record-backed doorbell (#5664)
* fix(calm): deliver Claude-bound operational input as a record-backed doorbell
Claude Code 2.1.280 removes U+2063 from every submitted prompt, so a typed
operational envelope reaches a Claude Code primary as plain text. The away
daemon now writes the envelope to a record under state/operational-inbox and
types only a plain doorbell naming it; the /afk return check and the Calm mod
recognize the doorbell only when that record holds a current envelope. Marker-
preserving harnesses keep the typed envelope. The live Calm guard accepts the
2.1.280 module-load log line, drives the doorbell, and asserts thinking stays
hidden.
* no-mistakes(review): Fix operational record retention at 7 days and document prune limit
* no-mistakes(document): Point Calm bounds at 2.1.280 evidence; fix afk-exit comment
* no-mistakes(lint): Pick newest Calm e2e transcript without parsing ls
* docs(calm): add a minimal turning-Calm-on step for Claude Code
* fix(spawn): deliver the Claude launch brief as a record-backed doorbell
Claude Code strips U+2063 from the launch-prompt argument too, so a
worker's launch brief arrived with its operational marker removed.
Publish the brief as a record in the receiving home's operational
inbox - a secondmate's own state, not the primary's - and pass only
the printable doorbell naming it, falling back to the typed envelope
when the record cannot be published so the brief body still delivers.
Unwrap doorbell-carried digests in the daemon digest tests that still
read the raw send log under the claude pin, and update the documented
bounds now that launch briefs hide like the other operational rows.
* test(spawn): cover a secondmate's launch-brief record landing in its own home
The record-backed doorbell resolves its state through the receiving
pane's home, so prove a claude secondmate launch publishes into the
seeded secondmate's operational inbox and never leaks a record into
the primary's.
* no-mistakes(review): Pass primary harness to daemon, tighten retention, refresh verdicts
* no-mistakes(review): Prune operational records by exact seven-day elapsed age
* no-mistakes(review): Batch record pruning so large inboxes still expire
* no-mistakes(review): Refuse Claude spawn when brief record cannot publish
* no-mistakes(review): Drop thinking probe from Claude Calm live test and docs
* no-mistakes(review): Record dated Claude Code 2.1.282 reproduction evidence
* no-mistakes(document): Clarify operational doorbell documentation and record expiry
* no-mistakes(document): Correct AFK escalation carrier guidance
* no-mistakes(review): Describe operational record retention as about seven days
* no-mistakes(document): Clarify Calm delivery and operational record retention
* no-mistakes(review): Remove out-of-scope Calm launch guide from Claude docs
* no-mistakes(document): Document Claude launch-brief delivery and refusal
* no-mistakes(document): Correct stale operational-input documentation
* no-mistakes(ci): Fixed the stale Claude trust test to verify that worker and secondmate launches deliver readable, record-backed briefs instead of expecting brief paths in their commands. Annotated the daemon’s output variable for ShellCheck without changing behavior. The affected tests, daemon tests, ShellCheck, and diff check pass locally
* no-mistakes(ci): parse rebased Claude launch after trailer hook prefix
* no-mistakes(review): Trust launch-brief record and restore thinking bound doc
* no-mistakes(review): Parse final Claude launch statement; drop Stop-hook docs
---------
Co-authored-by: Mike Sewell <maikunari@protonmail.com>
Co-authored-by: no-mistakes <no-mistakes@localhost>
* test: isolate lint fixture from tracked suite (#5727)
* fix(bin): republish parent metadata after a remote secondmate relaunch (#5583)
A host-local relaunch rewrote only the far endpoint, so this home kept the old harness, model, and effort, and appending those keys after pr= broke pull-request poll authentication.
* fix(bin): give slow watcher suites headroom under the changed-suite bound (#5516)
tests/fm-watch-triage.test.sh finishes in about 434s alone and about 698s
under CI load, so the 900s bound the changed-suite runner applies produced
a false timeout under ordinary concurrent validation. Raise the automatic
bound to 1500s, which keeps every measured script under it while staying
below the 30-minute normal CI tier so a genuinely hung script still fails
here with its output before the job cap cancels the lane.
Fixes #3869
Refs #3565
* fix(bin): stop nested steal-lock recursion and mid-steal watcher TERM (#5728)
* Fix nested watcher lock reclaim
* no-mistakes(review): Elect a single steal-mutex reaper and bound arm TERM wait
* no-mistakes(review): Reclaim self-held steal mutex and unify autoarm steal reaping
* no-mistakes(review): Resume own interrupted steal reap from its tombstone
* test: make supervision-host park-boundary tests deterministic (#5710)
* test: hold the back-to-back boundary close on the host's own clock
test_park_boundary_holds_under_back_to_back_closes assumed two engine
turns fit in the ~16s pre-refusal window and that the stub finished a
turn in 3s. Under load the stub's real drain, report, and
acknowledgement take ~13s, so the turn either died at its bound (which
hands the wake to main, no boundary line) or the second close landed
past the window and the fixture failed while the boundary held. 3
failures in 5 runs at a load average near 11.
Hold the first turn on a release file instead: once the engine is in
flight, a second close is appended mid-turn and the turn is released as
the refusal window opens (park bound minus turn bound and grace, read
off the host's own start record). The queued close can then only wait
for the boundary on any machine speed, which is what the test asserts:
the boundary line ends the output, the demo.status row stays queued for
main, and no second engine turn ever starts. A host too loaded to start
the turn at all hands the first close to the same boundary exit.
After: 12/12 at load ~15-42.
* no-mistakes(review): Print boundary test deadline as a decimal integer
* no-mistakes(review): Hold boundary test turn on a FIFO, require full sequence
* no-mistakes(review): Remove stray before/after supervision-host test copies
* test: hold the late close's render until the refusal window opens
The boundary recheck test's node shim slept a fixed 10s, which assumed
the first close was read before the host's refusal window opened. Under
load the close arrived after the refusal check, so the host correctly
refused it before the successor started and the render snapshot never
appeared. Block the wake-prompt render on a FIFO released at the
refusal-open instant read from the host's own start record, so the
pre-turn recheck must refuse on any machine speed.
* no-mistakes(review): Derive minimal park bounds and refresh supervision-host shard hint
* no-mistakes(review): Drive park-boundary tests from a seam-gated host test clock
* fix: stage remote home clones before publication (#5733)
* fix(bin): stage remote home clones before publishing them
A remote home provision cloned the code root directly into the public
FM_HOME path while rollback() claimed rm -rf of that same path on any
failure. Bash defers trapped signals past a foreground child, but any
other cleanup or lifecycle path that removes the home directory races
the live clone's object copy, producing the CI flake "fatal: failed to
copy file to .../.git/objects/...: No such file or directory".
Clone into a private staging directory beside the home and publish with
an atomic rename once complete, so no cleanup can remove a directory a
live clone is still writing; a home that appears mid-provision now dies
cleanly instead of inheriting torn state. The regression coverage holds
a real clone mid-copy, removes the public path, and requires the
provision to finish and publish intact.
* no-mistakes(review): Prove home ownership by sentinel and hold only a live clone
* no-mistakes(review): Assert raced provision publishes a complete, intact clone
* no-mistakes(document): Document remote home staging and publication safety
* no-mistakes(lint): Fix ShellCheck warning in clone integrity assertion
* no-mistakes(document): Clarify remote home publication and rollback guarantees
* fix: preserve Herdr status on Pi relaunch (#5161)
* fix(control): keep a relaunched Pi worker's herdr pane status authority alive
Defect: after `bin/fm-control.sh <id> relaunch` (observed live on a herdr
Pi crewmate whose pane read idle while it ran its validation pipeline),
the pane froze at whatever its previous agent had last reported.
Cause, measured on herdr 0.9.1 against a real Pi: a pane has one status
authority, and for Pi with its integration installed that authority is
the lifecycle hooks, so herdr also skips screen detection for the pane.
In the crew shape the registration outlives its agent process (upstream
issue #4115; docs/herdr-backend.md "Restart and liveness behavior"), and
herdr applies only reports carrying the session identity it bound. A
replacement started fresh in that pane reports a NEW session, so its
state reports are ignored and the pane stays frozen. Nothing from
outside repairs it: `pane report-agent-session` and `pane report-agent`
for `herdr:pi` are accepted (rc=0) without being applied unless the
reporter is the registered pane agent, and `pane release-agent` on the
stale record changes nothing.
Fix: a relaunch preserves the binding instead of fighting it. The launch
owner reads the session reference the endpoint's own runtime recorded
(`fm_backend_herdr_pane_agent_session_ref`) and passes it back as Pi's
own `--session <path-or-id>` (`relaunch_resume_args`;
`fm_control_relaunch_resume_flag` owns which adapters and which
registered-agent labels qualify). That is the same reference herdr
itself resumes Pi panes with after a server restart, and the resumed
session's reports land again, which the live check confirmed: the pane
returned to working while the replacement worked and idle when it
settled, on the same session identity.
Safety: relaunch-only (a fresh spawn binds nothing), herdr-only (the one
adapter that records a per-pane session), Pi-family only, and only when
the registration's own agent label matches - so no other adapter's
conversation can be handed to a Pi launch. An unreadable, missing, or
malformed reference degrades to exactly the fresh-session launch that
existed before. No lifecycle, liveness, isolation, or merge guard is
touched, and an empty result leaves every non-Pi launch byte-identical.
`resume` remains a refused verb; docs/agent-control.md and the
harness-adapters references are corrected where they claimed Pi had no
verified resume form at all.
* no-mistakes(document): docs: correct relaunch session-authority ownership and skill paths
* no-mistakes(document): docs: correct stale control-plane ownership claim
* no-mistakes(document): docs: drop unverified Herdr restart resume claim
* no-mistakes(test): Added offline Herdr Pi session-authority relaunch coverage
* no-mistakes(document): Document Herdr Pi relaunch session continuity
* no-mistakes(ci): The failing remote relaunch test tried to arm a PR poll for a secondmate, which `fm-pr-check.sh` correctly refuses. Removed that invalid test scenario; the remaining remote relaunch tests pass, and `git diff --check` is clean
* Seed the relaunch-ordering PR poll fixture without fm-pr-check.sh (#5758)
Main has been red since fm-pr-check.sh began refusing to arm a merge poll
on a kind=secondmate record (#5696): the relaunch-ordering case in
tests/fm-remote-secondmate-relaunch.test.sh armed its fixture through that
entry point and could no longer be set up.
The ordering guarantee still matters: a secondmate record armed before the
refusal can legitimately carry a trailing pr=/pr_head= identity block until
the watcher retires it, and fm-remote-secondmate-relaunch.sh must still keep
that block last when republishing harness/model/effort. Seed the fixture the
way such a record was really written - pr= appended last to the meta, then
the poll artifacts published throug…
… DevOps delivery (#15) * fix(bin): surface unrecognized status prefixes instead of reading them as silence (#5588) * fix(bin): surface an unrecognized status prefix instead of dropping it A parked or holding declaration, and a verb whose correlation token did not parse, never became an event, so the supervisor still saw the earlier line. * no-mistakes(review): Require verb-shaped unrecognized status prefixes, add continuation tests * no-mistakes(document): Document unrecognized status prefix escalation in afk skill * no-mistakes(ci): I reproduced the "Behavior portable serial 6" failure locally and fixed it by changing the test data in one test. No product code changed. **What failed:** `tests/fm-session-start.test.sh`, in `test_orphan_status_logs_are_printed`, with "matched status log was printed 2 times". **Why:** the test writes status lines with made-up prefixes, `matched: surfaced once` and `orphan: step N`. The test only uses them as placeholder text. It checks that the session-start digest prints each task's status tail exactly once. This PR (#4763) deliberately makes an unrecognized one-word lowercase prefix a status event. So those lines now surface as captain-relevant events, and the wake queue's STATUS OUTCOME BACKSTOP section prints them a second time. The code under review is behaving as the issue asks. Only the test's placeholder data had become meaningful. **Rule the test depends on:** its status lines must not be captain-relevant, so the digest is the only place they are printed. Both lines in this test broke that rule. The orphan line would have failed the same count check right after the matched line did. **Fix:** in that test only, I switched both lines to the recognized, non-captain verb `working:`: `working: surfaced once` and `working: orphan step 1..6`. I updated the matching assertions and counts to use the new text. What the test checks is unchanged: orphan logs are labelled, the tail is bounded, the log path is printed, and each tail appears once. **Verification:** before the fix, the test failed locally the same way as in CI. After it, `bash tests/fm-session-start.test.sh` reports "all assertions passed * no-mistakes(review): Detect unrecognized prefixes on unstamped lines; share verb list --------- Co-authored-by: Kun's firstmate <kunchenguid+firstmate@users.noreply.github.com> * fix(bin): report the failing item when remote inheritance fails (#5658) Fixes #5295 Session start now reports a remote inheritance failure using the push's own error line instead of the first unchanged item that happened to print before it, and the shared captain preferences header check now names the first required phrase it did not find, on both the local and remote inheritance paths. * fix(bin): stand down the Claude Stop auto-arm on pi-code-delivered payloads (#5657) * fix(bin): stand down the Claude Stop auto-arm on pi-code-delivered payloads pi-code loads the tracked Claude settings but has no asyncRewake, so it awaits every Stop hook; without a stand-down the auto-arm runs synchronously inside Pi's turn end and holds it open for the declared multi-hour timeout. Stand down when the payload's transcript_path contains a /.pi/ path component, the same discriminator the closed-but- unmerged fix in #3352 used, with an explicit string-type check on the jq filter. Fixes #3343 * no-mistakes(document): document pi-code stand-down in harness integrations reference * fix(bin): match whole multi-word project names in the registry lookup (#5659) * fix(bin): match whole multi-word project names in the registry lookup bin/fm-project-mode.sh matched a registered project name against only the first whitespace-delimited token of a registry row, so a name containing a space never matched, silently defaulting the project to no-mistakes off instead of its declared posture. The lookup now matches the whole registered name against the raw line text, so a name is compared literally (never as a regex) and a name that is a leading prefix of another registered name still resolves to its own row. * no-mistakes(document): docs already accurate for multiword registry name match * chore: drop accidental empty err file Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com> --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com> * fix(bin): classify shell stdin payloads after -s operands (#5546) * fix(bin): classify the stdin program of `bash -s` with operands in the arm policy With -s, sh/bash/zsh read the program from stdin even when operands follow; the operands are only positional parameters. The arm policy treated the first operand as a script path, so heredoc and here-string payloads were never classified and a hidden bin/fm-watch.sh execution was allowed. A protected path in the operand position still fails closed as before. Fixes #1489 * no-mistakes(document): Clarify stdin shell operand documentation * no-mistakes(ci): Captain, fixed `shellInvocation` so `bash -- -s` treats `-s` as a script name, and updated R21 to test the exact command. The targeted policy suite, lint, documentation check, and diff check pass. Both hosted workflows show `action_required` before any jobs ran; that external approval state remains unresolved * fix(bin): keep main's handling of words after a leading `--` Revert the pipeline CI-step change that made the first word after a leading `--` always a script. It turned forms that main denies today into allow (for example `bash -- -c 'bin/fm-watch.sh'`), which is outside #1489 and loosens a fail-closed policy. `--` after `-s` still ends option parsing. * fix(bin): strip AI co-author trailers from fleet-launched commits (#5695) * fix(bin): strip AI co-author trailers from fleet-launched commits Cursor and other non-Claude runtimes append the trailer after the typed message. A per-task commit-msg hook removes it and leaves human co-authors and the author identity untouched. * no-mistakes(review): Export pane hooksPath override and drop generated-with stripping * no-mistakes(ci): This PR caused all three CI failures, and the fix is test-only: 4 test files change, no product code. **Cause.** `fm-spawn.sh` now installs the AI-trailer strip hooks for every spawn, secondmates included. The installer refuses a worktree that is not a git repository, and the PR deliberately keeps that fail-closed rule because real secondmate homes are firstmate clones. Four test fixtures still gave secondmates a plain directory as their home, so each spawn failed with "not a git worktree ... could not install the AI-trailer strip hooks": - serial 5: `tests/fm-backlog-atomicity.test.sh` ("secondmate spawn failed"). - serial 8: `tests/fm-secondmate-harness.test.sh` ("split: no meta written"). - Herdr: `tests/fm-backend-herdr-launcher-workspace-e2e.test.sh` and `tests/fm-backend-herdr-workspace-per-home-e2e.test.sh`. **Rule that must hold.** Every home a test spawns as a secondmate must be a git worktree. I checked the other places in the changed area: the only secondmate spawns in these tests are the ones listed. The earlier rounds already fixed the other fixtures (`fm-secondmate-liveness`, `fm-secondmate-safety`) the same way. **Fix.** - Each of those four secondmate homes now gets the same `.gitignore` plus `git init -q -b main` that the liveness and safety tests already use. - The two Herdr tests clean up with their own plain `rm -rf "$TMP_ROOT"`, not the shared `tests/lib.sh` helper. Because the installer leaves each `state/<id>.git-hooks` directory read-only, that cleanup printed "Permission denied" and left the directories behind. Both cleanups now restore the owner's write bit on every directory before removing (`find ... -exec chmod u+rwx`), which is what `fm_test_remove_tree` in `tests/lib.sh` does. **Verification.** - `tests/fm-secondmate-harness.test.sh` passes. - `tests/fm-backlog-atomicity.test.sh` passes (99 ok, exit 0). - shellcheck is clean on all four files. - I could not run the two real-Herdr tests locally: the Herdr lab on this host refuses to start because it needs exactly one running default session, and I did not change the host's Herdr state to get around that. Instead I checked their two changed steps directly: the installer succeeds on a home set up the new way, and the new cleanup removes the read-only hooks directory completely. Those two tests will only be proven on CI * fix(bin): treat Pi's dollar-first cost footer as furniture (#5683) * fix(bin): treat Pi's dollar-first cost footer as furniture An idle Pi status row opening with $0.000 was read as a dead-shell prompt, so exit and relaunch refused on an empty composer. * test: wait for the draining holder to exec sleep before reading its identity The procevent drain fixture read fm_pid_identity immediately after backgrounding setsid sleep, racing the child's exec chain. Mid-exec the cmdline can read empty, failing the fixture on a loaded CI runner. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * fix(bin): refuse merges with unreported required checks (#5534) * fix(bin): refuse a merge when a required check never reported fm-pr-merge.sh built its GitHub refusals only from checks present in statusCheckRollup, so a required check that never ran was simply absent and the merge proceeded on the subset that reported, contradicting its own "every required check green" claim. The GitHub verify now reads the base branch's required contexts from the forge itself - the classic branch protection summary on GET repos/{o}/{r}/branches/{b} and the active ruleset rules on GET repos/{o}/{r}/rules/branches/{b} - and refuses when a required context has no entry in the same rollup, at the same head, that the merge is bound to. Absence reads as unknown, never green. The required-set read joins the existing refusal list, so a draft, a red check, and an unreported required check are all reported together. Could not read vs nothing required: both endpoints need only repository read access. The admin-only GET .../branches/{b}/protection endpoint is deliberately not used: it answers a non-admin token with the same 404 an unprotected branch gets (observed live on kunchenguid/firstmate main with this token), which would read a missing permission as "nothing required". Any failed or malformed read of either source (auth, missing fine-grained permission, rate limit, network, 404, unexpected shape) refuses the merge with a line naming the unreadable source. The one exception is GitHub's plan-gated 403 on the rules endpoint ("Upgrade to GitHub Pro or make this repository public"), which already means "this repository has no branch rules" for the merge-queue reader; that check moves into one shared helper and the classic summary still decides for such a repository. Attended waiver: --allow-missing <check-name> is the twin of --allow-red and follows the same design and recording path: once, separate name argument, waives only that exact unreported required check, still requires every other required check reported and every check green, never waives an unreadable required set, refused while the away-posture record exists, and refused on GitLab. Merge-state BLOCKED policy is unchanged. How this differs from the withdrawn #5353 (read from its diff): - #5353 read the admin-only branches/{b}/protection endpoint and treated its 404 as "no required checks", so for any non-admin token the required set silently read as empty; this change reads the read-access branch summary and treats every failure as unreadable. - #5353 ignored rulesets; this change also reads required_status_checks rules from the effective branch rules. - #5353 made separate per-head REST reads of statuses and check-runs capped at per_page=100 with no pagination; this change checks presence in the same statusCheckRollup view the red-check gate already reads at the verified head. - #5353 stopped at the first unreadable read; this change reports it as one refusal among all the others. - #5353 also claimed #5345 (lock stealing) and changed 39 files, most unrelated; this change is #5344 only. Live proof, read-only (a gh wrapper refused every merge and mutating call): - cli/cli#14474 (trunk requires 3 classic build contexts, none ran): refused, naming build (macos-latest), build (ubuntu-latest), build (windows-latest); with --allow-missing "build (macos-latest)" it still refused, naming the other two. - cli/cli#13665 (ran build (ubuntu-24.04-firewall) instead): refused, naming build (ubuntu-latest). - hashicorp/terraform#39262 (ruleset-required checks absent): refused, naming Code Consistency Checks, End-to-end Tests, Race Tests, Unit Tests. - cli/cli#14485 (all required reported and green): verified; the wrapper blocked the merge call and the pull request read back open. Fixes #5344 * fix(review): Preserve required-check producers and aggregate independent read failures * fix(document): Clarify required-check verification and waiver documentation * fix(bin): match an app-bound required commit status by name The producer-identity check resolved an app-bound required context only against check runs, so a required context that the required app reports as a commit status could never match and always read as "has not reported". A commit status carries no app id to compare, so an app-bound requirement that arrives as a status now matches by name, as before producer binding; check runs keep requiring the configured producer app. Live, read-only: hashicorp/terraform#39262 requires license/cla from integration 865473, reported green as a commit status by the CLA app. The previous head refused it as unreported; this head no longer does, while still naming the four required check runs that never ran there. Refs #5344 * fix(document): Clarify accepted commit-status producer verification limitation --------- Co-authored-by: firstmate-oss <firstmate-oss@kunchenguid.local> * fix: keep persistent secondmates out of landed-work cleanup (#5696) * fix(bin): never offer a persistent secondmate for teardown The return brief's "Landed, cleanup due" scan listed every state/*.meta record carrying a pr= and a merge-notified marker without regard to kind, so a secondmate record holding a relayed child's merged PR put the mate itself up for "bin/fm-teardown.sh <mate>" cleanup. A secondmate is a persistent worker, never landed work. - bin/fm-afk-return.sh: skip kind=secondmate in the landed-cleanup scan. - bin/fm-pr-check.sh: refuse to record pr= or arm a merge watch on a kind=secondmate record before any side effect; a PR reported on its routed status channel belongs to a task in the mate's own home, which arms its own watch. - bin/fm-watch.sh: a merged result from a poll already armed on a secondmate retires the poll silently - no merge outcome, marker, or wake. * no-mistakes(document): Document secondmate merge-watch and return-brief exclusions * no-mistakes(ci): CI failed in an unchanged watcher-shutdown test whose three-second wait was sensitive to runner load. Increased the wait for both state- and home-deletion cases without changing watcher behavior. The full fm-watch-arm suite passed locally; syntax and diff checks passed * fix: reject invalid X reply and follow-up arguments before posting (#5702) * fix(bin): refuse unknown dash-leading args in public-posting fm-x scripts fm-x-reply.sh collected any unrecognized argument into the positional pool and took the first one as the reply text, so an invocation like "fm-x-reply.sh <id> --followup --final <text>" posted the literal string "--final" to X and silently dropped the real text. Make argument parsing strict in every script that can post publicly: an unknown dash-leading argument, a dash-leading request_id/task id, a dash-leading option value, or a surplus positional now exits 2 with a usage error before any config load, outbox write, or network call. Reply text starting with '-' is still accepted via --text-file or stdin, and --help is honored wherever it appears instead of becoming text (a --help forwarded through fm-x-followup.sh would have counted as a posted follow-up and mutated the link). fm-x-link.sh and the fm-public-followup scripts already refuse unknown arguments; fm-x-poll.sh takes none. * no-mistakes(review): Refuse surplus follow-up text sources; drop post-ID help branches * no-mistakes(document): Clarify reply and follow-up argument usage * no-mistakes(review): Refuse dash-leading --text-file operands in fm-x-reply * no-mistakes(document): Correct follow-up argument parsing comment * no-mistakes(document): Document dismiss argument rejection in script header * fix: pause broken supervision-host sessions between engine probes (#5701) * feat(bin): latch the supervision host after repeated engine errors Rung 3c-1 of the PR 5631 re-cut: the host copies the Pi branch's broken-session policy. Two consecutive engine errors latch the session; every away wake then reaches main with one supervision-host line for a five-minute cooldown, after which one wake probes the engine, and each failed probe doubles the cooldown up to one hour. A reported turn without an engine error clears it. The latch is kept per main session, engine, and model in state/.supervision-host-health, and the engine conversation now uses the same main-session key, which includes the lock holder's process identity so a recycled pid never shares either. Lifted from the validated 5631 tree and adapted to main's away-only host: the attended recovery line and attended cooldown pass-through are left for the attended core, so a recovery is only logged. * no-mistakes(document): Consolidate supervision-host latch documentation * fix(bin): absorb routine secondmate working and paused status appends (#5535) * fix(bin): absorb routine second-mate progress while surfacing routed replies Fixes #2959 A kind=secondmate task's status signal was never absorbable, so a healthy mate's routine working: and paused: appends woke the primary every time. signal_crew_provably_working now reads the mate's lines new since the watcher's classified position: a decision, blocker, terminal outcome, note:, correlation-marked line, or unknown verb still surfaces regardless of busy evidence, while unmarked working:, paused:, and resolved: fall through to the same provably-working absorb an ordinary crewmate gets. * no-mistakes(review): narrow secondmate routine absorb to working and paused * fix(bin): use gh-axi for the ship DoD draft check (#5519) * fix(bin): use gh-axi for the ship DoD draft check * no-mistakes(review): use PR number not URL in gh-axi draft check * fix(bin): bind inactive-outcome receipt identity to structured fields only (#5520) * fix(bin): drop status prose from the inactive-outcome dedupe identity The inactive-outcome receipt fingerprint included the child's sanitized last status line, so a persistent child appending routine prose after one terminal outcome minted a fresh parent event per sentence. Bind the identity to incarnation, task id, terminal state, and PR only, keeping the last line in the record as status_head evidence. Fixes #2960 * no-mistakes(document): note structured-only inactive receipt identity in regression coverage * feat: record the Claude and Cursor dialog for the supervision host (#5707) * feat(bin): record the supervision host's dialog mirror on Claude and Cursor Add bin/fm-host-mirror.sh, the one owner of the supervision host's dialog mirror file, cursor, lock, and feed, plus the main-session key it keys entries to. The tracked Claude UserPromptSubmit and Stop hooks and the Cursor beforeSubmitPrompt and afterAgentResponse hooks record the captain's prompt and main's reply, only on a home with config/supervision-host, from a genuine primary checkout, for the lock-owning session. The mirror lands inert: writers record and nothing reads it yet; attended supervision on the host is the later step that consumes the feed. Codex, Grok, OpenCode, and omp have no writer here. * no-mistakes(review): Scope mirror dedup to session, atomic appends, marker-inclusive caps * no-mistakes(document): Clarify dialog mirror scope and remove duplicate contract details * no-mistakes(document): Correct Cursor hook documentation for dialog mirror registration * no-mistakes(review): Pass mirrored dialog text to jq via stdin * no-mistakes(document): Clarify dialog mirror documentation and remove duplicate claims * no-mistakes(review): Preserve internal dialog whitespace; drop mirror check and verified modes * no-mistakes(review): Drop only identical mirror repeats; remove redundant chmod guard * fix(bin): retire check-row receipts on branch acknowledgement so away escalations are not repeated (#5731) * fix(bin): retire check-row receipts on branch acks and report an unchanged situation once * fix(bin): scope a branch acknowledgement's check-row receipt retirement to its granted sequences The away posture lifts the attended partition's check/decision exclusions, so a branch grant can name check-kind rows - but the branch-actor ack still assumed check rows were main-only and skipped every receipt scan. The queue row was consumed while its terminal-outcome .pending receipt stayed behind, and each inactive-reconcile cadence scan re-queued the same fingerprint. In the first real away window on the supervision host that re-escalated one unchanged held-PR situation on every cycle (~1,734 of 4,149 outcomes). A branch ack now scans inactive-outcome and inactive-reconcile receipts and commits secondmate stall receipts against exactly the sequences in its eligible-row snapshot - the same rows it consumes - instead of none. Attended grants still name no check row, so the scans find nothing. * fix(bin): store a repeated captain verdict as routine while the task's durable situation is provably unchanged fm-branch-outcome.sh append computes a mechanical situation key per captain row - metadata bytes, captured status-log endpoint and identity, live crew-state verb, worktree head - and anchors it in state/.<task>.branch-captain-key. A later captain verdict whose recomputed key matches is stored as routine with "unchanged since seq <N>:" prefixed to its summary, so one situation escalates once until something provably changes. A task with no readable status ledger is never demoted, an unreadable record fails toward reporting, and teardown removes the sidecar with the task's other branch records. The append-only store schema is unchanged. This covers both hosts: the Pi supervision branch and the supervision host both funnel reports through append. * docs: check rows are main-owned only while attended; the away posture grants them to the branch, whose ack retires their receipts exactly * test: the away-flood reproduction as a regression test (branch ack retires the receipt and later scans stay quiet), store-level dedupe coverage, and a branch-ack secondmate stall receipt case * fix(bin): restore the secondmate child devin-config cleanup path The branch-captain-key sidecar addition mistyped the sibling entry as .$child_id.devin-config.json, so a forced secondmate teardown would have stopped removing each child's real <id>.devin-config.json. Restore the original path and add a behavioral test that stops the child sweep mid-loop on a refused close, proving the cleaned child's devin config and captain anchor are both removed while the unconsumed child's records are retained. * no-mistakes(review): Key captain dedupe on the covered wake rows' fingerprint * no-mistakes(review): Drop captain-key demotion; prove one escalation on both surfaces * no-mistakes(review): Drop unrelated teardown test; cite both receipt test files * no-mistakes(document): Docs already match branch-ack check-receipt retirement * fix(bin): deliver Claude-bound operational input as a record-backed doorbell (#5664) * fix(calm): deliver Claude-bound operational input as a record-backed doorbell Claude Code 2.1.280 removes U+2063 from every submitted prompt, so a typed operational envelope reaches a Claude Code primary as plain text. The away daemon now writes the envelope to a record under state/operational-inbox and types only a plain doorbell naming it; the /afk return check and the Calm mod recognize the doorbell only when that record holds a current envelope. Marker- preserving harnesses keep the typed envelope. The live Calm guard accepts the 2.1.280 module-load log line, drives the doorbell, and asserts thinking stays hidden. * no-mistakes(review): Fix operational record retention at 7 days and document prune limit * no-mistakes(document): Point Calm bounds at 2.1.280 evidence; fix afk-exit comment * no-mistakes(lint): Pick newest Calm e2e transcript without parsing ls * docs(calm): add a minimal turning-Calm-on step for Claude Code * fix(spawn): deliver the Claude launch brief as a record-backed doorbell Claude Code strips U+2063 from the launch-prompt argument too, so a worker's launch brief arrived with its operational marker removed. Publish the brief as a record in the receiving home's operational inbox - a secondmate's own state, not the primary's - and pass only the printable doorbell naming it, falling back to the typed envelope when the record cannot be published so the brief body still delivers. Unwrap doorbell-carried digests in the daemon digest tests that still read the raw send log under the claude pin, and update the documented bounds now that launch briefs hide like the other operational rows. * test(spawn): cover a secondmate's launch-brief record landing in its own home The record-backed doorbell resolves its state through the receiving pane's home, so prove a claude secondmate launch publishes into the seeded secondmate's operational inbox and never leaks a record into the primary's. * no-mistakes(review): Pass primary harness to daemon, tighten retention, refresh verdicts * no-mistakes(review): Prune operational records by exact seven-day elapsed age * no-mistakes(review): Batch record pruning so large inboxes still expire * no-mistakes(review): Refuse Claude spawn when brief record cannot publish * no-mistakes(review): Drop thinking probe from Claude Calm live test and docs * no-mistakes(review): Record dated Claude Code 2.1.282 reproduction evidence * no-mistakes(document): Clarify operational doorbell documentation and record expiry * no-mistakes(document): Correct AFK escalation carrier guidance * no-mistakes(review): Describe operational record retention as about seven days * no-mistakes(document): Clarify Calm delivery and operational record retention * no-mistakes(review): Remove out-of-scope Calm launch guide from Claude docs * no-mistakes(document): Document Claude launch-brief delivery and refusal * no-mistakes(document): Correct stale operational-input documentation * no-mistakes(ci): Fixed the stale Claude trust test to verify that worker and secondmate launches deliver readable, record-backed briefs instead of expecting brief paths in their commands. Annotated the daemon’s output variable for ShellCheck without changing behavior. The affected tests, daemon tests, ShellCheck, and diff check pass locally * no-mistakes(ci): parse rebased Claude launch after trailer hook prefix * no-mistakes(review): Trust launch-brief record and restore thinking bound doc * no-mistakes(review): Parse final Claude launch statement; drop Stop-hook docs --------- Co-authored-by: Mike Sewell <maikunari@protonmail.com> Co-authored-by: no-mistakes <no-mistakes@localhost> * test: isolate lint fixture from tracked suite (#5727) * fix(bin): republish parent metadata after a remote secondmate relaunch (#5583) A host-local relaunch rewrote only the far endpoint, so this home kept the old harness, model, and effort, and appending those keys after pr= broke pull-request poll authentication. * fix(bin): give slow watcher suites headroom under the changed-suite bound (#5516) tests/fm-watch-triage.test.sh finishes in about 434s alone and about 698s under CI load, so the 900s bound the changed-suite runner applies produced a false timeout under ordinary concurrent validation. Raise the automatic bound to 1500s, which keeps every measured script under it while staying below the 30-minute normal CI tier so a genuinely hung script still fails here with its output before the job cap cancels the lane. Fixes #3869 Refs #3565 * fix(bin): stop nested steal-lock recursion and mid-steal watcher TERM (#5728) * Fix nested watcher lock reclaim * no-mistakes(review): Elect a single steal-mutex reaper and bound arm TERM wait * no-mistakes(review): Reclaim self-held steal mutex and unify autoarm steal reaping * no-mistakes(review): Resume own interrupted steal reap from its tombstone * test: make supervision-host park-boundary tests deterministic (#5710) * test: hold the back-to-back boundary close on the host's own clock test_park_boundary_holds_under_back_to_back_closes assumed two engine turns fit in the ~16s pre-refusal window and that the stub finished a turn in 3s. Under load the stub's real drain, report, and acknowledgement take ~13s, so the turn either died at its bound (which hands the wake to main, no boundary line) or the second close landed past the window and the fixture failed while the boundary held. 3 failures in 5 runs at a load average near 11. Hold the first turn on a release file instead: once the engine is in flight, a second close is appended mid-turn and the turn is released as the refusal window opens (park bound minus turn bound and grace, read off the host's own start record). The queued close can then only wait for the boundary on any machine speed, which is what the test asserts: the boundary line ends the output, the demo.status row stays queued for main, and no second engine turn ever starts. A host too loaded to start the turn at all hands the first close to the same boundary exit. After: 12/12 at load ~15-42. * no-mistakes(review): Print boundary test deadline as a decimal integer * no-mistakes(review): Hold boundary test turn on a FIFO, require full sequence * no-mistakes(review): Remove stray before/after supervision-host test copies * test: hold the late close's render until the refusal window opens The boundary recheck test's node shim slept a fixed 10s, which assumed the first close was read before the host's refusal window opened. Under load the close arrived after the refusal check, so the host correctly refused it before the successor started and the render snapshot never appeared. Block the wake-prompt render on a FIFO released at the refusal-open instant read from the host's own start record, so the pre-turn recheck must refuse on any machine speed. * no-mistakes(review): Derive minimal park bounds and refresh supervision-host shard hint * no-mistakes(review): Drive park-boundary tests from a seam-gated host test clock * fix: stage remote home clones before publication (#5733) * fix(bin): stage remote home clones before publishing them A remote home provision cloned the code root directly into the public FM_HOME path while rollback() claimed rm -rf of that same path on any failure. Bash defers trapped signals past a foreground child, but any other cleanup or lifecycle path that removes the home directory races the live clone's object copy, producing the CI flake "fatal: failed to copy file to .../.git/objects/...: No such file or directory". Clone into a private staging directory beside the home and publish with an atomic rename once complete, so no cleanup can remove a directory a live clone is still writing; a home that appears mid-provision now dies cleanly instead of inheriting torn state. The regression coverage holds a real clone mid-copy, removes the public path, and requires the provision to finish and publish intact. * no-mistakes(review): Prove home ownership by sentinel and hold only a live clone * no-mistakes(review): Assert raced provision publishes a complete, intact clone * no-mistakes(document): Document remote home staging and publication safety * no-mistakes(lint): Fix ShellCheck warning in clone integrity assertion * no-mistakes(document): Clarify remote home publication and rollback guarantees * fix: preserve Herdr status on Pi relaunch (#5161) * fix(control): keep a relaunched Pi worker's herdr pane status authority alive Defect: after `bin/fm-control.sh <id> relaunch` (observed live on a herdr Pi crewmate whose pane read idle while it ran its validation pipeline), the pane froze at whatever its previous agent had last reported. Cause, measured on herdr 0.9.1 against a real Pi: a pane has one status authority, and for Pi with its integration installed that authority is the lifecycle hooks, so herdr also skips screen detection for the pane. In the crew shape the registration outlives its agent process (upstream issue #4115; docs/herdr-backend.md "Restart and liveness behavior"), and herdr applies only reports carrying the session identity it bound. A replacement started fresh in that pane reports a NEW session, so its state reports are ignored and the pane stays frozen. Nothing from outside repairs it: `pane report-agent-session` and `pane report-agent` for `herdr:pi` are accepted (rc=0) without being applied unless the reporter is the registered pane agent, and `pane release-agent` on the stale record changes nothing. Fix: a relaunch preserves the binding instead of fighting it. The launch owner reads the session reference the endpoint's own runtime recorded (`fm_backend_herdr_pane_agent_session_ref`) and passes it back as Pi's own `--session <path-or-id>` (`relaunch_resume_args`; `fm_control_relaunch_resume_flag` owns which adapters and which registered-agent labels qualify). That is the same reference herdr itself resumes Pi panes with after a server restart, and the resumed session's reports land again, which the live check confirmed: the pane returned to working while the replacement worked and idle when it settled, on the same session identity. Safety: relaunch-only (a fresh spawn binds nothing), herdr-only (the one adapter that records a per-pane session), Pi-family only, and only when the registration's own agent label matches - so no other adapter's conversation can be handed to a Pi launch. An unreadable, missing, or malformed reference degrades to exactly the fresh-session launch that existed before. No lifecycle, liveness, isolation, or merge guard is touched, and an empty result leaves every non-Pi launch byte-identical. `resume` remains a refused verb; docs/agent-control.md and the harness-adapters references are corrected where they claimed Pi had no verified resume form at all. * no-mistakes(document): docs: correct relaunch session-authority ownership and skill paths * no-mistakes(document): docs: correct stale control-plane ownership claim * no-mistakes(document): docs: drop unverified Herdr restart resume claim * no-mistakes(test): Added offline Herdr Pi session-authority relaunch coverage * no-mistakes(document): Document Herdr Pi relaunch session continuity * no-mistakes(ci): The failing remote relaunch test tried to arm a PR poll for a secondmate, which `fm-pr-check.sh` correctly refuses. Removed that invalid test scenario; the remaining remote relaunch tests pass, and `git diff --check` is clean * Seed the relaunch-ordering PR poll fixture without fm-pr-check.sh (#5758) Main has been red since fm-pr-check.sh began refusing to arm a merge poll on a kind=secondmate record (#5696): the relaunch-ordering case in tests/fm-remote-secondmate-relaunch.test.sh armed its fixture through that entry point and could no longer be set up. The ordering guarantee still matters: a secondmate record armed before the refusal can legitimately carry a trailing pr=/pr_head= identity block until the watcher retires it, and fm-remote-secondmate-relaunch.sh must still keep that block last when republishing harness/model/effort. Seed the fixture the way such a record was really written - pr= appended last to the meta, then the poll artifacts published through the same fm_pr_poll_prepare/fm_pr_poll_publish_prepared pair fm-pr-check.sh uses, a pattern tests/fm-pr-check-security.test.sh already follows - and drop the now-unused fake gh fixture. The #5696 refusal itself stays pinned by the security suite's secondmate-record case. * feat: add attended supervision for Claude and Cursor hosts (#5748) * feat: run attended supervision on the host for Claude and Cursor On a home opted into config/supervision-host with a Claude or Cursor primary, the supervision host now takes the attended wakes the Pi branch would take: routine outcomes stay off main, and a captain outcome wakes main once with a branch-outcome line and waits in the drain's new BRANCH OUTCOMES section until main acknowledges it with mark-processed. - The offer rule moves into branchOfferForWake, shared by the Pi watcher and the host through bin/fm-branch-dispatch.mjs offer. - The host feeds the dialog mirror at the head of each attended wake and passes a close through unchanged when it is main-only, the engine or a tool is missing, the primary has no verified mirror, the main session cannot be identified, or the session is cooling down. - The drain presents captain outcomes first, one line per task, never behind older routine outcomes, and collapses routine overflow into a count that is marked read. - The return advances the store's read cursor through the away window once the brief has rendered, so the first drain does not replay it. - The branch prompt's mirror wording is host-neutral, and the rule to report what main must act on as captain, once per unchanged situation, applies only to the attended posture on the host. * docs: record the attended supervision host live check * no-mistakes(review): Present pre-window unread outcomes and contiguous captain prefix * no-mistakes(review): Return brief presents every row it marks read * no-mistakes(review): Return brief lists every unread outcome in one list * no-mistakes(review): Keep return list in store order and gate cursor failures * no-mistakes(review): Make the drain the only branch-outcome presenter after return * no-mistakes(review): Gate return on drain outcome failures; byte-count outcome budgets * no-mistakes(review): Gate drain on projection failures; UTF-8-safe byte cuts * no-mistakes(review): Fail drain without jq; hand unreadable prompt mirror to main * no-mistakes(document): Correct supervision-host return and drain documentation * no-mistakes(review): Recheck attended offer at turn start; honest failed-drain brief * no-mistakes(document): Correct supervision-host posture and drain documentation * no-mistakes(document): Documentation remains accurate for attended supervision * fix(bin): dedup directed source expansions in fm-pending-reply-lib (#5753) Each '# shellcheck source=' directive makes ShellCheck's external-source traversal expand that library's whole transitive graph again at the site. fm-pending-reply-lib carried three directed lazy sources of fm-wake-lib and two of fm-parent-channel-lib on identical per-call re-source sites, so one file analysis peaked above 4 GiB and every caller (fm-watch, fm-teardown) inherited the multiplier - the root cause of the PR #5732 Lint 1 OOM kill. Keep the runtime '.' commands byte-identical: the lazy re-source under 'local STATE FM_WAKE_QUEUE FM_WAKE_QUEUE_LOCK' is real behavior. Drop the duplicate directives so each library expands once per unit, and drop the tmux/classify directives since classify already arrives through the kept fm-wake-lib expansion and no tmux symbol is referenced here. The directive above the lib-dir assignment is kept - it binds the bin/ prefix so the undirected sites still resolve without SC1091. Measured peak RSS, ShellCheck 0.11.0 -x on Linux arm64: bin/fm-pending-reply-lib.sh 4.06 GiB -> 1.96 GiB, zero findings * fix(bin): avoid bash 5.2 sibling $() in recovery mint and delivery log (#5773) * fix: split bash 5.2 sibling $() in recovery mint and delivery log Sibling command substitutions on one line can empty a recovery generation under bash 5.2 when a CHLD trap is set. Mint pid/epoch sequentially, refuse empty tokens before write, and clean delivery fields before printf. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: split bash 5.2 sibling $() in recovery mint and delivery log Sibling command substitutions on one line can empty a recovery generation under bash 5.2 when a CHLD trap is set. Mint pid/epoch sequentially, refuse empty tokens before write, and clean delivery fields before printf. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: keep recovery mint failure semantics after sibling $() split Remove the new pid/date refusal and grammar guard so a mint miss still yields a grammar-valid token and a durable wake row, matching accepted review intent. Drop the fake-failing-date case that locked in the refuse. Co-authored-by: Cursor <cursoragent@cursor.com> * no-mistakes(document): Point recovery-mint hazard comment at its regression test --------- Co-authored-by: Cursor <cursoragent@cursor.com> * fix(bin): name the recovery for a declined Claude imports dialog and stop calling Escape safe there (#5791) Fixes #4520 * fix(bin): report the newest status event in the voice status reader (#5790) Fixes #4756 The voice status reader in bin/fm_voice_records.py reports each worker's state from the last non-blank line of its status log. When a worker appends a status line and then a line of plain prose, the reader reported "note" with the prose line instead of the declared state, diverging from bin/fm-classify-lib.sh's shell scan. Scan back through the tail for the newest line whose prefix is a single lowercase verb-shaped word (letters and hyphens), and report that event's verb instead of always taking the last line. An unrecognised verb-shaped prefix still reports "note" rather than letting an earlier recognised line answer for it, and free text with no colon is skipped as prose. When the tail holds no such event, the last line is reported exactly as before. * fix(bin): gate a self-announcing tool's update-available report on a newer version (#5786) * fix(bin): stop reporting an already-installed version as an available update An update announcement named its version first ("current -> new"), so reading the first dotted number as the announced version compared the current version against itself and always looked newer. Read the last dotted number instead, and only report an available update when that announced version is newer than the newest installed copy found; when that version is already installed, report only PATH skew. Fixes #5151 * no-mistakes(document): docs: gate announce update-available report on newer-than-installed * fix(bin): pass the dispatch profile effort to OpenCode workers through their launch config (#5799) * fix(bin): pass the profile effort to OpenCode workers through their launch config The dispatch profile's effort axis was recorded in task metadata but never reached an OpenCode worker: the launch wrote only a permission grant into the config it constructs. OpenCode 1.18.32's config schema carries per-model reasoning effort as agent.<name>.variant, so the chosen effort is now merged into the same OPENCODE_CONFIG_CONTENT JSON as the default build agent's variant, keyed to the resolved model. With no effort chosen the launch stays byte-identical. Fixes #1373 * no-mistakes(review): gate OpenCode effort variant by model provider family * no-mistakes(document): docs(opencode): note provider-family gating for effort variant * fix: stop cancelled validation runs from reporting false failures (#5815) * fix: preserve cancellation as no verdict in crew state Reuse the green-delivery safeguard for cancelled CI monitors and permit a skipped rebase. Other cancelled outcomes and coarse ledger records use the existing unknown state. Four delivered-PR regressions failed before the fix and pass afterward. The isolated public resolver and fleet-summary tests prove that undelivered cancellation no longer creates a failure contradiction, while preserving historical records and the terminal_in_flight invariant. Evidence uses fixture no-mistakes responses, not a live daemon cancellation. Update the existing coarse cancellation assertion from failed to unknown because it encoded this defect; retain its newest-run precedence check. Full fm-crew-state suite and pinned lint pass. * fix(review): Verify PR disposition before reclassifying terminal validation runs * fix(test): Add captured cancellation replay coverage for resolver and fleet * fix(document): Clarify cancellation and terminal delivery documentation * fix: require declared waits for workers awaiting their own work (#5812) * fix: declare worker background and pipeline waits Require ship and scout workers to declare owned-work waits with the existing paused verb before ending a turn or waiting on a pipeline or long command. Keep the first-sight alert and existing liveness classification unchanged; subsequent inspection follows the existing long pause cadence. Validation: emitted brief regression failed before the instruction change and passes afterward. Public watcher/drain regressions cover the first alert, repeated wedge suppression, bounded rechecks, and undeclared idle alarms using isolated backend fixtures. Brief suite, pinned lint, Bash syntax, documentation inventory, and whitespace checks pass. No real worker harness was exercised for wait behavior. * fix(document): Clarify declared worker waits and documentation ownership * fix(ci): Captain, fixed the cadence test to age both the declaration and first-alert throttle while preserving declaration identity. Reproduced the CI failure using stable identity; corrected tests pass with both stable identity and native macOS behavior. Focused ShellCheck and diff checks pass. Production behavior is unchanged; Linux CI was not rerun locally * fix: bound ShellCheck to one canonical root per process (#5770) * fix(bin): bound each lint root in its own ShellCheck process CI job "Lint 1" died twice at about ten minutes because the two shard workers each packed about 110 canonical roots into one unbounded ShellCheck process, and a byte-weight rebalance moved the analysis-heavy fm-watch.sh into a partition with other heavy roots, so the pair outgrew the 16 GiB runner before anything could name a culprit. Run one canonical root per ShellCheck process under an enforced envelope: a wall deadline plus terminate-then-kill grace via the shared fm-timeout-lib.sh watchdog, and a per-root rlimit spec applied inside the child before exec (default a 4 GiB address-space cap, so two workers stay inside a 16 GiB job with headroom). A root that exceeds the envelope fails by name with a recorded reason - timeout, memory, signal, or limit-unavailable - instead of taking the runner down. The per-root watchdog runs in its own process group so the owner's group sweep cannot orphan the bounded subtree, and fm_exec_timed now starts the same escalation when its parent dies before it can be signalled. FM_LINT_REQUIRE_BOUNDS=1, set in CI, refuses the run outright when a configured bound cannot be enforced on the host rather than lint uncapped. Each root's begin/end, reason, duration, and peak RSS stream to stderr in partition mode and append to a retained <telemetry>.roots.tsv sidecar uploaded beside the partition telemetry. Coverage is unchanged: pinned ShellCheck 0.11.0, --norc, --external-sources full analysis, complete and disjoint partition inventory, workflow lint, and the backend-purity check, with byte-identical diagnostics across jobs=1/2 proven by tests/fm-lint.test.sh. * fix(bin): fail closed on unenforceable lint bounds and size the cap Required-bounds mode (FM_LINT_REQUIRE_BOUNDS=1, set by CI) now refuses the run with named errors before any root starts: a missing fm-timeout-lib.sh, a watchdog that cannot actually bound a probe command, or a host that rejects the address-space limit all stop the run rather than lint uncapped. The generalized FM_LINT_ROOT_RLIMITS flag:value interface is replaced by a single FM_LINT_ROOT_MEMORY_KIB, and the roots sidecar and telemetry record the run's final exit status after backend-purity and workflow checks instead of the pre-check lint status. The default cap is 6 GiB of address space per root, not 4 GiB: ulimit -v bounds virtual address space rather than resident memory, and ShellCheck's GHC runtime keeps roughly a third of that space as reservation, so 6 GiB yields about a 4 GiB working heap budget. A Linux measurement during this change showed eleven real canonical roots running out of memory under the earlier 4 GiB cap while the largest passing root peaked near 2.8 GiB resident; two 6 GiB roots plus runner overhead still fit the 16 GiB job. Roots that still exceed the cap keep failing by name, and the sidecar's per-root peak RSS keeps roots approaching the budget visible. tests/fm-lint.test.sh now proves the memory primitive where it can be proven: on hosts that accept ulimit -v a perl allocator is refused under a 256 MiB limit and reported by name as a memory death, the pinned ShellCheck lints a small file under the configured cap and is named when a far smaller cap binds it, and a watchdog-less copy refuses under REQUIRE_BOUNDS; the bounded cases skip on macOS, which cannot enforce the address-space limit. * no-mistakes(review): Prove memory cap binds, pass watchdog owner, drop unused modes * no-mistakes(review): Capture watchdog owner before startup for every fm_exec_timed caller * no-mistakes(document): Clarify bounded lint documentation and telemetry * no-mistakes(document): Correct bounded lint documentation and sidecar path * docs(bin): restore the per-root memory cap sizing rationale The pipeline's document step rewrote the ROOT_MEMORY_KIB comment and dropped the sizing reasoning the change is required to record: address space vs resident memory, the GHC reservation share, the measured 4 GiB failures and ~2.8 GiB peak, and the two-roots-plus-runner capacity arithmetic. Restore it beside the default while keeping the corrected "not a resident-memory ceiling" framing. * no-mistakes(review): Document memory cap RSS reduction threshold and first candidate * no-mistakes(review): Scope owner-death escalation docs to the perl watchdog * no-mistakes(document): Clarify bounded lint and timeout documentation * no-mistakes(review): Install perl watchdog signal handlers before forking the command * no-mistakes(document): Correct bounded lint documentation and stale watcher comments * no-mistakes(ci): Fixed the supervision-host test’s obsolete expectation: the watchdog now reaps an engine when its host dies. The timeout and supervision-host tests pass locally; the watcher test also passes locally. Lint 1 and 2 remain unresolved: seven canonical roots exceeded the required 6 GiB address-space cap in CI. I did not raise the cap, exempt roots, or reduce source-following coverage to make those failures disappear * no-mistakes(ci): The two lint checks failed when eight canonical roots hit the enforced memory cap. I reduced repeated ShellCheck source-graph expansion while keeping runtime imports and the canonical root inventory intact. Pinned ShellCheck passes for all changed roots; the relevant local tests pass. The 6 GiB Linux CI run remains unverified * no-mistakes(review): Restore source directives, raise cap to 8 GiB, classify OOM * no-mistakes(review): Classify memory deaths from root stderr, not source excerpts * no-mistakes(review): Match only whole runtime memory-error lines for memory reason * no-mistakes(document): Clarify lint memory classification in script documentation * no-mistakes(ci): Fixed both lint checks’ memory-limit failure: each CI lint job now runs one root at a time with a 12 GiB address-space cap. Kept the local two-worker default and updated the sizing comment and test expectation. The lint tests and workflow validation pass locally; Linux CI remains to confirm the heavy roots * fix: bound watcher cleanup wait on the downtime-marker lock (#5732) * fix(bin): bound the watcher cleanup marker-lock wait tests/fm-watch-triage.test.sh intermittently failed serial CI shard 1 with "watcher pid <pid> did not exit within 10s of TERM". The watcher had processed the TERM and was inside watcher_cleanup, where the recovery-marker publish waits on state/.watcher-down.lock through an unbounded fm_lock_acquire_wait. A live foreign holder of that lock leaves the TERM'd watcher spinning in its own EXIT trap until the lock frees or a second signal short-circuits the trap. fm_recovery_transition now takes an optional bound and both release-lock paths plus publish honour it through a new in-process fm_lock_acquire_wait_max. watcher_cleanup passes FM_WATCHER_CLEANUP_LOCK_BOUND (default 2s); on timeout the publish is skipped, the singleton stays behind as ordinary dead-pid evidence, and the next arm's clear-stale-lock still republishes it. Regression test drives a real watcher with .watcher-down.lock held by a live foreign process and asserts a single TERM still stops it. * no-mistakes(review): Parse watcher cleanup lock bound as decimal, zero defaults * no-mistakes(review): Pin cleanup bound tests to observed marker-lock contention * no-mistakes(document): Document bounded watcher cleanup and recovery * no-mistakes(document): Clarify bounded watcher cleanup and recovery documentation * no-mistakes: apply agent fixes * no-mistakes(review): Arm marker-lock FIFO before TERM; drop FM_TEST_ONLY_LATE * no-mistakes(review): Hold marker lock through a failed cleanup acquire * test: stop the remote secondmate e2e watcher before temp-root cleanup (#5845) * test: stop the leaked unreachable watcher before remote e2e cleanup The remote secondmate lifecycle e2e backgrounded fm-watch.sh through the remote_env shell function, so $! named the function's subshell rather than the watcher. Killing that subshell left the unreachable-leg watcher running, and its one-second liveness probe kept invoking the fake ssh, which rewrites ssh.count in the temp root. When a probe landed while the EXIT trap was removing the root, rm failed with "Directory not empty" after every assertion had passed. Exec the watcher from the backgrounded function so the recorded pid is the watcher itself, and assert the stopped watcher stops probing and writing its state. Cleanup also stops a watcher left running by a failed assertion and removes the root through fm_test_remove_tree, so a run that fails before retirement does not strand the read-only spawn hooks directory. Closes #5836 * no-mistakes(review): Clear reaped watcher PIDs and restore plain temp-root removal * no-mistakes(review): Let in-flight probe settle before stopped-watcher baseline * fix(bin): stop reporting untouched shared-captain copies as drift (#4806) * fix: stop quarantining ordinary shared-captain source updates * no-mistakes(document): Rewrap remote inherit header so usage prints fully * docs: restructure calm.md for readability (#5604) * docs: make calm easier to read Restructure the Calm mode prose into sections, lists, and tables without changing documented behavior. Every original heading, anchor, fenced code block, inline-code span, link target, number, and quoted string is preserved. * docs: restore reload case in calm override lead-in The Built-in tool override collisions lead-in covers a session that reloads with Calm already on, as the original text did. * docs: restructure turnend-guard.md for readability (#5611) * docs: make turnend-guard easier to read Restructure the turn-end guard doc's prose into shorter sections, lists, and tables without changing documented behavior. Every original heading, anchor, inline identifier, link target, and number is kept. * no-mistakes(review): Restore legacy-only scope on TERM retirement sentence * no-mistakes(review): Name Cursor park behavior in live e2e test line * docs: move situational AGENTS.md sections into on-demand skills (#5872) * docs: move situational AGENTS.md sections into on-demand skills Backpass memory optimization: shrink the always-loaded AGENTS.md by moving situational contracts (home layout, session-start recovery, validation and landing supervision, scout completion, away/quiet supervision, Relay ownership) into agent-only skills loaded at their triggers, with a trigger index skill. * docs: classify the new on-demand skills' documentation audience Register the seven new agent-only skills as agent-runtime docs and fix a link in validation-supervision that kept its AGENTS.md-relative path. * docs: close load-timing gaps found by the live regression check - load validation-supervision whenever an ask-user finding is decided or answered, so forbid --yes and process-every-return reach the worker - keep the mid-task captain-ask rule, the unconfirmed network-checks rule, and the worker account pin rule inline in AGENTS.md - fix cross-references that still pointed at moved AGENTS.md sections * fix: route second-mate signal wakes by presented status span (#5879) * fix: route second-mate signal wakes by their new status span A second mate's status log is a shared channel carrying many independently keyed decisions, so judging its signal rows by every decision still open in the whole log pinned each routine update to main behind any unrelated parked hold. scopeForUnreadWake (the one owner for Pi and the attended supervision host) now judges a second-mate signal row by the lines presented since the last drain, bounded by the existing status-presentation cursor: a decision, blocked, resolution, or captain-held line, or a line declaring the key of a still-open decision, keeps the whole row on main, and any cursor problem falls back to the whole log. Keys are read only at the status parser's declared positions, with readable time stamps stripped as bin/fm-classify-lib.sh does. Single-task crewmate and stale routing are unchanged, and stale and signal rows for one mate keep independent verdicts. The supervision branch now treats a second mate's done and merged lines as relayed child outcomes, and fm-teardown refuses the branch actor second-mate retirement through the existing role-partition helper in both postures. * no-mistakes(review): Route second-mate resolutions to main only when closing open decision * no-mistakes(review): Guard bare-verb fold lines and fold resolution spans incrementally * no-mistakes(document): Clarify second-mate wake routing and retirement documentation * no-mistakes(ci): The CI failure was a timing-sensitive watcher teardown test, not the PR’s signal-scope code. Extended the bounded wait for both state-directory and home removal. The watcher suite passes locally, and git diff --check is clean * fix(bin): take the source lock before the lifecycle lock in register-extension (#5882) register-extension took the extension lifecycle lock and then the source lock, while reconcile republishing an unhandled extension result holds the source lock and reaches the lifecycle lock through the extension host's process-event path. Both waits are unbounded and both owners stay alive, so the two could wait on each other forever and freeze the home's monitoring cycle. register-extension now takes the source lock first, matching every other path that holds both. The lifecycle lock still spans binding resolution through registration publication, so binding retirement stays serialized. A new lifecycle-order section in the extension-binding suite, run in the default aggregate, holds a re-registration inside binding resolution while reconcile republishes that source's unhandled result and requires both to finish within a bound. Fixes #5866 * fix: chain repository hooks under git -c overrides (#5877) * fix(bin): run the repository's own hooks when git -c carries the per-task hooksPath The per-task hooks wrappers cleared only the GIT_CONFIG_COUNT override before looking up the repository's own hooks directory. When core.hooksPath reached git through git -c (GIT_CONFIG_PARAMETERS), directly or inherited by a child process, the lookup found the wrapper directory again and exited 0, so the repository's real hook - such as a pre-push publish guard - never ran and the push succeeded. The lookup now ignores GIT_CONFIG_PARAMETERS too, so only the repository's config files decide its hooks directory, and a failed lookup exits nonzero instead of skipping the hook. AI-trailer stripping is unchanged. Fixes #5871 * no-mistakes(document): Document git hook chaining and lookup failure behavior * fix(bin): withhold never-send values from dispatch resolver requests (#5744) * feat(bin): add an optional never-send list to typed dispatch resolution * Added config/dispatch-never-send, an optional local list of literal values and re: regular expressions checked against every string of the resolver request before it is sent to typesafe.ai * A match, an unreadable list, or an empty or invalid pattern now stops the request and falls back to the off path, so firstmate dispatches through its existing intake; the one stderr diagnostic names at most the list line number and never the value * No list, or a list with no match, leaves resolution unchanged * no-mistakes(review): Match never-send literals across whitespace, drop regex mode * no-mistakes(review): Inherit the never-send list into secondmate homes * no-mistakes(ci): ci-2 (Lint 2), caused by this PR and now fixed. The rule that broke: the test script must not share shell variables with a library it sources. The new secondmate-inheritance test in tests/fm-dispatch-resolve.test.sh sourced bin/fm-config-inherit-lib.sh inside a `( ... )` subshell. That lib assigns `out`, so ShellCheck flagged every later `$out` in the test with SC2031. The subshell was the only place this PR sources that lib. The fix runs the propagation in a child shell instead (`bash -c '. "$1" && propagate_inheritable_config "$2" "$3"' _ lib from to`), so the test shell never sources the lib. `bin/fm-lint.sh tests/fm-dispatch-resolve.test.sh` now exits 0, and `bash tests/fm-dispatch-resolve.test.sh` passes. That includes the check that an inherited list is enforced in a secondmate home. ci-1 (Behavior portable serial 8), not caused by this PR, so no code change for it. The only failing test is tests/fm-remote-secondmate-relaunch.test.sh, which fails with "not ok - could not arm the PR poll fixture for the relaunch-ordering test". This PR doesn't touch that test or the code it runs. I reproduced the same failure locally on the base commit ea7c7f7. Main's own CI run on ea7c7f7 (run 36212602588) fails only this job, with the same message. The recent main runs before it also concluded failure * feat(jev): add the guard framework contract and the memory RSS/swap thrashing guard (#5903) * feat(jev): add the guard framework and the memory RSS/swap thrashing guard A Jev guard is a bounded read-only host diagnostic that turns one class of resource pressure into a machine-readable audit record and a one-line verdict. This lands the framework contract (docs/jev-guards.md) with one representative family (Pattern 46, memory RSS and swap thrashing) as the pilot: thin bash wrapper, stdlib-only python engine, and a behavioral test through the CLI. * no-mistakes(review): fix jev mem guard fail-open unknown and contract * no-mistakes(review): fix mem guard rss pairing, memfree, tests, docs * no-mistakes(review): fix inverted UNKNOWN branch in fail-forcing test * no-mistakes(review): register docs/jev-guards.md in audience inventory * no-mistakes(review): simplify wrapper, drop dead top_n, single-source thresholds * no-mistakes(review): assert exact exit code in fail-forcing test leg * no-mistakes(review): tolerate any stdout encoding in text output * no-mistakes(document): Fix guard contract dash style and output wording * fix: prevent contribution poll starvation on slow GitHub reads (#5900) * fix(bin): stop slow GitHub reads from starving and waking the contributions poll The poll's fixed budget cut the tail URL's reads short, and any read killed at the bound was recorded as a forge failure, so routine GitHub slowness produced an observation-unavailable wake. A configured FM_CONTRIBUTIONS_BUDGET now rides the generated check shim and is cut down to the watcher's per-check bound with a margin, a read killed at the bound or the deadline is budget refusal that leaves records untouched, and the poll moves to the next URL that still has a full observation reserve instead of ending. * fix(bin): report the bound when a signal death leaks through fm_run_timed fm_run_external_timeout trusted the wrapper's recorded status over the runner's own bound verdict, so a read killed by the bound's TERM could surface as 143 and callers such as the contributions poll classified a budget-bound read as a forge failure. When the runner reports the bound, a signal-death status is the bound's own TERM racing the wrapper's bookkeeping and is reported as 124; a natural exit still passes through. The replace-shell probe also moves to the repo's portable BASHPID fallbac…
kunchenguid#5583) A host-local relaunch rewrote only the far endpoint, so this home kept the old harness, model, and effort, and appending those keys after pr= broke pull-request poll authentication.
…ure DevOps delivery (#16) * fix(bin): surface unrecognized status prefixes instead of reading them as silence (#5588) * fix(bin): surface an unrecognized status prefix instead of dropping it A parked or holding declaration, and a verb whose correlation token did not parse, never became an event, so the supervisor still saw the earlier line. * no-mistakes(review): Require verb-shaped unrecognized status prefixes, add continuation tests * no-mistakes(document): Document unrecognized status prefix escalation in afk skill * no-mistakes(ci): I reproduced the "Behavior portable serial 6" failure locally and fixed it by changing the test data in one test. No product code changed. **What failed:** `tests/fm-session-start.test.sh`, in `test_orphan_status_logs_are_printed`, with "matched status log was printed 2 times". **Why:** the test writes status lines with made-up prefixes, `matched: surfaced once` and `orphan: step N`. The test only uses them as placeholder text. It checks that the session-start digest prints each task's status tail exactly once. This PR (#4763) deliberately makes an unrecognized one-word lowercase prefix a status event. So those lines now surface as captain-relevant events, and the wake queue's STATUS OUTCOME BACKSTOP section prints them a second time. The code under review is behaving as the issue asks. Only the test's placeholder data had become meaningful. **Rule the test depends on:** its status lines must not be captain-relevant, so the digest is the only place they are printed. Both lines in this test broke that rule. The orphan line would have failed the same count check right after the matched line did. **Fix:** in that test only, I switched both lines to the recognized, non-captain verb `working:`: `working: surfaced once` and `working: orphan step 1..6`. I updated the matching assertions and counts to use the new text. What the test checks is unchanged: orphan logs are labelled, the tail is bounded, the log path is printed, and each tail appears once. **Verification:** before the fix, the test failed locally the same way as in CI. After it, `bash tests/fm-session-start.test.sh` reports "all assertions passed * no-mistakes(review): Detect unrecognized prefixes on unstamped lines; share verb list --------- Co-authored-by: Kun's firstmate <kunchenguid+firstmate@users.noreply.github.com> * fix(bin): report the failing item when remote inheritance fails (#5658) Fixes #5295 Session start now reports a remote inheritance failure using the push's own error line instead of the first unchanged item that happened to print before it, and the shared captain preferences header check now names the first required phrase it did not find, on both the local and remote inheritance paths. * fix(bin): stand down the Claude Stop auto-arm on pi-code-delivered payloads (#5657) * fix(bin): stand down the Claude Stop auto-arm on pi-code-delivered payloads pi-code loads the tracked Claude settings but has no asyncRewake, so it awaits every Stop hook; without a stand-down the auto-arm runs synchronously inside Pi's turn end and holds it open for the declared multi-hour timeout. Stand down when the payload's transcript_path contains a /.pi/ path component, the same discriminator the closed-but- unmerged fix in #3352 used, with an explicit string-type check on the jq filter. Fixes #3343 * no-mistakes(document): document pi-code stand-down in harness integrations reference * fix(bin): match whole multi-word project names in the registry lookup (#5659) * fix(bin): match whole multi-word project names in the registry lookup bin/fm-project-mode.sh matched a registered project name against only the first whitespace-delimited token of a registry row, so a name containing a space never matched, silently defaulting the project to no-mistakes off instead of its declared posture. The lookup now matches the whole registered name against the raw line text, so a name is compared literally (never as a regex) and a name that is a leading prefix of another registered name still resolves to its own row. * no-mistakes(document): docs already accurate for multiword registry name match * chore: drop accidental empty err file Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com> --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com> * fix(bin): classify shell stdin payloads after -s operands (#5546) * fix(bin): classify the stdin program of `bash -s` with operands in the arm policy With -s, sh/bash/zsh read the program from stdin even when operands follow; the operands are only positional parameters. The arm policy treated the first operand as a script path, so heredoc and here-string payloads were never classified and a hidden bin/fm-watch.sh execution was allowed. A protected path in the operand position still fails closed as before. Fixes #1489 * no-mistakes(document): Clarify stdin shell operand documentation * no-mistakes(ci): Captain, fixed `shellInvocation` so `bash -- -s` treats `-s` as a script name, and updated R21 to test the exact command. The targeted policy suite, lint, documentation check, and diff check pass. Both hosted workflows show `action_required` before any jobs ran; that external approval state remains unresolved * fix(bin): keep main's handling of words after a leading `--` Revert the pipeline CI-step change that made the first word after a leading `--` always a script. It turned forms that main denies today into allow (for example `bash -- -c 'bin/fm-watch.sh'`), which is outside #1489 and loosens a fail-closed policy. `--` after `-s` still ends option parsing. * fix(bin): strip AI co-author trailers from fleet-launched commits (#5695) * fix(bin): strip AI co-author trailers from fleet-launched commits Cursor and other non-Claude runtimes append the trailer after the typed message. A per-task commit-msg hook removes it and leaves human co-authors and the author identity untouched. * no-mistakes(review): Export pane hooksPath override and drop generated-with stripping * no-mistakes(ci): This PR caused all three CI failures, and the fix is test-only: 4 test files change, no product code. **Cause.** `fm-spawn.sh` now installs the AI-trailer strip hooks for every spawn, secondmates included. The installer refuses a worktree that is not a git repository, and the PR deliberately keeps that fail-closed rule because real secondmate homes are firstmate clones. Four test fixtures still gave secondmates a plain directory as their home, so each spawn failed with "not a git worktree ... could not install the AI-trailer strip hooks": - serial 5: `tests/fm-backlog-atomicity.test.sh` ("secondmate spawn failed"). - serial 8: `tests/fm-secondmate-harness.test.sh` ("split: no meta written"). - Herdr: `tests/fm-backend-herdr-launcher-workspace-e2e.test.sh` and `tests/fm-backend-herdr-workspace-per-home-e2e.test.sh`. **Rule that must hold.** Every home a test spawns as a secondmate must be a git worktree. I checked the other places in the changed area: the only secondmate spawns in these tests are the ones listed. The earlier rounds already fixed the other fixtures (`fm-secondmate-liveness`, `fm-secondmate-safety`) the same way. **Fix.** - Each of those four secondmate homes now gets the same `.gitignore` plus `git init -q -b main` that the liveness and safety tests already use. - The two Herdr tests clean up with their own plain `rm -rf "$TMP_ROOT"`, not the shared `tests/lib.sh` helper. Because the installer leaves each `state/<id>.git-hooks` directory read-only, that cleanup printed "Permission denied" and left the directories behind. Both cleanups now restore the owner's write bit on every directory before removing (`find ... -exec chmod u+rwx`), which is what `fm_test_remove_tree` in `tests/lib.sh` does. **Verification.** - `tests/fm-secondmate-harness.test.sh` passes. - `tests/fm-backlog-atomicity.test.sh` passes (99 ok, exit 0). - shellcheck is clean on all four files. - I could not run the two real-Herdr tests locally: the Herdr lab on this host refuses to start because it needs exactly one running default session, and I did not change the host's Herdr state to get around that. Instead I checked their two changed steps directly: the installer succeeds on a home set up the new way, and the new cleanup removes the read-only hooks directory completely. Those two tests will only be proven on CI * fix(bin): treat Pi's dollar-first cost footer as furniture (#5683) * fix(bin): treat Pi's dollar-first cost footer as furniture An idle Pi status row opening with $0.000 was read as a dead-shell prompt, so exit and relaunch refused on an empty composer. * test: wait for the draining holder to exec sleep before reading its identity The procevent drain fixture read fm_pid_identity immediately after backgrounding setsid sleep, racing the child's exec chain. Mid-exec the cmdline can read empty, failing the fixture on a loaded CI runner. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * fix(bin): refuse merges with unreported required checks (#5534) * fix(bin): refuse a merge when a required check never reported fm-pr-merge.sh built its GitHub refusals only from checks present in statusCheckRollup, so a required check that never ran was simply absent and the merge proceeded on the subset that reported, contradicting its own "every required check green" claim. The GitHub verify now reads the base branch's required contexts from the forge itself - the classic branch protection summary on GET repos/{o}/{r}/branches/{b} and the active ruleset rules on GET repos/{o}/{r}/rules/branches/{b} - and refuses when a required context has no entry in the same rollup, at the same head, that the merge is bound to. Absence reads as unknown, never green. The required-set read joins the existing refusal list, so a draft, a red check, and an unreported required check are all reported together. Could not read vs nothing required: both endpoints need only repository read access. The admin-only GET .../branches/{b}/protection endpoint is deliberately not used: it answers a non-admin token with the same 404 an unprotected branch gets (observed live on kunchenguid/firstmate main with this token), which would read a missing permission as "nothing required". Any failed or malformed read of either source (auth, missing fine-grained permission, rate limit, network, 404, unexpected shape) refuses the merge with a line naming the unreadable source. The one exception is GitHub's plan-gated 403 on the rules endpoint ("Upgrade to GitHub Pro or make this repository public"), which already means "this repository has no branch rules" for the merge-queue reader; that check moves into one shared helper and the classic summary still decides for such a repository. Attended waiver: --allow-missing <check-name> is the twin of --allow-red and follows the same design and recording path: once, separate name argument, waives only that exact unreported required check, still requires every other required check reported and every check green, never waives an unreadable required set, refused while the away-posture record exists, and refused on GitLab. Merge-state BLOCKED policy is unchanged. How this differs from the withdrawn #5353 (read from its diff): - #5353 read the admin-only branches/{b}/protection endpoint and treated its 404 as "no required checks", so for any non-admin token the required set silently read as empty; this change reads the read-access branch summary and treats every failure as unreadable. - #5353 ignored rulesets; this change also reads required_status_checks rules from the effective branch rules. - #5353 made separate per-head REST reads of statuses and check-runs capped at per_page=100 with no pagination; this change checks presence in the same statusCheckRollup view the red-check gate already reads at the verified head. - #5353 stopped at the first unreadable read; this change reports it as one refusal among all the others. - #5353 also claimed #5345 (lock stealing) and changed 39 files, most unrelated; this change is #5344 only. Live proof, read-only (a gh wrapper refused every merge and mutating call): - cli/cli#14474 (trunk requires 3 classic build contexts, none ran): refused, naming build (macos-latest), build (ubuntu-latest), build (windows-latest); with --allow-missing "build (macos-latest)" it still refused, naming the other two. - cli/cli#13665 (ran build (ubuntu-24.04-firewall) instead): refused, naming build (ubuntu-latest). - hashicorp/terraform#39262 (ruleset-required checks absent): refused, naming Code Consistency Checks, End-to-end Tests, Race Tests, Unit Tests. - cli/cli#14485 (all required reported and green): verified; the wrapper blocked the merge call and the pull request read back open. Fixes #5344 * fix(review): Preserve required-check producers and aggregate independent read failures * fix(document): Clarify required-check verification and waiver documentation * fix(bin): match an app-bound required commit status by name The producer-identity check resolved an app-bound required context only against check runs, so a required context that the required app reports as a commit status could never match and always read as "has not reported". A commit status carries no app id to compare, so an app-bound requirement that arrives as a status now matches by name, as before producer binding; check runs keep requiring the configured producer app. Live, read-only: hashicorp/terraform#39262 requires license/cla from integration 865473, reported green as a commit status by the CLA app. The previous head refused it as unreported; this head no longer does, while still naming the four required check runs that never ran there. Refs #5344 * fix(document): Clarify accepted commit-status producer verification limitation --------- Co-authored-by: firstmate-oss <firstmate-oss@kunchenguid.local> * fix: keep persistent secondmates out of landed-work cleanup (#5696) * fix(bin): never offer a persistent secondmate for teardown The return brief's "Landed, cleanup due" scan listed every state/*.meta record carrying a pr= and a merge-notified marker without regard to kind, so a secondmate record holding a relayed child's merged PR put the mate itself up for "bin/fm-teardown.sh <mate>" cleanup. A secondmate is a persistent worker, never landed work. - bin/fm-afk-return.sh: skip kind=secondmate in the landed-cleanup scan. - bin/fm-pr-check.sh: refuse to record pr= or arm a merge watch on a kind=secondmate record before any side effect; a PR reported on its routed status channel belongs to a task in the mate's own home, which arms its own watch. - bin/fm-watch.sh: a merged result from a poll already armed on a secondmate retires the poll silently - no merge outcome, marker, or wake. * no-mistakes(document): Document secondmate merge-watch and return-brief exclusions * no-mistakes(ci): CI failed in an unchanged watcher-shutdown test whose three-second wait was sensitive to runner load. Increased the wait for both state- and home-deletion cases without changing watcher behavior. The full fm-watch-arm suite passed locally; syntax and diff checks passed * fix: reject invalid X reply and follow-up arguments before posting (#5702) * fix(bin): refuse unknown dash-leading args in public-posting fm-x scripts fm-x-reply.sh collected any unrecognized argument into the positional pool and took the first one as the reply text, so an invocation like "fm-x-reply.sh <id> --followup --final <text>" posted the literal string "--final" to X and silently dropped the real text. Make argument parsing strict in every script that can post publicly: an unknown dash-leading argument, a dash-leading request_id/task id, a dash-leading option value, or a surplus positional now exits 2 with a usage error before any config load, outbox write, or network call. Reply text starting with '-' is still accepted via --text-file or stdin, and --help is honored wherever it appears instead of becoming text (a --help forwarded through fm-x-followup.sh would have counted as a posted follow-up and mutated the link). fm-x-link.sh and the fm-public-followup scripts already refuse unknown arguments; fm-x-poll.sh takes none. * no-mistakes(review): Refuse surplus follow-up text sources; drop post-ID help branches * no-mistakes(document): Clarify reply and follow-up argument usage * no-mistakes(review): Refuse dash-leading --text-file operands in fm-x-reply * no-mistakes(document): Correct follow-up argument parsing comment * no-mistakes(document): Document dismiss argument rejection in script header * fix: pause broken supervision-host sessions between engine probes (#5701) * feat(bin): latch the supervision host after repeated engine errors Rung 3c-1 of the PR 5631 re-cut: the host copies the Pi branch's broken-session policy. Two consecutive engine errors latch the session; every away wake then reaches main with one supervision-host line for a five-minute cooldown, after which one wake probes the engine, and each failed probe doubles the cooldown up to one hour. A reported turn without an engine error clears it. The latch is kept per main session, engine, and model in state/.supervision-host-health, and the engine conversation now uses the same main-session key, which includes the lock holder's process identity so a recycled pid never shares either. Lifted from the validated 5631 tree and adapted to main's away-only host: the attended recovery line and attended cooldown pass-through are left for the attended core, so a recovery is only logged. * no-mistakes(document): Consolidate supervision-host latch documentation * fix(bin): absorb routine secondmate working and paused status appends (#5535) * fix(bin): absorb routine second-mate progress while surfacing routed replies Fixes #2959 A kind=secondmate task's status signal was never absorbable, so a healthy mate's routine working: and paused: appends woke the primary every time. signal_crew_provably_working now reads the mate's lines new since the watcher's classified position: a decision, blocker, terminal outcome, note:, correlation-marked line, or unknown verb still surfaces regardless of busy evidence, while unmarked working:, paused:, and resolved: fall through to the same provably-working absorb an ordinary crewmate gets. * no-mistakes(review): narrow secondmate routine absorb to working and paused * fix(bin): use gh-axi for the ship DoD draft check (#5519) * fix(bin): use gh-axi for the ship DoD draft check * no-mistakes(review): use PR number not URL in gh-axi draft check * fix(bin): bind inactive-outcome receipt identity to structured fields only (#5520) * fix(bin): drop status prose from the inactive-outcome dedupe identity The inactive-outcome receipt fingerprint included the child's sanitized last status line, so a persistent child appending routine prose after one terminal outcome minted a fresh parent event per sentence. Bind the identity to incarnation, task id, terminal state, and PR only, keeping the last line in the record as status_head evidence. Fixes #2960 * no-mistakes(document): note structured-only inactive receipt identity in regression coverage * feat: record the Claude and Cursor dialog for the supervision host (#5707) * feat(bin): record the supervision host's dialog mirror on Claude and Cursor Add bin/fm-host-mirror.sh, the one owner of the supervision host's dialog mirror file, cursor, lock, and feed, plus the main-session key it keys entries to. The tracked Claude UserPromptSubmit and Stop hooks and the Cursor beforeSubmitPrompt and afterAgentResponse hooks record the captain's prompt and main's reply, only on a home with config/supervision-host, from a genuine primary checkout, for the lock-owning session. The mirror lands inert: writers record and nothing reads it yet; attended supervision on the host is the later step that consumes the feed. Codex, Grok, OpenCode, and omp have no writer here. * no-mistakes(review): Scope mirror dedup to session, atomic appends, marker-inclusive caps * no-mistakes(document): Clarify dialog mirror scope and remove duplicate contract details * no-mistakes(document): Correct Cursor hook documentation for dialog mirror registration * no-mistakes(review): Pass mirrored dialog text to jq via stdin * no-mistakes(document): Clarify dialog mirror documentation and remove duplicate claims * no-mistakes(review): Preserve internal dialog whitespace; drop mirror check and verified modes * no-mistakes(review): Drop only identical mirror repeats; remove redundant chmod guard * fix(bin): retire check-row receipts on branch acknowledgement so away escalations are not repeated (#5731) * fix(bin): retire check-row receipts on branch acks and report an unchanged situation once * fix(bin): scope a branch acknowledgement's check-row receipt retirement to its granted sequences The away posture lifts the attended partition's check/decision exclusions, so a branch grant can name check-kind rows - but the branch-actor ack still assumed check rows were main-only and skipped every receipt scan. The queue row was consumed while its terminal-outcome .pending receipt stayed behind, and each inactive-reconcile cadence scan re-queued the same fingerprint. In the first real away window on the supervision host that re-escalated one unchanged held-PR situation on every cycle (~1,734 of 4,149 outcomes). A branch ack now scans inactive-outcome and inactive-reconcile receipts and commits secondmate stall receipts against exactly the sequences in its eligible-row snapshot - the same rows it consumes - instead of none. Attended grants still name no check row, so the scans find nothing. * fix(bin): store a repeated captain verdict as routine while the task's durable situation is provably unchanged fm-branch-outcome.sh append computes a mechanical situation key per captain row - metadata bytes, captured status-log endpoint and identity, live crew-state verb, worktree head - and anchors it in state/.<task>.branch-captain-key. A later captain verdict whose recomputed key matches is stored as routine with "unchanged since seq <N>:" prefixed to its summary, so one situation escalates once until something provably changes. A task with no readable status ledger is never demoted, an unreadable record fails toward reporting, and teardown removes the sidecar with the task's other branch records. The append-only store schema is unchanged. This covers both hosts: the Pi supervision branch and the supervision host both funnel reports through append. * docs: check rows are main-owned only while attended; the away posture grants them to the branch, whose ack retires their receipts exactly * test: the away-flood reproduction as a regression test (branch ack retires the receipt and later scans stay quiet), store-level dedupe coverage, and a branch-ack secondmate stall receipt case * fix(bin): restore the secondmate child devin-config cleanup path The branch-captain-key sidecar addition mistyped the sibling entry as .$child_id.devin-config.json, so a forced secondmate teardown would have stopped removing each child's real <id>.devin-config.json. Restore the original path and add a behavioral test that stops the child sweep mid-loop on a refused close, proving the cleaned child's devin config and captain anchor are both removed while the unconsumed child's records are retained. * no-mistakes(review): Key captain dedupe on the covered wake rows' fingerprint * no-mistakes(review): Drop captain-key demotion; prove one escalation on both surfaces * no-mistakes(review): Drop unrelated teardown test; cite both receipt test files * no-mistakes(document): Docs already match branch-ack check-receipt retirement * fix(bin): deliver Claude-bound operational input as a record-backed doorbell (#5664) * fix(calm): deliver Claude-bound operational input as a record-backed doorbell Claude Code 2.1.280 removes U+2063 from every submitted prompt, so a typed operational envelope reaches a Claude Code primary as plain text. The away daemon now writes the envelope to a record under state/operational-inbox and types only a plain doorbell naming it; the /afk return check and the Calm mod recognize the doorbell only when that record holds a current envelope. Marker- preserving harnesses keep the typed envelope. The live Calm guard accepts the 2.1.280 module-load log line, drives the doorbell, and asserts thinking stays hidden. * no-mistakes(review): Fix operational record retention at 7 days and document prune limit * no-mistakes(document): Point Calm bounds at 2.1.280 evidence; fix afk-exit comment * no-mistakes(lint): Pick newest Calm e2e transcript without parsing ls * docs(calm): add a minimal turning-Calm-on step for Claude Code * fix(spawn): deliver the Claude launch brief as a record-backed doorbell Claude Code strips U+2063 from the launch-prompt argument too, so a worker's launch brief arrived with its operational marker removed. Publish the brief as a record in the receiving home's operational inbox - a secondmate's own state, not the primary's - and pass only the printable doorbell naming it, falling back to the typed envelope when the record cannot be published so the brief body still delivers. Unwrap doorbell-carried digests in the daemon digest tests that still read the raw send log under the claude pin, and update the documented bounds now that launch briefs hide like the other operational rows. * test(spawn): cover a secondmate's launch-brief record landing in its own home The record-backed doorbell resolves its state through the receiving pane's home, so prove a claude secondmate launch publishes into the seeded secondmate's operational inbox and never leaks a record into the primary's. * no-mistakes(review): Pass primary harness to daemon, tighten retention, refresh verdicts * no-mistakes(review): Prune operational records by exact seven-day elapsed age * no-mistakes(review): Batch record pruning so large inboxes still expire * no-mistakes(review): Refuse Claude spawn when brief record cannot publish * no-mistakes(review): Drop thinking probe from Claude Calm live test and docs * no-mistakes(review): Record dated Claude Code 2.1.282 reproduction evidence * no-mistakes(document): Clarify operational doorbell documentation and record expiry * no-mistakes(document): Correct AFK escalation carrier guidance * no-mistakes(review): Describe operational record retention as about seven days * no-mistakes(document): Clarify Calm delivery and operational record retention * no-mistakes(review): Remove out-of-scope Calm launch guide from Claude docs * no-mistakes(document): Document Claude launch-brief delivery and refusal * no-mistakes(document): Correct stale operational-input documentation * no-mistakes(ci): Fixed the stale Claude trust test to verify that worker and secondmate launches deliver readable, record-backed briefs instead of expecting brief paths in their commands. Annotated the daemon’s output variable for ShellCheck without changing behavior. The affected tests, daemon tests, ShellCheck, and diff check pass locally * no-mistakes(ci): parse rebased Claude launch after trailer hook prefix * no-mistakes(review): Trust launch-brief record and restore thinking bound doc * no-mistakes(review): Parse final Claude launch statement; drop Stop-hook docs --------- Co-authored-by: Mike Sewell <maikunari@protonmail.com> Co-authored-by: no-mistakes <no-mistakes@localhost> * test: isolate lint fixture from tracked suite (#5727) * fix(bin): republish parent metadata after a remote secondmate relaunch (#5583) A host-local relaunch rewrote only the far endpoint, so this home kept the old harness, model, and effort, and appending those keys after pr= broke pull-request poll authentication. * fix(bin): give slow watcher suites headroom under the changed-suite bound (#5516) tests/fm-watch-triage.test.sh finishes in about 434s alone and about 698s under CI load, so the 900s bound the changed-suite runner applies produced a false timeout under ordinary concurrent validation. Raise the automatic bound to 1500s, which keeps every measured script under it while staying below the 30-minute normal CI tier so a genuinely hung script still fails here with its output before the job cap cancels the lane. Fixes #3869 Refs #3565 * fix(bin): stop nested steal-lock recursion and mid-steal watcher TERM (#5728) * Fix nested watcher lock reclaim * no-mistakes(review): Elect a single steal-mutex reaper and bound arm TERM wait * no-mistakes(review): Reclaim self-held steal mutex and unify autoarm steal reaping * no-mistakes(review): Resume own interrupted steal reap from its tombstone * test: make supervision-host park-boundary tests deterministic (#5710) * test: hold the back-to-back boundary close on the host's own clock test_park_boundary_holds_under_back_to_back_closes assumed two engine turns fit in the ~16s pre-refusal window and that the stub finished a turn in 3s. Under load the stub's real drain, report, and acknowledgement take ~13s, so the turn either died at its bound (which hands the wake to main, no boundary line) or the second close landed past the window and the fixture failed while the boundary held. 3 failures in 5 runs at a load average near 11. Hold the first turn on a release file instead: once the engine is in flight, a second close is appended mid-turn and the turn is released as the refusal window opens (park bound minus turn bound and grace, read off the host's own start record). The queued close can then only wait for the boundary on any machine speed, which is what the test asserts: the boundary line ends the output, the demo.status row stays queued for main, and no second engine turn ever starts. A host too loaded to start the turn at all hands the first close to the same boundary exit. After: 12/12 at load ~15-42. * no-mistakes(review): Print boundary test deadline as a decimal integer * no-mistakes(review): Hold boundary test turn on a FIFO, require full sequence * no-mistakes(review): Remove stray before/after supervision-host test copies * test: hold the late close's render until the refusal window opens The boundary recheck test's node shim slept a fixed 10s, which assumed the first close was read before the host's refusal window opened. Under load the close arrived after the refusal check, so the host correctly refused it before the successor started and the render snapshot never appeared. Block the wake-prompt render on a FIFO released at the refusal-open instant read from the host's own start record, so the pre-turn recheck must refuse on any machine speed. * no-mistakes(review): Derive minimal park bounds and refresh supervision-host shard hint * no-mistakes(review): Drive park-boundary tests from a seam-gated host test clock * fix: stage remote home clones before publication (#5733) * fix(bin): stage remote home clones before publishing them A remote home provision cloned the code root directly into the public FM_HOME path while rollback() claimed rm -rf of that same path on any failure. Bash defers trapped signals past a foreground child, but any other cleanup or lifecycle path that removes the home directory races the live clone's object copy, producing the CI flake "fatal: failed to copy file to .../.git/objects/...: No such file or directory". Clone into a private staging directory beside the home and publish with an atomic rename once complete, so no cleanup can remove a directory a live clone is still writing; a home that appears mid-provision now dies cleanly instead of inheriting torn state. The regression coverage holds a real clone mid-copy, removes the public path, and requires the provision to finish and publish intact. * no-mistakes(review): Prove home ownership by sentinel and hold only a live clone * no-mistakes(review): Assert raced provision publishes a complete, intact clone * no-mistakes(document): Document remote home staging and publication safety * no-mistakes(lint): Fix ShellCheck warning in clone integrity assertion * no-mistakes(document): Clarify remote home publication and rollback guarantees * fix: preserve Herdr status on Pi relaunch (#5161) * fix(control): keep a relaunched Pi worker's herdr pane status authority alive Defect: after `bin/fm-control.sh <id> relaunch` (observed live on a herdr Pi crewmate whose pane read idle while it ran its validation pipeline), the pane froze at whatever its previous agent had last reported. Cause, measured on herdr 0.9.1 against a real Pi: a pane has one status authority, and for Pi with its integration installed that authority is the lifecycle hooks, so herdr also skips screen detection for the pane. In the crew shape the registration outlives its agent process (upstream issue #4115; docs/herdr-backend.md "Restart and liveness behavior"), and herdr applies only reports carrying the session identity it bound. A replacement started fresh in that pane reports a NEW session, so its state reports are ignored and the pane stays frozen. Nothing from outside repairs it: `pane report-agent-session` and `pane report-agent` for `herdr:pi` are accepted (rc=0) without being applied unless the reporter is the registered pane agent, and `pane release-agent` on the stale record changes nothing. Fix: a relaunch preserves the binding instead of fighting it. The launch owner reads the session reference the endpoint's own runtime recorded (`fm_backend_herdr_pane_agent_session_ref`) and passes it back as Pi's own `--session <path-or-id>` (`relaunch_resume_args`; `fm_control_relaunch_resume_flag` owns which adapters and which registered-agent labels qualify). That is the same reference herdr itself resumes Pi panes with after a server restart, and the resumed session's reports land again, which the live check confirmed: the pane returned to working while the replacement worked and idle when it settled, on the same session identity. Safety: relaunch-only (a fresh spawn binds nothing), herdr-only (the one adapter that records a per-pane session), Pi-family only, and only when the registration's own agent label matches - so no other adapter's conversation can be handed to a Pi launch. An unreadable, missing, or malformed reference degrades to exactly the fresh-session launch that existed before. No lifecycle, liveness, isolation, or merge guard is touched, and an empty result leaves every non-Pi launch byte-identical. `resume` remains a refused verb; docs/agent-control.md and the harness-adapters references are corrected where they claimed Pi had no verified resume form at all. * no-mistakes(document): docs: correct relaunch session-authority ownership and skill paths * no-mistakes(document): docs: correct stale control-plane ownership claim * no-mistakes(document): docs: drop unverified Herdr restart resume claim * no-mistakes(test): Added offline Herdr Pi session-authority relaunch coverage * no-mistakes(document): Document Herdr Pi relaunch session continuity * no-mistakes(ci): The failing remote relaunch test tried to arm a PR poll for a secondmate, which `fm-pr-check.sh` correctly refuses. Removed that invalid test scenario; the remaining remote relaunch tests pass, and `git diff --check` is clean * Seed the relaunch-ordering PR poll fixture without fm-pr-check.sh (#5758) Main has been red since fm-pr-check.sh began refusing to arm a merge poll on a kind=secondmate record (#5696): the relaunch-ordering case in tests/fm-remote-secondmate-relaunch.test.sh armed its fixture through that entry point and could no longer be set up. The ordering guarantee still matters: a secondmate record armed before the refusal can legitimately carry a trailing pr=/pr_head= identity block until the watcher retires it, and fm-remote-secondmate-relaunch.sh must still keep that block last when republishing harness/model/effort. Seed the fixture the way such a record was really written - pr= appended last to the meta, then the poll artifacts published through the same fm_pr_poll_prepare/fm_pr_poll_publish_prepared pair fm-pr-check.sh uses, a pattern tests/fm-pr-check-security.test.sh already follows - and drop the now-unused fake gh fixture. The #5696 refusal itself stays pinned by the security suite's secondmate-record case. * feat: add attended supervision for Claude and Cursor hosts (#5748) * feat: run attended supervision on the host for Claude and Cursor On a home opted into config/supervision-host with a Claude or Cursor primary, the supervision host now takes the attended wakes the Pi branch would take: routine outcomes stay off main, and a captain outcome wakes main once with a branch-outcome line and waits in the drain's new BRANCH OUTCOMES section until main acknowledges it with mark-processed. - The offer rule moves into branchOfferForWake, shared by the Pi watcher and the host through bin/fm-branch-dispatch.mjs offer. - The host feeds the dialog mirror at the head of each attended wake and passes a close through unchanged when it is main-only, the engine or a tool is missing, the primary has no verified mirror, the main session cannot be identified, or the session is cooling down. - The drain presents captain outcomes first, one line per task, never behind older routine outcomes, and collapses routine overflow into a count that is marked read. - The return advances the store's read cursor through the away window once the brief has rendered, so the first drain does not replay it. - The branch prompt's mirror wording is host-neutral, and the rule to report what main must act on as captain, once per unchanged situation, applies only to the attended posture on the host. * docs: record the attended supervision host live check * no-mistakes(review): Present pre-window unread outcomes and contiguous captain prefix * no-mistakes(review): Return brief presents every row it marks read * no-mistakes(review): Return brief lists every unread outcome in one list * no-mistakes(review): Keep return list in store order and gate cursor failures * no-mistakes(review): Make the drain the only branch-outcome presenter after return * no-mistakes(review): Gate return on drain outcome failures; byte-count outcome budgets * no-mistakes(review): Gate drain on projection failures; UTF-8-safe byte cuts * no-mistakes(review): Fail drain without jq; hand unreadable prompt mirror to main * no-mistakes(document): Correct supervision-host return and drain documentation * no-mistakes(review): Recheck attended offer at turn start; honest failed-drain brief * no-mistakes(document): Correct supervision-host posture and drain documentation * no-mistakes(document): Documentation remains accurate for attended supervision * fix(bin): dedup directed source expansions in fm-pending-reply-lib (#5753) Each '# shellcheck source=' directive makes ShellCheck's external-source traversal expand that library's whole transitive graph again at the site. fm-pending-reply-lib carried three directed lazy sources of fm-wake-lib and two of fm-parent-channel-lib on identical per-call re-source sites, so one file analysis peaked above 4 GiB and every caller (fm-watch, fm-teardown) inherited the multiplier - the root cause of the PR #5732 Lint 1 OOM kill. Keep the runtime '.' commands byte-identical: the lazy re-source under 'local STATE FM_WAKE_QUEUE FM_WAKE_QUEUE_LOCK' is real behavior. Drop the duplicate directives so each library expands once per unit, and drop the tmux/classify directives since classify already arrives through the kept fm-wake-lib expansion and no tmux symbol is referenced here. The directive above the lib-dir assignment is kept - it binds the bin/ prefix so the undirected sites still resolve without SC1091. Measured peak RSS, ShellCheck 0.11.0 -x on Linux arm64: bin/fm-pending-reply-lib.sh 4.06 GiB -> 1.96 GiB, zero findings * fix(bin): avoid bash 5.2 sibling $() in recovery mint and delivery log (#5773) * fix: split bash 5.2 sibling $() in recovery mint and delivery log Sibling command substitutions on one line can empty a recovery generation under bash 5.2 when a CHLD trap is set. Mint pid/epoch sequentially, refuse empty tokens before write, and clean delivery fields before printf. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: split bash 5.2 sibling $() in recovery mint and delivery log Sibling command substitutions on one line can empty a recovery generation under bash 5.2 when a CHLD trap is set. Mint pid/epoch sequentially, refuse empty tokens before write, and clean delivery fields before printf. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: keep recovery mint failure semantics after sibling $() split Remove the new pid/date refusal and grammar guard so a mint miss still yields a grammar-valid token and a durable wake row, matching accepted review intent. Drop the fake-failing-date case that locked in the refuse. Co-authored-by: Cursor <cursoragent@cursor.com> * no-mistakes(document): Point recovery-mint hazard comment at its regression test --------- Co-authored-by: Cursor <cursoragent@cursor.com> * fix(bin): name the recovery for a declined Claude imports dialog and stop calling Escape safe there (#5791) Fixes #4520 * fix(bin): report the newest status event in the voice status reader (#5790) Fixes #4756 The voice status reader in bin/fm_voice_records.py reports each worker's state from the last non-blank line of its status log. When a worker appends a status line and then a line of plain prose, the reader reported "note" with the prose line instead of the declared state, diverging from bin/fm-classify-lib.sh's shell scan. Scan back through the tail for the newest line whose prefix is a single lowercase verb-shaped word (letters and hyphens), and report that event's verb instead of always taking the last line. An unrecognised verb-shaped prefix still reports "note" rather than letting an earlier recognised line answer for it, and free text with no colon is skipped as prose. When the tail holds no such event, the last line is reported exactly as before. * fix(bin): gate a self-announcing tool's update-available report on a newer version (#5786) * fix(bin): stop reporting an already-installed version as an available update An update announcement named its version first ("current -> new"), so reading the first dotted number as the announced version compared the current version against itself and always looked newer. Read the last dotted number instead, and only report an available update when that announced version is newer than the newest installed copy found; when that version is already installed, report only PATH skew. Fixes #5151 * no-mistakes(document): docs: gate announce update-available report on newer-than-installed * fix(bin): pass the dispatch profile effort to OpenCode workers through their launch config (#5799) * fix(bin): pass the profile effort to OpenCode workers through their launch config The dispatch profile's effort axis was recorded in task metadata but never reached an OpenCode worker: the launch wrote only a permission grant into the config it constructs. OpenCode 1.18.32's config schema carries per-model reasoning effort as agent.<name>.variant, so the chosen effort is now merged into the same OPENCODE_CONFIG_CONTENT JSON as the default build agent's variant, keyed to the resolved model. With no effort chosen the launch stays byte-identical. Fixes #1373 * no-mistakes(review): gate OpenCode effort variant by model provider family * no-mistakes(document): docs(opencode): note provider-family gating for effort variant * fix: stop cancelled validation runs from reporting false failures (#5815) * fix: preserve cancellation as no verdict in crew state Reuse the green-delivery safeguard for cancelled CI monitors and permit a skipped rebase. Other cancelled outcomes and coarse ledger records use the existing unknown state. Four delivered-PR regressions failed before the fix and pass afterward. The isolated public resolver and fleet-summary tests prove that undelivered cancellation no longer creates a failure contradiction, while preserving historical records and the terminal_in_flight invariant. Evidence uses fixture no-mistakes responses, not a live daemon cancellation. Update the existing coarse cancellation assertion from failed to unknown because it encoded this defect; retain its newest-run precedence check. Full fm-crew-state suite and pinned lint pass. * fix(review): Verify PR disposition before reclassifying terminal validation runs * fix(test): Add captured cancellation replay coverage for resolver and fleet * fix(document): Clarify cancellation and terminal delivery documentation * fix: require declared waits for workers awaiting their own work (#5812) * fix: declare worker background and pipeline waits Require ship and scout workers to declare owned-work waits with the existing paused verb before ending a turn or waiting on a pipeline or long command. Keep the first-sight alert and existing liveness classification unchanged; subsequent inspection follows the existing long pause cadence. Validation: emitted brief regression failed before the instruction change and passes afterward. Public watcher/drain regressions cover the first alert, repeated wedge suppression, bounded rechecks, and undeclared idle alarms using isolated backend fixtures. Brief suite, pinned lint, Bash syntax, documentation inventory, and whitespace checks pass. No real worker harness was exercised for wait behavior. * fix(document): Clarify declared worker waits and documentation ownership * fix(ci): Captain, fixed the cadence test to age both the declaration and first-alert throttle while preserving declaration identity. Reproduced the CI failure using stable identity; corrected tests pass with both stable identity and native macOS behavior. Focused ShellCheck and diff checks pass. Production behavior is unchanged; Linux CI was not rerun locally * fix: bound ShellCheck to one canonical root per process (#5770) * fix(bin): bound each lint root in its own ShellCheck process CI job "Lint 1" died twice at about ten minutes because the two shard workers each packed about 110 canonical roots into one unbounded ShellCheck process, and a byte-weight rebalance moved the analysis-heavy fm-watch.sh into a partition with other heavy roots, so the pair outgrew the 16 GiB runner before anything could name a culprit. Run one canonical root per ShellCheck process under an enforced envelope: a wall deadline plus terminate-then-kill grace via the shared fm-timeout-lib.sh watchdog, and a per-root rlimit spec applied inside the child before exec (default a 4 GiB address-space cap, so two workers stay inside a 16 GiB job with headroom). A root that exceeds the envelope fails by name with a recorded reason - timeout, memory, signal, or limit-unavailable - instead of taking the runner down. The per-root watchdog runs in its own process group so the owner's group sweep cannot orphan the bounded subtree, and fm_exec_timed now starts the same escalation when its parent dies before it can be signalled. FM_LINT_REQUIRE_BOUNDS=1, set in CI, refuses the run outright when a configured bound cannot be enforced on the host rather than lint uncapped. Each root's begin/end, reason, duration, and peak RSS stream to stderr in partition mode and append to a retained <telemetry>.roots.tsv sidecar uploaded beside the partition telemetry. Coverage is unchanged: pinned ShellCheck 0.11.0, --norc, --external-sources full analysis, complete and disjoint partition inventory, workflow lint, and the backend-purity check, with byte-identical diagnostics across jobs=1/2 proven by tests/fm-lint.test.sh. * fix(bin): fail closed on unenforceable lint bounds and size the cap Required-bounds mode (FM_LINT_REQUIRE_BOUNDS=1, set by CI) now refuses the run with named errors before any root starts: a missing fm-timeout-lib.sh, a watchdog that cannot actually bound a probe command, or a host that rejects the address-space limit all stop the run rather than lint uncapped. The generalized FM_LINT_ROOT_RLIMITS flag:value interface is replaced by a single FM_LINT_ROOT_MEMORY_KIB, and the roots sidecar and telemetry record the run's final exit status after backend-purity and workflow checks instead of the pre-check lint status. The default cap is 6 GiB of address space per root, not 4 GiB: ulimit -v bounds virtual address space rather than resident memory, and ShellCheck's GHC runtime keeps roughly a third of that space as reservation, so 6 GiB yields about a 4 GiB working heap budget. A Linux measurement during this change showed eleven real canonical roots running out of memory under the earlier 4 GiB cap while the largest passing root peaked near 2.8 GiB resident; two 6 GiB roots plus runner overhead still fit the 16 GiB job. Roots that still exceed the cap keep failing by name, and the sidecar's per-root peak RSS keeps roots approaching the budget visible. tests/fm-lint.test.sh now proves the memory primitive where it can be proven: on hosts that accept ulimit -v a perl allocator is refused under a 256 MiB limit and reported by name as a memory death, the pinned ShellCheck lints a small file under the configured cap and is named when a far smaller cap binds it, and a watchdog-less copy refuses under REQUIRE_BOUNDS; the bounded cases skip on macOS, which cannot enforce the address-space limit. * no-mistakes(review): Prove memory cap binds, pass watchdog owner, drop unused modes * no-mistakes(review): Capture watchdog owner before startup for every fm_exec_timed caller * no-mistakes(document): Clarify bounded lint documentation and telemetry * no-mistakes(document): Correct bounded lint documentation and sidecar path * docs(bin): restore the per-root memory cap sizing rationale The pipeline's document step rewrote the ROOT_MEMORY_KIB comment and dropped the sizing reasoning the change is required to record: address space vs resident memory, the GHC reservation share, the measured 4 GiB failures and ~2.8 GiB peak, and the two-roots-plus-runner capacity arithmetic. Restore it beside the default while keeping the corrected "not a resident-memory ceiling" framing. * no-mistakes(review): Document memory cap RSS reduction threshold and first candidate * no-mistakes(review): Scope owner-death escalation docs to the perl watchdog * no-mistakes(document): Clarify bounded lint and timeout documentation * no-mistakes(review): Install perl watchdog signal handlers before forking the command * no-mistakes(document): Correct bounded lint documentation and stale watcher comments * no-mistakes(ci): Fixed the supervision-host test’s obsolete expectation: the watchdog now reaps an engine when its host dies. The timeout and supervision-host tests pass locally; the watcher test also passes locally. Lint 1 and 2 remain unresolved: seven canonical roots exceeded the required 6 GiB address-space cap in CI. I did not raise the cap, exempt roots, or reduce source-following coverage to make those failures disappear * no-mistakes(ci): The two lint checks failed when eight canonical roots hit the enforced memory cap. I reduced repeated ShellCheck source-graph expansion while keeping runtime imports and the canonical root inventory intact. Pinned ShellCheck passes for all changed roots; the relevant local tests pass. The 6 GiB Linux CI run remains unverified * no-mistakes(review): Restore source directives, raise cap to 8 GiB, classify OOM * no-mistakes(review): Classify memory deaths from root stderr, not source excerpts * no-mistakes(review): Match only whole runtime memory-error lines for memory reason * no-mistakes(document): Clarify lint memory classification in script documentation * no-mistakes(ci): Fixed both lint checks’ memory-limit failure: each CI lint job now runs one root at a time with a 12 GiB address-space cap. Kept the local two-worker default and updated the sizing comment and test expectation. The lint tests and workflow validation pass locally; Linux CI remains to confirm the heavy roots * fix: bound watcher cleanup wait on the downtime-marker lock (#5732) * fix(bin): bound the watcher cleanup marker-lock wait tests/fm-watch-triage.test.sh intermittently failed serial CI shard 1 with "watcher pid <pid> did not exit within 10s of TERM". The watcher had processed the TERM and was inside watcher_cleanup, where the recovery-marker publish waits on state/.watcher-down.lock through an unbounded fm_lock_acquire_wait. A live foreign holder of that lock leaves the TERM'd watcher spinning in its own EXIT trap until the lock frees or a second signal short-circuits the trap. fm_recovery_transition now takes an optional bound and both release-lock paths plus publish honour it through a new in-process fm_lock_acquire_wait_max. watcher_cleanup passes FM_WATCHER_CLEANUP_LOCK_BOUND (default 2s); on timeout the publish is skipped, the singleton stays behind as ordinary dead-pid evidence, and the next arm's clear-stale-lock still republishes it. Regression test drives a real watcher with .watcher-down.lock held by a live foreign process and asserts a single TERM still stops it. * no-mistakes(review): Parse watcher cleanup lock bound as decimal, zero defaults * no-mistakes(review): Pin cleanup bound tests to observed marker-lock contention * no-mistakes(document): Document bounded watcher cleanup and recovery * no-mistakes(document): Clarify bounded watcher cleanup and recovery documentation * no-mistakes: apply agent fixes * no-mistakes(review): Arm marker-lock FIFO before TERM; drop FM_TEST_ONLY_LATE * no-mistakes(review): Hold marker lock through a failed cleanup acquire * test: stop the remote secondmate e2e watcher before temp-root cleanup (#5845) * test: stop the leaked unreachable watcher before remote e2e cleanup The remote secondmate lifecycle e2e backgrounded fm-watch.sh through the remote_env shell function, so $! named the function's subshell rather than the watcher. Killing that subshell left the unreachable-leg watcher running, and its one-second liveness probe kept invoking the fake ssh, which rewrites ssh.count in the temp root. When a probe landed while the EXIT trap was removing the root, rm failed with "Directory not empty" after every assertion had passed. Exec the watcher from the backgrounded function so the recorded pid is the watcher itself, and assert the stopped watcher stops probing and writing its state. Cleanup also stops a watcher left running by a failed assertion and removes the root through fm_test_remove_tree, so a run that fails before retirement does not strand the read-only spawn hooks directory. Closes #5836 * no-mistakes(review): Clear reaped watcher PIDs and restore plain temp-root removal * no-mistakes(review): Let in-flight probe settle before stopped-watcher baseline * fix(bin): stop reporting untouched shared-captain copies as drift (#4806) * fix: stop quarantining ordinary shared-captain source updates * no-mistakes(document): Rewrap remote inherit header so usage prints fully * docs: restructure calm.md for readability (#5604) * docs: make calm easier to read Restructure the Calm mode prose into sections, lists, and tables without changing documented behavior. Every original heading, anchor, fenced code block, inline-code span, link target, number, and quoted string is preserved. * docs: restore reload case in calm override lead-in The Built-in tool override collisions lead-in covers a session that reloads with Calm already on, as the original text did. * docs: restructure turnend-guard.md for readability (#5611) * docs: make turnend-guard easier to read Restructure the turn-end guard doc's prose into shorter sections, lists, and tables without changing documented behavior. Every original heading, anchor, inline identifier, link target, and number is kept. * no-mistakes(review): Restore legacy-only scope on TERM retirement sentence * no-mistakes(review): Name Cursor park behavior in live e2e test line * docs: move situational AGENTS.md sections into on-demand skills (#5872) * docs: move situational AGENTS.md sections into on-demand skills Backpass memory optimization: shrink the always-loaded AGENTS.md by moving situational contracts (home layout, session-start recovery, validation and landing supervision, scout completion, away/quiet supervision, Relay ownership) into agent-only skills loaded at their triggers, with a trigger index skill. * docs: classify the new on-demand skills' documentation audience Register the seven new agent-only skills as agent-runtime docs and fix a link in validation-supervision that kept its AGENTS.md-relative path. * docs: close load-timing gaps found by the live regression check - load validation-supervision whenever an ask-user finding is decided or answered, so forbid --yes and process-every-return reach the worker - keep the mid-task captain-ask rule, the unconfirmed network-checks rule, and the worker account pin rule inline in AGENTS.md - fix cross-references that still pointed at moved AGENTS.md sections * fix: route second-mate signal wakes by presented status span (#5879) * fix: route second-mate signal wakes by their new status span A second mate's status log is a shared channel carrying many independently keyed decisions, so judging its signal rows by every decision still open in the whole log pinned each routine update to main behind any unrelated parked hold. scopeForUnreadWake (the one owner for Pi and the attended supervision host) now judges a second-mate signal row by the lines presented since the last drain, bounded by the existing status-presentation cursor: a decision, blocked, resolution, or captain-held line, or a line declaring the key of a still-open decision, keeps the whole row on main, and any cursor problem falls back to the whole log. Keys are read only at the status parser's declared positions, with readable time stamps stripped as bin/fm-classify-lib.sh does. Single-task crewmate and stale routing are unchanged, and stale and signal rows for one mate keep independent verdicts. The supervision branch now treats a second mate's done and merged lines as relayed child outcomes, and fm-teardown refuses the branch actor second-mate retirement through the existing role-partition helper in both postures. * no-mistakes(review): Route second-mate resolutions to main only when closing open decision * no-mistakes(review): Guard bare-verb fold lines and fold resolution spans incrementally * no-mistakes(document): Clarify second-mate wake routing and retirement documentation * no-mistakes(ci): The CI failure was a timing-sensitive watcher teardown test, not the PR’s signal-scope code. Extended the bounded wait for both state-directory and home removal. The watcher suite passes locally, and git diff --check is clean * fix(bin): take the source lock before the lifecycle lock in register-extension (#5882) register-extension took the extension lifecycle lock and then the source lock, while reconcile republishing an unhandled extension result holds the source lock and reaches the lifecycle lock through the extension host's process-event path. Both waits are unbounded and both owners stay alive, so the two could wait on each other forever and freeze the home's monitoring cycle. register-extension now takes the source lock first, matching every other path that holds both. The lifecycle lock still spans binding resolution through registration publication, so binding retirement stays serialized. A new lifecycle-order section in the extension-binding suite, run in the default aggregate, holds a re-registration inside binding resolution while reconcile republishes that source's unhandled result and requires both to finish within a bound. Fixes #5866 * fix: chain repository hooks under git -c overrides (#5877) * fix(bin): run the repository's own hooks when git -c carries the per-task hooksPath The per-task hooks wrappers cleared only the GIT_CONFIG_COUNT override before looking up the repository's own hooks directory. When core.hooksPath reached git through git -c (GIT_CONFIG_PARAMETERS), directly or inherited by a child process, the lookup found the wrapper directory again and exited 0, so the repository's real hook - such as a pre-push publish guard - never ran and the push succeeded. The lookup now ignores GIT_CONFIG_PARAMETERS too, so only the repository's config files decide its hooks directory, and a failed lookup exits nonzero instead of skipping the hook. AI-trailer stripping is unchanged. Fixes #5871 * no-mistakes(document): Document git hook chaining and lookup failure behavior * fix(bin): withhold never-send values from dispatch resolver requests (#5744) * feat(bin): add an optional never-send list to typed dispatch resolution * Added config/dispatch-never-send, an optional local list of literal values and re: regular expressions checked against every string of the resolver request before it is sent to typesafe.ai * A match, an unreadable list, or an empty or invalid pattern now stops the request and falls back to the off path, so firstmate dispatches through its existing intake; the one stderr diagnostic names at most the list line number and never the value * No list, or a list with no match, leaves resolution unchanged * no-mistakes(review): Match never-send literals across whitespace, drop regex mode * no-mistakes(review): Inherit the never-send list into secondmate homes * no-mistakes(ci): ci-2 (Lint 2), caused by this PR and now fixed. The rule that broke: the test script must not share shell variables with a library it sources. The new secondmate-inheritance test in tests/fm-dispatch-resolve.test.sh sourced bin/fm-config-inherit-lib.sh inside a `( ... )` subshell. That lib assigns `out`, so ShellCheck flagged every later `$out` in the test with SC2031. The subshell was the only place this PR sources that lib. The fix runs the propagation in a child shell instead (`bash -c '. "$1" && propagate_inheritable_config "$2" "$3"' _ lib from to`), so the test shell never sources the lib. `bin/fm-lint.sh tests/fm-dispatch-resolve.test.sh` now exits 0, and `bash tests/fm-dispatch-resolve.test.sh` passes. That includes the check that an inherited list is enforced in a secondmate home. ci-1 (Behavior portable serial 8), not caused by this PR, so no code change for it. The only failing test is tests/fm-remote-secondmate-relaunch.test.sh, which fails with "not ok - could not arm the PR poll fixture for the relaunch-ordering test". This PR doesn't touch that test or the code it runs. I reproduced the same failure locally on the base commit ea7c7f7. Main's own CI run on ea7c7f7 (run 36212602588) fails only this job, with the same message. The recent main runs before it also concluded failure * feat(jev): add the guard framework contract and the memory RSS/swap thrashing guard (#5903) * feat(jev): add the guard framework and the memory RSS/swap thrashing guard A Jev guard is a bounded read-only host diagnostic that turns one class of resource pressure into a machine-readable audit record and a one-line verdict. This lands the framework contract (docs/jev-guards.md) with one representative family (Pattern 46, memory RSS and swap thrashing) as the pilot: thin bash wrapper, stdlib-only python engine, and a behavioral test through the CLI. * no-mistakes(review): fix jev mem guard fail-open unknown and contract * no-mistakes(review): fix mem guard rss pairing, memfree, tests, docs * no-mistakes(review): fix inverted UNKNOWN branch in fail-forcing test * no-mistakes(review): register docs/jev-guards.md in audience inventory * no-mistakes(review): simplify wrapper, drop dead top_n, single-source thresholds * no-mistakes(review): assert exact exit code in fail-forcing test leg * no-mistakes(review): tolerate any stdout encoding in text output * no-mistakes(document): Fix guard contract dash style and output wording * fix: prevent contribution poll starvation on slow GitHub reads (#5900) * fix(bin): stop slow GitHub reads from starving and waking the contributions poll The poll's fixed budget cut the tail URL's reads short, and any read killed at the bound was recorded as a forge failure, so routine GitHub slowness produced an observation-unavailable wake. A configured FM_CONTRIBUTIONS_BUDGET now rides the generated check shim and is cut down to the watcher's per-check bound with a margin, a read killed at the bound or the deadline is budget refusal that leaves records untouched, and the poll moves to the next URL that still has a full observation reserve instead of ending. * fix(bin): report the bound when a signal death leaks through fm_run_timed fm_run_external_timeout trusted the wrapper's recorded status over the runner's own bound verdict, so a read killed by the bound's TERM could surface as 143 and callers such as the contributions poll classified a budget-bound read as a forge failure. When the runner reports the bound, a signal-death status is the bound's own TERM racing the wrapper's bookkeeping and is reported as 124; a natural exit still passes through. The replace-shell probe also moves to the repo's portable BASHPID fall…
Intent
Fixes #5332.
Implement ready-for-pr issue #5332: a remote second-mate relaunch leaves the parent's record on the old harness, model and effort.
The fix already exists on the captain's fork as merged fork pull requests tiago-peixoto#61 (21e8008, republish the parent's own record after a remote secondmate relaunch) and tiago-peixoto#63 (c12ed36, keep pr= last when a remote relaunch republishes parent metadata). Upstream pull request #5321 was closed and must not be reopened.
After a successful remote relaunch, the parent's record for the mate names the harness, model, and effort the host actually launched, the same way it does after the local relaunch path.
A refused or failed remote relaunch leaves the parent's record untouched.
A record that carries pr= and pr_head= still passes fm_pr_metadata_identity_parse after a remote relaunch, and its armed PR merge poll still validates, because any rewrite must keep the pr= block last.
The parent's record reflects what the host confirms it launched, not what the parent asked for.
The documented operator path in secondmate-provisioning and docs/remote-secondmates.md leads to a correct parent record.
The existing remote relaunch callers, including bin/fm-secondmate-restart.sh and startup liveness, keep working.
What Changed
bin/fm-remote-secondmate-relaunch.sh. It runs the host-sidefm-remote-secondmate-control.sh relaunchthroughfm-on.sh, then reads the harness, model, and effort back from the route block the host prints. It rewrites the parent'sstate/<id>.metaunder the metadata lock with those values first, and keeps every other line in its original order so anypr=block stays last andfm_pr_metadata_identity_parsestill accepts it. If the relaunch fails, is refused, or prints no route confirmation, the parent record is not touched.fm-remote-secondmate-control.shnow includesmodel=andeffort=in its route block.relaunchprints that block at the end, read from the host's republished endpoint record, so adefaultpin shows up as the value the host actually resolved.bin/fm-secondmate-restart.shnow restarts remote mates through the new wrapper.secondmate-provisioning,docs/agent-control.md, anddocs/remote-secondmates.mdnow point to the wrapper as the way to relaunch a remote mate, and warn against callingrelaunchthroughfm-on.shdirectly. The branch addstests/fm-remote-secondmate-relaunch.test.sh, registers it infm-test-run.sh, and updates the restart test.Fixes #5332.
Risk Assessment
✅ Low: The change is a small, well-bounded wrapper that meets every intent criterion. On failure it exits before touching the record, because fm-on's non-zero exit returns and set -e on the host aborts before print_route runs. It reads the identity back from the host endpoint's own record, keeps the pr= block last under the per-record metadata lock, and leaves the restart caller's output parsing and the startup-liveness launch path (which uses the launch verb, not relaunch) unaffected.
Testing
I ran the new regression test, the existing restart test and the full remote lifecycle e2e; all passed. The e2e includes the startup and watch liveness auto-relaunch, so the existing remote relaunch callers still work. I also wrote an operator-style transcript script and ran it against the base commit and the target commit. At base, the documented relaunch path left the parent record stale; at the target, the wrapper records what the host confirmed (not what was requested), keeps pr= last so the PR poll stays valid, leaves the record unchanged on refusal, and refuses local mates. A separate check showed a record with pr=, pr_head= and x_* fields still passes fm_pr_metadata_identity_parse after a relaunch. Every run used a stubbed ssh or the repository's SSH/Herdr fixtures, so no scenario is live. A real remote leg needs a second SSH host running the Firstmate remote job worker, and the host-side relaunch always uses the Herdr session named fm-remote, which the runbook forbids outside fm-lab-* sessions. There is no UI surface, so there are no screenshots.
Evidence: Bug reproduced at base 31c47af (parent record stays on pi/gpt-5.6-sol/medium after relaunch)
Source: Bug reproduced at base 31c47af (parent record stays on pi/gpt-5.6-sol/medium after relaunch)
Evidence: Fixed at e21b3e1 (record takes the host-confirmed claude/claude-opus-5-5/high, pr= stays last, refusal leaves record unchanged)
Source: Fixed at e21b3e1 (record takes the host-confirmed claude/claude-opus-5-5/high, pr= stays last, refusal leaves record unchanged)
Evidence: pr/pr_head identity still parses after remote relaunch
Source: pr/pr_head identity still parses after remote relaunch
Evidence: Transcript driver script
Source: Transcript driver script
Evidence: New regression test log
Source: New regression test log
Evidence: Restart caller test log
Source: Restart caller test log
Evidence: Remote lifecycle e2e log (passed, including the liveness auto-relaunch)
Source: Remote lifecycle e2e log (passed, including the liveness auto-relaunch)
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
bin/fm-remote-secondmate-relaunch.sh:76- After a successful relaunch the wrapper refreshes only harness=, model= and effort= in the parent record. The route block it already parses also carries target= (and traceparent=), but the parent's remote_target= is left as it was at launch. bin/fm-fleet-snapshot.sh:770 reads remote_target to show a remote mate's endpoint. If the host-local fm-control relaunch opens a new Herdr pane, the snapshot keeps showing the pre-relaunch pane. This is pre-existing and outside the stated intent (which names only harness, model and effort), so no change is required here. Note that docs/remote-secondmates.md:247 says the wrapper republishes the route metadata 'the same way launch already records a fresh route', which slightly overstates what it refreshes.bash tests/fm-remote-secondmate-relaunch.test.sh(new regression test: confirmed, host-resolved, refused, local-refused, and armed-PR-poll cases)bash tests/fm-secondmate-restart.test.sh(existing restart caller, now routed through the wrapper for remote mates)bash tests/fm-remote-secondmate-lifecycle-e2e.test.sh(real fm-on, remote entrypoint, job worker and control scripts over the repository's SSH/Herdr fixtures; includes the startup and watch liveness auto-relaunch)$E/demo.sh <base worktree at 31c47af>: operator transcript showing the documented fm-on relaunch path leaving the parent record stale (bug reproduced)$E/demo.sh <this worktree at e21b3e1>: same transcript through bin/fm-remote-secondmate-relaunch.sh showing the record updated from the host's confirmation, PR poll still valid, refused relaunch leaving the record untouched, local mate refusedManual check: a record with pr=, pr_head= and x_platform= passed through the wrapper, thenfm_pr_metadata_identity_parserun on it before and after, plus file mode and no leftover temp files✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.