feat(bin): merge upstream Firstmate through 1b82b77a - #37
Merged
Merged
Conversation
…ning (kunchenguid#5566) * fix(bin): report a Lavish source armed only after its listener is running Registration alone was treated as ready, so arm could succeed before anything was collecting from the board. * no-mistakes(review): Guard Lavish arm launches, keep retire refusals, report live prior listener * no-mistakes(review): Keep polling through window before reporting a still-live prior listener * test: wait for a capture's claim to drop before the next arm The result is stored before the runner exits, so a re-arm in that gap was meeting a live claim. * no-mistakes(document): Record Lavish arm readiness evidence in verification doc * no-mistakes(ci): Both failures were caused by this PR, and both are fixed with test-only edits. Lint 2 (ShellCheck SC2034): this branch removed the only use of `reply_id` (a `start "$reply_id"` call) from tests/fm-procevent.test.sh, which left the assignment at line 1450 unused. I deleted that assignment. It was the only `reply_id` in the file. ShellCheck is now clean on both test files. Behavior portable serial 4: the failing test was tests/fm-bearings-board.test.sh, in the check "registration consumed its answer before the any-origin binding existed". I reproduced it locally: the hold was still `state: queued` when the test checked it. - What must hold: the test's check that the hold is closed must run after the listener has captured the answer. - Why it broke: the test used a stand-in adapter that ran `fm-procevent.sh start` in the foreground after `arm`, so capture finished before build returned. On this branch, `arm` starts the listener itself in the background, so the real listener captures the answer and closes the hold a moment after build returns. - Fix: removed the now-redundant stand-in adapter, the copied runtime directory, and its extra environment variables. The test now runs the real build through the existing `run_board` helper and waits up to about 10s for the hold to reach `state: done`. The checks that follow are unchanged: `Resolution mode: answered` and the any-origin binding. - Other tests: this was the only test in the file that stood in for the adapter this way. The shard's other pure-contract-unit test (tests/fm-trace-context-lib.test.sh) passed unchanged. Verification: - tests/fm-bearings-board.test.sh passed 3 times in a row via bin/fm-test-run.sh, all 18 checks, about 53s per run. - tests/fm-procevent.test.sh was not rerun, because the lint fix only removed an unused assignment
…elivered (kunchenguid#5599) * fix(bin): acknowledge a delivered unknown-wake escalation The same unrecognized wake was escalated again after it had already been handled, because delivery never recorded that identity. * no-mistakes(review): Scope unknown-wake acknowledgements to one away session * no-mistakes(review): Clear delivered digest when unknown-wake ack write fails * no-mistakes(review): Limit unknown-wake suppression to acknowledged lines * no-mistakes(document): List unknown-wake ack file among away-session artifacts
…ess wait (kunchenguid#5587) * fix(bin): keep a stated default retraction from cancelling a keyless wait A resolved line that names the shared default decision bucket was closing the keyless live wait that only prints as that same key. Keyless self-retraction still closes the keyless wait. * no-mistakes(review): Keep declared waits standing past foreign-key resolved lines * no-mistakes(review): Bound declared-wait read and share one decision-key parser * no-mistakes(document): Document supervisors' key-aware declared-wait read
…unchenguid#5544) * fix(bin): terminate a remote job worker that lost ownership when it receives TERM A serving worker whose lock directory is gone can no longer quarantine shutdown, and resuming service publishes a false ready heartbeat. Exit after stopping only that worker's own command tree, without removing a replacement owner's lock. * no-mistakes(review): Check worker lock ownership before publishing shutdown quarantine * no-mistakes(document): Correct worker shutdown comment on replacement-owned lock * fix(bin): keep an ousted remote job worker off the replacement quarantine Shutdown can lose the lock after the first ownership check and before it writes or clears quarantine. Bind both operations to the directory object this process still owns so a replacement's quarantine stays untouched. * no-mistakes(review): Make ousted-worker shutdown test reliably reach quarantine clear * no-mistakes(document): Reattach worker_shutdown doc comment to its function * no-mistakes(ci): Fixed the failing check (Behavior portable serial 7) with a test-only change to the stall test in tests/fm-remote-job.test.sh. Product code is unchanged; no other test changed. Cause: after the decoy dies, both workers run the same check-exists, read, delete sequence on the job records. On the CI runner the replacement deleted a record between the ousted worker's check and its read. The ousted worker exited 125, and because the file runs under set -e the unguarded `wait` ended the test with 125. The exit trap then killed the replacement, which produced the "Killed" line. Reproduction: a temporary 0.3 s delay between the check and the read, applied to the ousted worker only, made the committed test fail exactly as in CI (exit 125 and the "Killed" line). The new test passed with the same delay. The delay is reverted, along with a similar debug hook that the timed-out attempt had left in bin/fm-remote-job-worker.sh. Test changes: - The replacement is frozen (and confirmed stopped) before the decoy is killed and resumed only after the ousted worker exits, so only one worker touches the job records at a time. - The ousted worker is stopped only once its quarantine exists and its lane is reaped, which places it inside its stop loop. - Every fixed poll loop is now a wait on a named condition with a 30 s deadline and an explicit failure message. Exit detection also handles zombies. - The exit trap kills and waits for the decoy and both workers on every path. - A non-zero exit from the ousted worker now fails with its exit code and stderr instead of silently ending the file. The test still proves that the resumed ousted worker exits 0 and leaves the replacement's lock, quarantine contents and quarantine inode unchanged. Verification: the full test file passed four times on its own and three times under nice -n 10 with four busy-loop CPU hogs; bin/fm-lint.sh passes. Changes are not committed * no-mistakes(ci): I fixed the failing check (Behavior portable serial 7) by changing only the stall test in tests/fm-remote-job.test.sh. Product code is unchanged. **What failed:** "an ousted worker in shutdown leaves the replacement quarantine untouched" failed on CI with the ousted worker exiting 125 ("could not stop the active command tree"). **Why:** during shutdown, the worker retries the still-running decoy command group a fixed 100 times, 0.01 s apart, then gives up and exits 125. The test tried to freeze the worker partway through those retries by sending SIGSTOP from outside. On a slow runner the retries ran out before the stop arrived, so the worker had already given up. The invariant is that the test must hold the ousted worker inside that retry loop until the replacement owns the lock. That was the only place the test depended on timing. The other waits already watch for a named state change with a 30 s deadline. **Fix:** - The ousted worker now starts with a small `sleep` wrapper at the front of its PATH, and the SIGSTOP race is gone. - The wrapper only holds a `sleep` called directly by that worker's own process (it checks its parent pid against a hold file) while its quarantine file exists. - The only such `sleep` is the first retry in the shutdown stop loop, so the worker waits there as long as needed. - The wrapper writes a marker when it starts holding. The test waits for that marker, then hands the lock to the replacement, freezes the replacement, and kills the decoy. - The test releases the worker by deleting the hold file. Deleting the whole temp directory also releases it, so a failed run cannot leave the wrapper looping. - A process leak: the test overwrites the job's command-group record with the decoy, so no worker ever stopped the job's real command. `fm-hold-job.sh` and its `sleep 30` stayed running for up to 30 s after the test. The test now records that group before overwriting it and kills it at the end of the test and in the exit cleanup. - The test still asserts the same things: the ousted worker exits 0, and the replacement's lock, quarantine contents and quarantine inode are unchanged. **Verification:** - The full file passed twice on its own, twice under `nice -n 10` with six busy-loop CPU hogs, and twice more after the leak fix. - `pgrep` found no leftover processes afterwards. - With the worker from just before the fix commit (cf45cb6^), the test still fails with "the ousted worker wrote or cleared the replacement quarantine during shutdown", so it still proves the fix. - `bin/fm-lint.sh` passes. - I did not reproduce the CI failure locally. The cause comes from the fixed retry limit and the CI error message. The changes are not committed * no-mistakes(ci): I changed only the stall test ("an ousted worker in shutdown leaves the replacement quarantine untouched") in tests/fm-remote-job.test.sh. Product code is unchanged, and so is every other test. **Invariant:** the pid written to the job's group record must be a process-group leader whose group dies when that one process is killed. Otherwise the worker's bounded stop loop never sees the group die, gives up, and exits 125 ("could not stop the active command tree") before it reaches the lost-ownership exit. The decoy is the only place in this test that depends on this. **Fix:** - The decoy used to be `set -m; sleep 30 &`. It now starts as `perl -MPOSIX=setsid -e 'setsid() >= 0 or exit 1; exec @argv' sleep 30 &`, which gets its own session and group without shell job control. tests/fm-procevent.test.sh already uses the same idiom. - The test now waits, with the file's usual 30 s deadline and a named failure, until `ps -o pgid=` of the decoy equals its pid before writing it into the group record. This way the worker can never read the record before `setsid` has run. - The existing steps are unchanged: the test kills the decoy, reaps it with `wait` before releasing the hold file, and the exit trap still kills and reaps the decoy and both workers. - The assertions are unchanged: the ousted worker exits 0, and the replacement's lock pid, quarantine text and quarantine inode stay the same. **Cleanup:** I reverted a debug `printf` hook that the timed-out previous attempt had left in bin/fm-remote-job-worker.sh, and deleted its untracked `.tmp-repro/` directory. Neither was committed. **Verification:** - The full tests/fm-remote-job.test.sh passed twice normally and once under `setsid -w` with stdin from /dev/null (no controlling terminal). - `bin/fm-lint.sh` passes. - No leftover `sleep 30` processes afterwards. **Not reproduced:** I could not reproduce the CI failure locally. On this host `set -m` made the decoy its own group leader even without a controlling terminal, so the cause on the runner is not confirmed. The change removes the test's reliance on shell job control, as the user asked. Changes are not committed * no-mistakes(ci): I changed only the stall test ("an ousted worker in shutdown leaves the replacement quarantine untouched") in tests/fm-remote-job.test.sh. Product code is unchanged, and so is every other test. **Invariant:** the group record the ousted worker checks in its stop loop must stay the job's own command group, and the test must stop that group before it releases the hold. Otherwise the bounded retry keeps seeing a live group, gives up, and exits 125 ("could not stop the active command tree") before it reaches the lost-ownership exit. The test overwrote this record in one place (the decoy) and stopped the group in one place (killing the decoy); both are changed. **Fix:** - I removed the setsid decoy and the overwrite of `.claim/group`. The record keeps the job's real command group, which the test still saves as `STALL_JOB_GROUP`. - The two-line `group_start` stays. It is still needed: without it the worker kills the real group on its first pass, before the replacement takes over, so the hold would never matter. - The `sleep` wrapper that holds the worker at its first stop-loop retry is unchanged. - After the replacement owns the lock, its quarantine is planted and it is frozen, the test runs `kill -KILL -- -$STALL_JOB_GROUP`. It then waits, with the file's usual 30 s deadline and a named failure, until `kill -0` on the group fails. Only then does it remove the hold file. The worker therefore always sees its own command already stopped and never races its retry budget. - The exit trap still kills the saved command group if the test fails. It can't `wait` on that group because the group is not a child of the test shell. The decoy variable and its cleanup entry are gone. - The assertions are unchanged: the ousted worker exits 0, and the replacement's lock pid, quarantine text and quarantine inode stay the same. **Verification:** - The full tests/fm-remote-job.test.sh passed twice normally. - It passed once under `setsid -w` with stdin from /dev/null (no controlling terminal). - It passed once under `nice -n 10` with six busy-loop CPU hogs. - With the worker from before the fix (cf45cb6^), the test still fails with "the ousted worker wrote or cleared the replacement quarantine during shutdown", so it still proves the fix. - No `fm-hold-job` or `sleep 30` processes were left afterwards. - `bin/fm-lint.sh` passes. **Not reproduced:** I couldn't reproduce the CI failure locally; the decoy version also passed on this host. So I can't confirm why the decoy group stayed alive on the runner. The new wait turns any leftover live group into a clear named failure instead of an exit 125. The changes are not committed * fix(bin): keep a dead command group dead on bash 5.2 A bare return inside the liveness check drops the failing kill status when the check runs in a conditional, so shutdown keeps treating a stopped group as alive and exits 125. * no-mistakes(review): Use bash 3.2 fd syntax and fix trap return comments
…henguid#5589) * docs: make configuration settings easier to find and understand * no-mistakes(review): Restore dropped qualifiers and fix misplaced config doc labels * no-mistakes(review): Restore three dropped qualifiers in configuration reference
) * fix: bound worker edits of project AGENTS.md/CLAUDE.md to factual corrections These files are loaded into every agent session of a project, so additions should be a deliberate human choice rather than automated task output. The ship brief's project-memory section and AGENTS.md section 6 previously invited workers to record durable knowledge, which let project AGENTS.md files accrete detail the codebase or README already carries. Workers now edit only to fix factually wrong content - including content their own change made wrong - and fm-ensure-agents-md.sh runs only alongside such a correction. Stow no longer routes project-memory additions through ship tasks, and the generated skeleton no longer invites discovery-driven additions. * no-mistakes(review): Stop running fm-ensure-agents-md.sh on memory-file corrections * no-mistakes(document): Clarify manual project-memory initialization and remove duplicate guidance
…enguid#5635) * fix(bin): let gate agents drive lifecycle against marked lab homes Part 2 of the kunchenguid#5615 split. A no-mistakes gate agent runs inside a checkout carrying the fleet-captain identity, so fm-gate-refuse-lib refuses fleet mutation on the gate signal. That refusal was absolute, which kept gate validation from ever exercising the real lifecycle. Stamp a disposable lab FM_HOME with a .fm-lab-home marker file that only bin/fm-lab-home.sh writes, and only onto a fresh empty dir, so no call path can mark a populated real home. fm_refuse_if_gate_agent then permits lifecycle only when FM_HOME carries the marker and is driven through its stock layout - any FM_*_OVERRIDE relocation stays refused so part of the "lab" cannot be split back onto the real fleet. The threat model is a confused agent touching the real fleet, not deliberate forgery, so the marker is a plain token file rather than a bound record. FM_GATE_REFUSE_BYPASS is unchanged: it still serves the test harness, which cannot mark hundreds of temp homes. Teardown's slot-ownership scan compared state-dir paths textually while fm_firstmate_root_home canonicalizes, so a lab home under a symlinked TMPDIR scanned its own record twice and self-collided; compare file identity (-ef) instead. * no-mistakes(review): Refuse unlistable lab homes and hardlinked slot records * no-mistakes(review): Mint lab markers only on verified-empty fresh dirs * no-mistakes(document): Clarify lab-home gate documentation and comment contracts * no-mistakes(document): Clarify lab-home gate documentation and remove stale claims * no-mistakes(document): Clarify gate lab-home documentation and boundary wording
…ad of refusing every re-arm (kunchenguid#5594) * fix(bin): replace a watcher whose beacon stalls past a hard bound instead of refusing every re-arm A fleet watcher that is alive but whose liveness beacon has gone stale could never be replaced: every re-arm was refused because the lock holder was a live pid, and the holder was never evicted because it was not dead. Add FM_WATCHER_STALL_BOUND (default 3x the stale grace): below it the refusal is unchanged; at or past it the arm re-verifies the holder against the lock's recorded identity, sends TERM, waits boundedly, and takes the lock the normal way, ledgering a stalled-holder-replaced row. A holder that survives TERM keeps the old refusal. Fixes kunchenguid#4400 * no-mistakes(test): poll for replacement message to fix watcher-lock test flake * no-mistakes(document): document FM_WATCHER_STALL_BOUND in config inventory
…n can keep them (kunchenguid#5563) * fix(pi): hide queued Firstmate notifications under Calm only when the session can keep them Calm now keeps authenticated Firstmate operational inputs out of Pi's queued-message listing, but only after proving the live session exposes every member needed to keep them across Escape. A session missing any of them keeps stock rows and Escape and shows one generic warning. Escape and the dequeue key return only captain-authored messages to the editor and re-queue hidden notifications in order; after an abort that kept any in Pi's agent queue, the adapter starts the delivery turn itself because Pi 0.87.1 does not continue an aborted run. Compaction-held notifications stay with Pi's compaction flush and never start or announce a turn. Fixes kunchenguid#1588 * docs(calm): record Pi 0.87.1 queued-row retention verification * no-mistakes(review): Deliver kept Calm notifications after tree-navigation aborts too * no-mistakes(review): Defer Calm notification turn until tree navigation finishes * no-mistakes(lint): Silence SC2016 for literal JavaScript in queue-retention e2e test
…uid#5548) * fix(bin): refuse teardown when a required source disappears A missing sibling was sourced after cleanup had started, so Bash 3.2 exited 0 from the EXIT trap and Bash 5 continued and reported success. * no-mistakes(review): Remove unused FM_TEST_ONLY hook from teardown tests * no-mistakes(review): Check task backend sources before any teardown cleanup * test(gotmp): give teardown fixtures every tmux adapter sibling Teardown now refuses when a sibling the recorded backend's adapter sources is missing, so the fake bin must carry fm-session-lock-lib.sh, fm-agent-process-lib.sh and fm-gemini-lib.sh.
Restructure the supervision host doc's prose into shorter sections, lists, and tables without changing documented behavior. Every original heading, anchor, identifier, number, quoted string, and link target is preserved.
* docs: make herdr-backend easier to read Restructure the Herdr backend doc's prose into shorter sections, lists, numbered procedures, and tables without changing documented behavior. Every original heading, anchor, fenced code block, link target, and documented fact is kept. * no-mistakes(document): Restore composer-proof reason and complete Herdr topic table
Restructure the prose into sections, lists, and tables without changing documented behavior. Every original heading and anchor, inline-code span, link target, number, and quoted string is kept, and each sentence sits on its own line. Adds a topic navigation table and short subsections under the existing headings.
* docs: make watcher-continuity easier to read Restructure the prose into sections, lists, and tables without changing documented behavior. Every original heading, anchor, identifier, link target, and number is kept. * no-mistakes(review): Fix actor and supervision-host scope in watcher-continuity doc * no-mistakes(review): Make readiness TERM and retry conditional on unready successor
* docs: make sessionstart-nudge easier to read Restructure the prose into sections, lists, and tables without changing documented behavior. Every original heading, inline-code span, link target, number, and fact is preserved, and a harness-to-tier table now sits near the top. * no-mistakes(review): Drop helm glossary line and dedupe exit-code lead-in
* docs: make captain-hold-lifecycle easier to read Restructure the captain-hold lifecycle prose into sections, lists, and tables without changing documented behavior. Every original heading, anchor, identifier, number, quoted string, and link target is kept. * no-mistakes(review): Fix verification record subjects and grouping headings * no-mistakes(review): Clarify task-body read-back cases belong to the suite
* docs: make remote-secondmates easier to read Restructure the remote second mates prose into sections, lists, numbered procedures, and tables without changing documented behavior. Every original heading, anchor, fenced code block, identifier, link target, and qualifier is preserved. * no-mistakes(review): Merge remote-home table cell into one sentence * no-mistakes(review): Tighten readiness lead-in, restore causal link, fix dangling reference
…nguid#5554) * fix(bin): bound the away digest and log why a delivery failed The away daemon joined every buffered escalation into one unbounded digest. A start-up catch-all span can exceed what one transport argument carries (tmux rejects the send-keys command; Linux refuses to exec any argument above 131,071 bytes, which is how herdr receives it), so the initial send failed on every housekeeping pass and was logged as an unconfirmed Enter with text possibly in the composer. escalate_flush now builds the injected digest under a fixed byte budget: each event is cut at a UTF-8 boundary with an omitted-bytes marker, the joined events stop with a "+K more event(s)" tail, and a bounded digest names a state/.subsuper-digests/ file that keeps every buffered event verbatim. The buffer itself is untouched, so the return catch-up stays complete. The tmux submit core and the herdr literal send now replay the transport's stderr on failure, and inject_msg logs the failing stage (initial send versus Enter confirmation) with the byte count and that stderr. The wedge alarm line and marker carry the last failure reason. Fixes kunchenguid#4382 * no-mistakes(review): Drop digest pruning; label send-failed as send-or-Enter stage * no-mistakes(review): Keep digest full text once submit ran; reuse on retry * no-mistakes(lint): Count digest files with find instead of ls --------- Co-authored-by: firstmate-oss <firstmate@kunchenguid.local>
…henguid#5638) * feat(tests): add FM_TEST_SEAM launch seam and gate lab-primary recipe Part 1 of the kunchenguid#5615 split: the pieces that let the no-mistakes pipeline live-validate firstmate changes, without the gate-refusal rescoping. - bin/fm-afk-launch.sh: FM_TEST_HARNESS pins the detected harness only alongside the FM_TEST_SEAM=1 marker test suites set, so a leaked variable in a real primary's environment stays inert and unknown tokens fall through to real detection. - tests/lib.sh: export FM_TEST_SEAM=1 for every suite. - .no-mistakes.yaml: per-harness recipe for running a real fixture primary from a gate run - a plain mktemp lab FM_HOME on a private tmux socket, with FM_GATE_REFUSE_BYPASS=1 scoped to it and NO_MISTAKES_GATE scrubbed. - tests/fm-wake-queue.test.sh: stop the owned watcher fixture with KILL and clear its lifecycle state so the next leg starts clean; TERM could leave bash waiting in a child on some runners. - tests/fm-remote-secondmate-lifecycle-e2e.test.sh: wait for the liveness lock holder's post-acquire marker instead of the lock dir, which is published before the claim finishes. * no-mistakes(review): Scrub lab home overrides and require FM_TEST_SEAM separately * no-mistakes(document): Clarify test seam and disposable lab bypass documentation * no-mistakes(document): Clarify lab isolation and test-seam documentation * no-mistakes(ci): Fixed the CI failure: test cleanup killed the remote worker child but left its supervisor able to restart it during fixture removal. Cleanup now stops the worker tree. The lifecycle test passed locally; ShellCheck and diff checks passed
… home is gone (kunchenguid#5552) * fix(bin): refuse watchers from disposable checkouts and exit when the home is gone Fixes kunchenguid#321 Fixes kunchenguid#4760 A watcher armed from a disposable no-mistakes validation checkout under .no-mistakes/worktrees/ outlived the validation step and kept writing the real home's state, and a running watcher never noticed when its home, state directory, or code root disappeared. The arm now refuses from such a checkout with the typed failure line, the watcher checks once per poll that its home, state directory (or its own lock holder record), and bin directory still exist and exits with a logged reason scoped to itself, and the shared test helpers reap every watcher a suite armed for a temporary home through the home-scoped stop. * no-mistakes(lint): fix SC1007 by assigning empty string in watch-arm test * no-mistakes(ci): Found and fixed a genuine, reproducible hang introduced by this branch's test-watcher reaper, which is what killed both CI checks (serial-2 cancelled at the 30-min cap; Lint 2 exit 143 = the suite's own TERM-trap code). Root cause: test_drain_asserts_watcher_liveness (tests/fm-wake-queue.test.sh) fabricates a .watch.lock whose pid is the test runner's own $$ with the runner's real identity, to make the drain believe a live watcher exists. The new make_case tracking registers that state dir for reaping, so at fm_test_cleanup the new fm_test_reap_watchers drives fm-watch-arm.sh --stop; its identity check matches (the fixture recorded the runner's identity) and it kill -TERMs the test runner. tests/lib.sh:231 is `trap 'fm_test_cleanup; exit 143' TERM`, so the TERM re-enters cleanup -> reap -> kills $$ again -> infinite loop until the runner cap. I reproduced this locally: the suite ran all tests then looped forever in cleanup spawning fm-watch-arm.sh --stop against a lock naming its own PID. Fix (tests/lib.sh, +5 lines): in fm_test_reap_watchers, skip any tracked lock whose pid equals our own $$ before driving --stop. This is the single shared reap boundary; seven $$-self-lock fixtures across four test files are all covered by the one guard, and real armed watchers (pid != $$) are still reaped. Invariant: the test reaper must only signal real armed watcher processes, never the test runner itself. Verified locally: tests/fm-wake-queue.test.sh -> EXIT 0 (63 ok, no hang); tests/fm-watch-arm.test.sh -> EXIT 0 (21 ok, including test_reaper_stops_a_tracked_watcher, confirming the guard does not over-skip). Lint 2's exit 143 was the same shard/cap signature; a fresh CI run on this new commit will re-evaluate it --------- Co-authored-by: firstmate-oss <firstmate@kunchenguid.local>
…m as silence (kunchenguid#5588) * fix(bin): surface an unrecognized status prefix instead of dropping it A parked or holding declaration, and a verb whose correlation token did not parse, never became an event, so the supervisor still saw the earlier line. * no-mistakes(review): Require verb-shaped unrecognized status prefixes, add continuation tests * no-mistakes(document): Document unrecognized status prefix escalation in afk skill * no-mistakes(ci): I reproduced the "Behavior portable serial 6" failure locally and fixed it by changing the test data in one test. No product code changed. **What failed:** `tests/fm-session-start.test.sh`, in `test_orphan_status_logs_are_printed`, with "matched status log was printed 2 times". **Why:** the test writes status lines with made-up prefixes, `matched: surfaced once` and `orphan: step N`. The test only uses them as placeholder text. It checks that the session-start digest prints each task's status tail exactly once. This PR (kunchenguid#4763) deliberately makes an unrecognized one-word lowercase prefix a status event. So those lines now surface as captain-relevant events, and the wake queue's STATUS OUTCOME BACKSTOP section prints them a second time. The code under review is behaving as the issue asks. Only the test's placeholder data had become meaningful. **Rule the test depends on:** its status lines must not be captain-relevant, so the digest is the only place they are printed. Both lines in this test broke that rule. The orphan line would have failed the same count check right after the matched line did. **Fix:** in that test only, I switched both lines to the recognized, non-captain verb `working:`: `working: surfaced once` and `working: orphan step 1..6`. I updated the matching assertions and counts to use the new text. What the test checks is unchanged: orphan logs are labelled, the tail is bounded, the log path is printed, and each tail appears once. **Verification:** before the fix, the test failed locally the same way as in CI. After it, `bash tests/fm-session-start.test.sh` reports "all assertions passed * no-mistakes(review): Detect unrecognized prefixes on unstamped lines; share verb list --------- Co-authored-by: Kun's firstmate <kunchenguid+firstmate@users.noreply.github.com>
…henguid#5658) Fixes kunchenguid#5295 Session start now reports a remote inheritance failure using the push's own error line instead of the first unchanged item that happened to print before it, and the shared captain preferences header check now names the first required phrase it did not find, on both the local and remote inheritance paths.
…yloads (kunchenguid#5657) * fix(bin): stand down the Claude Stop auto-arm on pi-code-delivered payloads pi-code loads the tracked Claude settings but has no asyncRewake, so it awaits every Stop hook; without a stand-down the auto-arm runs synchronously inside Pi's turn end and holds it open for the declared multi-hour timeout. Stand down when the payload's transcript_path contains a /.pi/ path component, the same discriminator the closed-but- unmerged fix in kunchenguid#3352 used, with an explicit string-type check on the jq filter. Fixes kunchenguid#3343 * no-mistakes(document): document pi-code stand-down in harness integrations reference
…kunchenguid#5659) * fix(bin): match whole multi-word project names in the registry lookup bin/fm-project-mode.sh matched a registered project name against only the first whitespace-delimited token of a registry row, so a name containing a space never matched, silently defaulting the project to no-mistakes off instead of its declared posture. The lookup now matches the whole registered name against the raw line text, so a name is compared literally (never as a regex) and a name that is a leading prefix of another registered name still resolves to its own row. * no-mistakes(document): docs already accurate for multiword registry name match * chore: drop accidental empty err file Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com> --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
…d#5546) * fix(bin): classify the stdin program of `bash -s` with operands in the arm policy With -s, sh/bash/zsh read the program from stdin even when operands follow; the operands are only positional parameters. The arm policy treated the first operand as a script path, so heredoc and here-string payloads were never classified and a hidden bin/fm-watch.sh execution was allowed. A protected path in the operand position still fails closed as before. Fixes kunchenguid#1489 * no-mistakes(document): Clarify stdin shell operand documentation * no-mistakes(ci): Captain, fixed `shellInvocation` so `bash -- -s` treats `-s` as a script name, and updated R21 to test the exact command. The targeted policy suite, lint, documentation check, and diff check pass. Both hosted workflows show `action_required` before any jobs ran; that external approval state remains unresolved * fix(bin): keep main's handling of words after a leading `--` Revert the pipeline CI-step change that made the first word after a leading `--` always a script. It turned forms that main denies today into allow (for example `bash -- -c 'bin/fm-watch.sh'`), which is outside kunchenguid#1489 and loosens a fail-closed policy. `--` after `-s` still ends option parsing.
…nchenguid#5695) * fix(bin): strip AI co-author trailers from fleet-launched commits Cursor and other non-Claude runtimes append the trailer after the typed message. A per-task commit-msg hook removes it and leaves human co-authors and the author identity untouched. * no-mistakes(review): Export pane hooksPath override and drop generated-with stripping * no-mistakes(ci): This PR caused all three CI failures, and the fix is test-only: 4 test files change, no product code. **Cause.** `fm-spawn.sh` now installs the AI-trailer strip hooks for every spawn, secondmates included. The installer refuses a worktree that is not a git repository, and the PR deliberately keeps that fail-closed rule because real secondmate homes are firstmate clones. Four test fixtures still gave secondmates a plain directory as their home, so each spawn failed with "not a git worktree ... could not install the AI-trailer strip hooks": - serial 5: `tests/fm-backlog-atomicity.test.sh` ("secondmate spawn failed"). - serial 8: `tests/fm-secondmate-harness.test.sh` ("split: no meta written"). - Herdr: `tests/fm-backend-herdr-launcher-workspace-e2e.test.sh` and `tests/fm-backend-herdr-workspace-per-home-e2e.test.sh`. **Rule that must hold.** Every home a test spawns as a secondmate must be a git worktree. I checked the other places in the changed area: the only secondmate spawns in these tests are the ones listed. The earlier rounds already fixed the other fixtures (`fm-secondmate-liveness`, `fm-secondmate-safety`) the same way. **Fix.** - Each of those four secondmate homes now gets the same `.gitignore` plus `git init -q -b main` that the liveness and safety tests already use. - The two Herdr tests clean up with their own plain `rm -rf "$TMP_ROOT"`, not the shared `tests/lib.sh` helper. Because the installer leaves each `state/<id>.git-hooks` directory read-only, that cleanup printed "Permission denied" and left the directories behind. Both cleanups now restore the owner's write bit on every directory before removing (`find ... -exec chmod u+rwx`), which is what `fm_test_remove_tree` in `tests/lib.sh` does. **Verification.** - `tests/fm-secondmate-harness.test.sh` passes. - `tests/fm-backlog-atomicity.test.sh` passes (99 ok, exit 0). - shellcheck is clean on all four files. - I could not run the two real-Herdr tests locally: the Herdr lab on this host refuses to start because it needs exactly one running default session, and I did not change the host's Herdr state to get around that. Instead I checked their two changed steps directly: the installer succeeds on a home set up the new way, and the new cleanup removes the read-only hooks directory completely. Those two tests will only be proven on CI
…id#5683) * fix(bin): treat Pi's dollar-first cost footer as furniture An idle Pi status row opening with $0.000 was read as a dead-shell prompt, so exit and relaunch refused on an empty composer. * test: wait for the draining holder to exec sleep before reading its identity The procevent drain fixture read fm_pid_identity immediately after backgrounding setsid sleep, racing the child's exec chain. Mid-exec the cmdline can read empty, failing the fixture on a loaded CI runner. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…5534) * fix(bin): refuse a merge when a required check never reported fm-pr-merge.sh built its GitHub refusals only from checks present in statusCheckRollup, so a required check that never ran was simply absent and the merge proceeded on the subset that reported, contradicting its own "every required check green" claim. The GitHub verify now reads the base branch's required contexts from the forge itself - the classic branch protection summary on GET repos/{o}/{r}/branches/{b} and the active ruleset rules on GET repos/{o}/{r}/rules/branches/{b} - and refuses when a required context has no entry in the same rollup, at the same head, that the merge is bound to. Absence reads as unknown, never green. The required-set read joins the existing refusal list, so a draft, a red check, and an unreported required check are all reported together. Could not read vs nothing required: both endpoints need only repository read access. The admin-only GET .../branches/{b}/protection endpoint is deliberately not used: it answers a non-admin token with the same 404 an unprotected branch gets (observed live on kunchenguid/firstmate main with this token), which would read a missing permission as "nothing required". Any failed or malformed read of either source (auth, missing fine-grained permission, rate limit, network, 404, unexpected shape) refuses the merge with a line naming the unreadable source. The one exception is GitHub's plan-gated 403 on the rules endpoint ("Upgrade to GitHub Pro or make this repository public"), which already means "this repository has no branch rules" for the merge-queue reader; that check moves into one shared helper and the classic summary still decides for such a repository. Attended waiver: --allow-missing <check-name> is the twin of --allow-red and follows the same design and recording path: once, separate name argument, waives only that exact unreported required check, still requires every other required check reported and every check green, never waives an unreadable required set, refused while the away-posture record exists, and refused on GitLab. Merge-state BLOCKED policy is unchanged. How this differs from the withdrawn kunchenguid#5353 (read from its diff): - kunchenguid#5353 read the admin-only branches/{b}/protection endpoint and treated its 404 as "no required checks", so for any non-admin token the required set silently read as empty; this change reads the read-access branch summary and treats every failure as unreadable. - kunchenguid#5353 ignored rulesets; this change also reads required_status_checks rules from the effective branch rules. - kunchenguid#5353 made separate per-head REST reads of statuses and check-runs capped at per_page=100 with no pagination; this change checks presence in the same statusCheckRollup view the red-check gate already reads at the verified head. - kunchenguid#5353 stopped at the first unreadable read; this change reports it as one refusal among all the others. - kunchenguid#5353 also claimed kunchenguid#5345 (lock stealing) and changed 39 files, most unrelated; this change is kunchenguid#5344 only. Live proof, read-only (a gh wrapper refused every merge and mutating call): - cli/cli#14474 (trunk requires 3 classic build contexts, none ran): refused, naming build (macos-latest), build (ubuntu-latest), build (windows-latest); with --allow-missing "build (macos-latest)" it still refused, naming the other two. - cli/cli#13665 (ran build (ubuntu-24.04-firewall) instead): refused, naming build (ubuntu-latest). - hashicorp/terraform#39262 (ruleset-required checks absent): refused, naming Code Consistency Checks, End-to-end Tests, Race Tests, Unit Tests. - cli/cli#14485 (all required reported and green): verified; the wrapper blocked the merge call and the pull request read back open. Fixes kunchenguid#5344 * fix(review): Preserve required-check producers and aggregate independent read failures * fix(document): Clarify required-check verification and waiver documentation * fix(bin): match an app-bound required commit status by name The producer-identity check resolved an app-bound required context only against check runs, so a required context that the required app reports as a commit status could never match and always read as "has not reported". A commit status carries no app id to compare, so an app-bound requirement that arrives as a status now matches by name, as before producer binding; check runs keep requiring the configured producer app. Live, read-only: hashicorp/terraform#39262 requires license/cla from integration 865473, reported green as a commit status by the CLA app. The previous head refused it as unreported; this head no longer does, while still naming the four required check runs that never ran there. Refs kunchenguid#5344 * fix(document): Clarify accepted commit-status producer verification limitation --------- Co-authored-by: firstmate-oss <firstmate-oss@kunchenguid.local>
…uid#5696) * fix(bin): never offer a persistent secondmate for teardown The return brief's "Landed, cleanup due" scan listed every state/*.meta record carrying a pr= and a merge-notified marker without regard to kind, so a secondmate record holding a relayed child's merged PR put the mate itself up for "bin/fm-teardown.sh <mate>" cleanup. A secondmate is a persistent worker, never landed work. - bin/fm-afk-return.sh: skip kind=secondmate in the landed-cleanup scan. - bin/fm-pr-check.sh: refuse to record pr= or arm a merge watch on a kind=secondmate record before any side effect; a PR reported on its routed status channel belongs to a task in the mate's own home, which arms its own watch. - bin/fm-watch.sh: a merged result from a poll already armed on a secondmate retires the poll silently - no merge outcome, marker, or wake. * no-mistakes(document): Document secondmate merge-watch and return-brief exclusions * no-mistakes(ci): CI failed in an unchanged watcher-shutdown test whose three-second wait was sensitive to runner load. Increased the wait for both state- and home-deletion cases without changing watcher behavior. The full fm-watch-arm suite passed locally; syntax and diff checks passed
…unchenguid#5702) * fix(bin): refuse unknown dash-leading args in public-posting fm-x scripts fm-x-reply.sh collected any unrecognized argument into the positional pool and took the first one as the reply text, so an invocation like "fm-x-reply.sh <id> --followup --final <text>" posted the literal string "--final" to X and silently dropped the real text. Make argument parsing strict in every script that can post publicly: an unknown dash-leading argument, a dash-leading request_id/task id, a dash-leading option value, or a surplus positional now exits 2 with a usage error before any config load, outbox write, or network call. Reply text starting with '-' is still accepted via --text-file or stdin, and --help is honored wherever it appears instead of becoming text (a --help forwarded through fm-x-followup.sh would have counted as a posted follow-up and mutated the link). fm-x-link.sh and the fm-public-followup scripts already refuse unknown arguments; fm-x-poll.sh takes none. * no-mistakes(review): Refuse surplus follow-up text sources; drop post-ID help branches * no-mistakes(document): Clarify reply and follow-up argument usage * no-mistakes(review): Refuse dash-leading --text-file operands in fm-x-reply * no-mistakes(document): Correct follow-up argument parsing comment * no-mistakes(document): Document dismiss argument rejection in script header
…nchenguid#4806) * fix: stop quarantining ordinary shared-captain source updates * no-mistakes(document): Rewrap remote inherit header so usage prints fully
* docs: make calm easier to read Restructure the Calm mode prose into sections, lists, and tables without changing documented behavior. Every original heading, anchor, fenced code block, inline-code span, link target, number, and quoted string is preserved. * docs: restore reload case in calm override lead-in The Built-in tool override collisions lead-in covers a session that reloads with Calm already on, as the original text did.
* docs: make turnend-guard easier to read Restructure the turn-end guard doc's prose into shorter sections, lists, and tables without changing documented behavior. Every original heading, anchor, inline identifier, link target, and number is kept. * no-mistakes(review): Restore legacy-only scope on TERM retirement sentence * no-mistakes(review): Name Cursor park behavior in live e2e test line
…henguid#5872) * docs: move situational AGENTS.md sections into on-demand skills Backpass memory optimization: shrink the always-loaded AGENTS.md by moving situational contracts (home layout, session-start recovery, validation and landing supervision, scout completion, away/quiet supervision, Relay ownership) into agent-only skills loaded at their triggers, with a trigger index skill. * docs: classify the new on-demand skills' documentation audience Register the seven new agent-only skills as agent-runtime docs and fix a link in validation-supervision that kept its AGENTS.md-relative path. * docs: close load-timing gaps found by the live regression check - load validation-supervision whenever an ask-user finding is decided or answered, so forbid --yes and process-every-return reach the worker - keep the mid-task captain-ask rule, the unconfirmed network-checks rule, and the worker account pin rule inline in AGENTS.md - fix cross-references that still pointed at moved AGENTS.md sections
…guid#5879) * fix: route second-mate signal wakes by their new status span A second mate's status log is a shared channel carrying many independently keyed decisions, so judging its signal rows by every decision still open in the whole log pinned each routine update to main behind any unrelated parked hold. scopeForUnreadWake (the one owner for Pi and the attended supervision host) now judges a second-mate signal row by the lines presented since the last drain, bounded by the existing status-presentation cursor: a decision, blocked, resolution, or captain-held line, or a line declaring the key of a still-open decision, keeps the whole row on main, and any cursor problem falls back to the whole log. Keys are read only at the status parser's declared positions, with readable time stamps stripped as bin/fm-classify-lib.sh does. Single-task crewmate and stale routing are unchanged, and stale and signal rows for one mate keep independent verdicts. The supervision branch now treats a second mate's done and merged lines as relayed child outcomes, and fm-teardown refuses the branch actor second-mate retirement through the existing role-partition helper in both postures. * no-mistakes(review): Route second-mate resolutions to main only when closing open decision * no-mistakes(review): Guard bare-verb fold lines and fold resolution spans incrementally * no-mistakes(document): Clarify second-mate wake routing and retirement documentation * no-mistakes(ci): The CI failure was a timing-sensitive watcher teardown test, not the PR’s signal-scope code. Extended the bounded wait for both state-directory and home removal. The watcher suite passes locally, and git diff --check is clean
…extension (kunchenguid#5882) register-extension took the extension lifecycle lock and then the source lock, while reconcile republishing an unhandled extension result holds the source lock and reaches the lifecycle lock through the extension host's process-event path. Both waits are unbounded and both owners stay alive, so the two could wait on each other forever and freeze the home's monitoring cycle. register-extension now takes the source lock first, matching every other path that holds both. The lifecycle lock still spans binding resolution through registration publication, so binding retirement stays serialized. A new lifecycle-order section in the extension-binding suite, run in the default aggregate, holds a re-registration inside binding resolution while reconcile republishes that source's unhandled result and requires both to finish within a bound. Fixes kunchenguid#5866
* fix(bin): run the repository's own hooks when git -c carries the per-task hooksPath The per-task hooks wrappers cleared only the GIT_CONFIG_COUNT override before looking up the repository's own hooks directory. When core.hooksPath reached git through git -c (GIT_CONFIG_PARAMETERS), directly or inherited by a child process, the lookup found the wrapper directory again and exited 0, so the repository's real hook - such as a pre-push publish guard - never ran and the push succeeded. The lookup now ignores GIT_CONFIG_PARAMETERS too, so only the repository's config files decide its hooks directory, and a failed lookup exits nonzero instead of skipping the hook. AI-trailer stripping is unchanged. Fixes kunchenguid#5871 * no-mistakes(document): Document git hook chaining and lookup failure behavior
…unchenguid#5744) * feat(bin): add an optional never-send list to typed dispatch resolution * Added config/dispatch-never-send, an optional local list of literal values and re: regular expressions checked against every string of the resolver request before it is sent to typesafe.ai * A match, an unreadable list, or an empty or invalid pattern now stops the request and falls back to the off path, so firstmate dispatches through its existing intake; the one stderr diagnostic names at most the list line number and never the value * No list, or a list with no match, leaves resolution unchanged * no-mistakes(review): Match never-send literals across whitespace, drop regex mode * no-mistakes(review): Inherit the never-send list into secondmate homes * no-mistakes(ci): ci-2 (Lint 2), caused by this PR and now fixed. The rule that broke: the test script must not share shell variables with a library it sources. The new secondmate-inheritance test in tests/fm-dispatch-resolve.test.sh sourced bin/fm-config-inherit-lib.sh inside a `( ... )` subshell. That lib assigns `out`, so ShellCheck flagged every later `$out` in the test with SC2031. The subshell was the only place this PR sources that lib. The fix runs the propagation in a child shell instead (`bash -c '. "$1" && propagate_inheritable_config "$2" "$3"' _ lib from to`), so the test shell never sources the lib. `bin/fm-lint.sh tests/fm-dispatch-resolve.test.sh` now exits 0, and `bash tests/fm-dispatch-resolve.test.sh` passes. That includes the check that an inherited list is enforced in a secondmate home. ci-1 (Behavior portable serial 8), not caused by this PR, so no code change for it. The only failing test is tests/fm-remote-secondmate-relaunch.test.sh, which fails with "not ok - could not arm the PR poll fixture for the relaunch-ordering test". This PR doesn't touch that test or the code it runs. I reproduced the same failure locally on the base commit ea7c7f7. Main's own CI run on ea7c7f7 (run 36212602588) fails only this job, with the same message. The recent main runs before it also concluded failure
…hrashing guard (kunchenguid#5903) * feat(jev): add the guard framework and the memory RSS/swap thrashing guard A Jev guard is a bounded read-only host diagnostic that turns one class of resource pressure into a machine-readable audit record and a one-line verdict. This lands the framework contract (docs/jev-guards.md) with one representative family (Pattern 46, memory RSS and swap thrashing) as the pilot: thin bash wrapper, stdlib-only python engine, and a behavioral test through the CLI. * no-mistakes(review): fix jev mem guard fail-open unknown and contract * no-mistakes(review): fix mem guard rss pairing, memfree, tests, docs * no-mistakes(review): fix inverted UNKNOWN branch in fail-forcing test * no-mistakes(review): register docs/jev-guards.md in audience inventory * no-mistakes(review): simplify wrapper, drop dead top_n, single-source thresholds * no-mistakes(review): assert exact exit code in fail-forcing test leg * no-mistakes(review): tolerate any stdout encoding in text output * no-mistakes(document): Fix guard contract dash style and output wording
…enguid#5900) * fix(bin): stop slow GitHub reads from starving and waking the contributions poll The poll's fixed budget cut the tail URL's reads short, and any read killed at the bound was recorded as a forge failure, so routine GitHub slowness produced an observation-unavailable wake. A configured FM_CONTRIBUTIONS_BUDGET now rides the generated check shim and is cut down to the watcher's per-check bound with a margin, a read killed at the bound or the deadline is budget refusal that leaves records untouched, and the poll moves to the next URL that still has a full observation reserve instead of ending. * fix(bin): report the bound when a signal death leaks through fm_run_timed fm_run_external_timeout trusted the wrapper's recorded status over the runner's own bound verdict, so a read killed by the bound's TERM could surface as 143 and callers such as the contributions poll classified a budget-bound read as a forge failure. When the runner reports the bound, a signal-death status is the bound's own TERM racing the wrapper's bookkeeping and is reported as 124; a natural exit still passes through. The replace-shell probe also moves to the repo's portable BASHPID fallback so the suite runs on the macOS system bash. * no-mistakes(review): Rotate contribution polling and verify generated budget behavior * no-mistakes(review): Stabilize contribution rotation across successful observation refreshes * no-mistakes(review): Exclude settled contributions from live observation rotation * no-mistakes(document): Document contribution poll rotation and observation reserves * no-mistakes(ci): Captain, fixed timeout handling with a one-line change preserving natural exit 137. Full session-start and contributions suites, targeted regressions, and lint passed. The timeout suite still fails on the known Bash 3.2 BASHPID issue, left unchanged as instructed. Logs retained in .no-mistakes/ci-evidence/. Remote CI was not rerun * no-mistakes(ci): Fixed the fixture’s lock-acquisition race with a one-line bounded wait. Forced contention reproduced the CI error before the fix and passed afterward; the ordinary held-lock case, lint, syntax, and diff checks passed. Local Bash 3.2 failures remain: “cleanup lock bound 08 gave up before the marker lock freed” (also reproduced without the fix) and “TERM did not stop a watcher blocked inside a poll”. Remote CI was not rerun
…railers (kunchenguid#5859) * feat: add keep AI trailers setting * no-mistakes(document): Drop duplicate keep-ai-trailers line, update Cursor attribution note * no-mistakes(document): Honor keep-ai-trailers for Devin worker attribution * no-mistakes(document): Qualify Devin attribution note with keep-ai-trailers flag * no-mistakes(review): Inherit keep-ai-trailers into secondmate homes
…sh-command popup cannot hide the composer (kunchenguid#5876) * Fix Herdr composer reads blinded by the slash-command popup Capture the full visible viewport for every Herdr composer state and content read instead of a bounded tail. Claude Code renders its slash-command popup between the composer and the pane bottom, which pushes the composer outside a tail window. The pre-Enter payload proof then read an empty composer, judged a typed command unsent, and cleared it without pressing Enter. The proof-lines value now bounds only the clear cost, not the capture size. Growing the window adds rows above the composer only, so bottom-most shape selection and prior verdicts are unchanged. The suffix refusal is kept, and the shared inbox pending-line read stays a bounded tail. Portable regressions cover the popup-below-composer layout, and the live submit-confirmation guard gains a third exit scenario. The runtime-backends verification record documents the Herdr 0.9.0 and Claude Code 2.1.283 run. Closes kunchenguid#5533 * no-mistakes(document): Clarify composer capture bound ownership * no-mistakes(review): Drop unrequired bracketed-paste Enter fallback from live guard * no-mistakes(review): Latch trust prompt, check idle composer arm first
…configurable bound (kunchenguid#5917) Each per-task endpoint read in the session-start fleet digest runs in its own crash-isolated child under the existing timeout helper with a configurable FM_SESSION_START_ENDPOINT_TIMEOUT bound (default 10s). A hung or killed read becomes that task's endpoint error while the digest continues, and the wrapper banners any abnormal child exit naming its status.
…nchenguid#5886) * Keep lab tmux sockets on short private paths * no-mistakes(review): Fix Linux stat checks and fail-closed lab tmux teardown * no-mistakes(document): Update lab helper documentation for isolated tmux sockets
* Allow silent task-level no-change outcomes * no-mistakes(review): Exclude silent outcomes from captain-return handoffs * fix: look up supervision receipts by exact sequence * no-mistakes(review): Suppress silent notes in away-return brief * no-mistakes(review): Clarify visible notes; remove unused mode * no-mistakes(review): Clarify silent outcomes and avoid false drain promises * no-mistakes(document): Clarify silent supervision outcome documentation
…chenguid#5928) * feat: make /quiet a statement where the attended supervision host runs On a home that opted into the supervision host, quiet mode is what the attended host already does, so /quiet now enters nothing there instead of launching the quiet daemon and writing a record that would park a present captain's main. - bin/fm-afk-launch.sh quiet-check says quiet mode needs nothing where the attended host runs, or that the session is paused while its broken-session latch holds; a quiet enter refuses there before writing anything. - Where the home opted in but the attended host lacks a part (engine, tools, verified mirror writer, identifiable main session, valid mirror), quiet-check names it and quiet mode falls back to the daemon. - Under a live away record on that home, quiet-check and a quiet enter refuse and name the record, so the return runs first, whatever state/.afk says. - A quiet enter records mode: quiet in the posture record, so start and start-native launch the quiet daemon without FM_AFK_MODE, and the away refusal wording fires only for away. - bin/fm-host-mirror.sh check validates the dialog mirror read-only and exits 1 on a missing, unreadable, or invalid mirror. - The quiet and afk skills and the supervision-host docs describe the new behavior; homes without the opt-in and Pi homes keep the daemon path. * no-mistakes(review): Archive the quiet record when a quiet daemon start fails * no-mistakes(document): Clarify quiet-mode documentation and remove stale duplicates
…uid#5884) * fix(bin): grant Claude workers their task-channel dirs via --add-dir Since Claude Code 2.1.257, a file-tool read (Read/Glob/Grep, and an Edit's mandatory prior read) of a path outside the working directories parks --permission-mode auto panes on a one-time interactive question, and a "Block" answer lands permissions.blockReadsOutsideWorkingDirectories in user settings, refusing the same reads even under bypass. Firstmate launches Claude with no --add-dir, so a secondmate's parent-home steering inbox and a ship or scout worker's launch record, steering inbox, brief dir, and code-root .agents/skills were all outside: workers wedged on the question the first time they read a steer. Every Claude launch, spawn and relaunch, in both permission modes, now grants exactly the task's channel directories: state/<id>.inbox for a secondmate (in the parent home), or state/operational-inbox, state/<id>.inbox, data/<id>, and the code root's .agents/skills for a ship or scout. Paths resolve to real paths and lazily created channel dirs are made before launch so the grant never names a not-yet-existing directory; the whole state/ is deliberately never granted. The grant keeps the bypass-mode launch argv changed on purpose: it also protects bypass workers against a machine-recorded Block answer. * no-mistakes(document): Consolidate Claude launch guidance in configuration reference * no-mistakes(document): Clarify Claude permission documentation reference
…nguid#4819) * fix(supervision): prevent idle recovery loops without stranding wakes * no-mistakes(review): Remove unused wake-append rollback helper
…id#5889) * fix(bin): stop the remote-job worker busy-polling an idle queue The serving loop slept 50ms between passes and re-ran state preparation (chmod on every queue directory), the heartbeat publish, and the stale sweep on every pass. It now blocks on a worker.wake FIFO that staging, cancellation, and lane exit nudge, keeps a short fast-poll window after activity, refreshes the heartbeat at most once a second, and runs the sweep (which re-applies the queue directories' 0700 modes) at startup and then on a bounded interval. Lane-owned records are no longer re-read every pass. Measured with a fork/execve-interposing counter on a --serve worker in a disposable HOME and queue, bash 3.2, 20-second windows (the counter slows the old loop to about 5 passes a second, so real-host rates were higher): idle worker 146 forks/s, 61 execs/s -> 4.8 forks/s, 3.1 execs/s one running long job 232 forks/s, 100 execs/s -> 15 forks/s, 11 execs/s Stage-to-result latency for a no-op job, idle and back to back, stayed at about 0.8-1.2s in both versions (dominated by job execution, not pickup). * perf(bin): drop per-cycle forks from watcher, drain, and lock helpers The watcher, drain, inactive-reconcile scan, and branch-outcome reads forked small external commands on every cycle where bash can do the same work. - fm-wake-lib.sh gains fm_dirname_to, fm_basename_to, and fm_epoch_seconds_to, exact stand-ins for $(dirname --), $(basename --), and $(date +%s); the clock uses printf %(%s)T on bash 4.2+ and still forks date exactly once on stock macOS bash 3.2. - fm_lock_abs_path, fm_wake_signal_seen_path, fm_path_age, the watcher's age_of and wedge timer, and the recovery-marker line count use them or plain reads instead of dirname/basename/tr/date/wc. - window_to_task reads a meta file once instead of two grep | tail -1 | cut -d= -f2- pipelines per file per call. - fm-classify-lib.sh reads uname -s once at source time instead of in every status stat helper. - Libraries sourced every cycle derive their own directory without forking dirname, including the backend adapter siblings a subshell re-sources on each probe. tests/fm-fork-free-helpers.test.sh pins each replacement against the command it replaces on edge-case inputs, under every available bash and both the C and a UTF-8 locale; CI's stock macOS bash lane runs it under /bin/bash 3.2. Measured with a fork/execve-interposing counter in a disposable home, one tmux crew task, FM_POLL=1 (forks and execs per watcher cycle, per run otherwise): watcher cycle bash 5.3 299/138 -> 199/66 bash 3.2 341/146 -> 224/80 drain bash 5.3 492/238 -> 430/200 bash 3.2 567/250 -> 491/212 inactive scan bash 5.3 27/14 -> 17/4 bash 3.2 37/14 -> 17/4 branch-outcome bash 5.3 40/21 -> 35/16 bash 3.2 48/24 -> 38/19 * test: note the interpreter-expanded version probe for shellcheck * no-mistakes(review): Fix worker idle bounds and fork-free contributions snapshot * no-mistakes(review): Coalesce buffered worker wake nudges into one wake * no-mistakes(review): Coalesce wake nudges via pending marker so publishers never block * no-mistakes(review): Claim wake nudges atomically via noclobber pending marker * no-mistakes(review): Release abandoned wake claims only after a 30-second bound * no-mistakes(review): Drop worker wake FIFO; load path helpers side-effect free * no-mistakes(document): Document remote worker polling and preemption cadence * no-mistakes(ci): Fixed both failing CI shards: isolated remote and teardown test fixtures now include fm-path-lib.sh, which fm-wake-lib.sh requires. The three affected tests, fm-lint.sh, and git diff --check pass locally
…uid#5941) * fix: keep a successor watcher and remote-reply listeners across the gaps that dropped them A main-only supervision pass-through exited without leaving a watcher, and each remote-reply poll released its claim until the next cycle, so short-lived listeners stayed down. * no-mistakes(document): Clarify listener and supervision continuity documentation * no-mistakes(ci): Fixed the three Greptile findings: failed ingestion leaves one durable capture, failed reads exit instead of relistening, and the disposable-checkout guard rejects state paths outside the marked lab. Added behavioral tests; the remote-reply and watcher-lock suites, shell syntax checks, and git diff checks passed * no-mistakes(ci): Fixed a race in the wake-queue interruption test: it now waits for the drain to own the lock and enter handling before signaling it. The wake-queue suite, shell syntax check, and diff check pass
…enguid#5925) * fix: date replayed branch outcomes and ask main to check current state first A captain outcome main never acknowledged is presented again, which after a harness or posture switch, or the first drain after the upgrade whose earlier presenter never advanced the read cursor, can be days after its situation settled. The replay read as fresh news, so a PR since merged looked ready. bin/fm-branch-outcome.sh now adds a "recordedAgo" age (minutes, hours, then days) to present and unprocessed rows, one owner of that wording for both presenters. The drain's BRANCH OUTCOMES captain lines and the Pi branch's processing request name that age and ask main to check the task's current state first; an outcome already settled needs only the acknowledgement, with nothing relayed to the captain. Nothing is adopted as processed, so a fresh home's first outcome is still presented until acknowledged. * no-mistakes(review): Absent processed marker reads 0; never adopt read cursor * no-mistakes(review): Require recordedAgo in Pi requests; report undated rows to main * no-mistakes(review): Keep recordedAgo on captain rows only in present output * no-mistakes(document): Correct cutover documentation and retire stale migration guidance * fix: keep settled branch outcomes out of main's reply to the captain A live Pi primary that took over a host-drain home received the carried-over outcomes dated and check-first, but its processing reply still told the captain about an outcome whose decision had since been answered. The request also claimed every outcome was already shown as an anchor entry in this transcript, which is false for an outcome carried over from before a restart or a switch of primary. The Pi processing request now says each outcome was recorded earlier and may already have been seen or handled, and that a settled outcome gets no captain-facing mention at all in the reply or any recap, not even that it is settled. The drain's BRANCH OUTCOMES header and the supervision docs state the same rule, and the tests check both delivered texts. * fix: scope main's outcome reply to what is still open Telling main what not to say about a settled outcome was not enough: in two live Pi trials the processing reply still told the captain that an answered decision was settled. Main now sorts the outcomes by current state first, and its reply to the captain covers only the still-open ones, written as if the settled ones had never been listed. With that framing three live Pi trials kept the settled outcome out of the reply and relayed the open one each time. The drain's BRANCH OUTCOMES header and the supervision docs use the same framing, and the tests check both delivered texts. * no-mistakes(review): Clarify that main acknowledges every presented captain outcome * no-mistakes(document): Clarify outcome cursor ownership across Pi and host * no-mistakes(ci): Fixed Pi replay by batching unprocessed captain outcomes oldest-first and acknowledging only through each batch. Verified a marker-less backlog over 1 MiB replays through all batches. A real-drain regression confirms an older keyed decision remains under OPEN DECISIONS after a newer branch row is acknowledged; the check-first instruction now names those decisions. Relevant targeted tests and branch-supervision tests passed; the full host suite timed out * no-mistakes(ci): Fixed the host drain’s check-first wording in bin/fm-wake-drain.sh; the CI fixture now passes. The full host suite passed the affected fixtures but timed out later. Syntax and diff checks passed * no-mistakes(ci): Fixed ci-1: abbreviated Pi outcome summaries now stay within 1,024 characters and include a row-specific full-outcome lookup command. The delivered instruction requires reading the full outcome before acting, relaying, or acknowledging it. The new extension-driver regression failed before the fix and passes now; the Pi and supervision-host suites pass * no-mistakes(ci): Corrected the batching sentence in docs/pi-supervision-branch.md. The cancelled CI check needs no code fix; its clean rerun passed. The Pi branch extension suite and git diff check passed
…guid#5961) * fix: wake an idle Claude primary for attended main-only hand-backs An attended main-only pass-through confirmed a handling handoff for the successor it leaves running, which flipped the recovery marker to handling. The Claude Stop hook only rewakes main while that marker reads downtime, so the close reached no one and an idle primary slept with wakes queued. The pass-through now leaves the marker at downtime, and a close that turns main-only at its turn hands the consumed handoff back to downtime before it reaches main. Regression tests drive the real Stop hook around the real host on both paths and for the successor's own later close, and a new opt-in live guard proves it against an idle interactive Claude primary with a pre-fix negative control. * no-mistakes(review): Assert live lab Stop-hook registration via parsed settings JSON * no-mistakes(document): Correct supervision hand-back documentation * no-mistakes(ci): Fixed the failed downtime-write path so the Stop hook notifies main instead of silently dropping the close. Corrected the live guard’s tracked-hook check and added the requested at-turn main-only scenario. The new regression failed before the fix and passed after it; the host suite, syntax checks, and diff check pass. The credentialed live guard was not run in this CI phase because it writes outside the worktree * no-mistakes(ci): Fixed the Stop hook’s retry ordering: a crashed host gets its bounded retry before a non-crash hand-back failure is reported. The Stop-hook suite passes, including the crash regression. The host suite passed the failed-marker-write regression but timed out before completing; syntax and diff checks pass
… header, docs, skill
…n/fm-teardown.sh. No lint limits, roots or other files changed. Cause: the CI log shows only one failing root, bin/fm-teardown.sh, which ended with reason=memory at about 8.39 GiB RSS ("shellcheck: out of memory"). Every other root ended reason=ok. The merge brought in upstream growth in fm-wake-lib.sh, fm-classify-lib.sh and fm-lease-lib.sh, plus fm-path-lib.sh, and that pushed teardown's full-analysis ShellCheck past the address-space limit. Measured locally, peak RSS was 8.25 GiB at base and 8.55 GiB at HEAD. Invariant: fm-teardown.sh should load each of its sourced libraries once, from its top-level source block, and should not carry a source path that can never run. teardown_herdr_require_prerequisites had a second `. "$SCRIPT_DIR/fm-wake-lib.sh"` behind `if ! declare -F fm_lock_try_acquire`. That branch could never run. fm_lock_try_acquire is defined only in fm-wake-lib.sh, which is sourced unconditionally at the top of the script (line 386). All three callers of teardown_herdr_preflight_target (lines 3125, 3333, 3447) run after that line. ShellCheck still followed this dead source and analysed the whole wake-lib closure a second time. I removed the dead branch, as the removal-first rule asks. The existing check that refuses teardown when the lock functions are missing still stands, so the path still fails closed. Source-following itself was not narrowed or redirected, as the fm-lint.sh:880 comment forbids. The same wake-lib follows inside fm-public-followup-lib.sh and fm-pending-reply-lib.sh were left alone, because those libraries are part of other files' lint runs too and the user limited this round to teardown. Verification: - I ran the real guard, `FM_LINT_REQUIRE_BOUNDS=1 FM_LINT_JOBS=1 bin/fm-lint.sh --telemetry ... bin/fm-teardown.sh`, with its default 12 GiB cap. - The original file reproduces the CI failure: reason=memory, rc=251, rss 8386716 KiB, "shellcheck: out of memory". - The fixed file passes: reason=ok, rc=0, no findings, rss 7152892 KiB. That is about 1.2 GiB below the limit where CI failed, and lower than base. - These tests all pass with no failures: tests/fm-teardown.test.sh, tests/fm-teardown-endpoint-safety.test.sh, tests/fm-remote-secondmate-parent-binding.test.sh and tests/fm-backend.test.sh. - fm-teardown.test.sh includes the Herdr preflight cases (lock contention, preflight refusal, forced secondmate Herdr child preflight), so the edited function is exercised. - `bash -n` is clean. - I did not add a test, because a lint-memory test would be new machinery and a test that reads source text is forbidden. The before/after guard runs above are the regression evidence. - I did not run a full local `--partition 1of2`. Only teardown failed in the CI log, and that root was checked directly under the same guard
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Update firstmate and all its dependencies from Kun's repo only (2026-09-27 22:39 and 22:51), before a Pi firstmate takes over this home.
Upstream Firstmate (kunchenguid/firstmate) has 32 commits after this fork's last catch-up at 72b63ee, including quieter no-change supervision (ade7133), a memory and swap thrashing guard (6839202) and restored primary rewakes (1b82b77). Merge upstream into this fork.
Fetch upstream and merge upstream/main at its current head into a branch from this fork's main as a merge commit, with no rebase of published history, force-push, or reset. Keep every fork-unique feature with no upstream equivalent: off-table crew spawn refusal and missed-reply closes, exit and relaunch for quota-exhausted agents, the journey-walk skill, the cross-harness push skill, the prep gate and nav-prep install, Claude worker settings merge, ordered crew dispatch profiles, the Claude permission mode, board fixes, and the advisor line for eligible Claude workers. Where upstream fixed the same defect differently, take upstream's version and drop ours, naming the replaced fork commit. Carry any fork-only AGENTS.md text moved by 49a218b into the matching new skill or keep it in AGENTS.md. The full test suite passes; bin/fm-session-start.sh runs clean in an isolated scratch home. Complete the no-mistakes delivery with green checks.
What Changed
kunchenguid/firstmatemain through1b82b77aas a merge commit (no rebase or force-push), with the fork'smainmerged back in on top. The upstream changes are acrossbin/,.pi/extensions/, docs and tests. They include quieter no-change supervision outcomes (ade71339),/quietas a statement once attended supervision is ready (8c5493a0), and a new Jev guard framework with a memory RSS/swap thrashing guard (bin/fm-jev-mem-guard.{py,sh},docs/jev-guards.md,6839202a). They also restore primary rewakes after attended main-only closes (1b82b77a), add theconfig/keep-ai-trailersopt-out and the sharedbin/fm-path-lib.sh, and fix watcher, remote-worker, Herdr composer, contribution-poll and dispatch-resolver behavior.AGENTS.mdsections into on-demand skills, following upstream49a218bb. The new skills areagent-skill-trigger-index,away-quiet-supervision,operational-home-layout,scout-completion,session-start-recovery,ship-landingandvalidation-supervision. Fork-only wording from those sections moves word for word into the matching skills. Doc and skill references that pointed at the moved sections now point to the skills, and the fork's Codex quiet-mode refusal is kept in the quiet check, the header, the docs and thequietskill.tests/fm-supervision-host.test.shto the named timeout budgets, so the 30s+2s successor-confirmation path no longer flakes. Adds or extends tests for the upstream features, includingfm-jev-mem-guard,fm-fork-free-helpers,fm-supervision-host-attended-live-e2e,fm-lint,fm-crew-stateandfm-pi-branch-extension.Merge conflict resolutions
The upstream merge resolved the following 50 conflicts against fork head
a3f9be78.combinedmeans both sides were carried forward;upstreamorforkidentifies the chosen implementation for that conflicting file..agents/skills/afk/SKILL.md.agents/skills/captain-hold-lifecycle/SKILL.md.agents/skills/harness-adapters/references/harness/cursor.md.agents/skills/quiet/SKILL.md.pi/extensions/lib/fm-branch-dispatch.tsAGENTS.mdREADME.mdbin/backends/herdr.shbin/fm-afk-launch.shbin/fm-afk-return.shbin/fm-branch-outcome.shbin/fm-branch-prompt.shbin/fm-branch-report.shbin/fm-config-inherit-lib.shbin/fm-git-strip-ai-trailers.shbin/fm-host-mirror.shbin/fm-lab-home.shbin/fm-spawn.shbin/fm-supervise-daemon.shbin/fm-supervision-host.shbin/fm-teardown.shbin/fm-test-run.shbin/fm-wake-drain.shbin/fm-watch-arm.shdocs/calm.mddocs/captain-hold-lifecycle.mddocs/configuration.mddocs/herdr-backend.mddocs/pi-supervision-branch.mddocs/remote-secondmates.mddocs/scripts.mddocs/sessionstart-nudge.mddocs/supervision-host.mddocs/supervision-protocols/supervision-host.mddocs/turnend-guard.mddocs/watcher-continuity.mdtests/fm-afk-launch.test.shtests/fm-afk-return.test.shtests/fm-backend-herdr-workspace-per-home-e2e.test.shtests/fm-backend-herdr.test.shtests/fm-branch-supervision.test.shtests/fm-daemon.test.shtests/fm-gate-refuse.test.shtests/fm-git-strip-ai-trailers.test.shtests/fm-remote-job.test.shtests/fm-remote-secondmate-lifecycle-e2e.test.shtests/fm-spawn-dispatch-profile.test.shtests/fm-supervision-host.test.shtests/fm-watch-arm.test.shtests/fm-watcher-lock.test.shFork
mainmoved while this branch was in preparation.Merge commit
838a8972carried its Codex supervision repair (e599911e) through these four additional conflicts:AGENTS.mdbin/fm-watch.shdocs/turnend-guard.mddocs/watcher-continuity.mdFork work retained:
c9a45c02(off-table dispatch and missed-reply close),9a3aacc5(quota exit and relaunch),cdb1b34f(journey walk),a3f9be78(cross-harness push), and the earlier5c0addefcatch-up features (prep gate, Claude worker settings and permission mode, ordered dispatch profiles, board repairs, and eligible-worker advisor line).No fork commit was dropped wholesale or replaced by one upstream commit; individual conflicting files chosen from upstream are recorded above.
No manual state migration was identified in this catch-up; Firstmate will sync this home after merge.
Risk Assessment
Testing
The final published head
1046735falso repairs CI lint inbin/fm-teardown.sh. The unchanged file reproduced ShellCheck's memory failure under the existing 12 GiB address-space guard; removing four lines of unreachable duplicate sourcing passed the same guarded root at 7,152,892 KiB peak RSS, with four focused teardown/backend tests andbash -npassing. Earlier validation re-drove key fleet paths live in disposable lab homes. All 19 PR checks passed on the published head. With the old 25s wait, the new retry regression fails the same way the reported flake did. With the named 67s budget it passed twice (35-36s each, above the old limit), and the original write-failure test passed once. The loop was stopped early to report, so repeat coverage is limited to those runs, and the adjacent Stop-hook batch was not run. The following ran live in disposable lab homes, which were then torn down: bin/fm-session-start.sh in a scratch home (clean, exit 0), a real Claude primary on the lab tmux socket running session start (lock acquired, silent bootstrap, empty wake queue), off-table crew spawn refusal, and the Claude permission-mode and worker-settings refusals. All passed. There is no UI surface; the evidence is CLI transcripts and tmux pane captures. The transient runner scripts and the orphaned test fixture processes and temp directory from the stopped loop were removed, and the worktree is clean.not ok - hook retry write failure: the Stop hook did not finishwith HOOK_RETRY_POLLS=250Evidence: Retry regression fails with the old 25s wait (reproduces the flake)
Source: Retry regression fails with the old 25s wait (reproduces the flake)
Evidence: Retry and original write-failure Stop-hook tests with the named budget
Source: Retry and original write-failure Stop-hook tests with the named budget
Evidence: fm-session-start.sh in a scratch lab home at 4a3542e
Source: fm-session-start.sh in a scratch lab home at 4a3542e9
Evidence: Real Claude primary running session start in a lab home
Source: Real Claude primary running session start in a lab home
Evidence: Off-table crew spawn refusal
Source: Off-table crew spawn refusal
Evidence: Claude permission-mode and worker-settings refusals
Source: Claude permission-mode and worker-settings refusals
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
bin/fm-afk-launch.sh:309- The merge took upstream's quiet-mode text but kept the fork's Codex rule thatenterrefuses quiet mode (line 399). That leaves the script's runtime output and its header contract wrong for Codex. On a Codex home with config/supervision-host, the attended host never counts as ready, because FM_HOST_MIRROR_VERIFIED='claude cursor'.quiet-checkthen prints "...quiet mode enters through the quiet daemon instead." and exits 1, and the following quietenterrefuses because Codex runs no daemon. Other places that still describe the old Codex behavior: bin/fm-afk-launch.sh:17-24, where the header now says only Pi/pi-signed refusestartand 'every other harness still runs the daemon' (the fork's Codex sentence about the legacy-daemon handoff to the foreground checkpoint was dropped, but the code at 335 and 788-800 still does that); bin/fm-afk-launch.sh:38-40, where the QUIET MODE header says quiet 'enters through the daemon' when the host is unready; bin/fm-afk-launch.sh:324, a comment saying 'no longer launched on Pi' although the case at 335 also covers codex; docs/supervision-host.md:137, which says '/quiet names the missing part and enters the quiet daemon'; and .agents/skills/quiet/SKILL.md:47-48, which dropped the fork's instruction to tell the captain that quiet mode is unavailable on Codex, that/afkis the away posture there, and to stop. Fix: make the quiet-check line Codex-aware and restore the fork's Codex sentences in the upstream-structured header, comment, doc and skill..agents/skills/agent-skill-trigger-index/SKILL.md:15- Intent: "Carry any fork-only AGENTS.md text moved by 49a218b into the matching new skill or keep it in AGENTS.md." Some fork-only AGENTS.md lines were shortened instead of carried. (1) agent-skill-trigger-index:15research-first-decisionsdrops the triggers 'before recording such a selection as decided', 'integrating a library, SDK, API, CLI, framework, or service', and 'its Context7 version check is mandatory whether or not a candidate is being chosen'. (2) Line 16wayfindingdrops 'the queue looks fully gated' and 'the ready frontier lists only umbrellas' (now just 'empty'). This index says these skills load 'only at their precise triggers', so the shorter lines now narrow when they load, although each skill's own frontmatter still lists the full triggers. (3) operational-home-layout/SKILL.md:22.envdrops '(presence-gates bin/fm-dispatch-resolve.sh)', which both upstream and the fork had. (4) away-quiet-supervision/SKILL.md:16 drops 'Codex inside its foreground checkpoint loop'. (5) operational-home-layout/SKILL.md:25 claude-worker-settings drops its 'switch off add-on servers' example. Restoring the fork wording verbatim is the smallest remedy. It is flagged ask-user because it concerns what the intent requires to be carried.tests/fm-branch-supervision.test.sh:430- The merge commit changes upstream's 47h age sample to 47h30m, with a comment about a clock tick crossing the 48-hour boundary. recordedAgo in bin/fm-branch-outcome.sh:140-142 floors the age, so a 1-second tick cannot move 47h to 48h. The edit is harmless, but its stated reason is wrong and it differs from upstream for no need; consider reverting to upstream's line.🔧 Fix applied.
3 issues (2 warnings, 1 info) still open:
.agents/skills/agent-skill-trigger-index/SKILL.md:15- Intent: "Carry any fork-only AGENTS.md text moved by 49a218b into the matching new skill or keep it in AGENTS.md." Some fork-only AGENTS.md lines were shortened instead of carried. (1) agent-skill-trigger-index:15research-first-decisionsdrops the triggers 'before recording such a selection as decided', 'integrating a library, SDK, API, CLI, framework, or service', and 'its Context7 version check is mandatory whether or not a candidate is being chosen'. (2) Line 16wayfindingdrops 'the queue looks fully gated' and 'the ready frontier lists only umbrellas' (now just 'empty'). This index says these skills load 'only at their precise triggers', so the shorter lines now narrow when they load, although each skill's own frontmatter still lists the full triggers. (3) operational-home-layout/SKILL.md:22.envdrops '(presence-gates bin/fm-dispatch-resolve.sh)', which both upstream and the fork had. (4) away-quiet-supervision/SKILL.md:16 drops 'Codex inside its foreground checkpoint loop'. (5) operational-home-layout/SKILL.md:25 claude-worker-settings drops its 'switch off add-on servers' example. Restoring the fork wording verbatim is the smallest remedy. It is flagged ask-user because it concerns what the intent requires to be carried.tests/fm-branch-supervision.test.sh:430- The merge commit changes upstream's 47h age sample to 47h30m, with a comment about a clock tick crossing the 48-hour boundary. recordedAgo in bin/fm-branch-outcome.sh:140-142 floors the age, so a 1-second tick cannot move 47h to 48h. The edit is harmless, but its stated reason is wrong and it differs from upstream for no need; consider reverting to upstream's line..agents/skills/agent-skill-trigger-index/SKILL.md:15- Still present after the round 1 fix; it was left unselected and is still waiting on a decision. The intent says: "Carry any fork-only AGENTS.md text moved by 49a218b into the matching new skill or keep it in AGENTS.md." The merge resolution shortened several fork-only AGENTS.md lines instead of carrying them over. (1) agent-skill-trigger-index/SKILL.md:15research-first-decisionsdrops 'before recording such a selection as decided', 'integrating a library, SDK, API, CLI, framework, or service', and 'its Context7 version check is mandatory whether or not a candidate is being chosen' (fork AGENTS.md:620). (2) Line 16wayfindingdrops 'such as a stage, a release, a migration, or a campaign', 'the queue looks fully gated', and 'the ready frontier lists only umbrellas or nothing' (fork AGENTS.md:621). This index says skills load 'only at their precise triggers', so the shorter lines narrow when these skills load. (3) operational-home-layout/SKILL.md:22.envdrops '(presence-gates bin/fm-dispatch-resolve.sh)' (fork AGENTS.md:71). (4) away-quiet-supervision/SKILL.md:16 drops 'Codex inside its foreground checkpoint loop' (fork AGENTS.md:505). (5) operational-home-layout/SKILL.md:25 claude-worker-settings drops 'e.g. to switch off add-on servers workers never use' (fork AGENTS.md:74). The smallest remedy is to restore the fork wording verbatim. This is marked ask-user because it turns on how much the carry requirement covers.🔧 Fix applied.
1 info still open:
tests/fm-branch-supervision.test.sh:430- The merge commit changes upstream's 47h age sample to 47h30m, with a comment about a clock tick crossing the 48-hour boundary. recordedAgo in bin/fm-branch-outcome.sh:140-142 floors the age, so a 1-second tick cannot move 47h to 48h. The edit is harmless, but its stated reason is wrong and it differs from upstream for no need; consider reverting to upstream's line.🔧 **Test** - 1 issue found → auto-fixed ✅
tests/fm-supervision-host.test.sh:1140- test_claude_stop_hook_notifies_when_at_turn_downtime_write_fails flakes with 'hook write failure: the Stop hook did not finish'. It failed once in the targeted batch, once in two solo reruns, and intermittently in focused loops. The test file, bin/fm-supervision-host.sh and bin/fm-claude-stop-autoarm.sh are byte-identical to upstream 1b82b77, and upstream's untouched tree fails the same assertion. So this was inherited from upstream, not introduced by the merge. With a longer wait, both trees always reach the correct outcome (exit 2 plus the failure notice to main). The Stop hook's exit time is bimodal on both trees: 0-5s or 34-57s, with nothing in between, while the test waits only 25s (wait_until 250). The ~32s slow path matches the handling-successor confirmation budget in start_handling_successor (FM_ARM_CONFIRM_TIMEOUT default 30, plus 2s) in bin/fm-claude-stop-autoarm.sh:339-367. That suggests the successor arm sometimes fails to confirm after the failed downtime write and the hook waits out the full budget. CI runs this file and the intent requires green checks, so this can flake a required shard. The fix is to either bound the wait on that budget or set FM_ARM_CONFIRM_TIMEOUT low in this case, ideally upstream, rather than silently widening the wait. Evidence: flake-branch-timing.log, flake-upstream-timing.log, flake-upstream-1b82b77a.log, flake-branch-diag.log.fm-control.sh relaunch --model claude-fable-5-1the meta read model=claude-fable-5-1 and the brief had 0 adviso…bin/fm-lab-home.sh create $LABthenFM_HOME=$LAB bin/fm-session-start.shin a scratch lab home, followed bybin/fm-startup-network.sh reportRealclaudeprimary started on the lab's private tmux socket (tmux -L fm-lab new-session ... -e FM_HOME=$LAB claude) from the gate worktree and prompted to runbin/fm-session-start.shbin/fm-jev-mem-guard.sh --help, bare run,--json, and--check --warn-mem-pct 1 --crit-mem-pct 2on the live hostbin/fm-dispatch-select.sh <lab>/config/crew-dispatch.json 0, plus two off-tablebin/fm-spawn.sh ... --scoutattempts (default rule and--dispatch-rule 0)bin/fm-spawn.shrefusals withconfig/claude-permission-mode=yoloand withconfig/claude-worker-settings.json=[1,2]Real Claude scout spawn withclaude-permission-mode=autoand a worker settings object, with the launched argv and child MCP processes inspectedbin/fm-control.sh scout-z1 relaunch --model claude-fable-5-1 --note ..., checking the launch-brief advisor section before and afterbin/fm-control.sh scout-z1 exitcalled twice (stopped, then already-stopped), thenbin/fm-teardown.sh scout-z1 --forcebin/fm-test-run.sh --per-script-timeout-secs 1500on 21 targeted suites: fm-jev-mem-guard, fm-supervision-host, fm-afk-return, fm-branch-supervision, fm-supervision-instructions, fm-claude-stop-autoarm, fm-control-relaunch, fm-control, fm-spawn-dispatch-profile, fm-dispatch-resolve, fm-prep-install, fm-brief, fm-bearings-board, fm-captain-hold-lifecycle, fm-spawn-compact-adviser-disable, fm-send-resolve-key, fm-pending-reply-retire-notice, fm-codex-session, fm-session-start, fm-watcher-lock, fm-afk-launchbash tests/fm-supervision-host.test.shrerun twice alone, plus an 8-iteration focused loop oftest_claude_stop_hook_notifies_when_at_turn_downtime_write_failson this branch and on agit archive 1b82b77atree (temporary drivers, since removed)A second lab withconfig/supervision-hostand a live Claude scout in flight, a real Claude primary turn end, and thefm_primary_scope_matchesscope checkgit log/git merge-base --is-ancestorchecks of the merge parents and of fork and upstream feature commits;git diff e599911e HEADon the fork skill and feature files🔧 Fix applied.
✅ Re-checked - no issues remain.
not ok - hook retry write failure: the Stop hook did not finishwith HOOK_RETRY_POLLS=250git diff --stat b4c245d2 HEAD(only tests/fm-supervision-host.test.sh changed since the round-1 live validation)Focused runner (the test file's definitions plus selected calls), with HOOK_RETRY_POLLS forced to 250 (the old 25s wait):test_claude_stop_hook_retry_notifies_when_at_turn_downtime_write_fails, which gavenot ok - hook retry write failure: the Stop hook did not finishFocused runner at HEAD (HOOK_RETRY_POLLS=670):test_claude_stop_hook_retry_notifies_when_at_turn_downtime_write_failspassed twice (36s, 35s) andtest_claude_stop_hook_notifies_when_at_turn_downtime_write_failspassed once (4s); the loop was stopped early after thatbin/fm-lab-home.sh create $LABthenFM_HOME=$LAB bin/fm-session-start.shat 4a3542e9 (lock acquired, BOOTSTRAP silent, no queued wakes, exit 0)Real Claude primary on the lab tmux socket (tmux -L fm-lab new-session ... -e FM_HOME=$LAB2 claude), driven to run bin/fm-session-start.sh, thentmux -L fm-lab kill-serverandrm -rf $LAB2bin/fm-spawn.sh offtable-z1 ... --harness claude --model opus --effort highand--dispatch-rule 0 --harness codex --model gpt-5 --effort lowagainst a lab config/crew-dispatch.json (both refused, no state written)bin/fm-spawn.sh scout-z1 ... --harness claudewith an invalid config/claude-permission-mode, then with a non-object config/claude-worker-settings.json (both refused, no state written)bin/fm-merge-outcome-lib.sh:66- The runtime merge-landed reminder still ends with 'see AGENTS.md section 7'. The merge moved the post-merge verification rule into the ship-landing skill, and section 7 now only points there, so the reminder still leads the reader to the right rule, one hop later. Changing the string is a behavior change: the pinned MERGE_LANDED_REMINDER test fixture would have to change with it, so this documentation phase left both alone. Follow-up: point the reminder and its fixture at ship-landing together.✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.