fix(bin): merge upstream firstmate through 72b63eee - #31
Merged
Merged
Conversation
…ning (kunchenguid#5566) * fix(bin): report a Lavish source armed only after its listener is running Registration alone was treated as ready, so arm could succeed before anything was collecting from the board. * no-mistakes(review): Guard Lavish arm launches, keep retire refusals, report live prior listener * no-mistakes(review): Keep polling through window before reporting a still-live prior listener * test: wait for a capture's claim to drop before the next arm The result is stored before the runner exits, so a re-arm in that gap was meeting a live claim. * no-mistakes(document): Record Lavish arm readiness evidence in verification doc * no-mistakes(ci): Both failures were caused by this PR, and both are fixed with test-only edits. Lint 2 (ShellCheck SC2034): this branch removed the only use of `reply_id` (a `start "$reply_id"` call) from tests/fm-procevent.test.sh, which left the assignment at line 1450 unused. I deleted that assignment. It was the only `reply_id` in the file. ShellCheck is now clean on both test files. Behavior portable serial 4: the failing test was tests/fm-bearings-board.test.sh, in the check "registration consumed its answer before the any-origin binding existed". I reproduced it locally: the hold was still `state: queued` when the test checked it. - What must hold: the test's check that the hold is closed must run after the listener has captured the answer. - Why it broke: the test used a stand-in adapter that ran `fm-procevent.sh start` in the foreground after `arm`, so capture finished before build returned. On this branch, `arm` starts the listener itself in the background, so the real listener captures the answer and closes the hold a moment after build returns. - Fix: removed the now-redundant stand-in adapter, the copied runtime directory, and its extra environment variables. The test now runs the real build through the existing `run_board` helper and waits up to about 10s for the hold to reach `state: done`. The checks that follow are unchanged: `Resolution mode: answered` and the any-origin binding. - Other tests: this was the only test in the file that stood in for the adapter this way. The shard's other pure-contract-unit test (tests/fm-trace-context-lib.test.sh) passed unchanged. Verification: - tests/fm-bearings-board.test.sh passed 3 times in a row via bin/fm-test-run.sh, all 18 checks, about 53s per run. - tests/fm-procevent.test.sh was not rerun, because the lint fix only removed an unused assignment
…elivered (kunchenguid#5599) * fix(bin): acknowledge a delivered unknown-wake escalation The same unrecognized wake was escalated again after it had already been handled, because delivery never recorded that identity. * no-mistakes(review): Scope unknown-wake acknowledgements to one away session * no-mistakes(review): Clear delivered digest when unknown-wake ack write fails * no-mistakes(review): Limit unknown-wake suppression to acknowledged lines * no-mistakes(document): List unknown-wake ack file among away-session artifacts
…ess wait (kunchenguid#5587) * fix(bin): keep a stated default retraction from cancelling a keyless wait A resolved line that names the shared default decision bucket was closing the keyless live wait that only prints as that same key. Keyless self-retraction still closes the keyless wait. * no-mistakes(review): Keep declared waits standing past foreign-key resolved lines * no-mistakes(review): Bound declared-wait read and share one decision-key parser * no-mistakes(document): Document supervisors' key-aware declared-wait read
…unchenguid#5544) * fix(bin): terminate a remote job worker that lost ownership when it receives TERM A serving worker whose lock directory is gone can no longer quarantine shutdown, and resuming service publishes a false ready heartbeat. Exit after stopping only that worker's own command tree, without removing a replacement owner's lock. * no-mistakes(review): Check worker lock ownership before publishing shutdown quarantine * no-mistakes(document): Correct worker shutdown comment on replacement-owned lock * fix(bin): keep an ousted remote job worker off the replacement quarantine Shutdown can lose the lock after the first ownership check and before it writes or clears quarantine. Bind both operations to the directory object this process still owns so a replacement's quarantine stays untouched. * no-mistakes(review): Make ousted-worker shutdown test reliably reach quarantine clear * no-mistakes(document): Reattach worker_shutdown doc comment to its function * no-mistakes(ci): Fixed the failing check (Behavior portable serial 7) with a test-only change to the stall test in tests/fm-remote-job.test.sh. Product code is unchanged; no other test changed. Cause: after the decoy dies, both workers run the same check-exists, read, delete sequence on the job records. On the CI runner the replacement deleted a record between the ousted worker's check and its read. The ousted worker exited 125, and because the file runs under set -e the unguarded `wait` ended the test with 125. The exit trap then killed the replacement, which produced the "Killed" line. Reproduction: a temporary 0.3 s delay between the check and the read, applied to the ousted worker only, made the committed test fail exactly as in CI (exit 125 and the "Killed" line). The new test passed with the same delay. The delay is reverted, along with a similar debug hook that the timed-out attempt had left in bin/fm-remote-job-worker.sh. Test changes: - The replacement is frozen (and confirmed stopped) before the decoy is killed and resumed only after the ousted worker exits, so only one worker touches the job records at a time. - The ousted worker is stopped only once its quarantine exists and its lane is reaped, which places it inside its stop loop. - Every fixed poll loop is now a wait on a named condition with a 30 s deadline and an explicit failure message. Exit detection also handles zombies. - The exit trap kills and waits for the decoy and both workers on every path. - A non-zero exit from the ousted worker now fails with its exit code and stderr instead of silently ending the file. The test still proves that the resumed ousted worker exits 0 and leaves the replacement's lock, quarantine contents and quarantine inode unchanged. Verification: the full test file passed four times on its own and three times under nice -n 10 with four busy-loop CPU hogs; bin/fm-lint.sh passes. Changes are not committed * no-mistakes(ci): I fixed the failing check (Behavior portable serial 7) by changing only the stall test in tests/fm-remote-job.test.sh. Product code is unchanged. **What failed:** "an ousted worker in shutdown leaves the replacement quarantine untouched" failed on CI with the ousted worker exiting 125 ("could not stop the active command tree"). **Why:** during shutdown, the worker retries the still-running decoy command group a fixed 100 times, 0.01 s apart, then gives up and exits 125. The test tried to freeze the worker partway through those retries by sending SIGSTOP from outside. On a slow runner the retries ran out before the stop arrived, so the worker had already given up. The invariant is that the test must hold the ousted worker inside that retry loop until the replacement owns the lock. That was the only place the test depended on timing. The other waits already watch for a named state change with a 30 s deadline. **Fix:** - The ousted worker now starts with a small `sleep` wrapper at the front of its PATH, and the SIGSTOP race is gone. - The wrapper only holds a `sleep` called directly by that worker's own process (it checks its parent pid against a hold file) while its quarantine file exists. - The only such `sleep` is the first retry in the shutdown stop loop, so the worker waits there as long as needed. - The wrapper writes a marker when it starts holding. The test waits for that marker, then hands the lock to the replacement, freezes the replacement, and kills the decoy. - The test releases the worker by deleting the hold file. Deleting the whole temp directory also releases it, so a failed run cannot leave the wrapper looping. - A process leak: the test overwrites the job's command-group record with the decoy, so no worker ever stopped the job's real command. `fm-hold-job.sh` and its `sleep 30` stayed running for up to 30 s after the test. The test now records that group before overwriting it and kills it at the end of the test and in the exit cleanup. - The test still asserts the same things: the ousted worker exits 0, and the replacement's lock, quarantine contents and quarantine inode are unchanged. **Verification:** - The full file passed twice on its own, twice under `nice -n 10` with six busy-loop CPU hogs, and twice more after the leak fix. - `pgrep` found no leftover processes afterwards. - With the worker from just before the fix commit (cf45cb6^), the test still fails with "the ousted worker wrote or cleared the replacement quarantine during shutdown", so it still proves the fix. - `bin/fm-lint.sh` passes. - I did not reproduce the CI failure locally. The cause comes from the fixed retry limit and the CI error message. The changes are not committed * no-mistakes(ci): I changed only the stall test ("an ousted worker in shutdown leaves the replacement quarantine untouched") in tests/fm-remote-job.test.sh. Product code is unchanged, and so is every other test. **Invariant:** the pid written to the job's group record must be a process-group leader whose group dies when that one process is killed. Otherwise the worker's bounded stop loop never sees the group die, gives up, and exits 125 ("could not stop the active command tree") before it reaches the lost-ownership exit. The decoy is the only place in this test that depends on this. **Fix:** - The decoy used to be `set -m; sleep 30 &`. It now starts as `perl -MPOSIX=setsid -e 'setsid() >= 0 or exit 1; exec @argv' sleep 30 &`, which gets its own session and group without shell job control. tests/fm-procevent.test.sh already uses the same idiom. - The test now waits, with the file's usual 30 s deadline and a named failure, until `ps -o pgid=` of the decoy equals its pid before writing it into the group record. This way the worker can never read the record before `setsid` has run. - The existing steps are unchanged: the test kills the decoy, reaps it with `wait` before releasing the hold file, and the exit trap still kills and reaps the decoy and both workers. - The assertions are unchanged: the ousted worker exits 0, and the replacement's lock pid, quarantine text and quarantine inode stay the same. **Cleanup:** I reverted a debug `printf` hook that the timed-out previous attempt had left in bin/fm-remote-job-worker.sh, and deleted its untracked `.tmp-repro/` directory. Neither was committed. **Verification:** - The full tests/fm-remote-job.test.sh passed twice normally and once under `setsid -w` with stdin from /dev/null (no controlling terminal). - `bin/fm-lint.sh` passes. - No leftover `sleep 30` processes afterwards. **Not reproduced:** I could not reproduce the CI failure locally. On this host `set -m` made the decoy its own group leader even without a controlling terminal, so the cause on the runner is not confirmed. The change removes the test's reliance on shell job control, as the user asked. Changes are not committed * no-mistakes(ci): I changed only the stall test ("an ousted worker in shutdown leaves the replacement quarantine untouched") in tests/fm-remote-job.test.sh. Product code is unchanged, and so is every other test. **Invariant:** the group record the ousted worker checks in its stop loop must stay the job's own command group, and the test must stop that group before it releases the hold. Otherwise the bounded retry keeps seeing a live group, gives up, and exits 125 ("could not stop the active command tree") before it reaches the lost-ownership exit. The test overwrote this record in one place (the decoy) and stopped the group in one place (killing the decoy); both are changed. **Fix:** - I removed the setsid decoy and the overwrite of `.claim/group`. The record keeps the job's real command group, which the test still saves as `STALL_JOB_GROUP`. - The two-line `group_start` stays. It is still needed: without it the worker kills the real group on its first pass, before the replacement takes over, so the hold would never matter. - The `sleep` wrapper that holds the worker at its first stop-loop retry is unchanged. - After the replacement owns the lock, its quarantine is planted and it is frozen, the test runs `kill -KILL -- -$STALL_JOB_GROUP`. It then waits, with the file's usual 30 s deadline and a named failure, until `kill -0` on the group fails. Only then does it remove the hold file. The worker therefore always sees its own command already stopped and never races its retry budget. - The exit trap still kills the saved command group if the test fails. It can't `wait` on that group because the group is not a child of the test shell. The decoy variable and its cleanup entry are gone. - The assertions are unchanged: the ousted worker exits 0, and the replacement's lock pid, quarantine text and quarantine inode stay the same. **Verification:** - The full tests/fm-remote-job.test.sh passed twice normally. - It passed once under `setsid -w` with stdin from /dev/null (no controlling terminal). - It passed once under `nice -n 10` with six busy-loop CPU hogs. - With the worker from before the fix (cf45cb6^), the test still fails with "the ousted worker wrote or cleared the replacement quarantine during shutdown", so it still proves the fix. - No `fm-hold-job` or `sleep 30` processes were left afterwards. - `bin/fm-lint.sh` passes. **Not reproduced:** I couldn't reproduce the CI failure locally; the decoy version also passed on this host. So I can't confirm why the decoy group stayed alive on the runner. The new wait turns any leftover live group into a clear named failure instead of an exit 125. The changes are not committed * fix(bin): keep a dead command group dead on bash 5.2 A bare return inside the liveness check drops the failing kill status when the check runs in a conditional, so shutdown keeps treating a stopped group as alive and exits 125. * no-mistakes(review): Use bash 3.2 fd syntax and fix trap return comments
…henguid#5589) * docs: make configuration settings easier to find and understand * no-mistakes(review): Restore dropped qualifiers and fix misplaced config doc labels * no-mistakes(review): Restore three dropped qualifiers in configuration reference
) * fix: bound worker edits of project AGENTS.md/CLAUDE.md to factual corrections These files are loaded into every agent session of a project, so additions should be a deliberate human choice rather than automated task output. The ship brief's project-memory section and AGENTS.md section 6 previously invited workers to record durable knowledge, which let project AGENTS.md files accrete detail the codebase or README already carries. Workers now edit only to fix factually wrong content - including content their own change made wrong - and fm-ensure-agents-md.sh runs only alongside such a correction. Stow no longer routes project-memory additions through ship tasks, and the generated skeleton no longer invites discovery-driven additions. * no-mistakes(review): Stop running fm-ensure-agents-md.sh on memory-file corrections * no-mistakes(document): Clarify manual project-memory initialization and remove duplicate guidance
…enguid#5635) * fix(bin): let gate agents drive lifecycle against marked lab homes Part 2 of the kunchenguid#5615 split. A no-mistakes gate agent runs inside a checkout carrying the fleet-captain identity, so fm-gate-refuse-lib refuses fleet mutation on the gate signal. That refusal was absolute, which kept gate validation from ever exercising the real lifecycle. Stamp a disposable lab FM_HOME with a .fm-lab-home marker file that only bin/fm-lab-home.sh writes, and only onto a fresh empty dir, so no call path can mark a populated real home. fm_refuse_if_gate_agent then permits lifecycle only when FM_HOME carries the marker and is driven through its stock layout - any FM_*_OVERRIDE relocation stays refused so part of the "lab" cannot be split back onto the real fleet. The threat model is a confused agent touching the real fleet, not deliberate forgery, so the marker is a plain token file rather than a bound record. FM_GATE_REFUSE_BYPASS is unchanged: it still serves the test harness, which cannot mark hundreds of temp homes. Teardown's slot-ownership scan compared state-dir paths textually while fm_firstmate_root_home canonicalizes, so a lab home under a symlinked TMPDIR scanned its own record twice and self-collided; compare file identity (-ef) instead. * no-mistakes(review): Refuse unlistable lab homes and hardlinked slot records * no-mistakes(review): Mint lab markers only on verified-empty fresh dirs * no-mistakes(document): Clarify lab-home gate documentation and comment contracts * no-mistakes(document): Clarify lab-home gate documentation and remove stale claims * no-mistakes(document): Clarify gate lab-home documentation and boundary wording
…ad of refusing every re-arm (kunchenguid#5594) * fix(bin): replace a watcher whose beacon stalls past a hard bound instead of refusing every re-arm A fleet watcher that is alive but whose liveness beacon has gone stale could never be replaced: every re-arm was refused because the lock holder was a live pid, and the holder was never evicted because it was not dead. Add FM_WATCHER_STALL_BOUND (default 3x the stale grace): below it the refusal is unchanged; at or past it the arm re-verifies the holder against the lock's recorded identity, sends TERM, waits boundedly, and takes the lock the normal way, ledgering a stalled-holder-replaced row. A holder that survives TERM keeps the old refusal. Fixes kunchenguid#4400 * no-mistakes(test): poll for replacement message to fix watcher-lock test flake * no-mistakes(document): document FM_WATCHER_STALL_BOUND in config inventory
…n can keep them (kunchenguid#5563) * fix(pi): hide queued Firstmate notifications under Calm only when the session can keep them Calm now keeps authenticated Firstmate operational inputs out of Pi's queued-message listing, but only after proving the live session exposes every member needed to keep them across Escape. A session missing any of them keeps stock rows and Escape and shows one generic warning. Escape and the dequeue key return only captain-authored messages to the editor and re-queue hidden notifications in order; after an abort that kept any in Pi's agent queue, the adapter starts the delivery turn itself because Pi 0.87.1 does not continue an aborted run. Compaction-held notifications stay with Pi's compaction flush and never start or announce a turn. Fixes kunchenguid#1588 * docs(calm): record Pi 0.87.1 queued-row retention verification * no-mistakes(review): Deliver kept Calm notifications after tree-navigation aborts too * no-mistakes(review): Defer Calm notification turn until tree navigation finishes * no-mistakes(lint): Silence SC2016 for literal JavaScript in queue-retention e2e test
…uid#5548) * fix(bin): refuse teardown when a required source disappears A missing sibling was sourced after cleanup had started, so Bash 3.2 exited 0 from the EXIT trap and Bash 5 continued and reported success. * no-mistakes(review): Remove unused FM_TEST_ONLY hook from teardown tests * no-mistakes(review): Check task backend sources before any teardown cleanup * test(gotmp): give teardown fixtures every tmux adapter sibling Teardown now refuses when a sibling the recorded backend's adapter sources is missing, so the fake bin must carry fm-session-lock-lib.sh, fm-agent-process-lib.sh and fm-gemini-lib.sh.
Restructure the supervision host doc's prose into shorter sections, lists, and tables without changing documented behavior. Every original heading, anchor, identifier, number, quoted string, and link target is preserved.
* docs: make herdr-backend easier to read Restructure the Herdr backend doc's prose into shorter sections, lists, numbered procedures, and tables without changing documented behavior. Every original heading, anchor, fenced code block, link target, and documented fact is kept. * no-mistakes(document): Restore composer-proof reason and complete Herdr topic table
Restructure the prose into sections, lists, and tables without changing documented behavior. Every original heading and anchor, inline-code span, link target, number, and quoted string is kept, and each sentence sits on its own line. Adds a topic navigation table and short subsections under the existing headings.
* docs: make watcher-continuity easier to read Restructure the prose into sections, lists, and tables without changing documented behavior. Every original heading, anchor, identifier, link target, and number is kept. * no-mistakes(review): Fix actor and supervision-host scope in watcher-continuity doc * no-mistakes(review): Make readiness TERM and retry conditional on unready successor
* docs: make sessionstart-nudge easier to read Restructure the prose into sections, lists, and tables without changing documented behavior. Every original heading, inline-code span, link target, number, and fact is preserved, and a harness-to-tier table now sits near the top. * no-mistakes(review): Drop helm glossary line and dedupe exit-code lead-in
* docs: make captain-hold-lifecycle easier to read Restructure the captain-hold lifecycle prose into sections, lists, and tables without changing documented behavior. Every original heading, anchor, identifier, number, quoted string, and link target is kept. * no-mistakes(review): Fix verification record subjects and grouping headings * no-mistakes(review): Clarify task-body read-back cases belong to the suite
* docs: make remote-secondmates easier to read Restructure the remote second mates prose into sections, lists, numbered procedures, and tables without changing documented behavior. Every original heading, anchor, fenced code block, identifier, link target, and qualifier is preserved. * no-mistakes(review): Merge remote-home table cell into one sentence * no-mistakes(review): Tighten readiness lead-in, restore causal link, fix dangling reference
…nguid#5554) * fix(bin): bound the away digest and log why a delivery failed The away daemon joined every buffered escalation into one unbounded digest. A start-up catch-all span can exceed what one transport argument carries (tmux rejects the send-keys command; Linux refuses to exec any argument above 131,071 bytes, which is how herdr receives it), so the initial send failed on every housekeeping pass and was logged as an unconfirmed Enter with text possibly in the composer. escalate_flush now builds the injected digest under a fixed byte budget: each event is cut at a UTF-8 boundary with an omitted-bytes marker, the joined events stop with a "+K more event(s)" tail, and a bounded digest names a state/.subsuper-digests/ file that keeps every buffered event verbatim. The buffer itself is untouched, so the return catch-up stays complete. The tmux submit core and the herdr literal send now replay the transport's stderr on failure, and inject_msg logs the failing stage (initial send versus Enter confirmation) with the byte count and that stderr. The wedge alarm line and marker carry the last failure reason. Fixes kunchenguid#4382 * no-mistakes(review): Drop digest pruning; label send-failed as send-or-Enter stage * no-mistakes(review): Keep digest full text once submit ran; reuse on retry * no-mistakes(lint): Count digest files with find instead of ls --------- Co-authored-by: firstmate-oss <firstmate@kunchenguid.local>
…henguid#5638) * feat(tests): add FM_TEST_SEAM launch seam and gate lab-primary recipe Part 1 of the kunchenguid#5615 split: the pieces that let the no-mistakes pipeline live-validate firstmate changes, without the gate-refusal rescoping. - bin/fm-afk-launch.sh: FM_TEST_HARNESS pins the detected harness only alongside the FM_TEST_SEAM=1 marker test suites set, so a leaked variable in a real primary's environment stays inert and unknown tokens fall through to real detection. - tests/lib.sh: export FM_TEST_SEAM=1 for every suite. - .no-mistakes.yaml: per-harness recipe for running a real fixture primary from a gate run - a plain mktemp lab FM_HOME on a private tmux socket, with FM_GATE_REFUSE_BYPASS=1 scoped to it and NO_MISTAKES_GATE scrubbed. - tests/fm-wake-queue.test.sh: stop the owned watcher fixture with KILL and clear its lifecycle state so the next leg starts clean; TERM could leave bash waiting in a child on some runners. - tests/fm-remote-secondmate-lifecycle-e2e.test.sh: wait for the liveness lock holder's post-acquire marker instead of the lock dir, which is published before the claim finishes. * no-mistakes(review): Scrub lab home overrides and require FM_TEST_SEAM separately * no-mistakes(document): Clarify test seam and disposable lab bypass documentation * no-mistakes(document): Clarify lab isolation and test-seam documentation * no-mistakes(ci): Fixed the CI failure: test cleanup killed the remote worker child but left its supervisor able to restart it during fixture removal. Cleanup now stops the worker tree. The lifecycle test passed locally; ShellCheck and diff checks passed
… home is gone (kunchenguid#5552) * fix(bin): refuse watchers from disposable checkouts and exit when the home is gone Fixes kunchenguid#321 Fixes kunchenguid#4760 A watcher armed from a disposable no-mistakes validation checkout under .no-mistakes/worktrees/ outlived the validation step and kept writing the real home's state, and a running watcher never noticed when its home, state directory, or code root disappeared. The arm now refuses from such a checkout with the typed failure line, the watcher checks once per poll that its home, state directory (or its own lock holder record), and bin directory still exist and exits with a logged reason scoped to itself, and the shared test helpers reap every watcher a suite armed for a temporary home through the home-scoped stop. * no-mistakes(lint): fix SC1007 by assigning empty string in watch-arm test * no-mistakes(ci): Found and fixed a genuine, reproducible hang introduced by this branch's test-watcher reaper, which is what killed both CI checks (serial-2 cancelled at the 30-min cap; Lint 2 exit 143 = the suite's own TERM-trap code). Root cause: test_drain_asserts_watcher_liveness (tests/fm-wake-queue.test.sh) fabricates a .watch.lock whose pid is the test runner's own $$ with the runner's real identity, to make the drain believe a live watcher exists. The new make_case tracking registers that state dir for reaping, so at fm_test_cleanup the new fm_test_reap_watchers drives fm-watch-arm.sh --stop; its identity check matches (the fixture recorded the runner's identity) and it kill -TERMs the test runner. tests/lib.sh:231 is `trap 'fm_test_cleanup; exit 143' TERM`, so the TERM re-enters cleanup -> reap -> kills $$ again -> infinite loop until the runner cap. I reproduced this locally: the suite ran all tests then looped forever in cleanup spawning fm-watch-arm.sh --stop against a lock naming its own PID. Fix (tests/lib.sh, +5 lines): in fm_test_reap_watchers, skip any tracked lock whose pid equals our own $$ before driving --stop. This is the single shared reap boundary; seven $$-self-lock fixtures across four test files are all covered by the one guard, and real armed watchers (pid != $$) are still reaped. Invariant: the test reaper must only signal real armed watcher processes, never the test runner itself. Verified locally: tests/fm-wake-queue.test.sh -> EXIT 0 (63 ok, no hang); tests/fm-watch-arm.test.sh -> EXIT 0 (21 ok, including test_reaper_stops_a_tracked_watcher, confirming the guard does not over-skip). Lint 2's exit 143 was the same shard/cap signature; a fresh CI run on this new commit will re-evaluate it --------- Co-authored-by: firstmate-oss <firstmate@kunchenguid.local>
…m as silence (kunchenguid#5588) * fix(bin): surface an unrecognized status prefix instead of dropping it A parked or holding declaration, and a verb whose correlation token did not parse, never became an event, so the supervisor still saw the earlier line. * no-mistakes(review): Require verb-shaped unrecognized status prefixes, add continuation tests * no-mistakes(document): Document unrecognized status prefix escalation in afk skill * no-mistakes(ci): I reproduced the "Behavior portable serial 6" failure locally and fixed it by changing the test data in one test. No product code changed. **What failed:** `tests/fm-session-start.test.sh`, in `test_orphan_status_logs_are_printed`, with "matched status log was printed 2 times". **Why:** the test writes status lines with made-up prefixes, `matched: surfaced once` and `orphan: step N`. The test only uses them as placeholder text. It checks that the session-start digest prints each task's status tail exactly once. This PR (kunchenguid#4763) deliberately makes an unrecognized one-word lowercase prefix a status event. So those lines now surface as captain-relevant events, and the wake queue's STATUS OUTCOME BACKSTOP section prints them a second time. The code under review is behaving as the issue asks. Only the test's placeholder data had become meaningful. **Rule the test depends on:** its status lines must not be captain-relevant, so the digest is the only place they are printed. Both lines in this test broke that rule. The orphan line would have failed the same count check right after the matched line did. **Fix:** in that test only, I switched both lines to the recognized, non-captain verb `working:`: `working: surfaced once` and `working: orphan step 1..6`. I updated the matching assertions and counts to use the new text. What the test checks is unchanged: orphan logs are labelled, the tail is bounded, the log path is printed, and each tail appears once. **Verification:** before the fix, the test failed locally the same way as in CI. After it, `bash tests/fm-session-start.test.sh` reports "all assertions passed * no-mistakes(review): Detect unrecognized prefixes on unstamped lines; share verb list --------- Co-authored-by: Kun's firstmate <kunchenguid+firstmate@users.noreply.github.com>
…henguid#5658) Fixes kunchenguid#5295 Session start now reports a remote inheritance failure using the push's own error line instead of the first unchanged item that happened to print before it, and the shared captain preferences header check now names the first required phrase it did not find, on both the local and remote inheritance paths.
…yloads (kunchenguid#5657) * fix(bin): stand down the Claude Stop auto-arm on pi-code-delivered payloads pi-code loads the tracked Claude settings but has no asyncRewake, so it awaits every Stop hook; without a stand-down the auto-arm runs synchronously inside Pi's turn end and holds it open for the declared multi-hour timeout. Stand down when the payload's transcript_path contains a /.pi/ path component, the same discriminator the closed-but- unmerged fix in kunchenguid#3352 used, with an explicit string-type check on the jq filter. Fixes kunchenguid#3343 * no-mistakes(document): document pi-code stand-down in harness integrations reference
…kunchenguid#5659) * fix(bin): match whole multi-word project names in the registry lookup bin/fm-project-mode.sh matched a registered project name against only the first whitespace-delimited token of a registry row, so a name containing a space never matched, silently defaulting the project to no-mistakes off instead of its declared posture. The lookup now matches the whole registered name against the raw line text, so a name is compared literally (never as a regex) and a name that is a leading prefix of another registered name still resolves to its own row. * no-mistakes(document): docs already accurate for multiword registry name match * chore: drop accidental empty err file Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com> --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
…d#5546) * fix(bin): classify the stdin program of `bash -s` with operands in the arm policy With -s, sh/bash/zsh read the program from stdin even when operands follow; the operands are only positional parameters. The arm policy treated the first operand as a script path, so heredoc and here-string payloads were never classified and a hidden bin/fm-watch.sh execution was allowed. A protected path in the operand position still fails closed as before. Fixes kunchenguid#1489 * no-mistakes(document): Clarify stdin shell operand documentation * no-mistakes(ci): Captain, fixed `shellInvocation` so `bash -- -s` treats `-s` as a script name, and updated R21 to test the exact command. The targeted policy suite, lint, documentation check, and diff check pass. Both hosted workflows show `action_required` before any jobs ran; that external approval state remains unresolved * fix(bin): keep main's handling of words after a leading `--` Revert the pipeline CI-step change that made the first word after a leading `--` always a script. It turned forms that main denies today into allow (for example `bash -- -c 'bin/fm-watch.sh'`), which is outside kunchenguid#1489 and loosens a fail-closed policy. `--` after `-s` still ends option parsing.
…nchenguid#5695) * fix(bin): strip AI co-author trailers from fleet-launched commits Cursor and other non-Claude runtimes append the trailer after the typed message. A per-task commit-msg hook removes it and leaves human co-authors and the author identity untouched. * no-mistakes(review): Export pane hooksPath override and drop generated-with stripping * no-mistakes(ci): This PR caused all three CI failures, and the fix is test-only: 4 test files change, no product code. **Cause.** `fm-spawn.sh` now installs the AI-trailer strip hooks for every spawn, secondmates included. The installer refuses a worktree that is not a git repository, and the PR deliberately keeps that fail-closed rule because real secondmate homes are firstmate clones. Four test fixtures still gave secondmates a plain directory as their home, so each spawn failed with "not a git worktree ... could not install the AI-trailer strip hooks": - serial 5: `tests/fm-backlog-atomicity.test.sh` ("secondmate spawn failed"). - serial 8: `tests/fm-secondmate-harness.test.sh` ("split: no meta written"). - Herdr: `tests/fm-backend-herdr-launcher-workspace-e2e.test.sh` and `tests/fm-backend-herdr-workspace-per-home-e2e.test.sh`. **Rule that must hold.** Every home a test spawns as a secondmate must be a git worktree. I checked the other places in the changed area: the only secondmate spawns in these tests are the ones listed. The earlier rounds already fixed the other fixtures (`fm-secondmate-liveness`, `fm-secondmate-safety`) the same way. **Fix.** - Each of those four secondmate homes now gets the same `.gitignore` plus `git init -q -b main` that the liveness and safety tests already use. - The two Herdr tests clean up with their own plain `rm -rf "$TMP_ROOT"`, not the shared `tests/lib.sh` helper. Because the installer leaves each `state/<id>.git-hooks` directory read-only, that cleanup printed "Permission denied" and left the directories behind. Both cleanups now restore the owner's write bit on every directory before removing (`find ... -exec chmod u+rwx`), which is what `fm_test_remove_tree` in `tests/lib.sh` does. **Verification.** - `tests/fm-secondmate-harness.test.sh` passes. - `tests/fm-backlog-atomicity.test.sh` passes (99 ok, exit 0). - shellcheck is clean on all four files. - I could not run the two real-Herdr tests locally: the Herdr lab on this host refuses to start because it needs exactly one running default session, and I did not change the host's Herdr state to get around that. Instead I checked their two changed steps directly: the installer succeeds on a home set up the new way, and the new cleanup removes the read-only hooks directory completely. Those two tests will only be proven on CI
…id#5683) * fix(bin): treat Pi's dollar-first cost footer as furniture An idle Pi status row opening with $0.000 was read as a dead-shell prompt, so exit and relaunch refused on an empty composer. * test: wait for the draining holder to exec sleep before reading its identity The procevent drain fixture read fm_pid_identity immediately after backgrounding setsid sleep, racing the child's exec chain. Mid-exec the cmdline can read empty, failing the fixture on a loaded CI runner. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…5534) * fix(bin): refuse a merge when a required check never reported fm-pr-merge.sh built its GitHub refusals only from checks present in statusCheckRollup, so a required check that never ran was simply absent and the merge proceeded on the subset that reported, contradicting its own "every required check green" claim. The GitHub verify now reads the base branch's required contexts from the forge itself - the classic branch protection summary on GET repos/{o}/{r}/branches/{b} and the active ruleset rules on GET repos/{o}/{r}/rules/branches/{b} - and refuses when a required context has no entry in the same rollup, at the same head, that the merge is bound to. Absence reads as unknown, never green. The required-set read joins the existing refusal list, so a draft, a red check, and an unreported required check are all reported together. Could not read vs nothing required: both endpoints need only repository read access. The admin-only GET .../branches/{b}/protection endpoint is deliberately not used: it answers a non-admin token with the same 404 an unprotected branch gets (observed live on kunchenguid/firstmate main with this token), which would read a missing permission as "nothing required". Any failed or malformed read of either source (auth, missing fine-grained permission, rate limit, network, 404, unexpected shape) refuses the merge with a line naming the unreadable source. The one exception is GitHub's plan-gated 403 on the rules endpoint ("Upgrade to GitHub Pro or make this repository public"), which already means "this repository has no branch rules" for the merge-queue reader; that check moves into one shared helper and the classic summary still decides for such a repository. Attended waiver: --allow-missing <check-name> is the twin of --allow-red and follows the same design and recording path: once, separate name argument, waives only that exact unreported required check, still requires every other required check reported and every check green, never waives an unreadable required set, refused while the away-posture record exists, and refused on GitLab. Merge-state BLOCKED policy is unchanged. How this differs from the withdrawn kunchenguid#5353 (read from its diff): - kunchenguid#5353 read the admin-only branches/{b}/protection endpoint and treated its 404 as "no required checks", so for any non-admin token the required set silently read as empty; this change reads the read-access branch summary and treats every failure as unreadable. - kunchenguid#5353 ignored rulesets; this change also reads required_status_checks rules from the effective branch rules. - kunchenguid#5353 made separate per-head REST reads of statuses and check-runs capped at per_page=100 with no pagination; this change checks presence in the same statusCheckRollup view the red-check gate already reads at the verified head. - kunchenguid#5353 stopped at the first unreadable read; this change reports it as one refusal among all the others. - kunchenguid#5353 also claimed kunchenguid#5345 (lock stealing) and changed 39 files, most unrelated; this change is kunchenguid#5344 only. Live proof, read-only (a gh wrapper refused every merge and mutating call): - cli/cli#14474 (trunk requires 3 classic build contexts, none ran): refused, naming build (macos-latest), build (ubuntu-latest), build (windows-latest); with --allow-missing "build (macos-latest)" it still refused, naming the other two. - cli/cli#13665 (ran build (ubuntu-24.04-firewall) instead): refused, naming build (ubuntu-latest). - hashicorp/terraform#39262 (ruleset-required checks absent): refused, naming Code Consistency Checks, End-to-end Tests, Race Tests, Unit Tests. - cli/cli#14485 (all required reported and green): verified; the wrapper blocked the merge call and the pull request read back open. Fixes kunchenguid#5344 * fix(review): Preserve required-check producers and aggregate independent read failures * fix(document): Clarify required-check verification and waiver documentation * fix(bin): match an app-bound required commit status by name The producer-identity check resolved an app-bound required context only against check runs, so a required context that the required app reports as a commit status could never match and always read as "has not reported". A commit status carries no app id to compare, so an app-bound requirement that arrives as a status now matches by name, as before producer binding; check runs keep requiring the configured producer app. Live, read-only: hashicorp/terraform#39262 requires license/cla from integration 865473, reported green as a commit status by the CLA app. The previous head refused it as unreported; this head no longer does, while still naming the four required check runs that never ran there. Refs kunchenguid#5344 * fix(document): Clarify accepted commit-status producer verification limitation --------- Co-authored-by: firstmate-oss <firstmate-oss@kunchenguid.local>
…uid#5696) * fix(bin): never offer a persistent secondmate for teardown The return brief's "Landed, cleanup due" scan listed every state/*.meta record carrying a pr= and a merge-notified marker without regard to kind, so a secondmate record holding a relayed child's merged PR put the mate itself up for "bin/fm-teardown.sh <mate>" cleanup. A secondmate is a persistent worker, never landed work. - bin/fm-afk-return.sh: skip kind=secondmate in the landed-cleanup scan. - bin/fm-pr-check.sh: refuse to record pr= or arm a merge watch on a kind=secondmate record before any side effect; a PR reported on its routed status channel belongs to a task in the mate's own home, which arms its own watch. - bin/fm-watch.sh: a merged result from a poll already armed on a secondmate retires the poll silently - no merge outcome, marker, or wake. * no-mistakes(document): Document secondmate merge-watch and return-brief exclusions * no-mistakes(ci): CI failed in an unchanged watcher-shutdown test whose three-second wait was sensitive to runner load. Increased the wait for both state- and home-deletion cases without changing watcher behavior. The full fm-watch-arm suite passed locally; syntax and diff checks passed
…unchenguid#5702) * fix(bin): refuse unknown dash-leading args in public-posting fm-x scripts fm-x-reply.sh collected any unrecognized argument into the positional pool and took the first one as the reply text, so an invocation like "fm-x-reply.sh <id> --followup --final <text>" posted the literal string "--final" to X and silently dropped the real text. Make argument parsing strict in every script that can post publicly: an unknown dash-leading argument, a dash-leading request_id/task id, a dash-leading option value, or a surplus positional now exits 2 with a usage error before any config load, outbox write, or network call. Reply text starting with '-' is still accepted via --text-file or stdin, and --help is honored wherever it appears instead of becoming text (a --help forwarded through fm-x-followup.sh would have counted as a posted follow-up and mutated the link). fm-x-link.sh and the fm-public-followup scripts already refuse unknown arguments; fm-x-poll.sh takes none. * no-mistakes(review): Refuse surplus follow-up text sources; drop post-ID help branches * no-mistakes(document): Clarify reply and follow-up argument usage * no-mistakes(review): Refuse dash-leading --text-file operands in fm-x-reply * no-mistakes(document): Correct follow-up argument parsing comment * no-mistakes(document): Document dismiss argument rejection in script header
…nchenguid#5701) * feat(bin): latch the supervision host after repeated engine errors Rung 3c-1 of the PR 5631 re-cut: the host copies the Pi branch's broken-session policy. Two consecutive engine errors latch the session; every away wake then reaches main with one supervision-host line for a five-minute cooldown, after which one wake probes the engine, and each failed probe doubles the cooldown up to one hour. A reported turn without an engine error clears it. The latch is kept per main session, engine, and model in state/.supervision-host-health, and the engine conversation now uses the same main-session key, which includes the lock holder's process identity so a recycled pid never shares either. Lifted from the validated 5631 tree and adapted to main's away-only host: the attended recovery line and attended cooldown pass-through are left for the attended core, so a recovery is only logged. * no-mistakes(document): Consolidate supervision-host latch documentation
…kunchenguid#5535) * fix(bin): absorb routine second-mate progress while surfacing routed replies Fixes kunchenguid#2959 A kind=secondmate task's status signal was never absorbable, so a healthy mate's routine working: and paused: appends woke the primary every time. signal_crew_provably_working now reads the mate's lines new since the watcher's classified position: a decision, blocker, terminal outcome, note:, correlation-marked line, or unknown verb still surfaces regardless of busy evidence, while unmarked working:, paused:, and resolved: fall through to the same provably-working absorb an ordinary crewmate gets. * no-mistakes(review): narrow secondmate routine absorb to working and paused
* fix(bin): use gh-axi for the ship DoD draft check * no-mistakes(review): use PR number not URL in gh-axi draft check
… only (kunchenguid#5520) * fix(bin): drop status prose from the inactive-outcome dedupe identity The inactive-outcome receipt fingerprint included the child's sanitized last status line, so a persistent child appending routine prose after one terminal outcome minted a fresh parent event per sentence. Bind the identity to incarnation, task id, terminal state, and PR only, keeping the last line in the record as status_head evidence. Fixes kunchenguid#2960 * no-mistakes(document): note structured-only inactive receipt identity in regression coverage
…unchenguid#5707) * feat(bin): record the supervision host's dialog mirror on Claude and Cursor Add bin/fm-host-mirror.sh, the one owner of the supervision host's dialog mirror file, cursor, lock, and feed, plus the main-session key it keys entries to. The tracked Claude UserPromptSubmit and Stop hooks and the Cursor beforeSubmitPrompt and afterAgentResponse hooks record the captain's prompt and main's reply, only on a home with config/supervision-host, from a genuine primary checkout, for the lock-owning session. The mirror lands inert: writers record and nothing reads it yet; attended supervision on the host is the later step that consumes the feed. Codex, Grok, OpenCode, and omp have no writer here. * no-mistakes(review): Scope mirror dedup to session, atomic appends, marker-inclusive caps * no-mistakes(document): Clarify dialog mirror scope and remove duplicate contract details * no-mistakes(document): Correct Cursor hook documentation for dialog mirror registration * no-mistakes(review): Pass mirrored dialog text to jq via stdin * no-mistakes(document): Clarify dialog mirror documentation and remove duplicate claims * no-mistakes(review): Preserve internal dialog whitespace; drop mirror check and verified modes * no-mistakes(review): Drop only identical mirror repeats; remove redundant chmod guard
… escalations are not repeated (kunchenguid#5731) * fix(bin): retire check-row receipts on branch acks and report an unchanged situation once * fix(bin): scope a branch acknowledgement's check-row receipt retirement to its granted sequences The away posture lifts the attended partition's check/decision exclusions, so a branch grant can name check-kind rows - but the branch-actor ack still assumed check rows were main-only and skipped every receipt scan. The queue row was consumed while its terminal-outcome .pending receipt stayed behind, and each inactive-reconcile cadence scan re-queued the same fingerprint. In the first real away window on the supervision host that re-escalated one unchanged held-PR situation on every cycle (~1,734 of 4,149 outcomes). A branch ack now scans inactive-outcome and inactive-reconcile receipts and commits secondmate stall receipts against exactly the sequences in its eligible-row snapshot - the same rows it consumes - instead of none. Attended grants still name no check row, so the scans find nothing. * fix(bin): store a repeated captain verdict as routine while the task's durable situation is provably unchanged fm-branch-outcome.sh append computes a mechanical situation key per captain row - metadata bytes, captured status-log endpoint and identity, live crew-state verb, worktree head - and anchors it in state/.<task>.branch-captain-key. A later captain verdict whose recomputed key matches is stored as routine with "unchanged since seq <N>:" prefixed to its summary, so one situation escalates once until something provably changes. A task with no readable status ledger is never demoted, an unreadable record fails toward reporting, and teardown removes the sidecar with the task's other branch records. The append-only store schema is unchanged. This covers both hosts: the Pi supervision branch and the supervision host both funnel reports through append. * docs: check rows are main-owned only while attended; the away posture grants them to the branch, whose ack retires their receipts exactly * test: the away-flood reproduction as a regression test (branch ack retires the receipt and later scans stay quiet), store-level dedupe coverage, and a branch-ack secondmate stall receipt case * fix(bin): restore the secondmate child devin-config cleanup path The branch-captain-key sidecar addition mistyped the sibling entry as .$child_id.devin-config.json, so a forced secondmate teardown would have stopped removing each child's real <id>.devin-config.json. Restore the original path and add a behavioral test that stops the child sweep mid-loop on a refused close, proving the cleaned child's devin config and captain anchor are both removed while the unconsumed child's records are retained. * no-mistakes(review): Key captain dedupe on the covered wake rows' fingerprint * no-mistakes(review): Drop captain-key demotion; prove one escalation on both surfaces * no-mistakes(review): Drop unrelated teardown test; cite both receipt test files * no-mistakes(document): Docs already match branch-ack check-receipt retirement
…oorbell (kunchenguid#5664) * fix(calm): deliver Claude-bound operational input as a record-backed doorbell Claude Code 2.1.280 removes U+2063 from every submitted prompt, so a typed operational envelope reaches a Claude Code primary as plain text. The away daemon now writes the envelope to a record under state/operational-inbox and types only a plain doorbell naming it; the /afk return check and the Calm mod recognize the doorbell only when that record holds a current envelope. Marker- preserving harnesses keep the typed envelope. The live Calm guard accepts the 2.1.280 module-load log line, drives the doorbell, and asserts thinking stays hidden. * no-mistakes(review): Fix operational record retention at 7 days and document prune limit * no-mistakes(document): Point Calm bounds at 2.1.280 evidence; fix afk-exit comment * no-mistakes(lint): Pick newest Calm e2e transcript without parsing ls * docs(calm): add a minimal turning-Calm-on step for Claude Code * fix(spawn): deliver the Claude launch brief as a record-backed doorbell Claude Code strips U+2063 from the launch-prompt argument too, so a worker's launch brief arrived with its operational marker removed. Publish the brief as a record in the receiving home's operational inbox - a secondmate's own state, not the primary's - and pass only the printable doorbell naming it, falling back to the typed envelope when the record cannot be published so the brief body still delivers. Unwrap doorbell-carried digests in the daemon digest tests that still read the raw send log under the claude pin, and update the documented bounds now that launch briefs hide like the other operational rows. * test(spawn): cover a secondmate's launch-brief record landing in its own home The record-backed doorbell resolves its state through the receiving pane's home, so prove a claude secondmate launch publishes into the seeded secondmate's operational inbox and never leaks a record into the primary's. * no-mistakes(review): Pass primary harness to daemon, tighten retention, refresh verdicts * no-mistakes(review): Prune operational records by exact seven-day elapsed age * no-mistakes(review): Batch record pruning so large inboxes still expire * no-mistakes(review): Refuse Claude spawn when brief record cannot publish * no-mistakes(review): Drop thinking probe from Claude Calm live test and docs * no-mistakes(review): Record dated Claude Code 2.1.282 reproduction evidence * no-mistakes(document): Clarify operational doorbell documentation and record expiry * no-mistakes(document): Correct AFK escalation carrier guidance * no-mistakes(review): Describe operational record retention as about seven days * no-mistakes(document): Clarify Calm delivery and operational record retention * no-mistakes(review): Remove out-of-scope Calm launch guide from Claude docs * no-mistakes(document): Document Claude launch-brief delivery and refusal * no-mistakes(document): Correct stale operational-input documentation * no-mistakes(ci): Fixed the stale Claude trust test to verify that worker and secondmate launches deliver readable, record-backed briefs instead of expecting brief paths in their commands. Annotated the daemon’s output variable for ShellCheck without changing behavior. The affected tests, daemon tests, ShellCheck, and diff check pass locally * no-mistakes(ci): parse rebased Claude launch after trailer hook prefix * no-mistakes(review): Trust launch-brief record and restore thinking bound doc * no-mistakes(review): Parse final Claude launch statement; drop Stop-hook docs --------- Co-authored-by: Mike Sewell <maikunari@protonmail.com> Co-authored-by: no-mistakes <no-mistakes@localhost>
kunchenguid#5583) A host-local relaunch rewrote only the far endpoint, so this home kept the old harness, model, and effort, and appending those keys after pr= broke pull-request poll authentication.
…ound (kunchenguid#5516) tests/fm-watch-triage.test.sh finishes in about 434s alone and about 698s under CI load, so the 900s bound the changed-suite runner applies produced a false timeout under ordinary concurrent validation. Raise the automatic bound to 1500s, which keeps every measured script under it while staying below the 30-minute normal CI tier so a genuinely hung script still fails here with its output before the job cap cancels the lane. Fixes kunchenguid#3869 Refs kunchenguid#3565
…kunchenguid#5728) * Fix nested watcher lock reclaim * no-mistakes(review): Elect a single steal-mutex reaper and bound arm TERM wait * no-mistakes(review): Reclaim self-held steal mutex and unify autoarm steal reaping * no-mistakes(review): Resume own interrupted steal reap from its tombstone
…nguid#5710) * test: hold the back-to-back boundary close on the host's own clock test_park_boundary_holds_under_back_to_back_closes assumed two engine turns fit in the ~16s pre-refusal window and that the stub finished a turn in 3s. Under load the stub's real drain, report, and acknowledgement take ~13s, so the turn either died at its bound (which hands the wake to main, no boundary line) or the second close landed past the window and the fixture failed while the boundary held. 3 failures in 5 runs at a load average near 11. Hold the first turn on a release file instead: once the engine is in flight, a second close is appended mid-turn and the turn is released as the refusal window opens (park bound minus turn bound and grace, read off the host's own start record). The queued close can then only wait for the boundary on any machine speed, which is what the test asserts: the boundary line ends the output, the demo.status row stays queued for main, and no second engine turn ever starts. A host too loaded to start the turn at all hands the first close to the same boundary exit. After: 12/12 at load ~15-42. * no-mistakes(review): Print boundary test deadline as a decimal integer * no-mistakes(review): Hold boundary test turn on a FIFO, require full sequence * no-mistakes(review): Remove stray before/after supervision-host test copies * test: hold the late close's render until the refusal window opens The boundary recheck test's node shim slept a fixed 10s, which assumed the first close was read before the host's refusal window opened. Under load the close arrived after the refusal check, so the host correctly refused it before the successor started and the render snapshot never appeared. Block the wake-prompt render on a FIFO released at the refusal-open instant read from the host's own start record, so the pre-turn recheck must refuse on any machine speed. * no-mistakes(review): Derive minimal park bounds and refresh supervision-host shard hint * no-mistakes(review): Drive park-boundary tests from a seam-gated host test clock
* fix(bin): stage remote home clones before publishing them A remote home provision cloned the code root directly into the public FM_HOME path while rollback() claimed rm -rf of that same path on any failure. Bash defers trapped signals past a foreground child, but any other cleanup or lifecycle path that removes the home directory races the live clone's object copy, producing the CI flake "fatal: failed to copy file to .../.git/objects/...: No such file or directory". Clone into a private staging directory beside the home and publish with an atomic rename once complete, so no cleanup can remove a directory a live clone is still writing; a home that appears mid-provision now dies cleanly instead of inheriting torn state. The regression coverage holds a real clone mid-copy, removes the public path, and requires the provision to finish and publish intact. * no-mistakes(review): Prove home ownership by sentinel and hold only a live clone * no-mistakes(review): Assert raced provision publishes a complete, intact clone * no-mistakes(document): Document remote home staging and publication safety * no-mistakes(lint): Fix ShellCheck warning in clone integrity assertion * no-mistakes(document): Clarify remote home publication and rollback guarantees
* fix(control): keep a relaunched Pi worker's herdr pane status authority alive Defect: after `bin/fm-control.sh <id> relaunch` (observed live on a herdr Pi crewmate whose pane read idle while it ran its validation pipeline), the pane froze at whatever its previous agent had last reported. Cause, measured on herdr 0.9.1 against a real Pi: a pane has one status authority, and for Pi with its integration installed that authority is the lifecycle hooks, so herdr also skips screen detection for the pane. In the crew shape the registration outlives its agent process (upstream issue kunchenguid#4115; docs/herdr-backend.md "Restart and liveness behavior"), and herdr applies only reports carrying the session identity it bound. A replacement started fresh in that pane reports a NEW session, so its state reports are ignored and the pane stays frozen. Nothing from outside repairs it: `pane report-agent-session` and `pane report-agent` for `herdr:pi` are accepted (rc=0) without being applied unless the reporter is the registered pane agent, and `pane release-agent` on the stale record changes nothing. Fix: a relaunch preserves the binding instead of fighting it. The launch owner reads the session reference the endpoint's own runtime recorded (`fm_backend_herdr_pane_agent_session_ref`) and passes it back as Pi's own `--session <path-or-id>` (`relaunch_resume_args`; `fm_control_relaunch_resume_flag` owns which adapters and which registered-agent labels qualify). That is the same reference herdr itself resumes Pi panes with after a server restart, and the resumed session's reports land again, which the live check confirmed: the pane returned to working while the replacement worked and idle when it settled, on the same session identity. Safety: relaunch-only (a fresh spawn binds nothing), herdr-only (the one adapter that records a per-pane session), Pi-family only, and only when the registration's own agent label matches - so no other adapter's conversation can be handed to a Pi launch. An unreadable, missing, or malformed reference degrades to exactly the fresh-session launch that existed before. No lifecycle, liveness, isolation, or merge guard is touched, and an empty result leaves every non-Pi launch byte-identical. `resume` remains a refused verb; docs/agent-control.md and the harness-adapters references are corrected where they claimed Pi had no verified resume form at all. * no-mistakes(document): docs: correct relaunch session-authority ownership and skill paths * no-mistakes(document): docs: correct stale control-plane ownership claim * no-mistakes(document): docs: drop unverified Herdr restart resume claim * no-mistakes(test): Added offline Herdr Pi session-authority relaunch coverage * no-mistakes(document): Document Herdr Pi relaunch session continuity * no-mistakes(ci): The failing remote relaunch test tried to arm a PR poll for a secondmate, which `fm-pr-check.sh` correctly refuses. Removed that invalid test scenario; the remaining remote relaunch tests pass, and `git diff --check` is clean
…nchenguid#5758) Main has been red since fm-pr-check.sh began refusing to arm a merge poll on a kind=secondmate record (kunchenguid#5696): the relaunch-ordering case in tests/fm-remote-secondmate-relaunch.test.sh armed its fixture through that entry point and could no longer be set up. The ordering guarantee still matters: a secondmate record armed before the refusal can legitimately carry a trailing pr=/pr_head= identity block until the watcher retires it, and fm-remote-secondmate-relaunch.sh must still keep that block last when republishing harness/model/effort. Seed the fixture the way such a record was really written - pr= appended last to the meta, then the poll artifacts published through the same fm_pr_poll_prepare/fm_pr_poll_publish_prepared pair fm-pr-check.sh uses, a pattern tests/fm-pr-check-security.test.sh already follows - and drop the now-unused fake gh fixture. The kunchenguid#5696 refusal itself stays pinned by the security suite's secondmate-record case.
…id#5748) * feat: run attended supervision on the host for Claude and Cursor On a home opted into config/supervision-host with a Claude or Cursor primary, the supervision host now takes the attended wakes the Pi branch would take: routine outcomes stay off main, and a captain outcome wakes main once with a branch-outcome line and waits in the drain's new BRANCH OUTCOMES section until main acknowledges it with mark-processed. - The offer rule moves into branchOfferForWake, shared by the Pi watcher and the host through bin/fm-branch-dispatch.mjs offer. - The host feeds the dialog mirror at the head of each attended wake and passes a close through unchanged when it is main-only, the engine or a tool is missing, the primary has no verified mirror, the main session cannot be identified, or the session is cooling down. - The drain presents captain outcomes first, one line per task, never behind older routine outcomes, and collapses routine overflow into a count that is marked read. - The return advances the store's read cursor through the away window once the brief has rendered, so the first drain does not replay it. - The branch prompt's mirror wording is host-neutral, and the rule to report what main must act on as captain, once per unchanged situation, applies only to the attended posture on the host. * docs: record the attended supervision host live check * no-mistakes(review): Present pre-window unread outcomes and contiguous captain prefix * no-mistakes(review): Return brief presents every row it marks read * no-mistakes(review): Return brief lists every unread outcome in one list * no-mistakes(review): Keep return list in store order and gate cursor failures * no-mistakes(review): Make the drain the only branch-outcome presenter after return * no-mistakes(review): Gate return on drain outcome failures; byte-count outcome budgets * no-mistakes(review): Gate drain on projection failures; UTF-8-safe byte cuts * no-mistakes(review): Fail drain without jq; hand unreadable prompt mirror to main * no-mistakes(document): Correct supervision-host return and drain documentation * no-mistakes(review): Recheck attended offer at turn start; honest failed-drain brief * no-mistakes(document): Correct supervision-host posture and drain documentation * no-mistakes(document): Documentation remains accurate for attended supervision
…unchenguid#5753) Each '# shellcheck source=' directive makes ShellCheck's external-source traversal expand that library's whole transitive graph again at the site. fm-pending-reply-lib carried three directed lazy sources of fm-wake-lib and two of fm-parent-channel-lib on identical per-call re-source sites, so one file analysis peaked above 4 GiB and every caller (fm-watch, fm-teardown) inherited the multiplier - the root cause of the PR kunchenguid#5732 Lint 1 OOM kill. Keep the runtime '.' commands byte-identical: the lazy re-source under 'local STATE FM_WAKE_QUEUE FM_WAKE_QUEUE_LOCK' is real behavior. Drop the duplicate directives so each library expands once per unit, and drop the tmux/classify directives since classify already arrives through the kept fm-wake-lib expansion and no tmux symbol is referenced here. The directive above the lib-dir assignment is kept - it binds the bin/ prefix so the undirected sites still resolve without SC1091. Measured peak RSS, ShellCheck 0.11.0 -x on Linux arm64: bin/fm-pending-reply-lib.sh 4.06 GiB -> 1.96 GiB, zero findings
Keep fork-local behavior while adopting upstream pending-reply deduplication, Claude worker doorbells, and the other 47 upstream commits since the last catch-up. Local validation: focused task delivery, pending-reply, daemon, operational-input, AFK, and Pi Calm tests passed; lint and documentation checks passed. The session-lock ancestry E2E pty-host reparent assertion fails identically on clean origin/main and upstream/main source snapshots in this WSL host, so it is not caused by this merge. CI on Linux will decide the full-suite result.
…ne test, tests/fm-spawn-dispatch-profile.test.sh, case "an absent worker settings file changed the launch". The merge brought in upstream's new signature for the test helper `claude_expected_launch`. It changed from `<home> <id> <flag>` to `<launch> <home> <id> <flag>`, because the expected launch now rebuilds the launch-brief doorbell from the real launch. Our fork-only test `test_claude_worker_settings_absent_keeps_launch` (line 1705) still called it the old 3-argument way. Its arguments shifted by one: `cd workersettings-absent-z24/state` failed and `$4` was unbound, so the expected string came out empty. Invariant: every caller of `claude_expected_launch` must pass the 4-argument `<launch> <home> <id> <flag>` form. I checked every caller. The five upstream call sites (lines 144, 1144, 1582, 1600) already do, no other test file uses the helper, and line 1705 was the only stale one. Fix: one line in tests/fm-spawn-dispatch-profile.test.sh. The call now passes "$launch" first, like its sibling calls. The assertion is unchanged, and no production code changed. Verification: I ran `bash tests/fm-spawn-dispatch-profile.test.sh` locally and it exited 0. All 72 cases print `ok` and none fail, including all five claude-worker-settings cases (absent, ship merge, scout/secondmate, invalid refusal, codex unaffected). `bash -n` also passes. I left the change uncommitted for the outer executor
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
The captain asked "whats the update with firstmate?" (2026-09-26 01:49). Our fork is current with its own changes, but the original project (remote upstream, kunchenguid/firstmate) has 47 new commits since our last catch-up merge (PR #25), up to 72b63ee, including fixes for problems seen tonight (repeated pending-reply pings to second mates; messages to Claude workers stuck behind text already in their box). Bring them in while keeping all of our own changes.
What Changed
kunchenguid/firstmatecommits (through72b63eee) into the fork. Among them are the fixes for repeated pending-reply pings to second mates (dedup of directed source expansions infm-pending-reply-lib, retired check-row receipts, no repeat of unknown-wake escalations that were already delivered) and for Claude-bound operational input getting stuck behind text already in the box, which is now delivered as a record-backed doorbell.bin/, Pi extension, and Claude Calm mod changes. New scripts includefm-host-mirror.sh(attended supervision for Claude/Cursor hosts),fm-git-strip-ai-trailers.sh,fm-remote-secondmate-relaunch.sh, andfm-lab-home.sh. Upstream also hardened watcher locks, merge and teardown refusals, remote job workers, and X reply argument checks, and added the matching new and updated tests.herdr-backend.md,sessionstart-nudge.md, a duplicated paragraph incaptain-hold-lifecycle.md), restore fork-specific facts inconfiguration.mdandcaptain-hold-lifecycle.md, and restoretests/fm-calm-pi-extension.test.shbyte-for-byte to upstream.Conflict resolutions
The merge had 12 conflicted files. Each resolution keeps the upstream change and the fork behavior:
.agents/skills/afk/SKILL.md: kept upstream record-backed carrier and alarm wording alongside the fork's keyed wake replacement rule.bin/fm-spawn.sh: kept the fork's Claude worker settings placeholder and upstream's launch-brief doorbell; settings substitution remains last.bin/fm-supervise-daemon.sh: combined the fork's keyed escalation refresh with upstream's acknowledgement of unknown wakes.docs/captain-hold-lifecycle.md: kept upstream's document structure and the fork's datedlaterdeferral and legacy semantics; a duplicate paragraph was removed and omitted fork facts restored in follow-up commits.docs/configuration.md: kept upstream's attended host and reorganized sections with the fork's Codex quiet, Herdr notification, ordered dispatch, Grok Bot, Vercel dispatch, and Relay cadence facts; omitted fork facts were restored in a follow-up commit.docs/herdr-backend.md: kept upstream formatting and the fork's sidebar pane, Claude composer, and Codex daemon exemption facts; wording damaged in the merge was repaired in a follow-up commit.docs/pi-supervision-branch.md: kept upstream structure and the fork's post-merge verification rule.docs/sessionstart-nudge.md: kept upstream's structure and lock behavior with the fork's Codex 0.154.0 interactive startup and--codex-hookfacts; wording damaged in the merge was repaired in a follow-up commit.docs/supervision-host.md: kept upstream's attended host behavior with the fork's Codex quiet refusal.tests/fm-afk-launch.test.sh: kept the fork's Codex cases and upstream's test seam.tests/fm-backend-herdr-workspace-per-home-e2e.test.sh: kept the fork's prep scaffold and upstream's secondmate Git initialization.tests/fm-daemon.test.sh: kept the fork's keyed escalation cases and upstream's doorbell cases.The merge also changed
tests/fm-calm-pi-extension.test.shoutside a conflict. Review restored that file byte-for-byte to upstream72b63eee. Its full suite passes on Pi 0.87.1 (15/15, no skips).Risk Assessment
✅ Low: Every file changed by only one side of the merge is byte-identical to that parent. A line-level audit found only three parent-added lines missing across all non-doc files that both sides changed: all three were in the afk skill doc and fm-supervise-daemon.sh's
escalate_add, and each was an intentional conflict resolution that keeps both behaviors. In escalate_add, keyed check replacement sits beside upstream's unknown-wake acknowledgement, and every check-wake producer uses acheck:payload. The fork facts in the docs were restored and checked against the code, and the Pi Calm test is byte-identical to upstream.Testing
This round was driven at HEAD fec95ee. The unchanged upstream Pi Calm suite was run against the installed Pi 0.87.1: 15/15 passed with no gate skips, including the "delivers it once in a new announced turn" Escape assertion, and the live Pi queue-retention guard also passed. With the default npm-root lookup, the package-dependent Pi Calm cases gate-skip on this host because Pi is installed under linuxbrew, not under nvm's npm root. Pointing FM_PI_PACKAGE_DIR at the installed package fixed that setup issue without touching the test. A real Claude Code worker was driven live on a private tmux socket with a disposable lab home. fm-send left the captain's draft untouched, recorded the steer durably, and did not type the doorbell over the draft. After the draft was cleared, the re-ring got the message acted on and acknowledged in 6s. The documented Grok Bot facts were driven through the real CLIs in a throwaway lab home: the quota-balanced refusal (exit 2), the bridge-absent refusal naming the primary (exit 2), and launch_failed and blocked facts moving ordered selection to the next candidate. Seven targeted suites passed, covering the pending-reply, send-inbox, operational-input, captain-hold (including later deferrals), and dispatch/Grok Bot behavior. The repeated-ping fix and captain-hold behavior are covered only by these suites, not live, so those two scenarios are recorded as untested. No Herdr lab was needed because no Herdr lifecycle is touched. The worktree is clean and all temporary lab dirs and tmux servers were removed.
Evidence: Upstream Pi Calm suite on Pi 0.87.1 (15/15, no skips)
Source: Upstream Pi Calm suite on Pi 0.87.1 (15/15, no skips)
Evidence: Pi Calm default-lookup run plus live queue-retention guard on Pi 0.87.1 (the header has a harmless package-path probe error from the evidence script)
Source: Pi Calm default-lookup run plus live queue-retention guard on Pi 0.87.1 (the header has a harmless package-path probe error from the evidence script)
Evidence: Live Claude Code worker: draft not overwritten, re-ring acted+acked
Source: Live Claude Code worker: draft not overwritten, re-ring acted+acked
Evidence: Live draft-guard driver script
Source: Live draft-guard driver script
Evidence: Grok Bot documented behaviors via real CLIs
Source: Grok Bot documented behaviors via real CLIs
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
🔧 **Review** - 1 issue found → auto-fixed (3) ✅
docs/captain-hold-lifecycle.md:237- The conflict resolution kept two versions of the same paragraph. Upstream's restructured three-line paragraph is at lines 232-234 ("bin/fm-procevent-lavish.sh answersis one such built-in adapter command. ... relays a card's declared close mode ..."). The fork's older one-line version is at line 237, after the trust-boundary sentence ("... one such adapter command; it reads only rows taggedchoice, relays the selected option's close mode ..."). So the doc describes the same adapter twice, with slightly different wording on where the close mode comes from. Fix: delete line 237 and change line 233 to say "relays the selected option's close mode". That keeps the fork's per-option close-mode meaning, which thedefer:YYYY-MM-DDrow in the mode table above relies on. No other doc conflict resolution has this problem; I checked architecture.md, configuration.md, herdr-backend.md, pi-supervision-branch.md, scripts.md, sessionstart-nudge.md, supervision-host.md and supervision-protocols/supervision-host.md.🔧 Fix applied.
✅ Re-checked - no issues remain.
tests/fm-calm-pi-extension.test.sh:2643- The merge commit rewrote an upstream test that neither parent changed. The intent is to bring upstream in and keep the fork's own changes, and it does not ask for upstream tests to be rewritten.git show --remerge-diff 8430bdbdlists this file with no CONFLICT header, andgit diff 72b63eee HEADshows edits that exist only in the merge commit. There are three edits. (1) The upstream assertion that Pi Calm announces "Firstmate supervision continues in a new turn." after Escape now runs only when the installed Pi is 0.87.1 or newer, which weakens the test on older Pi. This host has Pi 0.85.1, so the assertion is skipped here. The version check also fails open: it splitspi --versionoutput on dots and passes it throughNumber(). Any output that is not a barex.y.z, such as apiprefix or a pre-release suffix, gives NaN, and the assertion is skipped without a message. CI installs an unpinned@earendil-works/pi-coding-agent(.github/workflows/ci.yml:111,215), so nothing guarantees the assertion runs anywhere. (2) At line 567,makeSessiongainssession.agent = { abort: () => session.abort() }. No code under .pi/ readssession.agent. (3) At line 2609, follow-up submission changes fromM-EntertoC-q, backed by a new keybindings.json at line 2508. The notice comes from .pi/extensions/lib/fm-calm-pending-operational-layout.ts:292 after the adapter drains a retained queue. The 0.85.1 carve-out is plausible, but it is a host-driven change to an upstream suite rather than part of the catch-up. The smallest remedy is to remove these merge-only edits and restore the upstream test. If the host adaptation is wanted, keep it as a separate, reviewed change with a version check that fails closed.docs/configuration.md:1064- The docs conflict resolution condensed the fork's text, and some fork facts were lost or changed. That goes against 'keeping all of our own changes'. (a) The fork documented that typed resolution reports a Grok Bot target as eligible but unranked and never emits it as aprofile:line. The merged Grok Bot paragraph (docs/configuration.md:1059-1065) no longer says this, but bin/fm-dispatch-resolve.sh:372 still does it. (b) docs/configuration.md:1064 now reads 'a Bot dispatch there refuses and supplies alaunch_failedfact'. bin/fm-grok-bot-dispatch.sh only exits 2 (die, line 41). The fork text said Firstmate supplies thelaunch_failedfact, so the new sentence gives that job to the script. (c) The fork's rule that bootstrap,fm-dispatch-select.sh, and typed resolution all refuse a quota-balanced Grok Bot rule is narrowed to 'Quota-balanced selection refuses' at line 1065. (d) Same class in docs/captain-hold-lifecycle.md. At lines 526-527, the verification record lost the fork's coverage of 'a boardlaterchoice dating the open hold under its unchanged reason whiledoneandreleasekeep their modes' and 'the refusal of undated, invalid, past, or non-laterdeferrals'. That behavior is still pinned in tests/fm-captain-hold-lifecycle.test.sh. At line 186 the doc still says 'a card-declared close', while the fork, and line 234 after the round-1 fix, say the close mode comes from the selected option. Fix: restore (a) through (c) in the Grok Bot paragraph with Firstmate as the subject of (b), add thelaterdeferral coverage to the Answers bullet at captain-hold-lifecycle.md:527, and change line 186 to 'the selected option's close'.🔧 Fix applied.
1 warning still open:
docs/configuration.md:1064- The docs conflict resolution condensed the fork's text, and some fork facts were lost or changed. That goes against 'keeping all of our own changes'. (a) The fork documented that typed resolution reports a Grok Bot target as eligible but unranked and never emits it as aprofile:line. The merged Grok Bot paragraph (docs/configuration.md:1059-1065) no longer says this, but bin/fm-dispatch-resolve.sh:372 still does it. (b) docs/configuration.md:1064 now reads 'a Bot dispatch there refuses and supplies alaunch_failedfact'. bin/fm-grok-bot-dispatch.sh only exits 2 (die, line 41). The fork text said Firstmate supplies thelaunch_failedfact, so the new sentence gives that job to the script. (c) The fork's rule that bootstrap,fm-dispatch-select.sh, and typed resolution all refuse a quota-balanced Grok Bot rule is narrowed to 'Quota-balanced selection refuses' at line 1065. (d) Same class in docs/captain-hold-lifecycle.md. At lines 526-527, the verification record lost the fork's coverage of 'a boardlaterchoice dating the open hold under its unchanged reason whiledoneandreleasekeep their modes' and 'the refusal of undated, invalid, past, or non-laterdeferrals'. That behavior is still pinned in tests/fm-captain-hold-lifecycle.test.sh. At line 186 the doc still says 'a card-declared close', while the fork, and line 234 after the round-1 fix, say the close mode comes from the selected option. Fix: restore (a) through (c) in the Grok Bot paragraph with Firstmate as the subject of (b), add thelaterdeferral coverage to the Answers bullet at captain-hold-lifecycle.md:527, and change line 186 to 'the selected option's close'.🔧 Fix applied.
✅ Re-checked - no issues remain.
✅ No issues found.
🔧 **Test** - 1 issue found → no changes applied (2) ✅
🔧 No changes applied.
1 warning still open:
tests/fm-watch-triage.test.sh- tests/fm-watch-triage.test.sh is timing-sensitive on this host. Run alone in round 2, it passed all 93 cases but did not finish within 580s. In round 1 it failed three different ways on HEAD (a heartbeat re-surface, TERM exit timing, and immediate stale surfacing). It also failed once on pure upstream 72b63ee and passed on the pre-merge fork 09c761f. The failures change from run to run and also show up on upstream alone, so this points to upstream suite timing under host load rather than a bad merge resolution. Remote CI's bounded lanes should settle it.bash live-claude-draft-guard.sh <worktree>: real Claude Code 2.1.283 in a private tmux socket with a throwaway home frombin/fm-lab-home.sh, driven throughbin/fm-send.sh, then the watcher re-ring viafm_task_inbox_ringbin/fm-test-run.sh tests/fm-operational-input.test.sh tests/fm-afk-inject-e2e.test.sh tests/fm-daemon.test.sh tests/fm-inactive-reconcile.test.sh tests/fm-wake-queue.test.sh tests/fm-pending-reply.test.sh tests/fm-dispatch-resolve.test.sh tests/fm-send-inbox.test.sh(8/8 pass)bin/fm-test-run.sh tests/fm-watch-triage.test.sh, run alone (timed out at 580s after 93 ok and 0 not ok)Fork-preservation diff: every non-blank line the fork added between merge-base 31c47af5 and 09c761f5 checked for presence in HEAD, with the leftovers inspectedgrepof docs/configuration.md comparing Vercel/gateway coverage in HEAD against 09c761f5tests/fm-calm-pi-extension.test.sh- Exact failure on this host (real Pi v0.85.1):not ok - Pi Calm restarted a turn silently after Escape (missing: 'Firstmate supervision continues in a new turn.'). The pane shows 'Error: This operation was aborted' and then MONITOR_HANDLED_queued_on, with no supervision notice. At HEAD the test file is byte-identical to upstream 72b63ee (git diff --quiet 72b63eee HEAD -- tests/fm-calm-pi-extension.test.sh), and.pi/has no diff from upstream. Pure upstream 72b63ee, run in a temporary tree against the same Pi, fails the same assertion with the same message. So the cause is the local Pi version, not the merge. It was not auto-fixed, for two reasons. The recorded decision forbids weakening or skipping the upstream assertion. The workspace boundary forbids upgrading the system Pi. CI installs an unpinned @earendil-works/pi-coding-agent, so the assertion should get a real run there on a newer Pi. The local run is also incomplete apart from this failure: several cases printedskip: installed @earendil-works/pi-coding-agent package not found. Decide whether to proceed and rely on CI, or upgrade the local Pi to 0.87.1 or newer and re-run.🚨 live validation verdict: no-go (4 of 8 scenarios were driven live against the product); failed: Pi Calm announces 'Firstmate supervision continues in a new turn.' after Escape (restored upstream test against real Pi 0.85.1)
Live validation: ❌ no-go - 4 of 8 scenarios driven live against the product
laterchoice dates the hold (restored doc facts d)bash $EVIDENCE/live-claude-draft-guard.sh $WORKTREE(real Claude Code worker on a private tmux socket plus a disposable fm-lab home, driving bin/fm-send.sh and fm_task_inbox_ring)bin/fm-test-run.sh tests/fm-calm-pi-extension.test.shat HEADgit diff --quiet 72b63eee HEAD -- tests/fm-calm-pi-extension.test.shandgit diff --stat 72b63eee HEAD -- .pi/bin/fm-test-run.sh tests/fm-calm-pi-extension.test.shon a temporarygit archive 72b63eeetree (removed afterwards)bin/fm-dispatch-select.shwith a quota-balanced Grok Bot rule, and with an ordered rule plus a launch_failed factFM_HOME=<fm-lab-home without bridge> bin/fm-grok-bot-dispatch.sh brief.md --bot fm-researcherbin/fm-test-run.sh tests/fm-pending-reply.test.sh tests/fm-send-inbox.test.sh tests/fm-dispatch-resolve.test.sh tests/fm-grok-bot-dispatch.test.sh tests/fm-captain-hold-lifecycle.test.shbin/fm-test-run.sh tests/fm-inactive-reconcile.test.sh tests/fm-daemon.test.sh🔧 No changes applied.
✅ Re-checked - no issues remain.
FM_PI_PACKAGE_DIR=~/.linuxbrew/lib/node_modules/@earendil-works/pi-coding-agent bash bin/fm-test-run.sh tests/fm-calm-pi-extension.test.sh(real Pi 0.87.1; 15 ok, 0 not ok, 0 skip)git rev-parse 72b63eee:tests/fm-calm-pi-extension.test.shvsHEAD:(both 5fde9877...)bash ~/.no-mistakes/evidence/01M3EMCMBKQC27KJDYRS6HBK68/live-claude-draft-guard.sh $PWD(real Claude Code worker on a private tmux socket, disposable lab FM_HOME, real bin/fm-send.sh + fm_task_inbox_ring)bash bin/fm-test-run.sh tests/fm-pending-reply.test.sh tests/fm-operational-input.test.sh tests/fm-send-inbox.test.sh tests/fm-inactive-reconcile.test.sh tests/fm-wake-queue.test.shbash bin/fm-test-run.sh tests/fm-dispatch-resolve.test.sh tests/fm-grok-bot-dispatch.test.shbin/fm-dispatch-select.sh qb.json 0with a quota-balanced rule that has a Grok Bot target✅ No issues found.
FM_PI_PACKAGE_DIR=~/.linuxbrew/lib/node_modules/@earendil-works/pi-coding-agent bin/fm-test-run.sh tests/fm-calm-pi-extension.test.sh(unchanged upstream test, real Pi 0.87.1, 15/15 ok, gate_skip=false)bin/fm-test-run.sh tests/fm-calm-pi-extension.test.sh tests/fm-calm-pi-queue-retention-live-e2e.test.sh(default npm-root lookup: the extension suite gate-skipped the package-dependent cases because Pi lives under linuxbrew, not under nvm's npm root; the live queue-retention guard passed against the running Pi 0.87.1)bash live-claude-draft-guard.sh <worktree>: real Claude Code 2.1.283 worker on a private tmux socket with a disposable fm-lab-home, steered through real bin/fm-send.sh while its composer held a draft, then re-rung by fm_task_inbox_ringbin/fm-dispatch-select.sh qb.json 0(quota-balanced rule with a Grok Bot target)FM_HOME=<lab home without bridge> bin/fm-grok-bot-dispatch.sh brief.md --bot fm-researcherbin/fm-dispatch-select.sh ord.json 0 --facts facts.jsonwith launch_failed and with blocked factsbin/fm-test-run.sh tests/fm-pending-reply.test.sh tests/fm-send-inbox.test.sh tests/fm-operational-input.test.sh tests/fm-inactive-reconcile.test.sh tests/fm-captain-hold-lifecycle.test.sh tests/fm-grok-bot-dispatch.test.sh tests/fm-dispatch-resolve.test.sh(7/7 exit 0)git rev-parse HEAD:tests/fm-calm-pi-extension.test.shvs72b63eee:(identical blob 5fde9877) andgit diff --quiet 72b63eee HEAD -- .pi✅ **Document** - passed
✅ No issues found.
✅ No issues found.
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ No issues found.
✅ No issues found.
✅ **Push** - passed
✅ No issues found.
✅ No issues found.
✅ No issues found.