Skip to content

fix(bin): merge upstream firstmate through 72b63eee - #31

Merged
ashkonmousavi merged 53 commits into
mainfrom
fm/firstmate-upstream-catchup-2
Sep 26, 2026
Merged

ashkonmousavi merged 53 commits into
mainfrom
fm/firstmate-upstream-catchup-2

Conversation

@ashkonmousavi

@ashkonmousavi ashkonmousavi commented Sep 26, 2026 •

Copy link
Copy Markdown
Owner

Intent

The captain asked "whats the update with firstmate?" (2026-09-26 01:49). Our fork is current with its own changes, but the original project (remote upstream, kunchenguid/firstmate) has 47 new commits since our last catch-up merge (PR #25), up to 72b63ee, including fixes for problems seen tonight (repeated pending-reply pings to second mates; messages to Claude workers stuck behind text already in their box). Bring them in while keeping all of our own changes.

What Changed

  • Merges 47 upstream kunchenguid/firstmate commits (through 72b63eee) into the fork. Among them are the fixes for repeated pending-reply pings to second mates (dedup of directed source expansions in fm-pending-reply-lib, retired check-row receipts, no repeat of unknown-wake escalations that were already delivered) and for Claude-bound operational input getting stuck behind text already in the box, which is now delivered as a record-backed doorbell.
  • Brings in the other upstream bin/, Pi extension, and Claude Calm mod changes. New scripts include fm-host-mirror.sh (attended supervision for Claude/Cursor hosts), fm-git-strip-ai-trailers.sh, fm-remote-secondmate-relaunch.sh, and fm-lab-home.sh. Upstream also hardened watcher locks, merge and teardown refusals, remote job workers, and X reply argument checks, and added the matching new and updated tests.
  • Resolves the merge conflicts while keeping the fork's own changes. The follow-up commits repair docs that the merge damaged (herdr-backend.md, sessionstart-nudge.md, a duplicated paragraph in captain-hold-lifecycle.md), restore fork-specific facts in configuration.md and captain-hold-lifecycle.md, and restore tests/fm-calm-pi-extension.test.sh byte-for-byte to upstream.

Conflict resolutions

The merge had 12 conflicted files. Each resolution keeps the upstream change and the fork behavior:

  1. .agents/skills/afk/SKILL.md: kept upstream record-backed carrier and alarm wording alongside the fork's keyed wake replacement rule.
  2. bin/fm-spawn.sh: kept the fork's Claude worker settings placeholder and upstream's launch-brief doorbell; settings substitution remains last.
  3. bin/fm-supervise-daemon.sh: combined the fork's keyed escalation refresh with upstream's acknowledgement of unknown wakes.
  4. docs/captain-hold-lifecycle.md: kept upstream's document structure and the fork's dated later deferral and legacy semantics; a duplicate paragraph was removed and omitted fork facts restored in follow-up commits.
  5. docs/configuration.md: kept upstream's attended host and reorganized sections with the fork's Codex quiet, Herdr notification, ordered dispatch, Grok Bot, Vercel dispatch, and Relay cadence facts; omitted fork facts were restored in a follow-up commit.
  6. docs/herdr-backend.md: kept upstream formatting and the fork's sidebar pane, Claude composer, and Codex daemon exemption facts; wording damaged in the merge was repaired in a follow-up commit.
  7. docs/pi-supervision-branch.md: kept upstream structure and the fork's post-merge verification rule.
  8. docs/sessionstart-nudge.md: kept upstream's structure and lock behavior with the fork's Codex 0.154.0 interactive startup and --codex-hook facts; wording damaged in the merge was repaired in a follow-up commit.
  9. docs/supervision-host.md: kept upstream's attended host behavior with the fork's Codex quiet refusal.
  10. tests/fm-afk-launch.test.sh: kept the fork's Codex cases and upstream's test seam.
  11. tests/fm-backend-herdr-workspace-per-home-e2e.test.sh: kept the fork's prep scaffold and upstream's secondmate Git initialization.
  12. tests/fm-daemon.test.sh: kept the fork's keyed escalation cases and upstream's doorbell cases.

The merge also changed tests/fm-calm-pi-extension.test.sh outside a conflict. Review restored that file byte-for-byte to upstream 72b63eee. Its full suite passes on Pi 0.87.1 (15/15, no skips).

Risk Assessment

✅ Low: Every file changed by only one side of the merge is byte-identical to that parent. A line-level audit found only three parent-added lines missing across all non-doc files that both sides changed: all three were in the afk skill doc and fm-supervise-daemon.sh's escalate_add, and each was an intentional conflict resolution that keeps both behaviors. In escalate_add, keyed check replacement sits beside upstream's unknown-wake acknowledgement, and every check-wake producer uses a check: payload. The fork facts in the docs were restored and checked against the code, and the Pi Calm test is byte-identical to upstream.

Testing

This round was driven at HEAD fec95ee. The unchanged upstream Pi Calm suite was run against the installed Pi 0.87.1: 15/15 passed with no gate skips, including the "delivers it once in a new announced turn" Escape assertion, and the live Pi queue-retention guard also passed. With the default npm-root lookup, the package-dependent Pi Calm cases gate-skip on this host because Pi is installed under linuxbrew, not under nvm's npm root. Pointing FM_PI_PACKAGE_DIR at the installed package fixed that setup issue without touching the test. A real Claude Code worker was driven live on a private tmux socket with a disposable lab home. fm-send left the captain's draft untouched, recorded the steer durably, and did not type the doorbell over the draft. After the draft was cleared, the re-ring got the message acted on and acknowledged in 6s. The documented Grok Bot facts were driven through the real CLIs in a throwaway lab home: the quota-balanced refusal (exit 2), the bridge-absent refusal naming the primary (exit 2), and launch_failed and blocked facts moving ordered selection to the next candidate. Seven targeted suites passed, covering the pending-reply, send-inbox, operational-input, captain-hold (including later deferrals), and dispatch/Grok Bot behavior. The repeated-ping fix and captain-hold behavior are covered only by these suites, not live, so those two scenarios are recorded as untested. No Herdr lab was needed because no Herdr lifecycle is touched. The worktree is clean and all temporary lab dirs and tmux servers were removed.

  • Live validation: ✅ go - 6 of 8 scenarios driven live against the product
Scenario Result Live Evidence
Pi Calm announces 'Firstmate supervision continues in a new turn.' after Escape, and delivers the queued Firstmate notification once (unchanged upstream test, real Pi 0.87.1) ✅ pass live round6-calm-pi-extension-pi-0.87.1-full.log: 'ok - Pi 0.87.1 with Calm on keeps a queued Firstmate notification unlisted, out of the editor on Escape, and delivers it once in a new announced turn'; te…
The whole restored upstream Pi Calm suite runs without skipping or weakening any assertion on Pi 0.87.1 ✅ pass live round6-calm-pi-extension-pi-0.87.1-full.log: 15 ok, gate_skip=false, exit=0
A running Pi 0.87.1 session exposes every queue-retention member Calm preflights ✅ pass live round6-calm-pi-pi-0.87.1.log: 'ok - Pi 0.87.1 exposes every queue-retention member Calm preflights'
Adversarial: a steer to a Claude worker whose box already holds captain draft text is not typed over the draft, is recorded durably, and is acted on and acknowledged after the draft clears and the wat… ✅ pass live round6-live-claude-draft-guard.log: 'doorbell skipped (composer visibly holds pending text)', draft still in composer, worker did not act; after re-ring acted=yes acked=yes (6s); RESULT: PASS
Adversarial: fm-dispatch-select refuses a quota-balanced rule that uses a Grok Bot target ✅ pass live round6-grok-bot-doc-facts-cli.txt: 'error: config: a quota-balanced rule cannot use a grok_bot profile' exit=2
In a home without the bridge (secondmate), fm-grok-bot-dispatch exits 2 and names the primary; a firstmate-supplied launch_failed or blocked fact makes ordered selection fall to the next candidate ✅ pass live round6-grok-bot-doc-facts-cli.txt: bridge-not-found exit=2; skipped[0] launch_failed / blocked, selected[1] claude sonnet
Second mates are not re-pinged repeatedly for the same pending reply, and the steering inbox, operational input, and inactive reconcile behave as upstream intends ⏸️ untested no The prior payload did not establish a live result. Only the fm-pending-reply, fm-send-inbox, fm-operational-input, and fm-inactive-reconcile suites passed (round6-targeted-suites.log). Driving repeate…
Captain hold board 'later' dates the open hold, 'done' and 'release' keep their modes, and undated, invalid, past, or non-later deferrals are refused (restored doc facts); typed resolution reports a G… ⏸️ untested no The prior payload did not establish a live result. Only the fm-captain-hold-lifecycle, fm-dispatch-resolve, and fm-grok-bot-dispatch suites passed (round6-targeted-suites.log). Typed resolution needs…
Evidence: Upstream Pi Calm suite on Pi 0.87.1 (15/15, no skips)

Source: Upstream Pi Calm suite on Pi 0.87.1 (15/15, no skips)

# pi --version: 0.87.1 (~/.linuxbrew/bin/pi)
# npm root -g (default lookup): ~/.nvm/versions/node/v22.22.0/lib/node_modules - Pi is not installed there, so the default lookup gate-skips
# FM_PI_PACKAGE_DIR=~/.linuxbrew/lib/node_modules/@earendil-works/pi-coding-agent (version 0.87.1)
# test file blob HEAD 5fde98777da0ced96b4af76094222f706519148a / upstream 72b63eee 5fde98777da0ced96b4af76094222f706519148a; HEAD fec95ee2
FM_TEST_BEGIN 2026-09-26T12:48:33Z tests/fm-calm-pi-extension.test.sh family=pure-contract-unit expected_gate_skip=none
ok - Pi calm resolves its persistent home independently of Pi's launch directory
ok - Pi calm compatibility evidence never rejects a Pi version for being newer than 0.82.0, and still fails closed on a missing or malformed version
ok - a missing collapsed-thinking presentation API degrades only that Calm adapter with a clear skip reason, while the rest of Calm still registers
ok - missing Pi presentation class exports reach the independent adapter degradation path
ok - Calm hides queued Firstmate rows only on a session that can keep them, keeps hidden ones out of the editor on Escape, delivers them once in order, and leaves unsupported sessions and Calm off stock
ok - Calm registers none of its 7 built-in tool wrappers at load while config/calm is off, and all 7 synchronously at load while config/calm is on
ok - Calm's first same-session /calm activation claims every uncontested built-in, leaves a foreign bash tool fully intact and callable, warns prominently and logs the contested name, and only rows constructed before that activation - the documented bound - fail to retroactively collapse
ok - Pi calm centralizes transcript visibility, preserves execution/export data, keeps Pi's stock working row visible while no run is active, and persists its choice across session starts
ok - Pi calm on collapses mid-turn assistant working notes to zero height while Calm off keeps them, leaves streaming, truncated-final, and genuine final replies untouched, never mutates the messages, ignores every /calm argument, and restores a legacy persisted max as ordinary Calm on
ok - Pi operational follow-up E2E processes exact user-role notifications once while Calm hides current and adjacent rows, Calm off and absent render them, and restart preserves semantics
ok - Pi 0.87.1 with Calm on keeps a queued Firstmate notification unlisted, out of the editor on Escape, and delivers it once in a new announced turn, while Calm off stays stock
ok - Pi Calm native /skill:ahoy geometry keeps every collapsed thinking and tool block at zero height while preserving expansion, history, restart, and Calm-off rendering
ok - Pi Calm working ship keeps its centered two-row asymmetric Unicode boat inside a deterministic long-wave trough, paints all water standard blue and the whole boat standard yellow with balanced resets, keeps ANSI-stripped width exact, reverses cleanly at both edges and every width, clamps visible and hidden resizes, falls back deterministically when narrow, freezes and resumes across settle/start without hidden-time jumps or duplicate timers, resets only on a fresh session, and leaves Calm-off visibility untouched
ok - the rendered-export-DOM guard renders in one pass, retries a bounded number of Chrome start-up failures, and reports the Chrome binary, Chrome version, Pi version, exit status, and Chrome diagnostic when every attempt fails
ok - Pi calm native E2E replaces the stock working row with a moving, resize-clamped working ship that freezes and resumes across two working periods in one Pi session, clears on abort, keeps captain turns visible, hides exact operational user rows without changing persistence, restores stock rendering Calm-off, survives restart, and preserves export plus Ctrl+O behavior
FM_TEST_END 2026-09-26T12:49:11Z tests/fm-calm-pi-extension.test.sh exit=0 duration_ms=38726 gate_skip=false
FM_TEST_SUMMARY total=1 failed=0 skipped_gate=0 duration_ms=38767
FM_TEST_SUMMARY_FAMILY family=pure-contract-unit count=1 duration_ms=38726 failed=0
FM_TEST_SLOWEST rank=1 script=tests/fm-calm-pi-extension.test.sh duration_ms=38726
runner exit=0
Evidence: Pi Calm default-lookup run plus live queue-retention guard on Pi 0.87.1 (the header has a harmless package-path probe error from the evidence script)

Source: Pi Calm default-lookup run plus live queue-retention guard on Pi 0.87.1 (the header has a harmless package-path probe error from the evidence script)

# pi --version: 0.87.1
# resolved pi package: ~/.linuxbrew/lib/node_modules/@earendil-works/pi-coding-agent/dist (node:internal/modules/cjs/loader:1386
  throw err;
  ^

Error: Cannot find module '~/.linuxbrew/lib/node_modules/@earendil-works/pi-coding-agent/dist/package.json'
Require stack:
- ~/.no-mistakes/worktrees/befb827bbae4/01M3EMCMBKQC27KJDYRS6HBK68/[eval]
    at Function._resolveFilename (node:internal/modules/cjs/loader:1383:15)
    at defaultResolveImpl (node:internal/modules/cjs/loader:1025:19)
    at resolveForCJSWithHooks (node:internal/modules/cjs/loader:1030:22)
    at Function._load (node:internal/modules/cjs/loader:1192:37)
    at TracingChannel.traceSync (node:diagnostics_channel:328:14)
    at wrapModuleLoad (node:internal/modules/cjs/loader:237:24)
    at Module.require (node:internal/modules/cjs/loader:1463:12)
    at require (node:internal/modules/helpers:147:16)
    at [eval]:1:1
    at runScriptInThisContext (node:internal/vm:209:10) {
  code: 'MODULE_NOT_FOUND',
  requireStack: [
    '~/.no-mistakes/worktrees/befb827bbae4/01M3EMCMBKQC27KJDYRS6HBK68/[eval]'
  ]
}

Node.js v22.22.0)
# test file blob HEAD 5fde98777da0ced96b4af76094222f706519148a / upstream 72b63eee 5fde98777da0ced96b4af76094222f706519148a
# HEAD fec95ee2
FM_TEST_BEGIN 2026-09-26T12:47:44Z tests/fm-calm-pi-extension.test.sh family=pure-contract-unit expected_gate_skip=none
skip: installed @earendil-works/pi-coding-agent package not found
ok - Pi calm compatibility evidence never rejects a Pi version for being newer than 0.82.0, and still fails closed on a missing or malformed version
skip: installed @earendil-works/pi-coding-agent package not found
ok - missing Pi presentation class exports reach the independent adapter degradation path
skip: installed @earendil-works/pi-coding-agent package not found
skip: installed @earendil-works/pi-coding-agent package not found
skip: installed @earendil-works/pi-coding-agent package not found
skip: installed @earendil-works/pi-coding-agent package not found
skip: installed @earendil-works/pi-coding-agent package not found
ok - Pi operational follow-up E2E processes exact user-role notifications once while Calm hides current and adjacent rows, Calm off and absent render them, and restart preserves semantics
ok - Pi 0.87.1 with Calm on keeps a queued Firstmate notification unlisted, out of the editor on Escape, and delivers it once in a new announced turn, while Calm off stays stock
ok - Pi Calm native /skill:ahoy geometry keeps every collapsed thinking and tool block at zero height while preserving expansion, history, restart, and Calm-off rendering
skip: installed @earendil-works/pi-coding-agent package not found
ok - the rendered-export-DOM guard renders in one pass, retries a bounded number of Chrome start-up failures, and reports the Chrome binary, Chrome version, Pi version, exit status, and Chrome diagnostic when every attempt fails
ok - Pi calm native E2E replaces the stock working row with a moving, resize-clamped working ship that freezes and resumes across two working periods in one Pi session, clears on abort, keeps captain turns visible, hides exact operational user rows without changing persistence, restores stock rendering Calm-off, survives restart, and preserves export plus Ctrl+O behavior
fm-test-run: gate skip: tests/fm-calm-pi-extension.test.sh: installed @earendil-works/pi-coding-agent package not found
FM_TEST_END 2026-09-26T12:48:11Z tests/fm-calm-pi-extension.test.sh exit=0 duration_ms=26286 gate_skip=true
FM_TEST_BEGIN 2026-09-26T12:48:11Z tests/fm-calm-pi-queue-retention-live-e2e.test.sh family=live-harness-optin expected_gate_skip=live-capability
ok - Pi 0.87.1 exposes every queue-retention member Calm preflights before hiding queued Firstmate rows
FM_TEST_END 2026-09-26T12:48:12Z tests/fm-calm-pi-queue-retention-live-e2e.test.sh exit=0 duration_ms=904 gate_skip=false
FM_TEST_SUMMARY total=2 failed=0 skipped_gate=1 duration_ms=27266
FM_TEST_SUMMARY_FAMILY family=live-harness-optin count=1 duration_ms=904 failed=0
FM_TEST_SUMMARY_FAMILY family=pure-contract-unit count=1 duration_ms=26286 failed=0
FM_TEST_SLOWEST rank=1 script=tests/fm-calm-pi-extension.test.sh duration_ms=26286
FM_TEST_SLOWEST rank=2 script=tests/fm-calm-pi-queue-retention-live-e2e.test.sh duration_ms=904
runner exit=0
Evidence: Live Claude Code worker: draft not overwritten, re-ring acted+acked

Source: Live Claude Code worker: draft not overwritten, re-ring acted+acked

# HEAD fec95ee2 claude 2.1.283 (Claude Code)
composer before draft: empty
composer with draft: pending
----- pane: draft typed, before fm-send
▝▜██████▀  Opus 5.5 with medium effort · Claude Max
 ▝▝   ▝▝   ~/.no-mistakes/worktrees/befb827bbae4/01M3EMCMBKQC27KJDYRS6HBK68 · /rc
                                                                                                                                                                                                        ◐ medium · /effort
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
❯ CAPTAIN_DRAFT_do_not_send_yet
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
  session · WORKTREE:01M3EMCMBKQC27KJDYRS... · opus/medium · default · -- · 5h--/7d-- · $0.00
  ⏵⏵ bypass permissions on (shift+tab to cycle)
fm-gate-refuse: gate agent lifecycle permitted only against lab home /tmp/fm-draft-live.7D3TCG/home
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
●  WATCHER DOWN - SUPERVISION IS OFF
●  1 task(s) in flight, but no watcher has a fresh beacon (last beat: never, grace 300s).
●  Trust the emitted supervision protocol for this harness; do not use shell & for watcher repair.
●  This is a supervision warning only; the requested message WILL still be sent.
●  watcher supervision needs Stop-owned automatic recovery; inspect the hook registration and startup status before ending the turn.
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
fm-send: doorbell skipped (composer visibly holds pending text); the steer is durably recorded at /tmp/fm-draft-live.7D3TCG/home/state/live-draft.inbox/001.msg and the watcher will re-ring
fm-send exit=0
durable record: present
----- pane: after fm-send with draft present
▝▜██████▀  Opus 5.5 with medium effort · Claude Max
 ▝▝   ▝▝   ~/.no-mistakes/worktrees/befb827bbae4/01M3EMCMBKQC27KJDYRS6HBK68 · /rc
                                                                                                                                                      Advisor Tool (experimental) is on and may use more tokens · /advisor
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
❯ CAPTAIN_DRAFT_do_not_send_yet
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
  session · WORKTREE:01M3EMCMBKQC27KJDYRS... · opus/medium · default · -- · 5h--/7d-- · $0.00
  ⏵⏵ bypass permissions on (shift+tab to cycle)
draft still in composer: yes
doorbell text visible in pane: no
worker acted while draft present: no
composer after clearing: empty
re-ring rc=0
----- pane: after re-ring
  ⎿  $ D='/tmp/fm-draft-live.7D3TCG/home/state/live-draft.inbox'; touch /tmp/fm-draft-live.7D3TCG/acted && mv "$D/001.msg" "$D/handled/" && ls "$D" "$D/handled"
* Ideating… (6s · ↓ 365 tokens · thought for 1s)
                                                                                                                                                      Advisor Tool (experimental) is on and may use more tokens · /advisor
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
❯ 
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
  Firstmate instruction queue... · WORKTREE:01M3EMCMBKQC27KJDYRS... · opus/medium · default · 8% · 5h36%/7d23% · $0.53
  ⏵⏵ bypass permissions on (shift+tab to cycle) · ← 2 agents
after re-ring: acted=yes acked=yes (6s)
RESULT: PASS
script exit=0
Evidence: Live draft-guard driver script

Source: Live draft-guard driver script

#!/usr/bin/env bash
# Live adversarial check: a real Claude Code worker whose composer already holds
# unsent draft text is steered through the real bin/fm-send.sh. The send must
# record the message durably but must NOT type the doorbell on top of the draft
# (no merged submission). After the draft is cleared, the watcher's re-ring
# (fm_task_inbox_ring) must get the message acted on and acknowledged.
# Isolation: private tmux socket, throwaway FM_HOME under a temp dir.
set -u
# Runs as a no-mistakes gate agent: fm-send is authorized only against a marked
# disposable lab home (bin/fm-lab-home.sh) with no FM_*_OVERRIDE.
ROOT=${1:?worktree root}
SOCKET="fm-draft-live-$$"
SESSION=draftlive
WIN=hx-claude
LAB=$(mktemp -d "${TMPDIR:-/tmp}/fm-draft-live.XXXXXX")
TIMEOUT=${TIMEOUT:-150}
cleanup() { tmux -L "$SOCKET" kill-server 2>/dev/null || true; rm -rf "$LAB"; }
trap cleanup EXIT
SHIM="$LAB/shim"; mkdir -p "$SHIM"
REAL_TMUX=$(command -v tmux)
printf '#!/usr/bin/env bash\nexec "%s" -L "%s" "$@"\n' "$REAL_TMUX" "$SOCKET" > "$SHIM/tmux"; chmod +x "$SHIM/tmux"
PATH="$SHIM:$PATH"
. "$ROOT/bin/fm-tmux-lib.sh"
. "$ROOT/bin/fm-task-inbox-lib.sh"
snap() { echo "----- pane: $1"; tmux -L "$SOCKET" capture-pane -p -t "$SESSION:$WIN" | grep '[^[:space:]]' | tail -8; }

tmux -L "$SOCKET" new-session -d -s "$SESSION" -x 220 -y 50 -c "$ROOT"
tmux -L "$SOCKET" new-window -d -t "$SESSION:" -n "$WIN" -c "$ROOT" -- bash -lc \
  "CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false CLAUDE_CODE_SEND_FEEDBACK=0 claude --dangerously-skip-permissions --settings '{\"feedbackDrafts\":\"off\"}'"
i=0; while [ $i -lt 60 ]; do [ "$(fm_tmux_composer_state "$SESSION:$WIN")" = empty ] && break; sleep 1; i=$((i+1)); done
echo "composer before draft: $(fm_tmux_composer_state "$SESSION:$WIN")"

DRAFT="CAPTAIN_DRAFT_do_not_send_yet"
tmux -L "$SOCKET" send-keys -t "$SESSION:$WIN" -l "$DRAFT"
sleep 2
echo "composer with draft: $(fm_tmux_composer_state "$SESSION:$WIN")"
snap "draft typed, before fm-send"

home=$("$ROOT/bin/fm-lab-home.sh" create "$LAB/home") || { echo "lab home create failed"; exit 1; }; mkdir -p "$home/state"; task=live-draft; acted="$LAB/acted"
printf 'window=%s:%s\nkind=ship\nharness=claude\n' "$SESSION" "$WIN" > "$home/state/$task.meta"
FM_HOME="$home" "$ROOT/bin/fm-send.sh" "$task" \
  "Firstmate live check: run exactly this shell command now: touch $acted - then follow the mv instruction you were given for this message. Reply with one short line."
echo "fm-send exit=$?"
rec="$home/state/$task.inbox/001.msg"; handled="$home/state/$task.inbox/handled/001.msg"
[ -f "$rec" ] && echo "durable record: present" || echo "durable record: MISSING"
sleep 8
snap "after fm-send with draft present"
pane=$(tmux -L "$SOCKET" capture-pane -p -S -200 -t "$SESSION:$WIN")
fail=0
printf '%s\n' "$pane" | grep -q "$DRAFT" && echo "draft still in composer: yes" || { echo "draft still in composer: NO"; fail=1; }
if printf '%s\n' "$pane" | grep -qE "inbox|\.msg"; then echo "doorbell text visible in pane: YES (typed over draft)"; fail=1; else echo "doorbell text visible in pane: no"; fi
[ -e "$acted" ] && { echo "worker acted while draft present: YES"; fail=1; } || echo "worker acted while draft present: no"

# Captain clears the draft; the watcher re-rings.
tmux -L "$SOCKET" send-keys -t "$SESSION:$WIN" C-u
sleep 1
tmux -L "$SOCKET" send-keys -t "$SESSION:$WIN" Escape
sleep 2
echo "composer after clearing: $(fm_tmux_composer_state "$SESSION:$WIN")"
fm_task_inbox_ring tmux "$SESSION:$WIN" "$rec"; echo "re-ring rc=$?"
i=0; while [ $i -lt "$TIMEOUT" ]; do [ -f "$handled" ] && [ -e "$acted" ] && break; sleep 1; i=$((i+1)); done
snap "after re-ring"
if [ -f "$handled" ] && [ -e "$acted" ]; then echo "after re-ring: acted=yes acked=yes (${i}s)"; else echo "after re-ring: acted=$([ -e "$acted" ] && echo yes || echo no) acked=$([ -f "$handled" ] && echo yes || echo no)"; fail=1; fi
echo "RESULT: $([ $fail -eq 0 ] && echo PASS || echo FAIL)"
exit $fail
Evidence: Grok Bot documented behaviors via real CLIs

Source: Grok Bot documented behaviors via real CLIs

# HEAD fec95ee2; lab dir /tmp/fm-grok6.mwlv9Z
$ fm-dispatch-select.sh qb.json 0   # doc: fm-dispatch-select refuses a quota-balanced Grok Bot rule
{"rules":[{"when":"Web research.","select":"quota-balanced","use":[{"grok_bot":"fm-researcher"},{"harness":"claude","model":"sonnet"}]}]}
error: config: a quota-balanced rule cannot use a grok_bot profile; Grok Bot targets have no quota evidence
exit=2

$ FM_HOME=<lab home without bridge> fm-grok-bot-dispatch.sh brief.md --bot fm-researcher   # doc: secondmate home refuses exit 2 and names it
error: Grok Bot bridge not found: /tmp/fm-grok6.mwlv9Z/home/data/grok-bot-bridge/grokbot.mjs; Grok Bot targets run only in the home that holds the bridge (the primary), so treat this candidate as unavailable and fall to the rule's other candidates
exit=2

$ fm-dispatch-select.sh ord.json 0 --facts facts.json   # doc: firstmate supplies launch_failed, rule falls to next candidate
[{"candidate":0,"kind":"launch_failed","detail":"fm-grok-bot-dispatch.sh exit 2: bridge absent in secondmate home"}]
selection: ordered
skipped[0]: launch_failed: fm-grok-bot-dispatch.sh exit 2: bridge absent in secondmate home; profile={"grok_bot":"fm-researcher"}
selected[1]: {"harness":"claude","model":"sonnet","effort":"high"}
exit=0

$ fm-dispatch-select.sh ord.json 0 --facts blocked.json   # doc: usage spent -> blocked fact moves ordered selection
selection: ordered
skipped[0]: blocked: Grok Bot usage spent; profile={"grok_bot":"fm-researcher"}
selected[1]: {"harness":"claude","model":"sonnet","effort":"high"}
exit=0

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 1 issue found → auto-fixed (3) ✅
  • ⚠️ docs/captain-hold-lifecycle.md:237 - The conflict resolution kept two versions of the same paragraph. Upstream's restructured three-line paragraph is at lines 232-234 ("bin/fm-procevent-lavish.sh answers is one such built-in adapter command. ... relays a card's declared close mode ..."). The fork's older one-line version is at line 237, after the trust-boundary sentence ("... one such adapter command; it reads only rows tagged choice, relays the selected option's close mode ..."). So the doc describes the same adapter twice, with slightly different wording on where the close mode comes from. Fix: delete line 237 and change line 233 to say "relays the selected option's close mode". That keeps the fork's per-option close-mode meaning, which the defer:YYYY-MM-DD row in the mode table above relies on. No other doc conflict resolution has this problem; I checked architecture.md, configuration.md, herdr-backend.md, pi-supervision-branch.md, scripts.md, sessionstart-nudge.md, supervision-host.md and supervision-protocols/supervision-host.md.

🔧 Fix applied.
✅ Re-checked - no issues remain.

  • ⚠️ tests/fm-calm-pi-extension.test.sh:2643 - The merge commit rewrote an upstream test that neither parent changed. The intent is to bring upstream in and keep the fork's own changes, and it does not ask for upstream tests to be rewritten. git show --remerge-diff 8430bdbd lists this file with no CONFLICT header, and git diff 72b63eee HEAD shows edits that exist only in the merge commit. There are three edits. (1) The upstream assertion that Pi Calm announces "Firstmate supervision continues in a new turn." after Escape now runs only when the installed Pi is 0.87.1 or newer, which weakens the test on older Pi. This host has Pi 0.85.1, so the assertion is skipped here. The version check also fails open: it splits pi --version output on dots and passes it through Number(). Any output that is not a bare x.y.z, such as a pi prefix or a pre-release suffix, gives NaN, and the assertion is skipped without a message. CI installs an unpinned @earendil-works/pi-coding-agent (.github/workflows/ci.yml:111,215), so nothing guarantees the assertion runs anywhere. (2) At line 567, makeSession gains session.agent = { abort: () =&gt; session.abort() }. No code under .pi/ reads session.agent. (3) At line 2609, follow-up submission changes from M-Enter to C-q, backed by a new keybindings.json at line 2508. The notice comes from .pi/extensions/lib/fm-calm-pending-operational-layout.ts:292 after the adapter drains a retained queue. The 0.85.1 carve-out is plausible, but it is a host-driven change to an upstream suite rather than part of the catch-up. The smallest remedy is to remove these merge-only edits and restore the upstream test. If the host adaptation is wanted, keep it as a separate, reviewed change with a version check that fails closed.
  • ⚠️ docs/configuration.md:1064 - The docs conflict resolution condensed the fork's text, and some fork facts were lost or changed. That goes against 'keeping all of our own changes'. (a) The fork documented that typed resolution reports a Grok Bot target as eligible but unranked and never emits it as a profile: line. The merged Grok Bot paragraph (docs/configuration.md:1059-1065) no longer says this, but bin/fm-dispatch-resolve.sh:372 still does it. (b) docs/configuration.md:1064 now reads 'a Bot dispatch there refuses and supplies a launch_failed fact'. bin/fm-grok-bot-dispatch.sh only exits 2 (die, line 41). The fork text said Firstmate supplies the launch_failed fact, so the new sentence gives that job to the script. (c) The fork's rule that bootstrap, fm-dispatch-select.sh, and typed resolution all refuse a quota-balanced Grok Bot rule is narrowed to 'Quota-balanced selection refuses' at line 1065. (d) Same class in docs/captain-hold-lifecycle.md. At lines 526-527, the verification record lost the fork's coverage of 'a board later choice dating the open hold under its unchanged reason while done and release keep their modes' and 'the refusal of undated, invalid, past, or non-later deferrals'. That behavior is still pinned in tests/fm-captain-hold-lifecycle.test.sh. At line 186 the doc still says 'a card-declared close', while the fork, and line 234 after the round-1 fix, say the close mode comes from the selected option. Fix: restore (a) through (c) in the Grok Bot paragraph with Firstmate as the subject of (b), add the later deferral coverage to the Answers bullet at captain-hold-lifecycle.md:527, and change line 186 to 'the selected option's close'.

🔧 Fix applied.
1 warning still open:

  • ⚠️ docs/configuration.md:1064 - The docs conflict resolution condensed the fork's text, and some fork facts were lost or changed. That goes against 'keeping all of our own changes'. (a) The fork documented that typed resolution reports a Grok Bot target as eligible but unranked and never emits it as a profile: line. The merged Grok Bot paragraph (docs/configuration.md:1059-1065) no longer says this, but bin/fm-dispatch-resolve.sh:372 still does it. (b) docs/configuration.md:1064 now reads 'a Bot dispatch there refuses and supplies a launch_failed fact'. bin/fm-grok-bot-dispatch.sh only exits 2 (die, line 41). The fork text said Firstmate supplies the launch_failed fact, so the new sentence gives that job to the script. (c) The fork's rule that bootstrap, fm-dispatch-select.sh, and typed resolution all refuse a quota-balanced Grok Bot rule is narrowed to 'Quota-balanced selection refuses' at line 1065. (d) Same class in docs/captain-hold-lifecycle.md. At lines 526-527, the verification record lost the fork's coverage of 'a board later choice dating the open hold under its unchanged reason while done and release keep their modes' and 'the refusal of undated, invalid, past, or non-later deferrals'. That behavior is still pinned in tests/fm-captain-hold-lifecycle.test.sh. At line 186 the doc still says 'a card-declared close', while the fork, and line 234 after the round-1 fix, say the close mode comes from the selected option. Fix: restore (a) through (c) in the Grok Bot paragraph with Firstmate as the subject of (b), add the later deferral coverage to the Answers bullet at captain-hold-lifecycle.md:527, and change line 186 to 'the selected option's close'.

🔧 Fix applied.
✅ Re-checked - no issues remain.

✅ No issues found.

🔧 **Test** - 1 issue found → no changes applied (2) ✅
  • ⚠️ The Test agent did not finish within its invocation budget. Reported: agent run tests timed out after 30m0s: agent last produced output 5s ago (113 observed); agent reported: claude parse events: context deadline exceeded. This is a budget or provider-slowness cut, not a code failure. Re-running the same request costs another full budget, so no further attempt is made automatically. If this repository's targeted tests or evidence gathering routinely approach the default 30m0s, raise test_agent_timeout in global config. Respond with fix to spend another budget: a repair turn runs only for selected findings other than this budget cut, then validation re-runs. Or abort and retry after raising the budget.

🔧 No changes applied.
1 warning still open:

  • ⚠️ tests/fm-watch-triage.test.sh - tests/fm-watch-triage.test.sh is timing-sensitive on this host. Run alone in round 2, it passed all 93 cases but did not finish within 580s. In round 1 it failed three different ways on HEAD (a heartbeat re-surface, TERM exit timing, and immediate stale surfacing). It also failed once on pure upstream 72b63ee and passed on the pre-merge fork 09c761f. The failures change from run to run and also show up on upstream alone, so this points to upstream suite timing under host load rather than a bad merge resolution. Remote CI's bounded lanes should settle it.
  • Live validation: ✅ go - 2 of 8 scenarios driven live against the product
Scenario Result Live Evidence
A captain's unsent draft in a real Claude worker's input box is not typed over: fm-send saves the message to the inbox and skips the doorbell; the worker does not act and the draft stays intact ✅ pass live round2-live-claude-draft-guard.log: 'doorbell skipped (composer visibly holds pending text)', 'draft still in composer: yes', 'doorbell text visible in pane: no', 'worker acted while draft present: no…
Once the draft is cleared, the watcher's re-ring delivers the stuck message to the Claude worker, which acts on it and moves 001.msg to handled/ ✅ pass live round2-live-claude-draft-guard.log: 'after re-ring: acted=yes acked=yes (6s)', RESULT: PASS
Claude-bound operational input (away-daemon injections) goes out as a saved inbox record plus doorbell rather than raw typing (upstream 1fe1a67) ⏸️ untested no The prior payload did not establish a live result. Coverage came only from tests/fm-operational-input.test.sh and tests/fm-afk-inject-e2e.test.sh with stub agents (both passed); a real away-daemon ses…
An escalation or wake that was already delivered is not re-sent: an acknowledged unknown wake does not escalate again, and branch acknowledgement retires check-row receipts (upstream 52e679b, c33b3f6… ⏸️ untested no The prior payload did not establish a live result. Coverage came only from tests/fm-daemon.test.sh, tests/fm-wake-queue.test.sh and tests/fm-inactive-reconcile.test.sh with fake tmux (all passed); rep…
Pending-reply library still works after the upstream shellcheck-directive dedup (72b63ee) ⏸️ untested no The prior payload did not establish a live result. Only tests/fm-pending-reply.test.sh was run (passed); the change touches only ShellCheck directives and keeps runtime '.' commands byte-identical, so…
A secondmate's routine working/paused status is absorbed instead of pinging firstmate again (upstream 97365aa) ⏸️ untested no The prior payload did not establish a live result. Coverage came only from the watcher-triage suite with fixture panes (relevant cases ok, suite timed out at 580s with 0 failures); no live fleet was r…
Fork's Vercel AI Gateway fallback for typed dispatch resolution survives the merge ⏸️ untested no The prior payload did not establish a live result. Only tests/fm-dispatch-resolve.test.sh and a docs grep were run (both consistent with 09c761f); a live call needs real TypeSafe and Vercel AI Gatewa…
No fork-owned code is lost in the merge (adversarial check of every line the fork added) ⏸️ untested no The prior payload did not establish a live result. This was a static diff check (round2-fork-lines-reworded.txt: 61 fork lines absent verbatim, 60 reworded doc sentences, one bin line gained an extra…
  • bash live-claude-draft-guard.sh &lt;worktree&gt;: real Claude Code 2.1.283 in a private tmux socket with a throwaway home from bin/fm-lab-home.sh, driven through bin/fm-send.sh, then the watcher re-ring via fm_task_inbox_ring

  • bin/fm-test-run.sh tests/fm-operational-input.test.sh tests/fm-afk-inject-e2e.test.sh tests/fm-daemon.test.sh tests/fm-inactive-reconcile.test.sh tests/fm-wake-queue.test.sh tests/fm-pending-reply.test.sh tests/fm-dispatch-resolve.test.sh tests/fm-send-inbox.test.sh (8/8 pass)

  • bin/fm-test-run.sh tests/fm-watch-triage.test.sh, run alone (timed out at 580s after 93 ok and 0 not ok)

  • Fork-preservation diff: every non-blank line the fork added between merge-base 31c47af5 and 09c761f5 checked for presence in HEAD, with the leftovers inspected

  • grep of docs/configuration.md comparing Vercel/gateway coverage in HEAD against 09c761f5

  • ⚠️ tests/fm-calm-pi-extension.test.sh - Exact failure on this host (real Pi v0.85.1): not ok - Pi Calm restarted a turn silently after Escape (missing: &#39;Firstmate supervision continues in a new turn.&#39;). The pane shows 'Error: This operation was aborted' and then MONITOR_HANDLED_queued_on, with no supervision notice. At HEAD the test file is byte-identical to upstream 72b63ee (git diff --quiet 72b63eee HEAD -- tests/fm-calm-pi-extension.test.sh), and .pi/ has no diff from upstream. Pure upstream 72b63ee, run in a temporary tree against the same Pi, fails the same assertion with the same message. So the cause is the local Pi version, not the merge. It was not auto-fixed, for two reasons. The recorded decision forbids weakening or skipping the upstream assertion. The workspace boundary forbids upgrading the system Pi. CI installs an unpinned @earendil-works/pi-coding-agent, so the assertion should get a real run there on a newer Pi. The local run is also incomplete apart from this failure: several cases printed skip: installed @earendil-works/pi-coding-agent package not found. Decide whether to proceed and rely on CI, or upgrade the local Pi to 0.87.1 or newer and re-run.

  • 🚨 live validation verdict: no-go (4 of 8 scenarios were driven live against the product); failed: Pi Calm announces 'Firstmate supervision continues in a new turn.' after Escape (restored upstream test against real Pi 0.85.1)

  • Live validation: ❌ no-go - 4 of 8 scenarios driven live against the product

Scenario Result Live Evidence
Claude worker with an unsent captain draft gets a steer via fm-send: the message is saved to the inbox, the draft is not overwritten or submitted, and after the draft is cleared the watcher re-ring ge… ✅ pass live round3-live-claude-draft-guard.log: 'doorbell skipped (composer visibly holds pending text)', durable record present, draft still in composer, no doorbell text, worker did not act; after the re-ring,…
Pi Calm announces 'Firstmate supervision continues in a new turn.' after Escape (restored upstream test against real Pi 0.85.1) ❌ fail live round3-calm-pi-extension.log: not ok - Pi Calm restarted a turn silently after Escape (missing: 'Firstmate supervision continues in a new turn.'), exit=1
HEAD's Pi Calm behavior matches pure upstream 72b63ee on the same Pi (test byte-identical, .pi/ unchanged, identical outcomes) ✅ pass live round3-calm-pi-extension.log vs round3-calm-pi-extension-upstream-72b63eee.log: both give the same ok and not ok lines
Restored Grok Bot doc facts hold when running the real CLIs: quota-balanced Grok Bot rule refused by fm-dispatch-select; fm-grok-bot-dispatch in a home without the bridge exits 2 and names the primary… ✅ pass live round3-grok-bot-doc-facts-cli.txt
Typed resolution reports a Grok Bot target as eligible but unranked and never emits it as a profile: line (restored doc fact a) ⏸️ untested no The prior payload did not establish a live result. Only tests/fm-dispatch-resolve.test.sh ran, against a stubbed TypeSafe response (round3-targeted-suites.log). Live typed resolution needs a real TYPE…
Repeated escalation and pending-reply pings stop after acknowledgement or delivery (upstream c33b3f6, 52e679b) ⏸️ untested no The prior payload did not establish a live result. Only the unit suites ran: tests/fm-inactive-reconcile.test.sh, tests/fm-daemon.test.sh and tests/fm-pending-reply.test.sh, all exit=0 (round3-escalat…
Captain hold lifecycle: the selected option's close mode applies, and a board later choice dates the hold (restored doc facts d) ⏸️ untested no The prior payload did not establish a live result. Only tests/fm-captain-hold-lifecycle.test.sh ran (exit=0, round3-targeted-suites.log), and it drives the public scripts in isolation. A live check ne…
Claude-bound record-backed doorbell over the Herdr backend (upstream 1fe1a67) ⏸️ untested no A Herdr lab session was not provisioned within the 30m budget, and only the tmux backend was driven live. Drive it with bin/fm-herdr-lab.sh using a named fm-lab-* session, following its prepare, provi…
  • bash $EVIDENCE/live-claude-draft-guard.sh $WORKTREE (real Claude Code worker on a private tmux socket plus a disposable fm-lab home, driving bin/fm-send.sh and fm_task_inbox_ring)
  • bin/fm-test-run.sh tests/fm-calm-pi-extension.test.sh at HEAD
  • git diff --quiet 72b63eee HEAD -- tests/fm-calm-pi-extension.test.sh and git diff --stat 72b63eee HEAD -- .pi/
  • bin/fm-test-run.sh tests/fm-calm-pi-extension.test.sh on a temporary git archive 72b63eee tree (removed afterwards)
  • bin/fm-dispatch-select.sh with a quota-balanced Grok Bot rule, and with an ordered rule plus a launch_failed fact
  • FM_HOME=<fm-lab-home without bridge> bin/fm-grok-bot-dispatch.sh brief.md --bot fm-researcher
  • bin/fm-test-run.sh tests/fm-pending-reply.test.sh tests/fm-send-inbox.test.sh tests/fm-dispatch-resolve.test.sh tests/fm-grok-bot-dispatch.test.sh tests/fm-captain-hold-lifecycle.test.sh
  • bin/fm-test-run.sh tests/fm-inactive-reconcile.test.sh tests/fm-daemon.test.sh

🔧 No changes applied.
✅ Re-checked - no issues remain.

  • Live validation: ✅ go - 4 of 6 scenarios driven live against the product
Scenario Result Live Evidence
On real Pi 0.87.1 with Calm on, pressing Escape on a queued Firstmate notification delivers it once in a new turn that shows 'Firstmate supervision continues in a new turn.' (unchanged upstream assert… ✅ pass live round5-calm-pi-extension-pi-0.87.1.log: 'ok - Pi 0.87.1 with Calm on keeps a queued Firstmate notification unlisted, out of the editor on Escape, and delivers it once in a new announced turn, while Ca…
The restored Pi Calm test is byte-identical to upstream 72b63ee and the whole suite passes on Pi 0.87.1 with no skips ✅ pass live round5-calm-pi-extension-pi-0.87.1.log header (blob 5fde9877 at both 72b63ee and HEAD), 15 ok / 0 not ok / 0 skip
Adversarial: fm-send to a real Claude worker whose input box already holds captain text does not type over it and does not get a merged submission; the message is stored durably and handled after the… ✅ pass live round5-live-claude-draft-guard.log: 'doorbell skipped (composer visibly holds pending text)', draft still in composer: yes, acted while draft present: no, after re-ring: acted=yes acked=yes, RESULT: P…
Pending replies to second mates resolve once and escalations are not repeated; Claude-bound operational input goes out as a record-backed doorbell ⏸️ untested no The earlier payload did not show a live result for this scenario. It rests only on repo suites (fm-pending-reply, fm-operational-input, fm-send-inbox, fm-inactive-reconcile, fm-wake-queue), which run…
Restored fork doc fact: a quota-balanced rule with a Grok Bot target is refused by fm-dispatch-select ✅ pass live round5-grok-bot-quota-balanced-refusal.txt: exit=2 with the quota-evidence refusal
Restored fork doc fact: typed resolution reports a Grok Bot target as eligible but unranked and never emits it as a profile: line ⏸️ untested no The earlier payload did not show a live result for this scenario. Typed resolution calls the external TypeSafe/Jev API, and the fm-dispatch-resolve suite stubs that response. Checking it live needs a…
  • FM_PI_PACKAGE_DIR=~/.linuxbrew/lib/node_modules/@earendil-works/pi-coding-agent bash bin/fm-test-run.sh tests/fm-calm-pi-extension.test.sh (real Pi 0.87.1; 15 ok, 0 not ok, 0 skip)
  • git rev-parse 72b63eee:tests/fm-calm-pi-extension.test.sh vs HEAD: (both 5fde9877...)
  • bash ~/.no-mistakes/evidence/01M3EMCMBKQC27KJDYRS6HBK68/live-claude-draft-guard.sh $PWD (real Claude Code worker on a private tmux socket, disposable lab FM_HOME, real bin/fm-send.sh + fm_task_inbox_ring)
  • bash bin/fm-test-run.sh tests/fm-pending-reply.test.sh tests/fm-operational-input.test.sh tests/fm-send-inbox.test.sh tests/fm-inactive-reconcile.test.sh tests/fm-wake-queue.test.sh
  • bash bin/fm-test-run.sh tests/fm-dispatch-resolve.test.sh tests/fm-grok-bot-dispatch.test.sh
  • bin/fm-dispatch-select.sh qb.json 0 with a quota-balanced rule that has a Grok Bot target

✅ No issues found.

  • Live validation: ✅ go - 6 of 8 scenarios driven live against the product
Scenario Result Live Evidence
Pi Calm announces 'Firstmate supervision continues in a new turn.' after Escape, and delivers the queued Firstmate notification once (unchanged upstream test, real Pi 0.87.1) ✅ pass live round6-calm-pi-extension-pi-0.87.1-full.log: 'ok - Pi 0.87.1 with Calm on keeps a queued Firstmate notification unlisted, out of the editor on Escape, and delivers it once in a new announced turn'; te…
The whole restored upstream Pi Calm suite runs without skipping or weakening any assertion on Pi 0.87.1 ✅ pass live round6-calm-pi-extension-pi-0.87.1-full.log: 15 ok, gate_skip=false, exit=0
A running Pi 0.87.1 session exposes every queue-retention member Calm preflights ✅ pass live round6-calm-pi-pi-0.87.1.log: 'ok - Pi 0.87.1 exposes every queue-retention member Calm preflights'
Adversarial: a steer to a Claude worker whose box already holds captain draft text is not typed over the draft, is recorded durably, and is acted on and acknowledged after the draft clears and the wat… ✅ pass live round6-live-claude-draft-guard.log: 'doorbell skipped (composer visibly holds pending text)', draft still in composer, worker did not act; after re-ring acted=yes acked=yes (6s); RESULT: PASS
Adversarial: fm-dispatch-select refuses a quota-balanced rule that uses a Grok Bot target ✅ pass live round6-grok-bot-doc-facts-cli.txt: 'error: config: a quota-balanced rule cannot use a grok_bot profile' exit=2
In a home without the bridge (secondmate), fm-grok-bot-dispatch exits 2 and names the primary; a firstmate-supplied launch_failed or blocked fact makes ordered selection fall to the next candidate ✅ pass live round6-grok-bot-doc-facts-cli.txt: bridge-not-found exit=2; skipped[0] launch_failed / blocked, selected[1] claude sonnet
Second mates are not re-pinged repeatedly for the same pending reply, and the steering inbox, operational input, and inactive reconcile behave as upstream intends ⏸️ untested no The prior payload did not establish a live result. Only the fm-pending-reply, fm-send-inbox, fm-operational-input, and fm-inactive-reconcile suites passed (round6-targeted-suites.log). Driving repeate…
Captain hold board 'later' dates the open hold, 'done' and 'release' keep their modes, and undated, invalid, past, or non-later deferrals are refused (restored doc facts); typed resolution reports a G… ⏸️ untested no The prior payload did not establish a live result. Only the fm-captain-hold-lifecycle, fm-dispatch-resolve, and fm-grok-bot-dispatch suites passed (round6-targeted-suites.log). Typed resolution needs…
  • FM_PI_PACKAGE_DIR=~/.linuxbrew/lib/node_modules/@earendil-works/pi-coding-agent bin/fm-test-run.sh tests/fm-calm-pi-extension.test.sh (unchanged upstream test, real Pi 0.87.1, 15/15 ok, gate_skip=false)
  • bin/fm-test-run.sh tests/fm-calm-pi-extension.test.sh tests/fm-calm-pi-queue-retention-live-e2e.test.sh (default npm-root lookup: the extension suite gate-skipped the package-dependent cases because Pi lives under linuxbrew, not under nvm's npm root; the live queue-retention guard passed against the running Pi 0.87.1)
  • bash live-claude-draft-guard.sh &lt;worktree&gt;: real Claude Code 2.1.283 worker on a private tmux socket with a disposable fm-lab-home, steered through real bin/fm-send.sh while its composer held a draft, then re-rung by fm_task_inbox_ring
  • bin/fm-dispatch-select.sh qb.json 0 (quota-balanced rule with a Grok Bot target)
  • FM_HOME=<lab home without bridge> bin/fm-grok-bot-dispatch.sh brief.md --bot fm-researcher
  • bin/fm-dispatch-select.sh ord.json 0 --facts facts.json with launch_failed and with blocked facts
  • bin/fm-test-run.sh tests/fm-pending-reply.test.sh tests/fm-send-inbox.test.sh tests/fm-operational-input.test.sh tests/fm-inactive-reconcile.test.sh tests/fm-captain-hold-lifecycle.test.sh tests/fm-grok-bot-dispatch.test.sh tests/fm-dispatch-resolve.test.sh (7/7 exit 0)
  • git rev-parse HEAD:tests/fm-calm-pi-extension.test.sh vs 72b63eee: (identical blob 5fde9877) and git diff --quiet 72b63eee HEAD -- .pi
✅ **Document** - passed

✅ No issues found.

✅ No issues found.

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ No issues found.

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

✅ No issues found.

✅ No issues found.

tiago-peixoto and others added 30 commits September 24, 2026 15:22
…ning (kunchenguid#5566)

* fix(bin): report a Lavish source armed only after its listener is running

Registration alone was treated as ready, so arm could succeed before anything was collecting from the board.

* no-mistakes(review): Guard Lavish arm launches, keep retire refusals, report live prior listener

* no-mistakes(review): Keep polling through window before reporting a still-live prior listener

* test: wait for a capture's claim to drop before the next arm

The result is stored before the runner exits, so a re-arm in that gap was meeting a live claim.

* no-mistakes(document): Record Lavish arm readiness evidence in verification doc

* no-mistakes(ci): Both failures were caused by this PR, and both are fixed with test-only edits. Lint 2 (ShellCheck SC2034): this branch removed the only use of `reply_id` (a `start "$reply_id"` call) from tests/fm-procevent.test.sh, which left the assignment at line 1450 unused. I deleted that assignment. It was the only `reply_id` in the file. ShellCheck is now clean on both test files. Behavior portable serial 4: the failing test was tests/fm-bearings-board.test.sh, in the check "registration consumed its answer before the any-origin binding existed". I reproduced it locally: the hold was still `state: queued` when the test checked it. - What must hold: the test's check that the hold is closed must run after the listener has captured the answer. - Why it broke: the test used a stand-in adapter that ran `fm-procevent.sh start` in the foreground after `arm`, so capture finished before build returned. On this branch, `arm` starts the listener itself in the background, so the real listener captures the answer and closes the hold a moment after build returns. - Fix: removed the now-redundant stand-in adapter, the copied runtime directory, and its extra environment variables. The test now runs the real build through the existing `run_board` helper and waits up to about 10s for the hold to reach `state: done`. The checks that follow are unchanged: `Resolution mode: answered` and the any-origin binding. - Other tests: this was the only test in the file that stood in for the adapter this way. The shard's other pure-contract-unit test (tests/fm-trace-context-lib.test.sh) passed unchanged. Verification: - tests/fm-bearings-board.test.sh passed 3 times in a row via bin/fm-test-run.sh, all 18 checks, about 53s per run. - tests/fm-procevent.test.sh was not rerun, because the lint fix only removed an unused assignment
…elivered (kunchenguid#5599)

* fix(bin): acknowledge a delivered unknown-wake escalation

The same unrecognized wake was escalated again after it had already been handled, because delivery never recorded that identity.

* no-mistakes(review): Scope unknown-wake acknowledgements to one away session

* no-mistakes(review): Clear delivered digest when unknown-wake ack write fails

* no-mistakes(review): Limit unknown-wake suppression to acknowledged lines

* no-mistakes(document): List unknown-wake ack file among away-session artifacts
…ess wait (kunchenguid#5587)

* fix(bin): keep a stated default retraction from cancelling a keyless wait

A resolved line that names the shared default decision bucket was closing the keyless live wait that only prints as that same key. Keyless self-retraction still closes the keyless wait.

* no-mistakes(review): Keep declared waits standing past foreign-key resolved lines

* no-mistakes(review): Bound declared-wait read and share one decision-key parser

* no-mistakes(document): Document supervisors' key-aware declared-wait read
…unchenguid#5544)

* fix(bin): terminate a remote job worker that lost ownership when it receives TERM

A serving worker whose lock directory is gone can no longer quarantine shutdown, and resuming service publishes a false ready heartbeat. Exit after stopping only that worker's own command tree, without removing a replacement owner's lock.

* no-mistakes(review): Check worker lock ownership before publishing shutdown quarantine

* no-mistakes(document): Correct worker shutdown comment on replacement-owned lock

* fix(bin): keep an ousted remote job worker off the replacement quarantine

Shutdown can lose the lock after the first ownership check and before it
writes or clears quarantine. Bind both operations to the directory object
this process still owns so a replacement's quarantine stays untouched.

* no-mistakes(review): Make ousted-worker shutdown test reliably reach quarantine clear

* no-mistakes(document): Reattach worker_shutdown doc comment to its function

* no-mistakes(ci): Fixed the failing check (Behavior portable serial 7) with a test-only change to the stall test in tests/fm-remote-job.test.sh. Product code is unchanged; no other test changed. Cause: after the decoy dies, both workers run the same check-exists, read, delete sequence on the job records. On the CI runner the replacement deleted a record between the ousted worker's check and its read. The ousted worker exited 125, and because the file runs under set -e the unguarded `wait` ended the test with 125. The exit trap then killed the replacement, which produced the "Killed" line. Reproduction: a temporary 0.3 s delay between the check and the read, applied to the ousted worker only, made the committed test fail exactly as in CI (exit 125 and the "Killed" line). The new test passed with the same delay. The delay is reverted, along with a similar debug hook that the timed-out attempt had left in bin/fm-remote-job-worker.sh. Test changes: - The replacement is frozen (and confirmed stopped) before the decoy is killed and resumed only after the ousted worker exits, so only one worker touches the job records at a time. - The ousted worker is stopped only once its quarantine exists and its lane is reaped, which places it inside its stop loop. - Every fixed poll loop is now a wait on a named condition with a 30 s deadline and an explicit failure message. Exit detection also handles zombies. - The exit trap kills and waits for the decoy and both workers on every path. - A non-zero exit from the ousted worker now fails with its exit code and stderr instead of silently ending the file. The test still proves that the resumed ousted worker exits 0 and leaves the replacement's lock, quarantine contents and quarantine inode unchanged. Verification: the full test file passed four times on its own and three times under nice -n 10 with four busy-loop CPU hogs; bin/fm-lint.sh passes. Changes are not committed

* no-mistakes(ci): I fixed the failing check (Behavior portable serial 7) by changing only the stall test in tests/fm-remote-job.test.sh. Product code is unchanged. **What failed:** "an ousted worker in shutdown leaves the replacement quarantine untouched" failed on CI with the ousted worker exiting 125 ("could not stop the active command tree"). **Why:** during shutdown, the worker retries the still-running decoy command group a fixed 100 times, 0.01 s apart, then gives up and exits 125. The test tried to freeze the worker partway through those retries by sending SIGSTOP from outside. On a slow runner the retries ran out before the stop arrived, so the worker had already given up. The invariant is that the test must hold the ousted worker inside that retry loop until the replacement owns the lock. That was the only place the test depended on timing. The other waits already watch for a named state change with a 30 s deadline. **Fix:** - The ousted worker now starts with a small `sleep` wrapper at the front of its PATH, and the SIGSTOP race is gone. - The wrapper only holds a `sleep` called directly by that worker's own process (it checks its parent pid against a hold file) while its quarantine file exists. - The only such `sleep` is the first retry in the shutdown stop loop, so the worker waits there as long as needed. - The wrapper writes a marker when it starts holding. The test waits for that marker, then hands the lock to the replacement, freezes the replacement, and kills the decoy. - The test releases the worker by deleting the hold file. Deleting the whole temp directory also releases it, so a failed run cannot leave the wrapper looping. - A process leak: the test overwrites the job's command-group record with the decoy, so no worker ever stopped the job's real command. `fm-hold-job.sh` and its `sleep 30` stayed running for up to 30 s after the test. The test now records that group before overwriting it and kills it at the end of the test and in the exit cleanup. - The test still asserts the same things: the ousted worker exits 0, and the replacement's lock, quarantine contents and quarantine inode are unchanged. **Verification:** - The full file passed twice on its own, twice under `nice -n 10` with six busy-loop CPU hogs, and twice more after the leak fix. - `pgrep` found no leftover processes afterwards. - With the worker from just before the fix commit (cf45cb6^), the test still fails with "the ousted worker wrote or cleared the replacement quarantine during shutdown", so it still proves the fix. - `bin/fm-lint.sh` passes. - I did not reproduce the CI failure locally. The cause comes from the fixed retry limit and the CI error message. The changes are not committed

* no-mistakes(ci): I changed only the stall test ("an ousted worker in shutdown leaves the replacement quarantine untouched") in tests/fm-remote-job.test.sh. Product code is unchanged, and so is every other test. **Invariant:** the pid written to the job's group record must be a process-group leader whose group dies when that one process is killed. Otherwise the worker's bounded stop loop never sees the group die, gives up, and exits 125 ("could not stop the active command tree") before it reaches the lost-ownership exit. The decoy is the only place in this test that depends on this. **Fix:** - The decoy used to be `set -m; sleep 30 &`. It now starts as `perl -MPOSIX=setsid -e 'setsid() >= 0 or exit 1; exec @argv' sleep 30 &`, which gets its own session and group without shell job control. tests/fm-procevent.test.sh already uses the same idiom. - The test now waits, with the file's usual 30 s deadline and a named failure, until `ps -o pgid=` of the decoy equals its pid before writing it into the group record. This way the worker can never read the record before `setsid` has run. - The existing steps are unchanged: the test kills the decoy, reaps it with `wait` before releasing the hold file, and the exit trap still kills and reaps the decoy and both workers. - The assertions are unchanged: the ousted worker exits 0, and the replacement's lock pid, quarantine text and quarantine inode stay the same. **Cleanup:** I reverted a debug `printf` hook that the timed-out previous attempt had left in bin/fm-remote-job-worker.sh, and deleted its untracked `.tmp-repro/` directory. Neither was committed. **Verification:** - The full tests/fm-remote-job.test.sh passed twice normally and once under `setsid -w` with stdin from /dev/null (no controlling terminal). - `bin/fm-lint.sh` passes. - No leftover `sleep 30` processes afterwards. **Not reproduced:** I could not reproduce the CI failure locally. On this host `set -m` made the decoy its own group leader even without a controlling terminal, so the cause on the runner is not confirmed. The change removes the test's reliance on shell job control, as the user asked. Changes are not committed

* no-mistakes(ci): I changed only the stall test ("an ousted worker in shutdown leaves the replacement quarantine untouched") in tests/fm-remote-job.test.sh. Product code is unchanged, and so is every other test. **Invariant:** the group record the ousted worker checks in its stop loop must stay the job's own command group, and the test must stop that group before it releases the hold. Otherwise the bounded retry keeps seeing a live group, gives up, and exits 125 ("could not stop the active command tree") before it reaches the lost-ownership exit. The test overwrote this record in one place (the decoy) and stopped the group in one place (killing the decoy); both are changed. **Fix:** - I removed the setsid decoy and the overwrite of `.claim/group`. The record keeps the job's real command group, which the test still saves as `STALL_JOB_GROUP`. - The two-line `group_start` stays. It is still needed: without it the worker kills the real group on its first pass, before the replacement takes over, so the hold would never matter. - The `sleep` wrapper that holds the worker at its first stop-loop retry is unchanged. - After the replacement owns the lock, its quarantine is planted and it is frozen, the test runs `kill -KILL -- -$STALL_JOB_GROUP`. It then waits, with the file's usual 30 s deadline and a named failure, until `kill -0` on the group fails. Only then does it remove the hold file. The worker therefore always sees its own command already stopped and never races its retry budget. - The exit trap still kills the saved command group if the test fails. It can't `wait` on that group because the group is not a child of the test shell. The decoy variable and its cleanup entry are gone. - The assertions are unchanged: the ousted worker exits 0, and the replacement's lock pid, quarantine text and quarantine inode stay the same. **Verification:** - The full tests/fm-remote-job.test.sh passed twice normally. - It passed once under `setsid -w` with stdin from /dev/null (no controlling terminal). - It passed once under `nice -n 10` with six busy-loop CPU hogs. - With the worker from before the fix (cf45cb6^), the test still fails with "the ousted worker wrote or cleared the replacement quarantine during shutdown", so it still proves the fix. - No `fm-hold-job` or `sleep 30` processes were left afterwards. - `bin/fm-lint.sh` passes. **Not reproduced:** I couldn't reproduce the CI failure locally; the decoy version also passed on this host. So I can't confirm why the decoy group stayed alive on the runner. The new wait turns any leftover live group into a clear named failure instead of an exit 125. The changes are not committed

* fix(bin): keep a dead command group dead on bash 5.2

A bare return inside the liveness check drops the failing kill status when the check runs in a conditional, so shutdown keeps treating a stopped group as alive and exits 125.

* no-mistakes(review): Use bash 3.2 fd syntax and fix trap return comments
…henguid#5589)

* docs: make configuration settings easier to find and understand

* no-mistakes(review): Restore dropped qualifiers and fix misplaced config doc labels

* no-mistakes(review): Restore three dropped qualifiers in configuration reference
)

* fix: bound worker edits of project AGENTS.md/CLAUDE.md to factual corrections

These files are loaded into every agent session of a project, so additions
should be a deliberate human choice rather than automated task output. The
ship brief's project-memory section and AGENTS.md section 6 previously invited
workers to record durable knowledge, which let project AGENTS.md files accrete
detail the codebase or README already carries. Workers now edit only to fix
factually wrong content - including content their own change made wrong - and
fm-ensure-agents-md.sh runs only alongside such a correction. Stow no longer
routes project-memory additions through ship tasks, and the generated skeleton
no longer invites discovery-driven additions.

* no-mistakes(review): Stop running fm-ensure-agents-md.sh on memory-file corrections

* no-mistakes(document): Clarify manual project-memory initialization and remove duplicate guidance
…enguid#5635)

* fix(bin): let gate agents drive lifecycle against marked lab homes

Part 2 of the kunchenguid#5615 split. A no-mistakes gate agent runs inside a
checkout carrying the fleet-captain identity, so fm-gate-refuse-lib
refuses fleet mutation on the gate signal. That refusal was absolute,
which kept gate validation from ever exercising the real lifecycle.

Stamp a disposable lab FM_HOME with a .fm-lab-home marker file that
only bin/fm-lab-home.sh writes, and only onto a fresh empty dir, so
no call path can mark a populated real home. fm_refuse_if_gate_agent
then permits lifecycle only when FM_HOME carries the marker and is
driven through its stock layout - any FM_*_OVERRIDE relocation stays
refused so part of the "lab" cannot be split back onto the real fleet.
The threat model is a confused agent touching the real fleet, not
deliberate forgery, so the marker is a plain token file rather than a
bound record. FM_GATE_REFUSE_BYPASS is unchanged: it still serves the
test harness, which cannot mark hundreds of temp homes.

Teardown's slot-ownership scan compared state-dir paths textually
while fm_firstmate_root_home canonicalizes, so a lab home under a
symlinked TMPDIR scanned its own record twice and self-collided;
compare file identity (-ef) instead.

* no-mistakes(review): Refuse unlistable lab homes and hardlinked slot records

* no-mistakes(review): Mint lab markers only on verified-empty fresh dirs

* no-mistakes(document): Clarify lab-home gate documentation and comment contracts

* no-mistakes(document): Clarify lab-home gate documentation and remove stale claims

* no-mistakes(document): Clarify gate lab-home documentation and boundary wording
…ad of refusing every re-arm (kunchenguid#5594)

* fix(bin): replace a watcher whose beacon stalls past a hard bound instead of refusing every re-arm

A fleet watcher that is alive but whose liveness beacon has gone stale could
never be replaced: every re-arm was refused because the lock holder was a live
pid, and the holder was never evicted because it was not dead. Add
FM_WATCHER_STALL_BOUND (default 3x the stale grace): below it the refusal is
unchanged; at or past it the arm re-verifies the holder against the lock's
recorded identity, sends TERM, waits boundedly, and takes the lock the normal
way, ledgering a stalled-holder-replaced row. A holder that survives TERM keeps
the old refusal.

Fixes kunchenguid#4400

* no-mistakes(test): poll for replacement message to fix watcher-lock test flake

* no-mistakes(document): document FM_WATCHER_STALL_BOUND in config inventory
…n can keep them (kunchenguid#5563)

* fix(pi): hide queued Firstmate notifications under Calm only when the session can keep them

Calm now keeps authenticated Firstmate operational inputs out of Pi's queued-message
listing, but only after proving the live session exposes every member needed to keep
them across Escape. A session missing any of them keeps stock rows and Escape and shows
one generic warning. Escape and the dequeue key return only captain-authored messages to
the editor and re-queue hidden notifications in order; after an abort that kept any in
Pi's agent queue, the adapter starts the delivery turn itself because Pi 0.87.1 does not
continue an aborted run. Compaction-held notifications stay with Pi's compaction flush and
never start or announce a turn.

Fixes kunchenguid#1588

* docs(calm): record Pi 0.87.1 queued-row retention verification

* no-mistakes(review): Deliver kept Calm notifications after tree-navigation aborts too

* no-mistakes(review): Defer Calm notification turn until tree navigation finishes

* no-mistakes(lint): Silence SC2016 for literal JavaScript in queue-retention e2e test
…uid#5548)

* fix(bin): refuse teardown when a required source disappears

A missing sibling was sourced after cleanup had started, so Bash 3.2
exited 0 from the EXIT trap and Bash 5 continued and reported success.

* no-mistakes(review): Remove unused FM_TEST_ONLY hook from teardown tests

* no-mistakes(review): Check task backend sources before any teardown cleanup

* test(gotmp): give teardown fixtures every tmux adapter sibling

Teardown now refuses when a sibling the recorded backend's adapter sources
is missing, so the fake bin must carry fm-session-lock-lib.sh,
fm-agent-process-lib.sh and fm-gemini-lib.sh.
Restructure the supervision host doc's prose into shorter sections, lists,
and tables without changing documented behavior. Every original heading,
anchor, identifier, number, quoted string, and link target is preserved.
* docs: make herdr-backend easier to read

Restructure the Herdr backend doc's prose into shorter sections, lists, numbered procedures, and tables without changing documented behavior. Every original heading, anchor, fenced code block, link target, and documented fact is kept.

* no-mistakes(document): Restore composer-proof reason and complete Herdr topic table
Restructure the prose into sections, lists, and tables without changing
documented behavior. Every original heading and anchor, inline-code span,
link target, number, and quoted string is kept, and each sentence sits on
its own line. Adds a topic navigation table and short subsections under
the existing headings.
* docs: make watcher-continuity easier to read

Restructure the prose into sections, lists, and tables without changing documented behavior. Every original heading, anchor, identifier, link target, and number is kept.

* no-mistakes(review): Fix actor and supervision-host scope in watcher-continuity doc

* no-mistakes(review): Make readiness TERM and retry conditional on unready successor
* docs: make sessionstart-nudge easier to read

Restructure the prose into sections, lists, and tables without changing documented behavior. Every original heading, inline-code span, link target, number, and fact is preserved, and a harness-to-tier table now sits near the top.

* no-mistakes(review): Drop helm glossary line and dedupe exit-code lead-in
* docs: make captain-hold-lifecycle easier to read

Restructure the captain-hold lifecycle prose into sections, lists, and tables without changing documented behavior. Every original heading, anchor, identifier, number, quoted string, and link target is kept.

* no-mistakes(review): Fix verification record subjects and grouping headings

* no-mistakes(review): Clarify task-body read-back cases belong to the suite
* docs: make remote-secondmates easier to read

Restructure the remote second mates prose into sections, lists, numbered procedures, and tables without changing documented behavior. Every original heading, anchor, fenced code block, identifier, link target, and qualifier is preserved.

* no-mistakes(review): Merge remote-home table cell into one sentence

* no-mistakes(review): Tighten readiness lead-in, restore causal link, fix dangling reference
…nguid#5554)

* fix(bin): bound the away digest and log why a delivery failed

The away daemon joined every buffered escalation into one unbounded
digest. A start-up catch-all span can exceed what one transport argument
carries (tmux rejects the send-keys command; Linux refuses to exec any
argument above 131,071 bytes, which is how herdr receives it), so the
initial send failed on every housekeeping pass and was logged as an
unconfirmed Enter with text possibly in the composer.

escalate_flush now builds the injected digest under a fixed byte budget:
each event is cut at a UTF-8 boundary with an omitted-bytes marker, the
joined events stop with a "+K more event(s)" tail, and a bounded digest
names a state/.subsuper-digests/ file that keeps every buffered event
verbatim. The buffer itself is untouched, so the return catch-up stays
complete.

The tmux submit core and the herdr literal send now replay the
transport's stderr on failure, and inject_msg logs the failing stage
(initial send versus Enter confirmation) with the byte count and that
stderr. The wedge alarm line and marker carry the last failure reason.

Fixes kunchenguid#4382

* no-mistakes(review): Drop digest pruning; label send-failed as send-or-Enter stage

* no-mistakes(review): Keep digest full text once submit ran; reuse on retry

* no-mistakes(lint): Count digest files with find instead of ls

---------

Co-authored-by: firstmate-oss <firstmate@kunchenguid.local>
…henguid#5638)

* feat(tests): add FM_TEST_SEAM launch seam and gate lab-primary recipe

Part 1 of the kunchenguid#5615 split: the pieces that let the no-mistakes pipeline
live-validate firstmate changes, without the gate-refusal rescoping.

- bin/fm-afk-launch.sh: FM_TEST_HARNESS pins the detected harness only
  alongside the FM_TEST_SEAM=1 marker test suites set, so a leaked variable
  in a real primary's environment stays inert and unknown tokens fall
  through to real detection.
- tests/lib.sh: export FM_TEST_SEAM=1 for every suite.
- .no-mistakes.yaml: per-harness recipe for running a real fixture primary
  from a gate run - a plain mktemp lab FM_HOME on a private tmux socket,
  with FM_GATE_REFUSE_BYPASS=1 scoped to it and NO_MISTAKES_GATE scrubbed.
- tests/fm-wake-queue.test.sh: stop the owned watcher fixture with KILL and
  clear its lifecycle state so the next leg starts clean; TERM could leave
  bash waiting in a child on some runners.
- tests/fm-remote-secondmate-lifecycle-e2e.test.sh: wait for the liveness
  lock holder's post-acquire marker instead of the lock dir, which is
  published before the claim finishes.

* no-mistakes(review): Scrub lab home overrides and require FM_TEST_SEAM separately

* no-mistakes(document): Clarify test seam and disposable lab bypass documentation

* no-mistakes(document): Clarify lab isolation and test-seam documentation

* no-mistakes(ci): Fixed the CI failure: test cleanup killed the remote worker child but left its supervisor able to restart it during fixture removal. Cleanup now stops the worker tree. The lifecycle test passed locally; ShellCheck and diff checks passed
… home is gone (kunchenguid#5552)

* fix(bin): refuse watchers from disposable checkouts and exit when the home is gone

Fixes kunchenguid#321
Fixes kunchenguid#4760

A watcher armed from a disposable no-mistakes validation checkout under
.no-mistakes/worktrees/ outlived the validation step and kept writing the
real home's state, and a running watcher never noticed when its home,
state directory, or code root disappeared. The arm now refuses from such
a checkout with the typed failure line, the watcher checks once per poll
that its home, state directory (or its own lock holder record), and bin
directory still exist and exits with a logged reason scoped to itself,
and the shared test helpers reap every watcher a suite armed for a
temporary home through the home-scoped stop.

* no-mistakes(lint): fix SC1007 by assigning empty string in watch-arm test

* no-mistakes(ci): Found and fixed a genuine, reproducible hang introduced by this branch's test-watcher reaper, which is what killed both CI checks (serial-2 cancelled at the 30-min cap; Lint 2 exit 143 = the suite's own TERM-trap code). Root cause: test_drain_asserts_watcher_liveness (tests/fm-wake-queue.test.sh) fabricates a .watch.lock whose pid is the test runner's own $$ with the runner's real identity, to make the drain believe a live watcher exists. The new make_case tracking registers that state dir for reaping, so at fm_test_cleanup the new fm_test_reap_watchers drives fm-watch-arm.sh --stop; its identity check matches (the fixture recorded the runner's identity) and it kill -TERMs the test runner. tests/lib.sh:231 is `trap 'fm_test_cleanup; exit 143' TERM`, so the TERM re-enters cleanup -> reap -> kills $$ again -> infinite loop until the runner cap. I reproduced this locally: the suite ran all tests then looped forever in cleanup spawning fm-watch-arm.sh --stop against a lock naming its own PID. Fix (tests/lib.sh, +5 lines): in fm_test_reap_watchers, skip any tracked lock whose pid equals our own $$ before driving --stop. This is the single shared reap boundary; seven $$-self-lock fixtures across four test files are all covered by the one guard, and real armed watchers (pid != $$) are still reaped. Invariant: the test reaper must only signal real armed watcher processes, never the test runner itself. Verified locally: tests/fm-wake-queue.test.sh -> EXIT 0 (63 ok, no hang); tests/fm-watch-arm.test.sh -> EXIT 0 (21 ok, including test_reaper_stops_a_tracked_watcher, confirming the guard does not over-skip). Lint 2's exit 143 was the same shard/cap signature; a fresh CI run on this new commit will re-evaluate it

---------

Co-authored-by: firstmate-oss <firstmate@kunchenguid.local>
…m as silence (kunchenguid#5588)

* fix(bin): surface an unrecognized status prefix instead of dropping it

A parked or holding declaration, and a verb whose correlation token did not parse, never became an event, so the supervisor still saw the earlier line.

* no-mistakes(review): Require verb-shaped unrecognized status prefixes, add continuation tests

* no-mistakes(document): Document unrecognized status prefix escalation in afk skill

* no-mistakes(ci): I reproduced the "Behavior portable serial 6" failure locally and fixed it by changing the test data in one test. No product code changed. **What failed:** `tests/fm-session-start.test.sh`, in `test_orphan_status_logs_are_printed`, with "matched status log was printed 2 times". **Why:** the test writes status lines with made-up prefixes, `matched: surfaced once` and `orphan: step N`. The test only uses them as placeholder text. It checks that the session-start digest prints each task's status tail exactly once. This PR (kunchenguid#4763) deliberately makes an unrecognized one-word lowercase prefix a status event. So those lines now surface as captain-relevant events, and the wake queue's STATUS OUTCOME BACKSTOP section prints them a second time. The code under review is behaving as the issue asks. Only the test's placeholder data had become meaningful. **Rule the test depends on:** its status lines must not be captain-relevant, so the digest is the only place they are printed. Both lines in this test broke that rule. The orphan line would have failed the same count check right after the matched line did. **Fix:** in that test only, I switched both lines to the recognized, non-captain verb `working:`: `working: surfaced once` and `working: orphan step 1..6`. I updated the matching assertions and counts to use the new text. What the test checks is unchanged: orphan logs are labelled, the tail is bounded, the log path is printed, and each tail appears once. **Verification:** before the fix, the test failed locally the same way as in CI. After it, `bash tests/fm-session-start.test.sh` reports "all assertions passed

* no-mistakes(review): Detect unrecognized prefixes on unstamped lines; share verb list

---------

Co-authored-by: Kun's firstmate <kunchenguid+firstmate@users.noreply.github.com>
…henguid#5658)

Fixes kunchenguid#5295

Session start now reports a remote inheritance failure using the
push's own error line instead of the first unchanged item that
happened to print before it, and the shared captain preferences
header check now names the first required phrase it did not find,
on both the local and remote inheritance paths.
…yloads (kunchenguid#5657)

* fix(bin): stand down the Claude Stop auto-arm on pi-code-delivered payloads

pi-code loads the tracked Claude settings but has no asyncRewake, so it
awaits every Stop hook; without a stand-down the auto-arm runs
synchronously inside Pi's turn end and holds it open for the declared
multi-hour timeout. Stand down when the payload's transcript_path
contains a /.pi/ path component, the same discriminator the closed-but-
unmerged fix in kunchenguid#3352 used, with an explicit string-type check on the
jq filter.

Fixes kunchenguid#3343

* no-mistakes(document): document pi-code stand-down in harness integrations reference
…kunchenguid#5659)

* fix(bin): match whole multi-word project names in the registry lookup

bin/fm-project-mode.sh matched a registered project name against only the
first whitespace-delimited token of a registry row, so a name containing a
space never matched, silently defaulting the project to no-mistakes off
instead of its declared posture.

The lookup now matches the whole registered name against the raw line text,
so a name is compared literally (never as a regex) and a name that is a
leading prefix of another registered name still resolves to its own row.

* no-mistakes(document): docs already accurate for multiword registry name match

* chore: drop accidental empty err file

Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
…d#5546)

* fix(bin): classify the stdin program of `bash -s` with operands in the arm policy

With -s, sh/bash/zsh read the program from stdin even when operands follow;
the operands are only positional parameters. The arm policy treated the first
operand as a script path, so heredoc and here-string payloads were never
classified and a hidden bin/fm-watch.sh execution was allowed.

A protected path in the operand position still fails closed as before.

Fixes kunchenguid#1489

* no-mistakes(document): Clarify stdin shell operand documentation

* no-mistakes(ci): Captain, fixed `shellInvocation` so `bash -- -s` treats `-s` as a script name, and updated R21 to test the exact command. The targeted policy suite, lint, documentation check, and diff check pass. Both hosted workflows show `action_required` before any jobs ran; that external approval state remains unresolved

* fix(bin): keep main's handling of words after a leading `--`

Revert the pipeline CI-step change that made the first word after a leading
`--` always a script. It turned forms that main denies today into allow
(for example `bash -- -c 'bin/fm-watch.sh'`), which is outside kunchenguid#1489 and
loosens a fail-closed policy. `--` after `-s` still ends option parsing.
…nchenguid#5695)

* fix(bin): strip AI co-author trailers from fleet-launched commits

Cursor and other non-Claude runtimes append the trailer after the typed
message. A per-task commit-msg hook removes it and leaves human co-authors
and the author identity untouched.

* no-mistakes(review): Export pane hooksPath override and drop generated-with stripping

* no-mistakes(ci): This PR caused all three CI failures, and the fix is test-only: 4 test files change, no product code. **Cause.** `fm-spawn.sh` now installs the AI-trailer strip hooks for every spawn, secondmates included. The installer refuses a worktree that is not a git repository, and the PR deliberately keeps that fail-closed rule because real secondmate homes are firstmate clones. Four test fixtures still gave secondmates a plain directory as their home, so each spawn failed with "not a git worktree ... could not install the AI-trailer strip hooks": - serial 5: `tests/fm-backlog-atomicity.test.sh` ("secondmate spawn failed"). - serial 8: `tests/fm-secondmate-harness.test.sh` ("split: no meta written"). - Herdr: `tests/fm-backend-herdr-launcher-workspace-e2e.test.sh` and `tests/fm-backend-herdr-workspace-per-home-e2e.test.sh`. **Rule that must hold.** Every home a test spawns as a secondmate must be a git worktree. I checked the other places in the changed area: the only secondmate spawns in these tests are the ones listed. The earlier rounds already fixed the other fixtures (`fm-secondmate-liveness`, `fm-secondmate-safety`) the same way. **Fix.** - Each of those four secondmate homes now gets the same `.gitignore` plus `git init -q -b main` that the liveness and safety tests already use. - The two Herdr tests clean up with their own plain `rm -rf "$TMP_ROOT"`, not the shared `tests/lib.sh` helper. Because the installer leaves each `state/<id>.git-hooks` directory read-only, that cleanup printed "Permission denied" and left the directories behind. Both cleanups now restore the owner's write bit on every directory before removing (`find ... -exec chmod u+rwx`), which is what `fm_test_remove_tree` in `tests/lib.sh` does. **Verification.** - `tests/fm-secondmate-harness.test.sh` passes. - `tests/fm-backlog-atomicity.test.sh` passes (99 ok, exit 0). - shellcheck is clean on all four files. - I could not run the two real-Herdr tests locally: the Herdr lab on this host refuses to start because it needs exactly one running default session, and I did not change the host's Herdr state to get around that. Instead I checked their two changed steps directly: the installer succeeds on a home set up the new way, and the new cleanup removes the read-only hooks directory completely. Those two tests will only be proven on CI
…id#5683)

* fix(bin): treat Pi's dollar-first cost footer as furniture

An idle Pi status row opening with $0.000 was read as a dead-shell prompt, so exit and relaunch refused on an empty composer.

* test: wait for the draining holder to exec sleep before reading its identity

The procevent drain fixture read fm_pid_identity immediately after
backgrounding setsid sleep, racing the child's exec chain. Mid-exec the
cmdline can read empty, failing the fixture on a loaded CI runner.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…5534)

* fix(bin): refuse a merge when a required check never reported

fm-pr-merge.sh built its GitHub refusals only from checks present in
statusCheckRollup, so a required check that never ran was simply absent and
the merge proceeded on the subset that reported, contradicting its own
"every required check green" claim.

The GitHub verify now reads the base branch's required contexts from the
forge itself - the classic branch protection summary on
GET repos/{o}/{r}/branches/{b} and the active ruleset rules on
GET repos/{o}/{r}/rules/branches/{b} - and refuses when a required context
has no entry in the same rollup, at the same head, that the merge is bound
to. Absence reads as unknown, never green. The required-set read joins the
existing refusal list, so a draft, a red check, and an unreported required
check are all reported together.

Could not read vs nothing required: both endpoints need only repository
read access. The admin-only GET .../branches/{b}/protection endpoint is
deliberately not used: it answers a non-admin token with the same 404 an
unprotected branch gets (observed live on kunchenguid/firstmate main with
this token), which would read a missing permission as "nothing required".
Any failed or malformed read of either source (auth, missing fine-grained
permission, rate limit, network, 404, unexpected shape) refuses the merge
with a line naming the unreadable source. The one exception is GitHub's
plan-gated 403 on the rules endpoint ("Upgrade to GitHub Pro or make this
repository public"), which already means "this repository has no branch
rules" for the merge-queue reader; that check moves into one shared helper
and the classic summary still decides for such a repository.

Attended waiver: --allow-missing <check-name> is the twin of --allow-red and
follows the same design and recording path: once, separate name argument,
waives only that exact unreported required check, still requires every
other required check reported and every check green, never waives an
unreadable required set, refused while the away-posture record exists, and
refused on GitLab. Merge-state BLOCKED policy is unchanged.

How this differs from the withdrawn kunchenguid#5353 (read from its diff):
- kunchenguid#5353 read the admin-only branches/{b}/protection endpoint and treated
  its 404 as "no required checks", so for any non-admin token the required
  set silently read as empty; this change reads the read-access branch
  summary and treats every failure as unreadable.
- kunchenguid#5353 ignored rulesets; this change also reads required_status_checks
  rules from the effective branch rules.
- kunchenguid#5353 made separate per-head REST reads of statuses and check-runs capped
  at per_page=100 with no pagination; this change checks presence in the
  same statusCheckRollup view the red-check gate already reads at the
  verified head.
- kunchenguid#5353 stopped at the first unreadable read; this change reports it as one
  refusal among all the others.
- kunchenguid#5353 also claimed kunchenguid#5345 (lock stealing) and changed 39 files, most
  unrelated; this change is kunchenguid#5344 only.

Live proof, read-only (a gh wrapper refused every merge and mutating call):
- cli/cli#14474 (trunk requires 3 classic build contexts, none ran):
  refused, naming build (macos-latest), build (ubuntu-latest),
  build (windows-latest); with --allow-missing "build (macos-latest)" it
  still refused, naming the other two.
- cli/cli#13665 (ran build (ubuntu-24.04-firewall) instead): refused,
  naming build (ubuntu-latest).
- hashicorp/terraform#39262 (ruleset-required checks absent): refused,
  naming Code Consistency Checks, End-to-end Tests, Race Tests, Unit Tests.
- cli/cli#14485 (all required reported and green): verified; the wrapper
  blocked the merge call and the pull request read back open.

Fixes kunchenguid#5344

* fix(review): Preserve required-check producers and aggregate independent read failures

* fix(document): Clarify required-check verification and waiver documentation

* fix(bin): match an app-bound required commit status by name

The producer-identity check resolved an app-bound required context only
against check runs, so a required context that the required app reports as
a commit status could never match and always read as "has not reported".
A commit status carries no app id to compare, so an app-bound requirement
that arrives as a status now matches by name, as before producer binding;
check runs keep requiring the configured producer app.

Live, read-only: hashicorp/terraform#39262 requires license/cla from
integration 865473, reported green as a commit status by the CLA app. The
previous head refused it as unreported; this head no longer does, while
still naming the four required check runs that never ran there.

Refs kunchenguid#5344

* fix(document): Clarify accepted commit-status producer verification limitation

---------

Co-authored-by: firstmate-oss <firstmate-oss@kunchenguid.local>
…uid#5696)

* fix(bin): never offer a persistent secondmate for teardown

The return brief's "Landed, cleanup due" scan listed every state/*.meta
record carrying a pr= and a merge-notified marker without regard to kind, so
a secondmate record holding a relayed child's merged PR put the mate itself
up for "bin/fm-teardown.sh <mate>" cleanup. A secondmate is a persistent
worker, never landed work.

- bin/fm-afk-return.sh: skip kind=secondmate in the landed-cleanup scan.
- bin/fm-pr-check.sh: refuse to record pr= or arm a merge watch on a
  kind=secondmate record before any side effect; a PR reported on its routed
  status channel belongs to a task in the mate's own home, which arms its
  own watch.
- bin/fm-watch.sh: a merged result from a poll already armed on a secondmate
  retires the poll silently - no merge outcome, marker, or wake.

* no-mistakes(document): Document secondmate merge-watch and return-brief exclusions

* no-mistakes(ci): CI failed in an unchanged watcher-shutdown test whose three-second wait was sensitive to runner load. Increased the wait for both state- and home-deletion cases without changing watcher behavior. The full fm-watch-arm suite passed locally; syntax and diff checks passed
…unchenguid#5702)

* fix(bin): refuse unknown dash-leading args in public-posting fm-x scripts

fm-x-reply.sh collected any unrecognized argument into the positional
pool and took the first one as the reply text, so an invocation like
"fm-x-reply.sh <id> --followup --final <text>" posted the literal
string "--final" to X and silently dropped the real text.

Make argument parsing strict in every script that can post publicly:
an unknown dash-leading argument, a dash-leading request_id/task id, a
dash-leading option value, or a surplus positional now exits 2 with a
usage error before any config load, outbox write, or network call. Reply
text starting with '-' is still accepted via --text-file or stdin, and
--help is honored wherever it appears instead of becoming text (a --help
forwarded through fm-x-followup.sh would have counted as a posted
follow-up and mutated the link).

fm-x-link.sh and the fm-public-followup scripts already refuse unknown
arguments; fm-x-poll.sh takes none.

* no-mistakes(review): Refuse surplus follow-up text sources; drop post-ID help branches

* no-mistakes(document): Clarify reply and follow-up argument usage

* no-mistakes(review): Refuse dash-leading --text-file operands in fm-x-reply

* no-mistakes(document): Correct follow-up argument parsing comment

* no-mistakes(document): Document dismiss argument rejection in script header
kunchenguid and others added 23 commits September 25, 2026 15:19
…nchenguid#5701)

* feat(bin): latch the supervision host after repeated engine errors

Rung 3c-1 of the PR 5631 re-cut: the host copies the Pi branch's
broken-session policy. Two consecutive engine errors latch the session;
every away wake then reaches main with one supervision-host line for a
five-minute cooldown, after which one wake probes the engine, and each
failed probe doubles the cooldown up to one hour. A reported turn without
an engine error clears it. The latch is kept per main session, engine, and
model in state/.supervision-host-health, and the engine conversation now
uses the same main-session key, which includes the lock holder's process
identity so a recycled pid never shares either.

Lifted from the validated 5631 tree and adapted to main's away-only host:
the attended recovery line and attended cooldown pass-through are left for
the attended core, so a recovery is only logged.

* no-mistakes(document): Consolidate supervision-host latch documentation
…kunchenguid#5535)

* fix(bin): absorb routine second-mate progress while surfacing routed replies

Fixes kunchenguid#2959

A kind=secondmate task's status signal was never absorbable, so a healthy
mate's routine working: and paused: appends woke the primary every time.
signal_crew_provably_working now reads the mate's lines new since the
watcher's classified position: a decision, blocker, terminal outcome, note:,
correlation-marked line, or unknown verb still surfaces regardless of busy
evidence, while unmarked working:, paused:, and resolved: fall through to the
same provably-working absorb an ordinary crewmate gets.

* no-mistakes(review): narrow secondmate routine absorb to working and paused
* fix(bin): use gh-axi for the ship DoD draft check

* no-mistakes(review): use PR number not URL in gh-axi draft check
… only (kunchenguid#5520)

* fix(bin): drop status prose from the inactive-outcome dedupe identity

The inactive-outcome receipt fingerprint included the child's sanitized
last status line, so a persistent child appending routine prose after one
terminal outcome minted a fresh parent event per sentence. Bind the
identity to incarnation, task id, terminal state, and PR only, keeping
the last line in the record as status_head evidence.

Fixes kunchenguid#2960

* no-mistakes(document): note structured-only inactive receipt identity in regression coverage
…unchenguid#5707)

* feat(bin): record the supervision host's dialog mirror on Claude and Cursor

Add bin/fm-host-mirror.sh, the one owner of the supervision host's dialog
mirror file, cursor, lock, and feed, plus the main-session key it keys
entries to. The tracked Claude UserPromptSubmit and Stop hooks and the
Cursor beforeSubmitPrompt and afterAgentResponse hooks record the captain's
prompt and main's reply, only on a home with config/supervision-host, from
a genuine primary checkout, for the lock-owning session. The mirror lands
inert: writers record and nothing reads it yet; attended supervision on the
host is the later step that consumes the feed.

Codex, Grok, OpenCode, and omp have no writer here.

* no-mistakes(review): Scope mirror dedup to session, atomic appends, marker-inclusive caps

* no-mistakes(document): Clarify dialog mirror scope and remove duplicate contract details

* no-mistakes(document): Correct Cursor hook documentation for dialog mirror registration

* no-mistakes(review): Pass mirrored dialog text to jq via stdin

* no-mistakes(document): Clarify dialog mirror documentation and remove duplicate claims

* no-mistakes(review): Preserve internal dialog whitespace; drop mirror check and verified modes

* no-mistakes(review): Drop only identical mirror repeats; remove redundant chmod guard
… escalations are not repeated (kunchenguid#5731)

* fix(bin): retire check-row receipts on branch acks and report an unchanged situation once

* fix(bin): scope a branch acknowledgement's check-row receipt retirement to
  its granted sequences

The away posture lifts the attended partition's check/decision exclusions, so
a branch grant can name check-kind rows - but the branch-actor ack still
assumed check rows were main-only and skipped every receipt scan. The queue
row was consumed while its terminal-outcome .pending receipt stayed behind,
and each inactive-reconcile cadence scan re-queued the same fingerprint. In
the first real away window on the supervision host that re-escalated one
unchanged held-PR situation on every cycle (~1,734 of 4,149 outcomes).

A branch ack now scans inactive-outcome and inactive-reconcile receipts and
commits secondmate stall receipts against exactly the sequences in its
eligible-row snapshot - the same rows it consumes - instead of none. Attended
grants still name no check row, so the scans find nothing.

* fix(bin): store a repeated captain verdict as routine while the task's
  durable situation is provably unchanged

fm-branch-outcome.sh append computes a mechanical situation key per captain
row - metadata bytes, captured status-log endpoint and identity, live
crew-state verb, worktree head - and anchors it in
state/.<task>.branch-captain-key. A later captain verdict whose recomputed key
matches is stored as routine with "unchanged since seq <N>:" prefixed to its
summary, so one situation escalates once until something provably changes. A
task with no readable status ledger is never demoted, an unreadable record
fails toward reporting, and teardown removes the sidecar with the task's
other branch records. The append-only store schema is unchanged.

This covers both hosts: the Pi supervision branch and the supervision host
both funnel reports through append.

* docs: check rows are main-owned only while attended; the away posture grants
  them to the branch, whose ack retires their receipts exactly

* test: the away-flood reproduction as a regression test (branch ack retires
  the receipt and later scans stay quiet), store-level dedupe coverage, and a
  branch-ack secondmate stall receipt case

* fix(bin): restore the secondmate child devin-config cleanup path

The branch-captain-key sidecar addition mistyped the sibling entry as
.$child_id.devin-config.json, so a forced secondmate teardown would have
stopped removing each child's real <id>.devin-config.json. Restore the
original path and add a behavioral test that stops the child sweep mid-loop
on a refused close, proving the cleaned child's devin config and captain
anchor are both removed while the unconsumed child's records are retained.

* no-mistakes(review): Key captain dedupe on the covered wake rows' fingerprint

* no-mistakes(review): Drop captain-key demotion; prove one escalation on both surfaces

* no-mistakes(review): Drop unrelated teardown test; cite both receipt test files

* no-mistakes(document): Docs already match branch-ack check-receipt retirement
…oorbell (kunchenguid#5664)

* fix(calm): deliver Claude-bound operational input as a record-backed doorbell

Claude Code 2.1.280 removes U+2063 from every submitted prompt, so a typed
operational envelope reaches a Claude Code primary as plain text. The away
daemon now writes the envelope to a record under state/operational-inbox and
types only a plain doorbell naming it; the /afk return check and the Calm mod
recognize the doorbell only when that record holds a current envelope. Marker-
preserving harnesses keep the typed envelope. The live Calm guard accepts the
2.1.280 module-load log line, drives the doorbell, and asserts thinking stays
hidden.

* no-mistakes(review): Fix operational record retention at 7 days and document prune limit

* no-mistakes(document): Point Calm bounds at 2.1.280 evidence; fix afk-exit comment

* no-mistakes(lint): Pick newest Calm e2e transcript without parsing ls

* docs(calm): add a minimal turning-Calm-on step for Claude Code

* fix(spawn): deliver the Claude launch brief as a record-backed doorbell

Claude Code strips U+2063 from the launch-prompt argument too, so a
worker's launch brief arrived with its operational marker removed.
Publish the brief as a record in the receiving home's operational
inbox - a secondmate's own state, not the primary's - and pass only
the printable doorbell naming it, falling back to the typed envelope
when the record cannot be published so the brief body still delivers.

Unwrap doorbell-carried digests in the daemon digest tests that still
read the raw send log under the claude pin, and update the documented
bounds now that launch briefs hide like the other operational rows.

* test(spawn): cover a secondmate's launch-brief record landing in its own home

The record-backed doorbell resolves its state through the receiving
pane's home, so prove a claude secondmate launch publishes into the
seeded secondmate's operational inbox and never leaks a record into
the primary's.

* no-mistakes(review): Pass primary harness to daemon, tighten retention, refresh verdicts

* no-mistakes(review): Prune operational records by exact seven-day elapsed age

* no-mistakes(review): Batch record pruning so large inboxes still expire

* no-mistakes(review): Refuse Claude spawn when brief record cannot publish

* no-mistakes(review): Drop thinking probe from Claude Calm live test and docs

* no-mistakes(review): Record dated Claude Code 2.1.282 reproduction evidence

* no-mistakes(document): Clarify operational doorbell documentation and record expiry

* no-mistakes(document): Correct AFK escalation carrier guidance

* no-mistakes(review): Describe operational record retention as about seven days

* no-mistakes(document): Clarify Calm delivery and operational record retention

* no-mistakes(review): Remove out-of-scope Calm launch guide from Claude docs

* no-mistakes(document): Document Claude launch-brief delivery and refusal

* no-mistakes(document): Correct stale operational-input documentation

* no-mistakes(ci): Fixed the stale Claude trust test to verify that worker and secondmate launches deliver readable, record-backed briefs instead of expecting brief paths in their commands. Annotated the daemon’s output variable for ShellCheck without changing behavior. The affected tests, daemon tests, ShellCheck, and diff check pass locally

* no-mistakes(ci): parse rebased Claude launch after trailer hook prefix

* no-mistakes(review): Trust launch-brief record and restore thinking bound doc

* no-mistakes(review): Parse final Claude launch statement; drop Stop-hook docs

---------

Co-authored-by: Mike Sewell <maikunari@protonmail.com>
Co-authored-by: no-mistakes <no-mistakes@localhost>
kunchenguid#5583)

A host-local relaunch rewrote only the far endpoint, so this home kept the old harness, model, and effort, and appending those keys after pr= broke pull-request poll authentication.
…ound (kunchenguid#5516)

tests/fm-watch-triage.test.sh finishes in about 434s alone and about 698s
under CI load, so the 900s bound the changed-suite runner applies produced
a false timeout under ordinary concurrent validation. Raise the automatic
bound to 1500s, which keeps every measured script under it while staying
below the 30-minute normal CI tier so a genuinely hung script still fails
here with its output before the job cap cancels the lane.

Fixes kunchenguid#3869
Refs kunchenguid#3565
…kunchenguid#5728)

* Fix nested watcher lock reclaim

* no-mistakes(review): Elect a single steal-mutex reaper and bound arm TERM wait

* no-mistakes(review): Reclaim self-held steal mutex and unify autoarm steal reaping

* no-mistakes(review): Resume own interrupted steal reap from its tombstone
…nguid#5710)

* test: hold the back-to-back boundary close on the host's own clock

test_park_boundary_holds_under_back_to_back_closes assumed two engine
turns fit in the ~16s pre-refusal window and that the stub finished a
turn in 3s. Under load the stub's real drain, report, and
acknowledgement take ~13s, so the turn either died at its bound (which
hands the wake to main, no boundary line) or the second close landed
past the window and the fixture failed while the boundary held. 3
failures in 5 runs at a load average near 11.

Hold the first turn on a release file instead: once the engine is in
flight, a second close is appended mid-turn and the turn is released as
the refusal window opens (park bound minus turn bound and grace, read
off the host's own start record). The queued close can then only wait
for the boundary on any machine speed, which is what the test asserts:
the boundary line ends the output, the demo.status row stays queued for
main, and no second engine turn ever starts. A host too loaded to start
the turn at all hands the first close to the same boundary exit.

After: 12/12 at load ~15-42.

* no-mistakes(review): Print boundary test deadline as a decimal integer

* no-mistakes(review): Hold boundary test turn on a FIFO, require full sequence

* no-mistakes(review): Remove stray before/after supervision-host test copies

* test: hold the late close's render until the refusal window opens

The boundary recheck test's node shim slept a fixed 10s, which assumed
the first close was read before the host's refusal window opened. Under
load the close arrived after the refusal check, so the host correctly
refused it before the successor started and the render snapshot never
appeared. Block the wake-prompt render on a FIFO released at the
refusal-open instant read from the host's own start record, so the
pre-turn recheck must refuse on any machine speed.

* no-mistakes(review): Derive minimal park bounds and refresh supervision-host shard hint

* no-mistakes(review): Drive park-boundary tests from a seam-gated host test clock
* fix(bin): stage remote home clones before publishing them

A remote home provision cloned the code root directly into the public
FM_HOME path while rollback() claimed rm -rf of that same path on any
failure. Bash defers trapped signals past a foreground child, but any
other cleanup or lifecycle path that removes the home directory races
the live clone's object copy, producing the CI flake "fatal: failed to
copy file to .../.git/objects/...: No such file or directory".

Clone into a private staging directory beside the home and publish with
an atomic rename once complete, so no cleanup can remove a directory a
live clone is still writing; a home that appears mid-provision now dies
cleanly instead of inheriting torn state. The regression coverage holds
a real clone mid-copy, removes the public path, and requires the
provision to finish and publish intact.

* no-mistakes(review): Prove home ownership by sentinel and hold only a live clone

* no-mistakes(review): Assert raced provision publishes a complete, intact clone

* no-mistakes(document): Document remote home staging and publication safety

* no-mistakes(lint): Fix ShellCheck warning in clone integrity assertion

* no-mistakes(document): Clarify remote home publication and rollback guarantees
* fix(control): keep a relaunched Pi worker's herdr pane status authority alive

Defect: after `bin/fm-control.sh <id> relaunch` (observed live on a herdr
Pi crewmate whose pane read idle while it ran its validation pipeline),
the pane froze at whatever its previous agent had last reported.

Cause, measured on herdr 0.9.1 against a real Pi: a pane has one status
authority, and for Pi with its integration installed that authority is
the lifecycle hooks, so herdr also skips screen detection for the pane.
In the crew shape the registration outlives its agent process (upstream
issue kunchenguid#4115; docs/herdr-backend.md "Restart and liveness behavior"), and
herdr applies only reports carrying the session identity it bound. A
replacement started fresh in that pane reports a NEW session, so its
state reports are ignored and the pane stays frozen. Nothing from
outside repairs it: `pane report-agent-session` and `pane report-agent`
for `herdr:pi` are accepted (rc=0) without being applied unless the
reporter is the registered pane agent, and `pane release-agent` on the
stale record changes nothing.

Fix: a relaunch preserves the binding instead of fighting it. The launch
owner reads the session reference the endpoint's own runtime recorded
(`fm_backend_herdr_pane_agent_session_ref`) and passes it back as Pi's
own `--session <path-or-id>` (`relaunch_resume_args`;
`fm_control_relaunch_resume_flag` owns which adapters and which
registered-agent labels qualify). That is the same reference herdr
itself resumes Pi panes with after a server restart, and the resumed
session's reports land again, which the live check confirmed: the pane
returned to working while the replacement worked and idle when it
settled, on the same session identity.

Safety: relaunch-only (a fresh spawn binds nothing), herdr-only (the one
adapter that records a per-pane session), Pi-family only, and only when
the registration's own agent label matches - so no other adapter's
conversation can be handed to a Pi launch. An unreadable, missing, or
malformed reference degrades to exactly the fresh-session launch that
existed before. No lifecycle, liveness, isolation, or merge guard is
touched, and an empty result leaves every non-Pi launch byte-identical.
`resume` remains a refused verb; docs/agent-control.md and the
harness-adapters references are corrected where they claimed Pi had no
verified resume form at all.

* no-mistakes(document): docs: correct relaunch session-authority ownership and skill paths

* no-mistakes(document): docs: correct stale control-plane ownership claim

* no-mistakes(document): docs: drop unverified Herdr restart resume claim

* no-mistakes(test): Added offline Herdr Pi session-authority relaunch coverage

* no-mistakes(document): Document Herdr Pi relaunch session continuity

* no-mistakes(ci): The failing remote relaunch test tried to arm a PR poll for a secondmate, which `fm-pr-check.sh` correctly refuses. Removed that invalid test scenario; the remaining remote relaunch tests pass, and `git diff --check` is clean
…nchenguid#5758)

Main has been red since fm-pr-check.sh began refusing to arm a merge poll
on a kind=secondmate record (kunchenguid#5696): the relaunch-ordering case in
tests/fm-remote-secondmate-relaunch.test.sh armed its fixture through that
entry point and could no longer be set up.

The ordering guarantee still matters: a secondmate record armed before the
refusal can legitimately carry a trailing pr=/pr_head= identity block until
the watcher retires it, and fm-remote-secondmate-relaunch.sh must still keep
that block last when republishing harness/model/effort. Seed the fixture the
way such a record was really written - pr= appended last to the meta, then
the poll artifacts published through the same
fm_pr_poll_prepare/fm_pr_poll_publish_prepared pair fm-pr-check.sh uses, a
pattern tests/fm-pr-check-security.test.sh already follows - and drop the
now-unused fake gh fixture. The kunchenguid#5696 refusal itself stays pinned by the
security suite's secondmate-record case.
…id#5748)

* feat: run attended supervision on the host for Claude and Cursor

On a home opted into config/supervision-host with a Claude or Cursor
primary, the supervision host now takes the attended wakes the Pi branch
would take: routine outcomes stay off main, and a captain outcome wakes
main once with a branch-outcome line and waits in the drain's new
BRANCH OUTCOMES section until main acknowledges it with mark-processed.

- The offer rule moves into branchOfferForWake, shared by the Pi watcher
  and the host through bin/fm-branch-dispatch.mjs offer.
- The host feeds the dialog mirror at the head of each attended wake and
  passes a close through unchanged when it is main-only, the engine or a
  tool is missing, the primary has no verified mirror, the main session
  cannot be identified, or the session is cooling down.
- The drain presents captain outcomes first, one line per task, never
  behind older routine outcomes, and collapses routine overflow into a
  count that is marked read.
- The return advances the store's read cursor through the away window
  once the brief has rendered, so the first drain does not replay it.
- The branch prompt's mirror wording is host-neutral, and the rule to
  report what main must act on as captain, once per unchanged situation,
  applies only to the attended posture on the host.

* docs: record the attended supervision host live check

* no-mistakes(review): Present pre-window unread outcomes and contiguous captain prefix

* no-mistakes(review): Return brief presents every row it marks read

* no-mistakes(review): Return brief lists every unread outcome in one list

* no-mistakes(review): Keep return list in store order and gate cursor failures

* no-mistakes(review): Make the drain the only branch-outcome presenter after return

* no-mistakes(review): Gate return on drain outcome failures; byte-count outcome budgets

* no-mistakes(review): Gate drain on projection failures; UTF-8-safe byte cuts

* no-mistakes(review): Fail drain without jq; hand unreadable prompt mirror to main

* no-mistakes(document): Correct supervision-host return and drain documentation

* no-mistakes(review): Recheck attended offer at turn start; honest failed-drain brief

* no-mistakes(document): Correct supervision-host posture and drain documentation

* no-mistakes(document): Documentation remains accurate for attended supervision
…unchenguid#5753)

Each '# shellcheck source=' directive makes ShellCheck's external-source
traversal expand that library's whole transitive graph again at the site.
fm-pending-reply-lib carried three directed lazy sources of fm-wake-lib and
two of fm-parent-channel-lib on identical per-call re-source sites, so one
file analysis peaked above 4 GiB and every caller (fm-watch, fm-teardown)
inherited the multiplier - the root cause of the PR kunchenguid#5732 Lint 1 OOM kill.

Keep the runtime '.' commands byte-identical: the lazy re-source under
'local STATE FM_WAKE_QUEUE FM_WAKE_QUEUE_LOCK' is real behavior. Drop the
duplicate directives so each library expands once per unit, and drop the
tmux/classify directives since classify already arrives through the kept
fm-wake-lib expansion and no tmux symbol is referenced here. The directive
above the lib-dir assignment is kept - it binds the bin/ prefix so the
undirected sites still resolve without SC1091.

Measured peak RSS, ShellCheck 0.11.0 -x on Linux arm64:
  bin/fm-pending-reply-lib.sh  4.06 GiB -> 1.96 GiB, zero findings
Keep fork-local behavior while adopting upstream pending-reply deduplication, Claude worker doorbells, and the other 47 upstream commits since the last catch-up.

Local validation: focused task delivery, pending-reply, daemon, operational-input, AFK, and Pi Calm tests passed; lint and documentation checks passed. The session-lock ancestry E2E pty-host reparent assertion fails identically on clean origin/main and upstream/main source snapshots in this WSL host, so it is not caused by this merge. CI on Linux will decide the full-suite result.
…ne test, tests/fm-spawn-dispatch-profile.test.sh, case "an absent worker settings file changed the launch". The merge brought in upstream's new signature for the test helper `claude_expected_launch`. It changed from `<home> <id> <flag>` to `<launch> <home> <id> <flag>`, because the expected launch now rebuilds the launch-brief doorbell from the real launch. Our fork-only test `test_claude_worker_settings_absent_keeps_launch` (line 1705) still called it the old 3-argument way. Its arguments shifted by one: `cd workersettings-absent-z24/state` failed and `$4` was unbound, so the expected string came out empty. Invariant: every caller of `claude_expected_launch` must pass the 4-argument `<launch> <home> <id> <flag>` form. I checked every caller. The five upstream call sites (lines 144, 1144, 1582, 1600) already do, no other test file uses the helper, and line 1705 was the only stale one. Fix: one line in tests/fm-spawn-dispatch-profile.test.sh. The call now passes "$launch" first, like its sibling calls. The assertion is unchanged, and no production code changed. Verification: I ran `bash tests/fm-spawn-dispatch-profile.test.sh` locally and it exited 0. All 72 cases print `ok` and none fail, including all five claude-worker-settings cases (absent, ship merge, scout/secondmate, invalid refusal, codex unaffected). `bash -n` also passes. I left the change uncommitted for the outer executor
@ashkonmousavi
ashkonmousavi merged commit 5c0adde into main Sep 26, 2026
19 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants