Skip to content

feat(bin): sync the fork with upstream firstmate main and reconcile the zcode harness - #10

Merged
d-ploutarchos merged 235 commits into
mainfrom
fm/fm-upstream-sync
Oct 1, 2026
Merged

d-ploutarchos merged 235 commits into
mainfrom
fm/fm-upstream-sync

Conversation

@d-ploutarchos

@d-ploutarchos d-ploutarchos commented Sep 30, 2026 •

Copy link
Copy Markdown

Intent

The captain (2026-09-29): "we were planning to update the firstmate repo to the latest version" and agreed to do it Wed 06:00Z. The OK-LG/firstmate fork is 224 commits behind upstream kunchenguid/firstmate main (c35b9a6) and carries 9 fork-only zcode harness commits (#1-#9) that must be kept.

What Changed

  • Merged upstream kunchenguid/firstmate main (c35b9a69) into the fork, bringing the supervision host (bin/fm-supervision-host.sh, bin/fm-supervision-engine-lib.sh, enabled by default for Claude primaries), the firstmate-calm Claude mod under .claude/mods/, new helpers for dispatch resolution, fleet ledger, contributions, host mirroring, worker accounts, PR state/reviewers, the live supervision lab, Devin and Gerrit forge support, plus the accompanying docs/, .agents/skills/, and tests/ updates.
  • Kept the fork-only zcode harness and registered it in the tables that arrived with the merge: zcode added to the arm-signal, press-gap, and hazard-signal interrupt cases in bin/fm-control-lib.sh, to the "no effort levels" rule in bin/fm-dispatch-resolve.sh, and fm-zcode-lib.sh added to the tmux and herdr backend sibling prechecks in bin/fm-backend.sh.
  • Extended tests and docs for the reconciled surface: a table-completeness guard driven off fm_control_harnesses in tests/fm-control.test.sh, zcode assertions in tests/fm-zcode-harness.test.sh and tests/fm-backend.test.sh, fm-zcode-lib.sh/fm-path-lib.sh staged in the session-lock and turn-end sandbox fixtures, and doc updates recording that zcode has no supervision-host arm owner and that docs/architecture.md owns the per-harness wait-shape mapping.

Merge report

Summary

Merges upstream kunchenguid/firstmate main at c35b9a69 into this fork (a merge, not a rebase).
The fork was 224 commits behind upstream and carries 9 fork-only zcode harness commits (#1-#9).
All of that zcode work is kept: crewmate and scout launches, the primary-session harness, the opt-in --zcode-tui variant, the turn-end hook with resume and quota evidence, herdr agent-view busy flips, and the variant-shaped interrupt record.

Upstream added a Devin adapter in most of the same places, so nearly every conflict is two new harnesses landing in one list.
In each case the resolution keeps upstream's change and puts zcode next to Devin.

Conflicts and how each was resolved

22 files conflicted.

Docs and skills

File Resolution
.agents/skills/harness-adapters/SKILL.md (3) Kept upstream's Devin entries and the fork's zcode entries in the description, the crew-only sentence, and the reference map.
AGENTS.md (2) Took upstream's move of the home-layout block into the new operational-home-layout skill, and moved the fork's two zcode state rows (<id>.zcode-turnend-token, <id>.zcode-session) into that skill. Took upstream's section 4 bullets (Devin plus the worker account-pin rule) with zcode kept in the verified list.
docs/architecture.md The spawn-compat list names both devin and zcode.
docs/configuration.md (2) Kept the zcode verification paragraph, the Devin paragraph, and upstream's new subsection headings. Kept the fork's choice for the per-harness wait-shape sentence: a pointer to docs/architecture.md, which owns that list, instead of upstream's duplicate. A future upstream sync will bring the duplicate back at this spot.
docs/documentation-audiences.json Both entries (supervision-protocols/zcode.md, supervision-host.md).
docs/sessionstart-nudge.md (2) Upstream split the transport table into one subsection per harness. Added a Zcode row to the new tier table and a ### Zcode subsection carrying the fork's text.
docs/turnend-guard.md Both bullets: zcode's missing turn-end integration and upstream's pi-code stand-down.
docs/verification/runtime-backends.md Both subsections: the zcode agent-view bridge record and upstream's pane-status-authority record.
docs/watcher-continuity.md Took upstream's bulleted layout. The Grok bullet now reads "Grok and Zcode retain their tracked background-task notification protocols".

Scripts

File Resolution
bin/fm-agent-process-lib.sh (2), bin/fm-session-lock-lib.sh Took upstream's _FM_*_LIB_DIR sourcing pattern and source fm-zcode-lib.sh the same way. The process-name classifier anchors both devin and the zcode names.
bin/fm-bootstrap.sh Took upstream's typed-dispatch validator. zcode is in the static verified list, and the typed path reads fm_control_harnesses, which now includes zcode.
bin/fm-busy-lib.sh (3) Merged the source-list header. Kept both devin-hook and zcode-hook. Kept the zcode rendered-tail fallback next to upstream's launch-prompt signatures.
bin/fm-control-lib.sh (8) Took upstream's fm_control_harnesses table (zcode added to it). Every per-harness table has both devin and zcode.
bin/fm-control.sh (2) Merged the header knobs (FM_CONTROL_ARM_WAIT plus the fork's settle wording). Merged do_exit locals: upstream's hazard/composer/absence checks plus the fork's signal-exit counters.
bin/fm-crew-state.sh Took upstream's unconditional tail capture, which covers the fork's grok|zcode capture.
bin/fm-harness.sh (2) Usage line and process verdict carry both devin and zcode.
bin/fm-spawn.sh (16) Upstream reindented the file and added --branch-prefix, Devin, worker account pins, and Claude launch changes. The fork's zcode pieces were re-applied in upstream's formatting: --zcode-tui parsing, the relaunch override guard and forwarding, the launch template, binary resolution, model and effort omission, the busy arm, hook wiring and session resume, the zcode_tui meta key, placeholder substitution, and the outer env wrap. --harness "$HARNESS" stays on every busy arm (the herdr bridge reads it), including Devin's.
bin/fm-test-run.sh (4) Family and live-guard lists include both adapters. Weight hints stay sorted, with the zcode hint added. The changed-file map keeps upstream's new entries and the zcode hook mapping.

Tests

File Resolution
tests/fm-control.test.sh (3) Fake sleep keeps upstream's Devin passthrough and the fork's interrupt-death hook. The env passthrough carries all three knobs. Family-resolution pairs include devin:devin and zcode:zcode.
tests/fm-cursor-primary.test.sh Fixture copies both fm-path-lib.sh (upstream) and fm-zcode-lib.sh (fork).
tests/fm-session-lock-ancestry.test.sh (4) Rebuilt from upstream's file plus the fork's zcode primary-ancestry test, lock-acquisition fixture, and runner lines.

Changes outside the conflicts

  • bin/fm-harness.sh supervision_primary_pin and bin/fm-supervision-engine-lib.sh's test seam gained zcode, because these harness-name lists are new upstream.
  • The operational-home-layout skill gained the zcode state rows (see AGENTS.md above).
  • tests/fm-zcode-harness.test.sh: the TUI fixture now reads upstream's staged launch file (a spawn types . '<file>' instead of the whole command), matching how upstream's kimi fixture handles it.
  • tests/fm-afk-return.test.sh, tests/fm-gotmp.test.sh, tests/fm-turnend-guard.test.sh, and tests/fm-session-lock-ancestry.test.sh: sandbox fixtures that copy fm-harness.sh, the agent-process lib, or the session-lock lib also copy fm-zcode-lib.sh, which those scripts source.

Upstream changes that alter running-fleet behaviour

  • Supervision host is on by default for Claude primaries (feat: enable supervision host by default for Claude primaries kunchenguid/firstmate#6124, fix: inherit supervision host opt-out across secondmates kunchenguid/firstmate#6154).
    With no config/supervision-host, a Claude primary now runs the supervision host: a headless engine session that owns the watcher cycle and handles routine wakes without waking main.
    This home runs a Claude primary with Claude secondmates and has no config/supervision-host file, so the host turns on after restart in the primary and in every Claude secondmate home.
    To opt out, create config/supervision-host-off in the primary. Secondmates inherit that opt-out.
    A zcode primary is not a host primary and keeps its existing background-notify protocol.
  • Attended supervision on Claude and Cursor (feat: add attended supervision for Claude and Cursor hosts kunchenguid/firstmate#5748), and the away posture on non-Pi primaries (feat: extend opt-in away supervision to non-Pi primaries kunchenguid/firstmate#5503, feat: add opt-in Claude away supervision host kunchenguid/firstmate#5488).
    /afk on a host home launches no away daemon, because the host is the away session there.
    /quiet enters nothing where the attended host runs.
  • AGENTS.md was restructured into load-on-trigger skills: operational-home-layout, session-start-recovery, ship-landing, validation-supervision, scout-completion, away-quiet-supervision, and agent-skill-trigger-index.
  • Worker launches changed.
    • Claude workers receive the brief as a published launch-brief record plus a printable doorbell, a task-worker trust statement through --append-system-prompt, and --add-dir task-channel grants.
    • Codex crewmates launch with --disable hooks.
    • Long launch commands are staged in a file and sourced.
    • Pi relaunches resume the herdr-bound session.
  • New opt-ins that are off unless configured: typed dispatch (TYPESAFE_API_KEY), Claude and Pi worker account pins (config/claude-account, config/pi-account), --branch-prefix, config/keep-ai-trailers, and the Devin crewmate adapter.
  • Tool version floors went up: tasks-axi 0.2.4 to 0.2.6, quota-axi 0.1.29 to 0.1.51, and lavish-axi 0.1.46 to 0.1.80 (board 0.1.77).
    This machine currently has tasks-axi 0.2.5, quota-axi 0.1.42, and lavish-axi 0.1.67, all below the new floors.
    Until tasks-axi reaches 0.2.6, task cleanup refuses because its automatic backlog transitions need it.
    Bootstrap reports the others at session start and asks before installing.

Restart impact

  • Running homes pick this up only after it lands on main and each home fast-forwards and restarts (/updatefirstmate).
    Live crewmates keep their current launch shape until they are relaunched.
  • Upgrade tasks-axi (to 0.2.6 or newer), quota-axi, and lavish-axi before or right after the restart, or cleanup and some reports will refuse.
  • After the restart, the Claude primary and every Claude secondmate run the supervision host unless config/supervision-host-off is created first.
  • A zcode primary and zcode crewmates are unaffected beyond the shared upstream changes.

Validation

  • PR CI at 469dd177: 21 of 22 checks pass. "Behavior portable serial 5" was cancelled at its 30-minute job limit twice (original run and one rerun) without any failure: all 13 tests it reached before the limit passed. The limit is consumed by tests/fm-supervision-host.test.sh, which takes about 17.5-18.5 minutes on these runners against about 10 minutes in upstream's CI and a 41-second timing hint in bin/fm-test-run.sh. Locally that suite takes the same time on pure upstream and on this merge (636s vs 634s), so the merge does not cause it. Fixing the shard's time budget is a follow-up.

  • CI on the first push then caught a real gap: upstream's fm-wake-lib.sh now sources fm-path-lib.sh, which the fork's zcode lock-acquisition fixture did not stage, so fm-lock.sh looped forever and the zcode lock tests hung their shard. The fixture now stages it. A later review registered fm-zcode-lib.sh in upstream's new backend sibling-readability precheck for tmux and herdr. It also recorded in docs/supervision-host.md that zcode is outside the supervision host.

  • Follow-up, not in this merge: upstream's herdr sibling list also omits fm-cursor-lib.sh.

  • The no-mistakes review then caught four more merge gaps, fixed in the pipeline commits on this branch. zcode was missing from upstream's three new interrupt tables, which would have broken interrupting or stopping a zcode worker. It was also missing from the no-effort rule in upstream's new dispatch resolver. Two new test assertions could pass vacuously. And the zcode primary protocol still said it was the "supervision host". The inert live-lab fixture line was dropped.

  • Follow-up, not in this merge: several per-harness lists predate the merge and still omit zcode (the /afk harness buckets, docs/agent-control.md, docs/trace-context.md, docs/remote-secondmates.md). Each needs its own verification.

  • bin/fm-test-run.sh --check-coverage passes (the coverage check needs LC_ALL=C on this host).

  • bin/fm-lint.sh and bin/fm-doc-audience-check.sh pass.

  • zcode suites pass: fm-zcode-harness, fm-zcode-herdr-bridge, fm-busy-adapter-wiring, fm-control-relaunch, fm-crew-state, fm-cursor-primary, fm-session-lock-ancestry (apart from the environment case below), fm-supervision-instructions, and fm-tmux-agent-liveness.

  • Full suite, run locally the way CI shards it (both portable-parallel lanes, all 9 portable-serial shards, and the real-Herdr family): 246 scripts, 204 passed, and 36 of those gate-skipped as opt-in live guards.
    The turn-end guard failure was a missing fixture copy and is fixed in the second commit; that suite now passes.
    The Stop auto-arm suite failed once under 4-way load and passes when run alone.
    Every other failure reproduces with the same assertion on a pure copy of upstream c35b9a69 on this host, so none comes from the merge.
    They come from this host: it runs as root (so permission-denied fixtures don't deny), it has tasks-axi 0.2.5 (below the new 0.2.6 floor), it has no ruby or Chrome, and it has lavish-axi 0.1.67.
    fm-remote-reply hangs at the same point on upstream too.
    CI on the PR is the clean-environment signal.

  • Opt-in zcode live guards against real zcode (zcode-app-cli 3.12.3-26, zcode-runtime 0.16.5) and herdr 0.9.1:

    • fm-zcode-primary-live-e2e passes.
    • fm-zcode-herdr-bridge-live-e2e passes.
    • fm-zcode-signals-live-e2e passes every check up to the credential-less probe. A key-stripped config there now exits 0 where exit 1 was pinned.
      That probe runs zcode directly, and its test and the zcode hook are byte-identical to the fork's main, so this is vendor drift since the recorded 3.11.2-24 pass, not a merge regression.
      It is left as a follow-up.

Risk Assessment

⚠️ Medium: The fork's own reconciliation delta is small, bounded, and verified faithful (95 deletions vs upstream, every one a zcode-augmented replacement, with no zcode coverage lost and no upstream behavior reverted), but the change wholesale adopts 224 unreviewed upstream commits and leaves one test-only fixture weakness, so it is safe to merge with that fixture fix as a follow-up.

Testing

I ran the six test files this change touches plus the zcode harness suite as baseline (all pass), then drove the change's real surfaces end to end. On a real tmux server with a real process whose kernel comm is zcode, the control plane interrupts and exits a zcode task correctly, the interrupt proof follows the recorded launch variant (headless ends with agent-ended-by-interrupt, the TUI variant stays agent-alive), the endpoint survives exit, and an unverified adapter is still refused with a diagnostic. The backend sibling registration shows a real difference: on a home missing fm-zcode-lib.sh, the merge commit's list let fm-control dot-source the tmux adapter anyway (27 sourcing failures, then it acted on the endpoint), while this branch refuses first with one diagnostic and sends no keys. A real zcode-cli-named session acquires the fleet lock through the staged bin/ only with fm-path-lib.sh present, and plain bash is refused; that is the pair the session-lock suite cannot reach on this host, since it still aborts earlier on the upstream pid-1 reparent assumption (this environment runs under a systemd --user subreaper) - the already-declined environment issue, unrelated to the change. The resolver rejects a zcode profile carrying an effort in both the rule and default positions, demands zcode's explicit provider, and accepts zcode with provider and no effort. The rendered zcode session-start supervision block now reads "Interactive TUI sessions are the supported Zcode primary surface" and contains no supervision-host claim even on a home that opts into the host, while a claude home still renders the full host protocol. There is no GUI, HTML, or other rendered surface in this change, so the reviewer-visible evidence is CLI transcripts and the generated supervision block rather than screenshots. Broad regression across the 224 merged upstream commits is CI's job, not this targeted step.

  • Live validation: ✅ go - 9 of 10 scenarios driven live against the product
Scenario Result Live Evidence
A zcode task is interrupted and then exited through bin/fm-control.sh on a real tmux endpoint; the agent is stopped and the endpoint is preserved ✅ pass live control-plane-zcode-live.txt - real tmux socket, a real process with comm=zcode; interrupt returns verified=agent-alive, exit returns stopped and tmux list-windows still shows fm-zc1
The exit verb's bounded second interrupt key stops a TUI-variant zcode worker that survived the first C-c ✅ pass live control-exit-second-key-zcode.txt - the pane records C-c press 1 (turn cancelled), then press 2 and the process exiting; fm-control reports stopped
The zcode interrupt postcondition follows the recorded launch variant: a headless worker ending is success, the opt-in TUI variant must stay alive ✅ pass live zcode-interrupt-variants-live.txt - verified=agent-ended-by-interrupt with no zcode_tui line, verified=agent-alive with zcode_tui=1
Adversarial: a task recording an unverified harness is refused with a diagnostic rather than acted on with guessed mechanics ✅ pass live control-plane-zcode-live.txt section C - error: task zc1 records harness &#39;someagent&#39;, which has no verified control mechanics, exit 1
Adversarial: a firstmate home whose bin/ is missing fm-zcode-lib.sh refuses the lifecycle before touching the endpoint instead of running on a partially loaded library ✅ pass live backend-sibling-precheck-live.txt - merge-commit list: 27 sourcing failures then keys delivered and an unconfirmed exit; this branch: one `endpoint reads 'unverified' ... refusing to send a lifecycle…
A real zcode-cli-named primary session acquires the fleet lock from a staged home bin/, and a session with no zcode ancestry is refused ✅ pass live zcode-fleet-lock-live.txt - lock acquired: harness pid &lt;session pid&gt; with state/.lock matching, exit 0; plain bash gets cannot locate harness process in ancestry, exit 1 and no lock file; without…
A crew-dispatch config that gives a zcode profile an effort is an actionable configuration error before any request, in both rule and default position ✅ pass live dispatch-resolve-zcode.txt cases 1 and 2 - each use profile effort must be supported by its harness and model and each default profile effort must be supported by its harness and model, exit 2 eac…
Adversarial: zcode must declare its provider, and a zcode profile with provider and no effort is accepted ✅ pass live dispatch-resolve-zcode.txt cases 3 and 4 - provider-less zcode exits 2 naming zcode; with provider and no effort the config validates and the tool proceeds to its request path
A zcode primary's session-start supervision block no longer claims the supervision host, while a claude primary on the same home still renders the host protocol ✅ pass live supervision-block-zcode.txt and supervision-host-gating.txt - the rendered zcode block reads "Interactive TUI sessions are the supported Zcode primary surface" and matches 'supervision host' 0 times e…
The reference-doc and configuration-doc wording changes (docs/agent-control.md interrupt-table paragraphs, docs/configuration.md wait-shape pointer, .agents/skills harness reference) ⏸️ untested no These three files have no runtime consumer: unlike docs/supervision-protocols/zcode.md (which bin/fm-supervision-instructions.sh renders into a live session and which I did drive), they are read by hu…
Evidence: zcode control plane on a real tmux endpoint (interrupt, exit, unverified-adapter refusal), before and after the interrupt-table registration

Source: zcode control plane on a real tmux endpoint (interrupt, exit, unverified-adapter refusal), before and after the interrupt-table registration

################ A. the merged-upstream tables BEFORE the fork's zcode registration
#  bin/ staged with bin/fm-control-lib.sh exactly as merge commit 3140240 left it
pane foreground process group, as ps sees it:
  2508679 2508679 2508718 bash
  2508718 2508718 2508718 zcode
  2509364 2508718 2508718 sleep

$ fm-control.sh zc1 interrupt
interrupt-delivered zc1 harness=zcode backend=tmux verified=agent-alive cancel=unconfirmed
exit=0   stdout/stderr bytes=90
pane after the command (the agent was never touched):
  root@vultr:/tmp/fm-live-zcode.t1E2HA# PATH=/tmp/fm-live-zcode.t1E2HA/fakebin:$PATH; zcode /tmp/fm-live-zcode.t1E2HA/agent.sh
  [zcode TUI] running a turn
  ^C[zcode TUI] turn cancelled by C-c (press 1)

$ fm-control.sh zc1 exit
stopped zc1 harness=zcode backend=tmux endpoint=fmses:fm-zc1 worktree=/tmp/fm-live-zcode.t1E2HA/wt
exit=0   stdout/stderr bytes=98

################ B. this branch's bin/ (zcode registered in every interrupt table)
pane foreground process group, as ps sees it:
  2511401 2511401 2511429 bash
  2511429 2511429 2511429 zcode
  2512090 2511429 2511429 sleep

$ fm-control.sh zc1 interrupt
interrupt-delivered zc1 harness=zcode backend=tmux verified=agent-alive cancel=unconfirmed
exit=0
pane after the interrupt (C-c reached the real process, the TUI survived):
  root@vultr:/tmp/fm-live-zcode.t1E2HA# PATH=/tmp/fm-live-zcode.t1E2HA/fakebin:$PATH; zcode /tmp/fm-live-zcode.t1E2HA/agent.sh
  [zcode TUI] running a turn
  ^C[zcode TUI] turn cancelled by C-c (press 1)

$ fm-control.sh zc1 exit    (signal-shaped exit: first C-c, then the second key)
stopped zc1 harness=zcode backend=tmux endpoint=fmses:fm-zc1 worktree=/tmp/fm-live-zcode.t1E2HA/wt
exit=0
pane after exit:
  ^C[zcode TUI] turn cancelled by C-c (press 1)
  ^C[zcode TUI] turn cancelled by C-c (press 2)
  [zcode TUI] exiting
  root@vultr:/tmp/fm-live-zcode.t1E2HA#
window still present (exit preserves the endpoint):
  fm-zc1

################ C. adversarial: an unverified adapter must still be refused, loudly
$ fm-control.sh zc1 interrupt   (state/zc1.meta records harness=someagent)
error: task zc1 records harness 'someagent', which has no verified control mechanics; fm-control refuses to guess an interrupt key or exit command
exit=1
Evidence: exit verb's second interrupt key on a TUI-variant zcode task, before and after

Source: exit verb's second interrupt key on a TUI-variant zcode task, before and after

################ A. BEFORE: merge commit 3140240's interrupt tables (no zcode row)
$ fm-control.sh zc1 exit
stopped zc1 harness=zcode backend=tmux endpoint=fmses:fm-zc1 worktree=/tmp/fm-live-zcx.MQeBib/wt
exit=0   captured output bytes=96
  ps: no zcode process left in the pane
  pane| ^C[zcode TUI] C-c press 2: turn cancelled
  pane| [zcode TUI] second key: exiting
  pane| root@vultr:/tmp/fm-live-zcx.MQeBib#

################ B. AFTER: this branch's tables (zcode registered)
$ fm-control.sh zc1 exit
stopped zc1 harness=zcode backend=tmux endpoint=fmses:fm-zc1 worktree=/tmp/fm-live-zcx.MQeBib/wt
exit=0
  ps: no zcode process left in the pane
  pane| ^C[zcode TUI] C-c press 1: turn cancelled
  pane| ^C[zcode TUI] C-c press 2: turn cancelled
  pane| [zcode TUI] second key: exiting
  pane| root@vultr:/tmp/fm-live-zcx.MQeBib#
  endpoint preserved:
  window| fm-zc1
Evidence: interrupt proof follows the recorded launch variant: headless process ends, TUI survives

Source: interrupt proof follows the recorded launch variant: headless process ends, TUI survives

\### headless spawn (state/zc1.meta records no zcode_tui line)
$ fm-control.sh zc1 interrupt
  interrupt-delivered zc1 harness=zcode backend=tmux verified=agent-ended-by-interrupt cancel=unconfirmed
  exit=0
  pane| ^C
  pane| root@vultr:/tmp/fm-live-hl.xHyjPz#

\### the opt-in TUI variant (state/zc1.meta records zcode_tui=1)
$ fm-control.sh zc1 interrupt
  interrupt-delivered zc1 harness=zcode backend=tmux verified=agent-alive cancel=unconfirmed
  exit=0
  pane| [zcode TUI] working
  pane| ^C[zcode TUI] turn cancelled, still here
Evidence: a home missing fm-zcode-lib.sh: partial source and a blind lifecycle action before the fix, one clean refusal after

Source: a home missing fm-zcode-lib.sh: partial source and a blind lifecycle action before the fix, one clean refusal after

################ A. BEFORE: fm_backend_source's tmux siblings do not name fm-zcode-lib.sh
   staged home bin/ is missing fm-zcode-lib.sh:
   files: 209
   ls: cannot access '/tmp/fm-live-sib.2DgAmT/bin-pre/fm-zcode-lib.sh': No such file or directory
$ fm-control.sh zc1 exit
  /tmp/fm-live-sib.2DgAmT/bin-pre/fm-session-lock-lib.sh: line 26: /tmp/fm-live-sib.2DgAmT/bin-pre/fm-zcode-lib.sh: No such file or directory
  /tmp/fm-live-sib.2DgAmT/bin-pre/fm-session-lock-lib.sh: line 26: /tmp/fm-live-sib.2DgAmT/bin-pre/fm-zcode-lib.sh: No such file or directory
  /tmp/fm-live-sib.2DgAmT/bin-pre/fm-agent-process-lib.sh: line 23: /tmp/fm-live-sib.2DgAmT/bin-pre/fm-zcode-lib.sh: No such file or directory
  /tmp/fm-live-sib.2DgAmT/bin-pre/fm-session-lock-lib.sh: line 26: /tmp/fm-live-sib.2DgAmT/bin-pre/fm-zcode-lib.sh: No such file or directory
  /tmp/fm-live-sib.2DgAmT/bin-pre/fm-session-lock-lib.sh: line 26: /tmp/fm-live-sib.2DgAmT/bin-pre/fm-zcode-lib.sh: No such file or directory
  /tmp/fm-live-sib.2DgAmT/bin-pre/fm-agent-process-lib.sh: line 23: /tmp/fm-live-sib.2DgAmT/bin-pre/fm-zcode-lib.sh: No such file or directory
  /tmp/fm-live-sib.2DgAmT/bin-pre/fm-session-lock-lib.sh: line 26: /tmp/fm-live-sib.2DgAmT/bin-pre/fm-zcode-lib.sh: No such file or directory
  /tmp/fm-live-sib.2DgAmT/bin-pre/fm-session-lock-lib.sh: line 26: /tmp/fm-live-sib.2DgAmT/bin-pre/fm-zcode-lib.sh: No such file or directory
  /tmp/fm-live-sib.2DgAmT/bin-pre/fm-agent-process-lib.sh: line 23: /tmp/fm-live-sib.2DgAmT/bin-pre/fm-zcode-lib.sh: No such file or directory
  /tmp/fm-live-sib.2DgAmT/bin-pre/fm-session-lock-lib.sh: line 26: /tmp/fm-live-sib.2DgAmT/bin-pre/fm-zcode-lib.sh: No such file or directory
  /tmp/fm-live-sib.2DgAmT/bin-pre/fm-session-lock-lib.sh: line 26: /tmp/fm-live-sib.2DgAmT/bin-pre/fm-zcode-lib.sh: No such file or directory
  /tmp/fm-live-sib.2DgAmT/bin-pre/fm-agent-process-lib.sh: line 23: /tmp/fm-live-sib.2DgAmT/bin-pre/fm-zcode-lib.sh: No such file or directory
  /tmp/fm-live-sib.2DgAmT/bin-pre/fm-session-lock-lib.sh: line 26: /tmp/fm-live-sib.2DgAmT/bin-pre/fm-zcode-lib.sh: No such file or directory
  /tmp/fm-live-sib.2DgAmT/bin-pre/fm-session-lock-lib.sh: line 26: /tmp/fm-live-sib.2DgAmT/bin-pre/fm-zcode-lib.sh: No such file or directory
  /tmp/fm-live-sib.2DgAmT/bin-pre/fm-agent-process-lib.sh: line 23: /tmp/fm-live-sib.2DgAmT/bin-pre/fm-zcode-lib.sh: No such file or directory
  /tmp/fm-live-sib.2DgAmT/bin-pre/fm-session-lock-lib.sh: line 26: /tmp/fm-live-sib.2DgAmT/bin-pre/fm-zcode-lib.sh: No such file or directory
  /tmp/fm-live-sib.2DgAmT/bin-pre/fm-session-lock-lib.sh: line 26: /tmp/fm-live-sib.2DgAmT/bin-pre/fm-zcode-lib.sh: No such file or directory
  /tmp/fm-live-sib.2DgAmT/bin-pre/fm-agent-process-lib.sh: line 23: /tmp/fm-live-sib.2DgAmT/bin-pre/fm-zcode-lib.sh: No such file or directory
  /tmp/fm-live-sib.2DgAmT/bin-pre/fm-session-lock-lib.sh: line 26: /tmp/fm-live-sib.2DgAmT/bin-pre/fm-zcode-lib.sh: No such file or directory
  /tmp/fm-live-sib.2DgAmT/bin-pre/fm-session-lock-lib.sh: line 26: /tmp/fm-live-sib.2DgAmT/bin-pre/fm-zcode-lib.sh: No such file or directory
  /tmp/fm-live-sib.2DgAmT/bin-pre/fm-agent-process-lib.sh: line 23: /tmp/fm-live-sib.2DgAmT/bin-pre/fm-zcode-lib.sh: No such file or directory
  /tmp/fm-live-sib.2DgAmT/bin-pre/fm-session-lock-lib.sh: line 26: /tmp/fm-live-sib.2DgAmT/bin-pre/fm-zcode-lib.sh: No such file or directory
  /tmp/fm-live-sib.2DgAmT/bin-pre/fm-session-lock-lib.sh: line 26: /tmp/fm-live-sib.2DgAmT/bin-pre/fm-zcode-lib.sh: No such file or directory
  /tmp/fm-live-sib.2DgAmT/bin-pre/fm-agent-process-lib.sh: line 23: /tmp/fm-live-sib.2DgAmT/bin-pre/fm-zcode-lib.sh: No such file or directory
  /tmp/fm-live-sib.2DgAmT/bin-pre/fm-session-lock-lib.sh: line 26: /tmp/fm-live-sib.2DgAmT/bin-pre/fm-zcode-lib.sh: No such file or directory
  /tmp/fm-live-sib.2DgAmT/bin-pre/fm-session-lock-lib.sh: line 26: /tmp/fm-live-sib.2DgAmT/bin-pre/fm-zcode-lib.sh: No such file or directory
  /tmp/fm-live-sib.2DgAmT/bin-pre/fm-agent-process-lib.sh: line 23: /tmp/fm-live-sib.2DgAmT/bin-pre/fm-zcode-lib.sh: No such file or directory
  error: exit-delivered zc1 interrupt=signal cancel=unconfirmed agent-state=alive exit=unconfirmed; the agent did not stop within 4s
  exit=1
  ps: the zcode agent is STILL RUNNING (no lifecycle step happened)

################ B. AFTER: fm-zcode-lib.sh registered in the tmux sibling precheck
$ fm-control.sh zc1 exit
  error: task zc1's endpoint reads 'unverified' rather than a positively classified state; refusing to send a lifecycle command into an unattributed endpoint
  exit=1
  ps: the zcode agent is STILL RUNNING (no lifecycle step happened)

################ C. control: the same bin/ WITH fm-zcode-lib.sh restored still works
$ fm-control.sh zc1 exit
  error: exit-delivered zc1 interrupt=signal cancel=unconfirmed agent-state=alive exit=unconfirmed; the agent did not stop within 4s
  exit=1
  ps: the zcode agent is STILL RUNNING (no lifecycle step happened)

################ reading this transcript
A: the precheck did not name fm-zcode-lib.sh, so fm-control.sh dot-sourced the
   tmux adapter anyway - 27 sourcing failures on stderr - and then ACTED on the
   endpoint with a partially loaded library before refusing.
B: the registered sibling makes the same incomplete bin/ refuse first, with one
   diagnostic and no keys delivered.
C: the lab agent script ignores SIGINT forever, so the bounded `exit=unconfirmed`
   refusal is the correct outcome for it; the point of C is that a complete bin/
   still reaches the normal lifecycle path with no sourcing failures.
Evidence: a real zcode-cli-named session acquiring the fleet lock, with and without the fm-path-lib.sh fixture staging, plus the plain-bash refusal

Source: a real zcode-cli-named session acquiring the fleet lock, with and without the fm-path-lib.sh fixture staging, plus the plain-bash refusal

################ A. BEFORE commit 094f8bb: the staged home carries no fm-path-lib.sh
  session process name: zcode-cli
/tmp/fm-live-zcode-lock.sh: line 32: /tmp/fm-live-lock.zSZZpB/a/state/lock.rc: No such file or directory
  fm-lock.sh exit: <never finished>
  fm-lock.sh output:
    /tmp/fm-live-lock.zSZZpB/a/bin/fm-wake-lib.sh: line 14: /tmp/fm-live-lock.zSZZpB/a/bin/fm-path-lib.sh: No such file or directory
    /tmp/fm-live-lock.zSZZpB/a/bin/fm-wake-lib.sh: line 497: fm_dirname_to: command not found
    /tmp/fm-live-lock.zSZZpB/a/bin/fm-wake-lib.sh: line 498: fm_basename_to: command not found
    /tmp/fm-live-lock.zSZZpB/a/bin/fm-wake-lib.sh: line 499: dir: unbound variable
    /tmp/fm-live-lock.zSZZpB/a/bin/fm-wake-lib.sh: line 497: fm_dirname_to: command not found
    /tmp/fm-live-lock.zSZZpB/a/bin/fm-wake-lib.sh: line 498: fm_basename_to: command not found
    /tmp/fm-live-lock.zSZZpB/a/bin/fm-wake-lib.sh: line 499: dir: unbound variable
    ... (the same three lines repeat without bound: the staged session never
        finishes, so fm-lock.sh writes no exit status and takes no lock)
  state/.lock: absent (no lock taken)

################ B. AFTER: fm-path-lib.sh staged alongside the other libs
  session process name: zcode-cli
  fm-lock.sh exit: 0
  fm-lock.sh output:
    lock acquired: harness pid 2629383
  state/.lock holder pid: 2629383  (session pid was 2629383)

################ C. adversarial: a plain bash session with no zcode ancestry
  session process name: bash
  fm-lock.sh exit: 1
  fm-lock.sh output:
    error: cannot locate harness process in ancestry
  state/.lock: absent (no lock taken)

################ reading this transcript
A: the fixture gap commit 094f8bb closed, driven live - a zcode-named session whose
   home bin/ lacks fm-path-lib.sh cannot acquire the fleet lock at all: fm-wake-lib.sh
   loses fm_dirname_to/fm_basename_to and the session spins instead of locking.
B: with fm-path-lib.sh staged, the real zcode-cli-named session acquires the fleet
   lock and state/.lock records that session's own pid.
C: a plain bash session with no zcode ancestry is refused and writes no lock file.
Evidence: crew-dispatch configuration outcomes for zcode profiles (effort refused, provider required, accepted shape)

Source: crew-dispatch configuration outcomes for zcode profiles (effort refused, provider required, accepted shape)

=== 1. zcode profile that declares an effort (zcode 0.16.5 rejects --effort) ===
$ cat config/crew-dispatch.json
{"rules":[{"when":"a docs-only fix","use":{"harness":"zcode","provider":"zai","effort":"high"}}]}
$ TYPESAFE_API_KEY=... bin/fm-dispatch-resolve.sh brief.md
error: malformed rules file: /tmp/fm-live-dispatch.JwTKcZ/home/config/crew-dispatch.json - each use profile effort must be supported by its harness and model
exit=2

=== 2. same rule with a default zcode profile carrying an effort ===
error: malformed rules file: /tmp/fm-live-dispatch.JwTKcZ/home/config/crew-dispatch.json - each default profile effort must be supported by its harness and model
exit=2

=== 3. accepted: zcode with its declared provider and no effort axis ===
dispatch-resolve: error (http 401 after 207 ms: {"detail":{"error_type":"authentication_error","message":"Cannot authenticate with the server. Please check your API key and try again."}})
dispatch-resolve:
  status: error
  reason: http 401 after 207 ms: {"detail":{"error_type":"authentication_error","message":"Cannot authenticate with the server. Please check your API key and try again."}}
exit=0

=== 4. adversarial: zcode without the required provider declaration ===
error: malformed rules file: /tmp/fm-live-dispatch.JwTKcZ/home/config/crew-dispatch.json - use profiles whose harness lacks one authoritative provider family require provider: zcode
exit=2
Evidence: rendered zcode session-start supervision block (the generated interface delivered to a zcode primary)

Source: rendered zcode session-start supervision block (the generated interface delivered to a zcode primary)

$ FM_HOME=<lab> bin/fm-supervision-instructions.sh --harness zcode

================================================================================
SUPERVISION OPERATING INSTRUCTIONS - primary harness: zcode
================================================================================
Current state:
- Lock: held by this session; this session owns normal supervision unless away mode says otherwise.
- Away/quiet mode: inactive.
- X mode: inactive; use the default watcher cadence.
- Ordinary wake: re-arm exactly one bin/fm-watch-arm.sh zcode Bash tool background task (run_in_background) as directed below.

Mode: Zcode background-notify supervision.

When this session owns supervision and away mode is not active:
1. Drain first with `bin/fm-wake-drain.sh`.
   After handling all emitted wakes and reconciling open decisions and unread status lines, run the exact `--ack-through` command printed as `WAKE_ACK_REQUIRED`; until then the work remains durable for idempotent re-handling after interruption.
2. Source `/tmp/fm-live-sup.Y8oDoR/home/config/x-mode.env` first when Relay is active.
3. First cycle: arm with the Bash tool's tracked background mechanism, as its own call with `run_in_background: true` on:

   `[ -f '/tmp/fm-live-sup.Y8oDoR/home/config/x-mode.env' ] && . '/tmp/fm-live-sup.Y8oDoR/home/config/x-mode.env'; exec bin/fm-watch-arm.sh`

4. Trust only the arm's one-line status.
5. `watcher: started ...` or `watcher: attached ...` means a live cycle exists.
   On attach, the background task follows verified identity-matched successors instead of exiting when the first cycle ends.
6. Failure or missing cycle only: `watcher: FAILED ...` means supervision is down; fix and re-arm.
7. After a successful start or attach status, end the turn.
   The background arm remains the live wait until it returns an actionable wake or failure.
8. Waiting is silent.
9. Never use shell `&` for firstmate supervision.
10. Never bundle the arm onto another command or run it in the foreground of an ordinary tool call.

Zcode re-invokes the model with a task notification when a background task completes, whether the arm exits with a wake reason or a FAILED line, so cycle completion is itself both the wake path and the failure path.
When you see a background-task-completed notification for the arm:
1. Run `bin/fm-wake-drain.sh` first.
2. Read the arm task's output for the reason line.
3. Handle `signal`, `stale`, `check`, or `heartbeat` using the harness-neutral contract in `AGENTS.md`.
4. Ordinary wake: re-arm the next cycle with the same background `bin/fm-watch-arm.sh` call if the home still needs supervision, as `bin/fm-supervision-lib.sh` defines it.
5. Do not invent a wake from an attach-status line alone.
   Drain the queue and act only on real wake records, the drain's `OPEN DECISIONS` and `UNREAD STATUS` entries, or a real watcher reason line.
   Re-arm attaches to an existing healthy cycle when one is already present and follows its verified successor chain.
   See [`watcher-continuity.md`](../watcher-continuity.md) for the arm-layer successor and clean-close failure contract.

Zcode has no turn-end guard backstop yet: the arm task's completion notification is the only cycle-end signal, so a healthy background arm must exist before every turn ends.
The crew-side Stop hook pair (`bin/fm-zcode-turnend-hook.sh`) is task-scoped and never arms primary supervision.

Interactive TUI sessions are the supported Zcode primary surface.
Headless `zcode --prompt` is the one-shot crewmate launch shape and cannot host the primary's supervision cycle.
Verified live on zcode-app-cli 3.11.2-24 wrapping zcode-runtime 0.16.5: the TUI Bash tool's background tasks survive the tool call and re-invoke the model on completion (observed in the live TUI primary session, 2026-09-15), and a session acquires the fleet lock through the `zcode-cli` engine in its own ancestry (`tests/fm-zcode-primary-live-e2e.test.sh`, opt-in, refreshes the lock and detection facts against the installed harness on demand).


exit=0
Evidence: supervision-host gating: zcode renders no host lines even on an opted-in home, claude renders four

Source: supervision-host gating: zcode renders no host lines even on an opted-in home, claude renders four

Adversarial: the lab home opts INTO the supervision host (config/supervision-host present).

$ bin/fm-supervision-instructions.sh --harness zcode | grep -ci 'supervision host'
0
$ bin/fm-supervision-instructions.sh --harness claude | grep -ci 'supervision host'
4

--- the claude host lines the zcode block correctly lacks ---
- Supervision host: on; it takes away-posture wakes and, where the dialog mirror is verified, eligible attended wakes itself, and hands the rest to you (protocol at the end of this block).
The watcher itself remains `bin/fm-watch.sh`, and `bin/fm-watch-arm.sh` remains the verified arm wrapper that the Stop hook foregrounds on a home that opted out of the [supervision host](../supervision-host.md) (`config/supervision-host-off`).
Supervision host: on for this home (`config/supervision-host-off` turns it off; [`supervision-host.md`](../supervision-host.md) owns the design).
The Stop hook runs the supervision host in the arm's place, and everything above still holds with these additions:
- Outcome: ⚠️ 1 info across 1 run (17m29s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - 1 info
  • ℹ️ tests/fm-backend.test.sh:577 - The new guard's present-case leg does not actually load the herdr adapter, so half of its control is vacuous. The fixture copies bin/ FLAT into $dir (cp -R &#34;$ROOT/bin/.&#34; &#34;$dir/&#34;, line 577) and points FM_BACKEND_LIB_DIR at $dir. That works for tmux, whose adapter sources siblings through $FM_BACKEND_LIB_DIR, but bin/backends/herdr.sh derives its own root at line 71 as cd &#34;${BASH_SOURCE[0]%/*}/../..&#34; = $dir/backends/../.. = $TMP_ROOT, then dot-sources "$FM_BACKEND_HERDR_ROOT/bin/fm-composer-lib.sh" (:79), fm-transition-lib.sh (:87) and fm-agent-process-lib.sh (:94). $TMP_ROOT/bin does not exist, so all three fail. The leg still passes only because fm_backend_source invokes the adapter as . &#34;$adapter&#34; || return 1 (bin/fm-backend.sh:658), and bash suppresses errexit for the whole dynamic extent of a command in an || list: the failed inner sources do not abort, herdr.sh runs to its last top-level statements (assignments at :3710-3711, status 0), so . &#34;$adapter&#34; returns 0 and the test writes its continuation file. Net effect: fm_backend_source herdr is asserted to 'continue the lifecycle on a complete bin/' while the adapter is in fact only partially loaded - the exact false-success the precheck exists to prevent. The absent-case leg still detects removal of either zcode registration, so the guard's protective assertion holds and no product path is affected (a real home has bin/ under its root, where both resolutions agree). Remedy is mechanical and test-only: stage into "$dir/bin" (mkdir -p "$dir/bin"; cp -R "$ROOT/bin/." "$dir/bin/") and pass "$dir/bin" as FM_BACKEND_LIB_DIR, so herdr's BASH_SOURCE-derived root resolves back into the staged tree. This is a narrow fixture-layout fix, not a reinstatement of the closure-walk machinery the round-2 decision removed.
⚠️ **Test** - 1 info
  • ℹ️ docs/agent-control.md:187 - The new completeness guard's stated failure mode does not reproduce live. docs/agent-control.md:187 and tests/fm-control.test.sh:466 both claim that an interrupt table with no row for a verified harness "aborts bin/fm-control.sh under errexit at the bare send_interrupt_keys the exit verb's second key uses - with no diagnostic". Driven live on a real tmux zcode pane with bin/fm-control-lib.sh restored to merge commit 3140240 (no zcode rows in the arm-signal, press-gap, and hazard tables), both fm-control.sh zc1 interrupt and fm-control.sh zc1 exit completed normally and stopped the agent, including the second-key path: bash does not apply errexit inside the $(do_exit) / $(do_interrupt) command substitutions those calls live in (shopt inherit_errexit is off and is set nowhere in bin/). For zcode specifically the missing rows were also value-equivalent - repeat is 1, so the empty press gap is never slept on, and the intended arm and hazard signals are empty anyway. The registrations and the owner-driven completeness test are still correct and make the tables total; only the rationale sentence overstates the consequence. No action needed on this branch; the document phase may want the wording narrowed.
  • Live validation: ✅ go - 9 of 10 scenarios driven live against the product
Scenario Result Live Evidence
A zcode task is interrupted and then exited through bin/fm-control.sh on a real tmux endpoint; the agent is stopped and the endpoint is preserved ✅ pass live control-plane-zcode-live.txt - real tmux socket, a real process with comm=zcode; interrupt returns verified=agent-alive, exit returns stopped and tmux list-windows still shows fm-zc1
The exit verb's bounded second interrupt key stops a TUI-variant zcode worker that survived the first C-c ✅ pass live control-exit-second-key-zcode.txt - the pane records C-c press 1 (turn cancelled), then press 2 and the process exiting; fm-control reports stopped
The zcode interrupt postcondition follows the recorded launch variant: a headless worker ending is success, the opt-in TUI variant must stay alive ✅ pass live zcode-interrupt-variants-live.txt - verified=agent-ended-by-interrupt with no zcode_tui line, verified=agent-alive with zcode_tui=1
Adversarial: a task recording an unverified harness is refused with a diagnostic rather than acted on with guessed mechanics ✅ pass live control-plane-zcode-live.txt section C - error: task zc1 records harness &#39;someagent&#39;, which has no verified control mechanics, exit 1
Adversarial: a firstmate home whose bin/ is missing fm-zcode-lib.sh refuses the lifecycle before touching the endpoint instead of running on a partially loaded library ✅ pass live backend-sibling-precheck-live.txt - merge-commit list: 27 sourcing failures then keys delivered and an unconfirmed exit; this branch: one `endpoint reads 'unverified' ... refusing to send a lifecycle…
A real zcode-cli-named primary session acquires the fleet lock from a staged home bin/, and a session with no zcode ancestry is refused ✅ pass live zcode-fleet-lock-live.txt - lock acquired: harness pid &lt;session pid&gt; with state/.lock matching, exit 0; plain bash gets cannot locate harness process in ancestry, exit 1 and no lock file; without…
A crew-dispatch config that gives a zcode profile an effort is an actionable configuration error before any request, in both rule and default position ✅ pass live dispatch-resolve-zcode.txt cases 1 and 2 - each use profile effort must be supported by its harness and model and each default profile effort must be supported by its harness and model, exit 2 eac…
Adversarial: zcode must declare its provider, and a zcode profile with provider and no effort is accepted ✅ pass live dispatch-resolve-zcode.txt cases 3 and 4 - provider-less zcode exits 2 naming zcode; with provider and no effort the config validates and the tool proceeds to its request path
A zcode primary's session-start supervision block no longer claims the supervision host, while a claude primary on the same home still renders the host protocol ✅ pass live supervision-block-zcode.txt and supervision-host-gating.txt - the rendered zcode block reads "Interactive TUI sessions are the supported Zcode primary surface" and matches 'supervision host' 0 times e…
The reference-doc and configuration-doc wording changes (docs/agent-control.md interrupt-table paragraphs, docs/configuration.md wait-shape pointer, .agents/skills harness reference) ⏸️ untested no These three files have no runtime consumer: unlike docs/supervision-protocols/zcode.md (which bin/fm-supervision-instructions.sh renders into a live session and which I did drive), they are read by hu…
  • bash tests/fm-zcode-harness.test.sh (23 cases, all pass)
  • bash tests/fm-control.test.sh (47 cases, includes the new test_every_verified_harness_answers_each_interrupt_table)
  • bash tests/fm-backend.test.sh (includes the new test_backend_source_requires_the_zcode_sibling)
  • bash tests/fm-dispatch-resolve.test.sh (includes the two new zcode effort rows)
  • bash tests/fm-supervision-instructions.test.sh (includes the reworded zcode TUI assertion)
  • bash tests/fm-turnend-guard.test.sh (89 cases, includes the fm-zcode-lib.sh fixture staging)
  • bash tests/fm-session-lock-ancestry.test.sh - aborts on this host at the upstream pid-1 reparent assertion (already-declined environment issue); its zcode lock cases were driven live instead
  • Live: real tmux + a real zcode-comm process, bin/fm-control.sh zc1 interrupt and bin/fm-control.sh zc1 exit, run against both this branch's bin/ and merge commit 3140240's bin/fm-control-lib.sh
  • Live: same lab with harness=someagent recorded, to confirm an unverified adapter is still refused loudly
  • Live: bin/fm-control.sh zc1 exit on a staged home bin/ missing fm-zcode-lib.sh, against 3140240's fm-backend.sh sibling lists and this branch's
  • Live: real zcode-cli-named session running bin/fm-lock.sh from a staged bin/ with and without fm-path-lib.sh, plus a plain-bash negative case
  • Live: bin/fm-dispatch-resolve.sh with four crew-dispatch.json shapes (zcode+effort in use, zcode+effort in default, zcode+provider, zcode with no provider)
  • Live: bin/fm-supervision-instructions.sh --harness zcode and --harness claude on a home that opts into config/supervision-host
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

Rangezi and others added 30 commits September 14, 2026 11:23
… asked, not declined (kunchenguid#4387)

* fix(bin): read Claude Code's default external-imports flags as never asked, not declined (kunchenguid#4378)

fm-claude-trust.sh refused the whole trust registration whenever the project-root entry
carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code
writes that value only on an explicit "No, disable". Claude Code's default project
entry carries Approved and WarningShown both false before the dialog is ever shown, so
every such project refused every spawn.

Only Approved === false with WarningShown === true — the pair the dialog writes on a
decline — now counts as a decline. false/false behaves like an absent flag: trust is
registered and no import consent is manufactured.

New case test_project_root_entry_default_import_flags_are_not_a_decline fails on
b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31,
bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(review): Correct harness doc's external-imports decline predicate

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…chenguid#4445)

* fix(brief): keep operator address out of composed intent

Teach raw-word authoring for intent sections and mid-task relays, with a neutral [captain] provenance marker for legacy mixed tasks. Keep headings and contract prose outside the serialized intent body.

The legacy selector already excluded the old speaker labels from its output; preserve that read compatibility. The reproduced leak comes from adding labels inside a modern intent body, not from the legacy selector. Do not scrub actual request content.

Add exact serialized-input and generated-contract regressions, retaining refusal of unmarked legacy tasks and coverage of scout promotion.

Fixes kunchenguid#3882

* no-mistakes(review): Refuse operator-address lines in Captain's intent body

* no-mistakes(document): Document operator-address refusal in intent contract comments
…as a proven empty composer (kunchenguid#4455)

* fix(composer): accept Grok title overhang

* no-mistakes(review): summary: named Grok overhang constant, doc caveat, restored tmux typed-title coverage
…ailure (kunchenguid#4474)

* fix(bin): recover Claude auto-arm after timeout

* no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list
* fix(spawn): establish Claude task channel authority

* no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference
…or pending text (kunchenguid#4458)

* fix: guard relaunch exit against pending input

* no-mistakes(review): Verifying test run in progress

* no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard

* no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now
…unchenguid#4460)

* fix: reconcile diverged secondmate updates

* no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md

* no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md
…d#4497)

* fix(dispatch): support Codex Luna max effort

* no-mistakes(review): use portable CODEX_HOME path in codex effort reference
kunchenguid#4498)

* feat(calm): render smooth Unicode swell

* feat(calm): make sails asymmetric

* feat(calm): use quarter sail glyph

* no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer

* no-mistakes(document): docs: sync calm wave phase doc comment

* no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。
kunchenguid#4491)

* fix: supersede scout delivery brief on promotion

* fix: preserve ship safety contract after promotion

* no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch
…d stop cleanup dropping accents from a held body (kunchenguid#4471)

* fix(bin): let captain holds work on hosts with an older JSON::PP

Holding a task for the captain, and the cleanup that keeps a captain-held row
open, both fail outright on any host whose JSON::PP defaults allow_nonref off -
2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`,
but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older
library rejects that whole value with "must be object or array".

The consequence is fleet-wide on such a host, not one broken command: a worker
there cannot formally record a decision for the captain at all. It can only
mention the decision in passing in a status line, where it can be missed - which
is how a real decision goes unrecorded. The hold reports that the task lost its
hold-set stamp; the cleanup cannot return the row to Queued.

Both call sites now ask for allow_nonref explicitly rather than inheriting
whatever the installed library defaults to. The second one is worth naming: its
`/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string
case that fails, so the guard selects for the failing input rather than
protecting against it.

The regression case forces the older default back off for every perl the commands
spawn, then drives both paths - holding a task that carries a body, and tearing
down a captain-held row whose deliverable must still be appended. It also probes
that the simulation genuinely rejects a bare scalar, so the case cannot pass
vacuously on a lenient host. Each half was verified failing on its own unfixed
call site with that site's real error message. Suites: fm-captain-hold-lifecycle
51 cases, fm-backlog-atomicity 99 cases, 0 failures.

Verification limit: the mechanism is reproduced and tested, but neither fix is
verified against a real JSON::PP 2.27202 host, because none is in the loop. This
laptop runs 4.06, where the bug does not manifest.

`bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a
brace-delimited object before decoding, so allow_nonref never applies.

* fix(bin): stop cleanup silently dropping accented characters from a held body

Cleanup rewrites a captain-held row's body to append the finished work's
deliverable, and the decoder it reads that body with printed decoded characters
to a stream with no `:raw` layer. A character at or below U+00FF then came out
as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the
accent. `fm_backlog_retain` writes that body straight back through
`--body-file`, and nothing reported an error - the character was simply gone
from a row still waiting on the captain.

The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus
`utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already
used.

Review of the parent commit found this on one of the lines that commit already
changed. It predates that change.

The test asserts bytes rather than decoded strings, because comparing strings
cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any
character above U+00FF makes perl print the whole string as UTF-8, so one body
carrying both an accent and an em dash passes even unfixed and proves nothing.
Verified failing before the fix on the accented row, passing after. Suites:
fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures.

* no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc

* no-mistakes(review): drop whole-file UTF-8 check from retained-body test

* no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc

* no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc
…furniture (kunchenguid#4532)

* fix(composer): read codex 0.154's idle starfield and status footer as furniture

codex-cli 0.154.0 animates a braille "starfield" around its idle composer:
on the row above the bold `›` prompt row, on the `›` row behind the SGR-2
dim `Ask Codex to do anything` placeholder, and on the row below it, then
draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`).
The cells are truecolor greys on both sides of the ghost luminance ceiling,
so the brighter ones survive ghost stripping, and the rows below the glyph
carry no structural edge. The shared classifier selected the bare `›` shape,
extended its wrap region over the two rows beneath the glyph, read the
survivors and the footer as wrapped typed input, and answered `pending`;
the steering doorbell defers on exactly that verdict, so no doorbell ever
reached an idle codex 0.154 pane.

bin/fm-composer-lib.sh now recognises that furniture by shape, declared
once next to the idle placeholders and reached from the two wrap-region
boundary points:
- a row whose non-whitespace content is entirely braille cells
  (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it
  never counts as wrapped typed content and bounds a bare composer's wrap
  region; braille behind the glyph row's content is stripped before the
  emptiness decision when nothing else follows the glyph; a row mixing
  braille with other text stays typed content;
- the codex status footer bounds the wrap region exactly as omp's status
  row does, anchored on the effort token, a spaced middle dot, and a `~` or
  `/` path cell, so a typed `fix · tests` stays composer input;
- `^Ask Codex to do anything$` joins the verified idle-placeholder set; the
  ghost strip remains what proves that row empty, and the bare-row rule that
  bright placeholder text is real input is unchanged.

Unchanged: the strict blank-row rule, the styled=0 degradation (a plain
cmux/orca capture of this screen still reads `unknown`, never `pending`),
FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape.

tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte
with the divergence (letters in place of the starfield read `pending`) and
the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh
is the default-on live guard (token-free, skips explicitly without codex or
tmux) that launches the installed codex idle and asserts `empty` through
both the tmux and the cursorless styled reads, naming codex --version on
failure. docs/verification/runtime-backends.md records the dated Herdr
evidence: `pending` before, `empty` after, on the captured screen.

* no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry

---------

Co-authored-by: Todd Billings <todd@usdvcapital.com>
* fix(bin): refuse empty text steers in fm-send

A marked secondmate request sent with an empty message delivered only
marker and correlation bytes and minted a pending-reply expectation the
parent could never see resolved, stalling the fleet with no loud error
(kunchenguid#4255). Fail closed on an empty or whitespace-only message on the text
path, mirroring the existing --resolve-key refusal.

* chore: retain ambient Pi-lens autoformat as its own commit

Formatting-only edits produced by ambient Pi-lens autoformat during the
msg-loss investigation, kept separate from the behavioural change in
c23acba so the fix stays reviewable on its own.

AGENTS.md is deliberately excluded: its only autoformat edit stripped the
trailing space from the documented FM_OPERATIONAL_PREFIX value, which
bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11
records as permanent compatibility. Documenting that constant without its
trailing space makes the doc wrong about the contract, so that one line was
restored rather than retained.
…chenguid#4554)

On rose-pine-moon the two-color water (cyan crests over blue troughs) read as
a pink stripe over aqua, the yellow left sail and mast clashed with the red
right sail, and the hull carried a blue interior run. Every water cell is now
blue so the swell reads through glyph height alone, and both sail halves, the
mast, and the whole hull are one yellow run. Geometry, cadence, animation,
direction flip, resize clamping, and the narrow fallback are unchanged.

Update the unit and real-TUI color assertions to the new palette and the Calm
docs that described the old one.
…chenguid#4270)

* fix(watch): stop aging a second mate's active turn from its launch

The parent watcher's second-mate wake-loop stall check exempts a mate that
is demonstrably inside an active turn, but secondmate_in_active_turn asked
busy_turn_over_age first and returned "not in a turn" whenever that said
the bound was crossed.

busy_turn_over_age ages from state/<task>.turn-ended, falling back to
state/<task>.meta. A second mate's turns end in its own home, so the
parent never gets a turn-ended mark for it and the fallback ages the
mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago
was therefore permanently "over age", the busy pane was never consulted,
and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false
wake-loop stall.

The gate now bounds the busy exemption by <idle> - how long the queue's
drain position has not moved - which is evidence this home actually
holds. A busy mate stays exempt while the queue has been frozen for less
than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so
the bound that stops a busy pane from proving liveness forever is kept
rather than removed. busy_turn_over_age is untouched; its remaining
callers are the ordinary crew busy-pane bound.

The regression pins the case that actually broke: a mate whose launch
record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn
must not escalate, while the same mate with its queue frozen past the
bound still publishes exactly one notification. The existing coverage
only exercised a freshly launched mate, which passes either way.

Reaching that alert now costs a pane capture inside the gate, so the
three checkpoints in this suite that assert an alert move from a 1s to a
4s bound - the value the neighbouring active-turn cases already use. The
bound is a ceiling, not a wait: the checkpoint returns on the first
actionable wake. On a loaded machine a 1s bound missed the alert
repeatedly; at 4s it did not miss in 20 runs under the same load.

* no-mistakes(review): scope the second-mate active-turn regression test's coverage claim

* no-mistakes(document): fix stale second-mate active-turn comments in fm-watch
…unchenguid#4278)

* feat(bin): add read-only PR blocker and reviewer-discovery commands

Two focused, opt-in commands that read GitHub and never write to it.

fm-pr-state.sh reports what still blocks one pull request from the
author's side: a closed or merged state, draft state, unknown or
conflicting mergeability, absent or failing required checks, and a
blocking CHANGES_REQUESTED decision explained by each reviewer's latest
verdict, marked STALE when it was left at a superseded head. A pull
request that only awaits an approval is not reported as blocked, and
advisory checks are omitted. Every reading is taken against one exact
head; a push that lands mid-read invalidates the whole result rather
than mixing two snapshots.

fm-pr-reviewers.sh suggests reviewers from the most recent commits to
the pull request's exact changed paths, counting each commit once,
resolving handles through GitHub's own commit author.login mapping, and
excluding the author and Bot accounts.

Both stay read-only: no review request, no approval, no merge.
Unresolved review-thread state is left unreported because the REST API
does not expose it and unattended commands may not use GraphQL.

Closes kunchenguid#3731

* no-mistakes(review): accept only PR URLs and stop at terminal state

* no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate

* no-mistakes(review): stop attributing readings to unverified heads

* no-mistakes(review): narrow readiness contract to checks that have reported

* no-mistakes(review): read the pull request once, drop the head guard

* no-mistakes(document): scope pr-forge isolation proof to its measured members

* no-mistakes(document): record uncovered pr-forge members and their pending proof

* docs(isolation-proof): re-prove pr-forge at its full membership

tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the
pr-forge family in this branch, and script_allows_concurrency grants
four workers by family membership alone, so both ran concurrently on a
proof measured before they existed.

Re-proved the family at all eight members: two consecutive runs, 0
failures, each begun with the one-minute load average below 6.0 so the
result measures isolation rather than contention. A third run taken
between them is disclosed rather than recorded, because it started
while the previous run's workers were still decaying.

The new durations are not comparable with the six-member measurement
above them, so they are not presented as evidence about the two new
members, and that record's 1.72x four-worker figure is left as a
statement about its own run rather than restated as current.

* no-mistakes(review): disclose gh error-text coupling at its matching site and tests
…uid#2752)

* fix(bin): teach validation-round pauses in briefs

* no-mistakes(document): Point classifier comments to authoritative pause examples
…guid#4510)

* fix(teardown): refuse a cleanup whose endpoint close failed

bin/fm-teardown.sh discarded both the exit status and the stderr of every
fm_backend_kill call, so a close that genuinely failed was indistinguishable
from one that succeeded. Teardown continued past it, deleted the task's durable
records, returned its worktree, and reported the cleanup as completed. The
deleted metadata is the only record of which endpoint belongs to the task, so
such a close did not merely leave a stray session behind, it stranded one:
nothing was left on disk naming it.

The adapters could not carry that signal either. Driven against the real code,
every backend arm returned 0 for a genuine failure exactly as it did for an
already-exited endpoint, so there was nothing for the four call sites to
propagate even once they stopped swallowing it.

The tmux arm now resolves a close that did not succeed against the window's
exact recorded identity, since kill-window fails the same way for a window that
is gone and one that is still there. The Orca arm reports a close its missing
CLI never attempted. Both stay silent for an endpoint that is already
legitimately gone, and the remaining arms are unchanged: their close-command
timing cannot be established without the real Zellij, Orca, and cmux binaries,
and a gate that refused ordinary cleanup of an already-exited session would be
worse than the defect. docs/verification/runtime-backends.md records what each
backend can prove.

A reported close failure now reaches teardown's existing retain-and-stop
refusal before the records naming the endpoint are removed, matching where the
Herdr confirmed-gone gates already sit for the same hazard, and the retained
records let a rerun finish once the close works.

* no-mistakes(review): refuse unreadable tmux close re-read; honor --force override

* no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close

* no-mistakes(document): document endpoint-close refusal in its backend and retirement owners

* no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree
* feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag

Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that
brings Calm to Claude Code: the sailboat replaces the stock working row through a
Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration,
and canonically classified operational user rows draw at zero height. /calm is
registered by the hooks module itself and toggles the same per-home config/calm
preference the Pi extension uses, so one choice applies on either harness; rows
redraw retroactively on toggle and stay hidden across claude --continue.

The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS
flag is on. Nothing sets that flag in any settings file, and the plugin carries no
command file, skill, agent, or classic hook, so it is a complete no-op while the
flag is off. The trusted project auto-loads it through an .agents/skills symlink,
the only path Claude Code scans for project plugins.

Extract the working-ship geometry, bounce track, cadences, and freeze/resume state
into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module
imports from outside the plugin folder) and have the Pi widget paint that core's
frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify
operational rows through a port of bin/fm-operational-input.sh's classify command
guarded by a corpus parity test against the shell owner.

Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering,
Raster packing, policy, classifier parity), the mod's own claude plugin test suites
behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op,
the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272.

Docs: record the version-scoped Claude Code evidence and the three bounded gaps in
docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md,
and make the shared preference, layout, and contributor notes harness-neutral.

* no-mistakes(review): Preserve colliding final replies and strengthen parser parity

* no-mistakes(review): Preserve final replies and strengthen canonical parity checks

* no-mistakes(review): Require exact function-hooks opt-in before Calm activation

* no-mistakes(review): Clarify Calm module loading and activation boundaries

* no-mistakes(review): Reset Calm presentation state across session starts

* no-mistakes(document): Refresh Calm session lifecycle documentation

* feat(calm): paint the Claude Code working ship in Claude's own theme colors

The captain picked the "Claude native" palette for the Claude Code mod's Raster:
every water cell takes the spinner blue of the active theme family (#93a5ff dark,
#5769f7 light) and the whole boat takes the Claude orange of the stock spinner
(#d77757), one water color and one boat color. The family follows the `theme`
setting's prefix, read at load through $.config.list and re-read on a
config.set of that row, with `auto` and custom themes falling back to the dark
set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte.

Rename the shared sprite's color classes from hue names to `water` and `boat`,
since each harness now maps them to its own colors; geometry, motion, cadence,
and the activation gate are untouched.

Tests cover both palettes' packing and the family rule under Node, and the
plugin kit drives every theme value, a theme change mid-session, the Calm-off
pass-through, and inertness of the menu read while the flag is off. The docs
describe the Claude Code colors and record the guard passing on 2.1.273.

* no-mistakes(review): Use light palette for unresolved Claude themes

* no-mistakes(document): Refresh Claude Calm verification evidence
…kunchenguid#4586)

* fix(watch): honour a declared wait before wedge-escalating a quiet pane

wedge_timer_check escalated on elapsed idle time alone. Nothing asked
whether the worker had already said why its pane was quiet, so a lane
that declared a bounded external wait climbed the escalation ladder for
as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every
repeat carried demand-deep-inspection - which by its own wording forbids
re-absorbing on the run-step or pane state, so the supervisor could not
use the evidence that was there either.

The generated brief promises that declaring `paused:` buys the long
recheck cadence instead of a wedge, but the timer was still reachable
while that declaration stood: a crew that declares a wait and then has an
active run or busy pane attributed to it is handed to the timer as
provably-working. The declaration is what the worker said about its own
silence, so it now outranks a liveness verdict that only says something
is running.

The consult runs in the at-threshold branch that was about to escalate,
beside the worktree walk already there, and costs one status-line read.
Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS
recheck the declared-wait absorber already uses, so the wait is still
rechecked and cannot rot invisibly. Which verb declared it decides the
wording, because the two block on different people: a `paused:` wait is
owed by an external dependency and asks the reader to confirm it still
holds, while a `captain-held:` transfer is owed by the captain reading
the recheck and asks them to answer or release the hold. A hold is not
rechecked at all while the away-posture record exists, as on every other
captain-held path, and that absorb arms no throttle so the recheck is
owed in full on return.

A declared clearing time that has already passed stops counting, and a
lane that never declared one keeps the identical escalation schedule,
reason, count and demand-deep-inspection wording, so detection and its
worst-case time are unchanged. The deferral restarts the idle timer
rather than cancelling it, so a lane that stops waiting escalates again
within one threshold.

A lane quiet because its own validation run is parked at a gate awaiting
a human decision is deliberately out of scope: reading that state needs a
signal carrying who the wait is on and what clears it, rather than one
inferred from a parked verdict that also covers gates awaiting the
crewmate itself.

Tests pin both directions for each case and were each confirmed to fail
with the consult removed.

* no-mistakes(document): docs: honour declared waits in stale-escalation docs
* fix(bin): derive passed PR state from PR record

A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists.

For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim.

Fixes kunchenguid#4607

* no-mistakes(review): Add bounded GitLab merge-request state reads

* no-mistakes(review): Preserve network-free inactive crew-state scans

* no-mistakes(document): Document PR record readers in shared library
kunchenguid#4627)

* fix: restore published contribution follow-up (Fixes kunchenguid#4469)

* fix(review): Fix contribution freshness and merge actor routing

* fix(review): Restore issue triage and scope contribution follow-up

* fix(test): test: assert one wake per contribution signal

* fix(document): Document contribution follow-up

* fix: restore truthful terminal delivery evidence

* fix(review): Disclose unsupported contributions and deduplicate watcher wakes

* fix(review): Preserve unmeasured unsupported contributions across Bearings

* fix(review): Deduplicate shared contribution wakes and isolate diagnostics

* fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower
…nguid#4658)

* fix(bin): make a remote-reply document gap self-clearing and re-attemptable

A remote mate's undelivered document raised a keyed `blocked` decision that
nothing could ever resolve, and any `data/*.md` substring in any mirrored line
was an unconditional fetch instruction. A mate announcing a report it had not
written yet therefore manufactured a permanent, factually false blocker, and
its own explanation of the false alarm manufactured more.

The reader has no permanence vocabulary: a report still being written refuses
exactly like a path that will never exist. So an undelivered document is now a
durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`,
re-attempted on the next delta and on the channel's own quiet poll, and retired
with a matching `resolved` line naming the local copy once it arrives. The
cursor still advances and no delta stalls on one bad pointer.

Only a structured `report=data/....md` pointer now offers a document, so a path
merely mentioned in prose - including one under another home's mirror tree,
which is provably not that mate's to serve - is never fetched. Offers are
deduplicated across the whole delta, the escalation names each missing document
once and carries the reader's own reason instead of discarding it, and a
strictly increasing notice ordinal keeps a later escalation from being
swallowed as duplicate bytes. A mirrored line still lands once whichever
pointer form it was first written under.

* no-mistakes(review): Require structured pointer token boundaries

* no-mistakes(review): Unify boundary-safe pointer extraction and rewriting

* fix(bin): identify a mirrored line independently of its delivery state

Two defects in the boundary-safe pointer work.

The at-most-once check compared only the all-remote and all-local renderings
of a line, so it could not recognize a mixed one. A line offering two documents
where only the first was deliverable mirrored as local-plus-remote; once the
second arrived, a cursor-loss whole-log recapture rendered the same line
all-local, matched neither alternate, and mirrored a second time. A line's
identity is now the canonical form every boundary-valid pointer would take once
delivered, derived by the same parser that does extraction and rewriting, so it
no longer depends on which documents happened to be deliverable at the time.

The pointer map was passed to awk through the process environment. A delta may
carry up to the configured 1 MiB bound, and an expanded map of delivered
pointers can exceed the platform's exec argument limit, so awk would fail to
start; because no caller checked, the empty result would have been appended as
blank lines while the cursor advanced past dropped status content. The map now
travels in a file, and every call site checks the exit status and stops the
ingest rather than committing a delta it could not render.

Both passes now run once per stream instead of twice per line.

* no-mistakes(review): Abort ingest when document pointer extraction fails

* no-mistakes(review): Exclude structured cross-home pointers from document transfer

* fix(bin): fail open on an undeliverable remote document instead of tracking it

Narrow the remote-reply document fix to the scope the diagnosis actually
requires, as decided after measuring a simpler alternative.

A document the reader cannot deliver now fails open. The mate's line is
mirrored with its own pointer, the cursor advances, and one unkeyed note
carries the reader's reason. A note never enters the open-decision fold, so it
cannot stand open the way the original keyed block did - which removes the
never-clearing false blocker by construction rather than by resolving it.

That makes the durable self-clearing obligation unnecessary, so it goes: the
per-mate pending-documents record, its notice ordinal and resolved
announcements, and the poll-side retry. Canonical line identity goes too, and
with it a way to silently drop a genuine status line; mirroring is back to
at-most-once on exact bytes. The cross-home exclusion goes as well: under
fail-open a cross-home report= either fails harmlessly or is a nested remote
report this mate genuinely holds, which is now relayed again.

Kept: fetching only on a structured report= pointer, the boundary-correct
parser, the file-based rewrite map, and checked extraction and rewrite exit
status. The parser now scans behind a sentinel byte so a rejected candidate can
no longer give the text right after it a false leading boundary.

The reported incident is covered end to end: a report path announced in prose
before it exists raises no decision, and the report still arrives through the
ledger publisher's structured offer once written.

* no-mistakes(review): Preserve source-line identity across remote reply replays

* no-mistakes(document): Document remote reply transfer and replay semantics

* no-mistakes(lint): Fix staging truncation lint checks
* Preserve substantive Calm mid-turn text

* no-mistakes(review): Distinguish newline-preserved replies from short narration

* no-mistakes(document): Document Calm mid-turn preservation boundaries

* no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`
…#4656)

* fix(bin): re-record PR poll identity after a volume device renumber (Fixes kunchenguid#4260)

A volume remount can renumber the state filesystem's st_dev while every
inode and byte stays the same; APFS does this across a reboot. A poll
registration records its sidecar and check as device:inode, so every poll
armed before the remount failed strict validation and the watcher refused
all of them as unauthenticated state checks until each was re-armed by hand.

There are two device comparisons. fm_pr_private_file_valid compares a live
file's device with the state directory's device read in the same invocation:
it refuses a file that is not on the state directory's own filesystem and
already survives a renumber, so it is unchanged. The registration's recorded
identity versus the live identity (from kunchenguid#556, reused by the kunchenguid#932 retirement
receipt) binds the registration to the exact files published in its own
transaction; its device part is what breaks.

When strict capture fails, the watcher now proves the device is the only
difference: every other artifact check passes (template bytes, both hashes,
private mode, single link, live device, metadata), both recorded identities
name one device, and each recorded inode equals its live inode. Only then,
under the task's control lock, does it rewrite the two identity lines,
repeating the whole proof and comparing the registration's file identity and
bytes just before the rename, and then capture strictly again. A swapped,
altered, re-moded, relinked, split-device, or foreign-device artifact still
fails a proof and is still refused, and a pending retirement receipt blocks
the rewrite.

Reproduction: on macOS a poll armed on an APFS disk image that was detached
and re-attached behind another image moved st_dev 16777239 -> 16777243 with
inodes, bytes, mode, and link count unchanged; the real watcher refused it on
main and reports its merge with this change. The portable regression test
rewrites a real registration's recorded device and drives the watcher.

Not changed here: the status presentation cursor keys rows by its own
device:inode identity in bin/fm-classify-lib.sh, a different helper that
needs its own fix; a retirement receipt left by a reboot between its
publication and removal still names the old device and stays refused; custom
check trust binds only a content hash and is unaffected.

* fix(review): Serialize PR poll publication writers

* fix(review): Bound PR poll publication lock scope
…llow-up to kunchenguid#4627) (kunchenguid#4661)

A budget that expires partway through an observation no longer records an
error or prints the unavailable wake; the URL keeps its prior record and is
observed first next poll. forge() flags budget exhaustion at the point it
refuses, or when a read is killed at the budget's own deadline, so a genuine
forge failure still records the error and wakes. Each distinct URL is now
observed once per poll and applied to every owning task.
…kunchenguid#4680)

* fix(bin): clear parent pending-replies on local secondmate retirement

Local secondmate teardown left resolved parent pending-reply records behind
after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced
retirement while any reply for that id is still unresolved, and delete every
matching record plus its delivery confirmation after a successful local or
remote retirement, matching the remote cleanup path.

* no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup

* no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung

* no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern

* no-mistakes(review): Pending-replies Basename und corr_id abgleichen

* no-mistakes(document): Clarify forced retirement pending-reply cleanup

---------

Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>
kunchenguid and others added 25 commits September 28, 2026 20:18
…kunchenguid#6043)

* fix(bin): name an in-window engine latch in the return brief and drop the false handling GAP line

The away return brief said nothing had failed after the supervision host
latched on engine errors during the window, and printed a GAP: watcher
downtime line whenever a wake was merely being handled or queued at return.

The failures section now reads the host ledger and latch record and names
the latch time, the window's engine-error count, and whether the session
is still paused or recovered. An open recovery episode is reported as
information, and as a gap only when a queued episode outlived the return
grace or the marker cannot be read.

* no-mistakes(review): Fix latch trip time, drop marker-age grace, bound error count

* no-mistakes(review): Report paused latch without ledger trip row; bound errors

* no-mistakes(review): Never report a failed probe's latch row as trip time

* no-mistakes(review): Only a retained trip row marks a pre-window latch

* no-mistakes(document): Clarify return-brief latch and watcher-gap documentation

* no-mistakes(ci): Fixed Lint 1 by marking the shared cooldown constant as used by sourcing scripts. The repository lint command and diff check pass; the return test run was stopped by a 180-second timeout after its completed cases passed

* no-mistakes(ci): Fixed the return brief so the trip time and error count come from the same initial latch row, and ledger rows before the current session’s lock boundary cannot affect its latch report. Added real-script regressions for both findings. The return test suite, repository lint, and diff check pass

* no-mistakes(ci): Fixed the return brief’s restart cutoff so it retains in-window failures, prints one line per initial-trip row, and omits zero-error count wording. Added real-script restart regressions. The return test suite, ShellCheck, and diff check pass

* no-mistakes(ci): Fixed the return brief so a recorded trip followed by recovery stays recovered, while a later pause with a lost trip append gets a separate “trip time unavailable” line. Added a real-script regression that failed before the fix. The return test suite, ShellCheck, syntax checks, and diff check pass

* no-mistakes(ci): Fixed the false second latch during recovery. A real-script regression failed before the fix and passes now; the lost-second-trip test still passes. The return test suite, ShellCheck, syntax checks, and diff check pass
* fix(calm): name the Claude Code Calm plugin fm so supervision notes read "fm: "

Claude Code labels every mod transcript line with the plugin name, so the
notes rendered as "firstmate-calm: ⚓ ...". Rename the plugin to fm, update
the live guard to assert the fm: label, and document the one-time replay for
sessions resumed across the rename.

* no-mistakes(document): Clarify Calm plugin rename in documentation
…#6037)

* feat(bin): add fm-live-lab.sh, a one-command live supervision lab builder

* fix(bin): exact lab windows, per-lab task ids, self-safe teardown

* fix(bin): target lab windows by id, stop lab descendants, add readiness tests

* fix(bin): keep Claude's auto-updater off in live labs; list fm-live-lab.sh

* fix(bin): start the lab tmux server without user config

* no-mistakes(review): Scope lab teardown to its store, root, and task ids

* no-mistakes(review): Record selected user stores at up for check and down

* no-mistakes(document): Clarify live lab documentation and remove stale narratives

* no-mistakes(ci): Fixed the CI failure by checking for an existing lab root before looking up the harness executable. The affected behavioral test and shell syntax check pass; the refusal also works with Claude absent from PATH

* no-mistakes(ci): Fixed all four Greptile findings: teardown signals only recorded lab processes and their descendants; the worker gate is in its granted task directory and its path is exposed; readiness uses current crew state; and mate and worker IDs use 12 nonce hex digits. The CLI behavior tests pass, as do shell syntax, ShellCheck, and diff checks. The Claude no-host path is unchanged

* no-mistakes(ci): Fixed the CI test’s dependence on an installed Claude binary by supplying a test-local stub. The full fm-live-lab test, shell syntax check, and diff check pass

* no-mistakes(ci): Fixed all three selected findings in bin/fm-live-lab.sh: down waits for recorded processes and escalates before cleanup, PID roots are checked against recorded start times, and Claude primary trust is rechecked after mate/worker readiness. Added behavioral tests in tests/fm-live-lab.test.sh. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass

* no-mistakes(ci): Fixed the pre-primary settle wait, worker gate instructions, unused retry variable, and teardown PID revalidation in bin/fm-live-lab.sh. Added behavioral tests in tests/fm-live-lab.test.sh. Both requested commands pass: tests/fm-live-lab.test.sh and bin/fm-lint.sh

* no-mistakes(ci): Fixed teardown to track pre-kill lab processes by PID and start time, including children orphaned when a root exits. Up now rejects an empty pane PID before calling ps. Added regression tests and a Linux-safe worker fixture. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass

* no-mistakes(ci): Fixed teardown tracking for children spawned during shutdown and made the worker fixture verify its exact window with a Linux-available shell. Both requested checks pass. The lab test takes about 66 seconds locally, so the under-one-minute target remains unmet

* no-mistakes(ci): Fixed ci-2 and ci-4 in bin/fm-live-lab.sh and tests/fm-live-lab.test.sh. Teardown now tracks identity-checked members of captured lab process groups, including children orphaned during shutdown, without signaling the caller’s group or unrelated processes. Lint passed, and the lab test passed four times

* no-mistakes(ci): Fixed teardown so an observed-empty process group is permanently dropped, preventing a reused group ID from signalling unrelated work. Added a ps-shim regression test. The lab test, lint, and diff checks pass

* no-mistakes(ci): Fixed ci-1 in bin/fm-live-lab.sh and tests/fm-live-lab.test.sh. The TERM-born-child fixture now waits until its handler is installed before calling down. Down sends SIGKILL to identity-valid survivors on every pass from pass 20 onward and includes survivor process details if it must refuse cleanup. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass locally; Linux CI remains to be verified

* no-mistakes(ci): Fixed down’s teardown wait to require two empty identity-checked scans separated by 0.5 seconds, and removed the unused test loop variable without changing the TERM-born-child test. The lab test, lint, and diff check pass locally
…unchenguid#6103)

* fix(bin): keep slow watcher cycles and preempted reply polls from breaking supervision

- fm_pending_reply_tick selects the records it has work for in one awk pass,
  so settled records cost no lock or fork and the walk no longer grows with
  the never-pruned store.
- An attached arm keeps following a live, identity-matched holder whose beacon
  went stale until the lock changes or the shared stall bound
  (fm_watcher_stall_bound), then fails with a typed stalled-holder line so the
  retry replaces the holder.
- The remote-reply adapter reports the job worker's preemption (exit 76) as a
  closed window, so the listener keeps its claim and polls again instead of
  being relaunched every watcher cycle.

* no-mistakes(document): Clarify watcher grace and attached-arm documentation
…geable is UNKNOWN (kunchenguid#6110)

* fix(bin): retry a bounded number of times when GitHub mergeable is UNKNOWN

Fixes kunchenguid#6020

bin/fm-pr-merge.sh refused a GitHub merge whenever the pull request's
mergeable field was not literally MERGEABLE. GitHub reports UNKNOWN for
a short while after a push or a base-branch change while it recomputes
mergeability, so a green, conflict-free pull request was refused as if
it could not be merged.

github_verify_mergeable now returns a distinct status when mergeable is
the only failing condition and reads UNKNOWN. The caller retries up to
5 times, 3 seconds apart (overridable in tests), re-reading and
re-checking every live condition on each attempt. Once the bound is
spent it reports mergeability as still being computed rather than
unmergeable, with the same nonzero exit as before. Every other refusal
(closed, draft, conflicting, red or missing checks, away authority,
queue protection) is unchanged and never retried.

* no-mistakes(ci): I fixed both review findings the way you asked. The full suite (`bash tests/fm-pr-merge.test.sh`) ran to completion. Its last lines showed all `ok`, and any failure would have stopped the run early. I watched the output through `tail`, so I didn't see the new test's own `ok` line directly. **ci-2 (`bin/fm-pr-merge.sh`), retry delay not validated.** What must hold: the retry wait is always a short, valid `sleep` argument, so a bad `FM_PR_GITHUB_MERGEABLE_RETRY_DELAY` can never trip `set -e` or hold the task lock for a long time. The retry loop is the only place that reads this variable. The script now reads the value once before the loop and accepts only whole numbers from 0 to 10. Anything else (empty, `abc`, `-1`, `1.5`, `11`, a huge number, leading spaces) falls back to 3. I ran those values through the check by hand and each came out as expected. The loop now sleeps on that checked value. **ci-1 (`tests/fm-pr-merge.test.sh`), no test for a check changing between UNKNOWN reads.** What must hold: every retry re-checks all live conditions, not just mergeable. The fake `gh pr view` in the test can now take an optional second word on each line of the mergeable sequence, which sets the first check's result. The new test `test_github_mergeable_unknown_retry_rechecks_checks` feeds `UNKNOWN`, then `UNKNOWN FAILURE`. It asserts: - exit code 1 after exactly 2 reads, - the refusal names `check 'ci' is not green`, - the message does not say mergeability is still being computed, - `pr merge` was never called. If a later change made the retry look only at mergeable, the loop would read UNKNOWN 5 times, end with the "still being computed" message, and this test would fail. I didn't run it against a deliberately broken script to confirm that. `bash -n` passes. `shellcheck` reports only the existing info-level notes about files it can't follow. Only `bin/fm-pr-merge.sh` and `tests/fm-pr-merge.test.sh` changed
kunchenguid#6112)

* fix(bin): converge every open owner onto a known terminal contribution

settle_final only cleared a stale error on retry, so an owner whose saved
row still said open kept projecting a merged or closed pull request as
open after another owner's row had already recorded the terminal
observation. Copy the known terminal observation to every owner whose
saved row is not itself terminal, keeping that owner's own pending and
notified state, and clear its error.

* no-mistakes(review): Carry terminal checked_at when converging existing owner rows

* no-mistakes(ci): I fixed Greptile finding ci-2 as you asked, with a change to tests/fm-contributions.test.sh only. The rule it enforces: when a retry converges an owner onto a URL that is already merged or closed, that owner gets the terminal owner's whole observation, not just its state. The same weak check appeared twice in test_interrupted_multi_owner_poll_settles_every_owner, so I fixed both: - **Open owner (line 784):** the check now also requires `.observation == $terminal[0].records[0].observation`. The existing checks for error, checked_at, pending and notified are unchanged. - **Errored owner (just below):** it only checked state and error before. It now reads the terminal owner's file and makes the same full-observation comparison. Adding the comparison alone would not have caught anything. The test fixtures gave both owners identical observations apart from `state`, so copying only the state would still have passed. In both cases I also set the terminal owner's observation head to HEAD_B, so the two observations now really differ. Verification: - The focused test passes against the current bin/fm-contributions.sh. - I temporarily changed `settle_final` so it copied only the state. The test then failed, reporting the owner still on the old head (HEAD_A). I restored the file afterwards, and `git status` shows only the test file modified. - The full tests/fm-contributions.test.sh suite exits 0. No product code changed. The other CI finding (ci-1, "Behavior portable serial 9") was left alone because you chose to ignore it
…nguid#6124)

* feat: run the supervision host by default on a Claude primary

An absent config/supervision-host on a Claude primary now reads as on with
the default engine, and a file holding `off` opts any home out. Cursor,
OpenCode, omp, Grok, and Codex stay file-gated, with `off` read as disabled
there too. Every reader asks fm_supervision_host_enabled instead of testing
the file, and non-bash readers query it through the lib's `enabled` entry.
A primary's `off` is not inherited by secondmates: each home keeps its own
supervision posture.

* test: pin the watcher-path posture in fixtures that assume no supervision host

Fixtures that drive the watcher arm or assert a non-host drain now write
an explicit off file, and fixtures that copy the Stop auto-arm or the
supervision instructions carry the engine lib they now source. The two
drain suites also stop reading the code root's config.

* fix: name the opt-out when an off home passes an attended wake to main

A host parked when the home writes off now logs that the home does not run
the supervision host, rather than claiming it has no engine.

* no-mistakes(document): Clarify Claude supervision defaults and historical evidence

* no-mistakes(ci): Fixed process leaks in the two added host tests. Each case now stops its recorded watcher and host/arm processes; fake hook sessions exit through session.stop. The full host suite passed before the final cleanup refinement, and both affected cases, bash syntax, ShellCheck, and diff checks passed afterward. CI runtime still needs confirmation
…start scope check (kunchenguid#6125)

* fix(bin): create the state dir on a fresh primary before the session-start scope check

fm_primary_scope_matches required an already-existing state directory, so
bin/fm-sessionstart-run.sh stood down on a fresh clone before anything could
create it. Split out fm_primary_root_matches so the run wrapper can confirm
primary-home identity first, create the gitignored state dir when it is
missing, and only then run the unchanged scope check.

* no-mistakes(document): Document session-start state dir creation on fresh clones

* no-mistakes(ci): I fixed the Greptile P1 the way you asked. When a fresh primary can't create `state/`, the run wrapper no longer stands down silently. **Invariant:** when an otherwise eligible fresh primary cannot create `state/`, startup must never fail silently. This path has only one site: the mkdir in `bin/fm-sessionstart-run.sh`. Other hooks and the nudge wrapper never create `state/`, so they have no equivalent failure. **What changed:** - **Run wrapper** (`bin/fm-sessionstart-run.sh`): it captures mkdir's error and prints one line to stderr before standing down as before (exit 0, or 3 for the Pi prerequisite). The line looks like `fm-sessionstart-run: startup could not create the state directory <path>: <reason>`. - **Test** (`tests/fm-sessionstart-nudge.test.sh`): the new case `test_run_reports_a_state_dir_it_cannot_create` uses a fresh primary with no `state/` and a read-only (0500) root. It checks four things: exit 0, no digest on stdout, no state dir created, and exactly one stderr line ending in "Permission denied". It fails without the fix and passes with it. - **Docs** (`docs/sessionstart-nudge.md`): I added one sentence describing the stderr line and one describing what the new test proves. **Verification:** I ran `tests/fm-sessionstart-nudge.test.sh`, and every test passes. `bin/fm-lint.sh` on the changed scripts (pinned ShellCheck 0.11.0) and `tests/fm-documentation-audiences.test.sh` also pass. As you asked, the wrapper still stands down with the ineligible-checkout status afterwards. It does not report this as a failed eligible startup, which is what the bot suggested
…ery (kunchenguid#6126)

* fix(bin): measure pending-reply grace from turn completion, not delivery

Fixes kunchenguid#6057

The pending-reply guard demanded a repost ("REPOST REQUIRED: previous
marked request had no correlated parent report") while the second
mate's correlated reply was already on its way.
fm_pending_reply_send_recovery measured its grace window from delivery
instead of from the request turn's completion, so any turn longer than
the grace fired the demand the moment the turn ended, before the reply
could have landed. The missed-report escalation had the same gap: it
fired the instant the recovery turn's completion was observed, with no
grace at all.

Both now measure grace from the relevant turn's completion (request
turn for the recovery repost, recovery turn for the escalation), and
both take one fresh, uncached read of the parent status file
immediately before firing, accepting a correlated line regardless of
its verb. Transport-failure escalations stay immediate, and the
one-repost limit is unchanged.

* no-mistakes(review): Document grace window as measured from turn completion

* no-mistakes(ci): Both Greptile findings were real and caused by this PR, so I fixed them. The full `tests/fm-pending-reply.test.sh` suite passes. **ci-1 (a reply could be overwritten by a repost).** The rule that must hold: a recovery send is recorded only if the record is still unresolved, checked under the same per-correlation lock that resolution uses. The escalation path already did this (`_fm_pending_reply_maybe_escalate_locked` reads fresh and publishes under one lock). The recovery path did not: `fm_pending_reply_send_recovery` did its fresh read through `fm_pending_reply_try_resolve`, which let go of the lock before the send was recorded. A reply landing in that gap could be overwritten, and the repost would go out anyway. Now `send_recovery` takes the lock once and, while holding it, re-checks that the phase is still `awaiting_report`, runs the fresh uncached read, and records the send (sender pid and identity, attempt time, phase `recovery_sending`). It releases the lock before actually sending, so the lock is not held during the send. It uses the same lock helpers the other lock wrappers use. Grace timing, the one-repost limit and the escalation path are unchanged. **ci-2 (the test would pass even without the fix).** In `test_recovery_fresh_status_read_resolves_before_firing`, the reply is still appended to the status file, but the stored file signature is then set to the file's new signature. That stands in for a same-size rewrite that the signature cache cannot see. The test first checks that a normal cached read misses the reply, then that the fresh read before sending catches it. I also added the same check for the fresh read before escalation, which the review said was uncovered. The test now sets its own send hook, so it no longer depends on one left over from an earlier test (that leftover had made failures exit silently). **Checks:** - I removed the fresh-read bypass at each site in turn and reran the suite. With it gone from recovery, the test fails with "recovery must not fire once a correlated reply has landed". With it gone from escalation, it fails with "the fresh pre-escalation read should have resolved the record, got escalated". With both in place, all tests pass. - Shellcheck with `-x` timed out locally. Without `-x` and ignoring SC1091, the only warnings are SC2034 on the existing `maybe_escalate` lock wrapper, which is not part of this change. The new code adds no warnings. Changes are in `bin/fm-pending-reply-lib.sh` and `tests/fm-pending-reply.test.sh`. Nothing is committed yet; a plain commit message such as "fix(bin): record the pending-reply recovery send under the fresh-read lock" fits the instruction

* no-mistakes(ci): ci-1 was real and caused by this PR. The same bug was also in the escalation path, so both are fixed. The full tests/fm-pending-reply.test.sh suite passes. The rule that must hold: a recovery repost or an escalation goes out only if the record's phase, read after the fresh-read resolve, is still what it was before. The resolver writes phase=resolved first and only then writes the other resolution fields. If one of those later writes fails, it returns an error even though the record is already resolved. Places this rule applies, both fixed: - Recovery (fm_pending_reply_send_recovery): the fresh-read resolve now runs first, and the phase is re-read right after it, whatever it returned. The send is recorded and made only if the phase is still exactly awaiting_report. This replaces the earlier phase check rather than adding a second one. - Escalation (_fm_pending_reply_maybe_escalate_locked): same bug. After a failed resolve it went on to publish the blocked line and set phase=escalated. One added line after the resolve call returns 1 without publishing if the phase has changed. Test: added test_partial_resolve_write_blocks_firing. It forces a failure on the resolved_epoch write after a correlated reply has landed. It checks that the recovery send hook is never called, that no escalation line is published, and that the phase stays resolved. The forced failure runs in a subshell so it can't affect later tests. Checks: - With the recovery fix reverted, the new test fails with "recovery must not fire after a partial resolve". - With the escalation fix reverted, it fails with "partial resolve should block escalation, got escalated". - With both fixes in, every test passes. - Shellcheck was run with SC1091 excluded and without -x, not through the repo's lint script. The only new message is one SC2329 info on the test's override function; other test overrides in the same file already get that same info, unsuppressed. Changed files: bin/fm-pending-reply-lib.sh and tests/fm-pending-reply.test.sh. Nothing is committed. Suggested plain commit message: "fix(bin): recheck pending-reply phase after the fresh read before sending
…to stderr (kunchenguid#6001)

* fix: provider-table lookup never writes a broken-pipe error to stderr

Fixes kunchenguid#5956

fm_quota_single_provider_for_harness returned from its while read loop
as soon as it found a match, closing the pipe while
fm_quota_single_provider_table's printf could still be writing.
Where SIGPIPE is ignored, as on GitHub Actions runners, bash then
prints "printf: write error: Broken pipe" on the resolver's stderr,
which intermittently broke the one-diagnostic-line assertions in
tests/fm-dispatch-resolve.test.sh.

Read the whole table before answering, the way
fm_control_harness_supported already does, so the writer always
finishes. Return values and output are unchanged.

Reproduced by running tests/fm-dispatch-resolve.test.sh with SIGPIPE
ignored on a single pinned core under CPU contention: 30 of 30 runs
failed before the fix, 0 of 30 after. Note: reproducing requires
setting the trap inside the tested shell because nice(1) resets an
inherited SIGPIPE ignore to SIG_DFL. tests/fm-quota-choose.test.sh
passes and bin/fm-lint.sh is clean.

* no-mistakes(ci): Fixed both Greptile findings the user chose to address. ci-1 (bin/fm-quota-axi-lib.sh:154). Invariant: looking up a harness must always end with status 0 and print the provider, even when the caller runs under `set -e`. The loop body `[ -z "$found" ] && [ "$harness" = "$1" ] && found=$provider` now ends in `|| :`. Every iteration succeeds and the whole table is still read. Only `fm_quota_single_provider_for_harness` loops over the table this way, so this is the one place the fix was needed. One caveat: on bash 5.3 the old code did not actually exit under `set -e`, because the `while` loop is not the function's last command, so the new `set -e` test would have passed before this fix too. The change makes the loop's success explicit, as the user asked. ci-2 (regression coverage). I added three cases to the existing `tests/fm-quota-choose.test.sh`, all calling the public lookup function after sourcing the library: 1. With SIGPIPE ignored (`trap "" PIPE`), it looks up every harness 200 times and checks that nothing reaches stderr. 2. A deterministic version of the race: the table function is wrapped so it writes the first row, pauses 0.2 s, then writes the rest. With SIGPIPE ignored, it checks that looking up `claude` prints `claude` and writes nothing to stderr. The stress loop alone reproduced the bug in only about 1 of 5 local runs, which is why this case exists. 3. A direct call under `set -e` prints `claude`. Verification: - `bash tests/fm-quota-choose.test.sh`: all pass. - Same test against the pre-PR library (fa48367, via `FM_ROOT_OVERRIDE`): fails with `printf: write error: Broken pipe`. The deterministic case failed in one run and the stress loop caught it in another. - `shellcheck` on both files: clean. - `tests/fm-dispatch-resolve.test.sh`: passes
…isioning (kunchenguid#6162)

* fix: survive Pi 0.99 rendering and Git 2.55 local-clone races

Pi 0.99 puts arguments on the stock tool header and leaves hidden custom messages in the export conversation column. Match that header, and keep Calm's boundary on the visible column. Clone a remote home with --no-local so a prune during Git's loose-object copy cannot fail the seed.

* no-mistakes(review): Stop SIGPIPE write errors; cover older Pi export and project clones

* no-mistakes(document): Clarify Calm export visibility and tool rendering

* no-mistakes(ci): Fixed the dispatch diagnostic to list every provider-less use/default profile in one line and added a multi-profile behavior test. Shortened supervision fixtures using the existing engine-grace and park-clock knobs; removed stray scratch files. Dispatch tests, syntax checks, and three targeted supervision cases passed. CI’s prior supervision duration was 751s; the single permitted local full-suite run timed out at 1200s, so an after-duration is not established. The cancelled serial check had no failure verdict. The outer executor should record the measured before/after duration in the PR body when available

* no-mistakes(review): Gate Pi 0.99 call headers by version; drop hidden-row assertion

* no-mistakes(review): Test stock call headers under Pi 0.87 and 0.99 stubs

* no-mistakes(test): Fix older-Pi queued-row test and verify park-boundary behavior

* no-mistakes(document): Clarify Pi Calm export and queued-turn documentation

* no-mistakes(ci): Fixed the stock macOS Bash 3.2 parse failure in tests/fm-calm-pi-extension.test.sh; its parse check passes. The watcher CI failure is in unchanged code: the isolated five-minute/66-minute case passes locally, but the CI log omits the drain error needed to establish its cause. No speculative watcher fix was made. The full local watcher suite timed out after 500 seconds
…6169)

* Prevent premature Lavish board handoffs

* Prove Lavish arm lacks reply acknowledgement

* Confirm Lavish replies before arming worker boards

* no-mistakes(review): Post Lavish reply only after locked arm eligibility checks

* no-mistakes(review): Fail Lavish reply closed on unknown version

* no-mistakes(document): Correct Lavish reply documentation and remove stale guidance

* no-mistakes(document): Clarify Lavish reply routing and remove duplicate version guidance
…#6154)

* feat: inherit the supervision-host opt-out from the primary

Move the supervision host's off opt-out out of config/supervision-host into
its own presence flag, config/supervision-host-off, and add that flag to the
primary-authoritative inherited config set. A primary that opts out now opts
every secondmate home out at spawn and convergence, and clearing it converges
them back. config/supervision-host stays the home-local engine choice.

Shape: config/supervision-host mixed two things, a fleet posture (off) and a
per-home engine and model. Only the posture should follow the primary, so it
becomes a separate presence flag that rides the existing inherited-config
mechanism (FM_INHERITABLE_CONFIG in bin/fm-config-inherit-lib.sh) with no new
machinery, while the engine line stays local. The parse stays in its one
owner, fm_supervision_host_enabled. There is no migration or compatibility
handling for a home that still holds off in config/supervision-host.

Primary off, mate on: inherited material is primary-authoritative by design,
so a mate cannot keep the host while the primary is opted out, and a mate's
own opt-out is removed at the next convergence while the primary has none.
Running the host on a mate is the primary's choice for the fleet; no override
mechanism is added.

Live validation (disposable bin/fm-live-lab.sh lab, Claude primary with a
real seeded secondmate, --supervision-host off):
- up: every readiness check ok, including "host: none running, as expected"
  and a live mate session; the spawned mate home held the inherited
  config/supervision-host-off and the gate read primary OFF, mate OFF.
- primary removed its opt-out, then bin/fm-config-push.sh reported
  "supervision-host-off: pushed - mirrored primary absence" and a config
  reread sent; the gate read primary ON, mate ON, and the live mate handled
  the reread.
- primary opted out again and pushed: "supervision-host-off: pushed", mate
  gate OFF.
- down stopped every lab process and left no lab process running.

Out of scope, follow-up: default-on for the other harnesses, away-daemon
retirement, rollout.

* no-mistakes(document): Document inherited supervision-host opt-out ownership

* no-mistakes(ci): Fixed ci-4: with `--supervision-host off --mate`, lab readiness now requires the inherited flag in the mate home and a disabled mate supervision-host gate. The focused behavior test, shellcheck, and diff checks pass. Left ci-1–ci-3 untouched as directed

* no-mistakes(test): Fix mate readiness HOST_OFF initialization in lab up

* no-mistakes(ci): Fixed Lint 2 by making the new test’s fixtures source resolvable to ShellCheck; its off/on readiness test and ShellCheck now pass locally. Behavior portable serial 5 failed in the unchanged remote-reply test at generation 7. That test passes locally, and no PR-caused defect was identified, so no remote-reply code was changed
…nguid#6179)

* fix(tests): cut the fixed sleeps in supervision-host cycles

The serial CI lane keeps brushing its 30-minute cap because
fm-supervision-host.test.sh spends ~903s of the job, and per the
run-36635306527 case profile the top nine cases are all multi-cycle
ones (3-10 park/close/turn cycles each): every close waits out the
host's sleep $POLL in await_close plus a watcher sleep $FM_POLL scan
cycle, and every engine turn waits out the fixed sleep 1 descendant
snapshot. That is ~3s of pure sleep per cycle before any real work.

The host poll now accepts positive decimal seconds through a new
seconds_or validator (FM_SUPERVISION_HOST_POLL), and the engine turn's
snapshot loop takes FM_SUPERVISION_ENGINE_SNAPSHOT_SECONDS, also a
positive decimal defaulting to one second - the smallest seam at each
wait's single owner. The suite drives them at 0.2 alongside the
existing FM_POLL=0.5 and FM_ARM_ATTACH_POLL=0.2 knobs, so the real
poll loops still run. The park-boundary case moves onto the injected
test clock instead of a real 3s wait, per-case cleanup polls the host
pid rather than sleeping a full second, and the proof-by-absence
windows (flood re-escalation, successor re-announce, watcher
persistence, recovery staying off main) shrink from 2-3s to 1s, which
still spans two watcher polls at the test cadence.

Every assertion, process lifecycle, and reaping path is unchanged;
production defaults stay at one second. Isolated case timings on a
contended host, base vs branch: attended-latch 54.3->34.6s,
undelivered-dialog 67.7->59.1s, away-latch 46.5->30.5s, held-cadence
47.9->21.6s, unreadable-mirror 39.2->38.5s, park-limit 18.2->12.3s,
registration-fallback 14.1->10.0s, first-cycle-status 12.6->8.4s,
latch-scope 16.7->16.3s. Full suite: 65/65 pass. fm-lint and
shellcheck clean.

* no-mistakes(review): Wait for scan lock release before duplicate check

* no-mistakes(document): Correct supervision snapshot cadence documentation

* fix(tests): keep production poll cadence, probe exits at 0.1s

The fractional poll cadences multiplied the cost of each loop body:
full process-table scans in the engine turn and process refreshes in
await_close ran five times more often, which swamped the thin CI runner
and nearly doubled every multi-cycle case (serial 5 was cancelled at its
30-minute limit on run 36635306527's successor). Restore the production
cadence and notice arm/engine exits with a cheap kill -0 probe at a
tenth of a second between the one-second bodies instead: strictly less
dead time than baseline with no added CPU.

Also hold each injected-clock park bound well past its case's
wall-clock checks so a host that ignored the test clock fails instead
of silently passing at a real-time boundary, and restore the shortened
proof windows (watcher liveness, recovery-off-main absence, first-cycle
stream) to their baseline depth.

* no-mistakes(document): Clarify supervision engine snapshot documentation
Bring the fork up to date with upstream's 224 commits while keeping the
fork-only zcode harness adapter: crewmate and scout launches, the primary
session harness, the opt-in TUI variant, the turn-end hook with resume and
quota evidence, herdr agent-view busy flips, and the variant-shaped
interrupt record.

Conflicts keep upstream's behaviour and formatting and re-apply the zcode
surface beside upstream's new devin adapter. Post-merge adaptations add
zcode to upstream's new harness-name lists (supervision-branch primary pin,
supervision-engine test seam), move the zcode state rows into the
operational-home-layout skill that now owns the home layout, teach the
zcode TUI test fixture upstream's staged launch file, and make sandbox
fixtures that copy fm-harness.sh carry fm-zcode-lib.sh.
…-lock lib

fm-session-lock-lib.sh sources fm-zcode-lib.sh, so the turn-end guard and
session-start hook fixtures that stage a private copy of the lock lib must
stage the zcode lib beside it.
Upstream's fm-wake-lib.sh now sources fm-path-lib.sh. Without it the
fixture's fm-lock.sh loops forever taking the claim lock, so the zcode and
plain-bash lock e2e cases never finish and hang the CI shard.
@d-ploutarchos d-ploutarchos changed the title feat(bin): merge upstream firstmate main into the fork and reconcile the zcode harness feat(bin): sync the fork with upstream firstmate main and reconcile the zcode harness Sep 30, 2026
@d-ploutarchos
d-ploutarchos merged commit be03cb5 into main Oct 1, 2026
59 of 61 checks passed
d-ploutarchos added a commit that referenced this pull request Oct 1, 2026
Records OK-LG/firstmate main (9bc7f86) as a parent so this branch merges
cleanly, while keeping this branch's tree unchanged: upstream main 8f756bb
plus the cherry-picked Bash 3.2 owner-capture fix (#11). The fork-only
zcode adapter (#1-#10) is intentionally dropped; #10's upstream content is
already contained in upstream main.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.