Skip to content

feat(bin): merge upstream Firstmate aec043c7 into the fork - #54

Merged
BenWilcox8 merged 161 commits into
mainfrom
fm/upstream-reconcile-r2
Sep 25, 2026
Merged

BenWilcox8 merged 161 commits into
mainfrom
fm/upstream-reconcile-r2

Conversation

@BenWilcox8

@BenWilcox8 BenWilcox8 commented Sep 24, 2026 •

Copy link
Copy Markdown
Owner

Intent

Captain's request: "I want you to look and see if the main firstmate repo has updates and if it does, reconcile my version of firstmate with the newest one. After, give me a report of what was changed and what was retained." Preserve the personal fork, local customizations, private operational records, and unfinished work during this update.

Accepted build requirements:

  • Integrate upstream kunchenguid/firstmate into the personal fork BenWilcox8/firstmate. The initial local and origin head was af823b7. The integration target is upstream main aec043c (fetched 2026-09-21; 321 local-only and 121 upstream-only commits; common ancestor 72bfdd0). Upstream commits after aec043c are deliberately NOT integrated here; MAIN will do a separate follow-up merge of newer upstream commits after this lands.
  • Start from local main, retain both histories with merge commits (integration merge 3432404 has exactly the two pinned parents), and resolve conflicts by behavior. Never replace all conflicted files with one side, and never discard a local contract because upstream changed its owner.
  • Fork main must be fully contained, merged with merge commits (never rebase this branch): 712aed8, 534a987, c91b03a, d5fbb6c, 0089320, eeb841b, cc9045e, d978c18 (fork PRs 44-57). The fork-only retire repair folds decisions without the task kind; PR 52 pane hooks step aside for upstream stopped-server recovery and refuse a recycled pane.
  • Preserve Codex operation, Herdr spaces, agent-axi placement, RTK policy, dashboard and Atlas integration, captain authority (explicit merge authority and yolo; away posture alone never authorizes a merge), and durable inbox and approval records. Preserve Claude account pins (account-pinned trust stays outside the supervisor credential store), complete accepted validation intent in briefs and promotion, archive-aware captain answers (a live-read timeout refuses and never falls back to a stale archived answer), remoteless pool freshness (local default-branch content is authoritative), bounded declared waits and duplicate suppression, exact Codex MainThread attribution, and endpoint retirement that never loses unfinished work.
  • Preserve supported upstream harnesses and new upstream behavior unless a concrete existing captain requirement conflicts: adopt AGY, native ultra effort, Codex hook isolation (--disable hooks), Claude permission mode, durable away records, mail, typed dispatch, staged launch files, Herdr stopped-server and destroyed/drifted endpoint recovery, teardown close-failure refusal, held-call serialization, non-draft PR requirements, and upstream's rule that a live worker's declared future wait time controls when the watcher rechecks it.
  • Keep the official upstream Pi files; do not revive the archived custom Pi. Rovo dispatch stays disabled and its verification doc stays deleted. New Orca task selection stays retired (existing-record handling kept).
  • Instruction policy (MAIN decision): retain the complete local AGENTS.md instruction policy and the existing tracked token ceiling mechanism (.agentsmd-ceiling). Upstream VISION's 9,000-word rule does not supersede it; VISION points to the fork ceiling instead. Do not delete instruction substance and do not arbitrarily raise limits. AGENTS.md is 29,219 estimated tokens against the ceiling 29,220; the fork main merge dropped the redundant words "to this default" from the Atlas doctrine so PR 47's verb list fits without raising the ceiling.
  • CI (MAIN decisions): bin/fm-lint.sh accepts 1of4..4of4 besides 1of2/2of2 (same least-loaded split; the two-way split is unchanged), and CI runs four lint partitions with FM_LINT_JOBS "1", and each root in bin/fm-lint-heavy-roots.list (measured single-file RSS near 4 GB or more) runs alone in its own ShellCheck invocation, because fork Lint 1 shards exceeded runner memory; tests/fm-pending-reply.test.sh points its source directives for bin/fm-pending-reply-lib.sh at /dev/null (fork 14.65 GB vs upstream 12.69 GB alone; 6.36 GB after), the library staying its own lint root. PR follow-up (not this change): find which fork library growth inflates that graph. Seams row upstream: generic lint memory fix. Pin upload-artifact ea165f8d (v4.6.2); Bearings count 60.
  • Park and resume (PR 47) combine with upstream relaunch recovery: a --resume-session relaunch always opens a new endpoint in the recorded worktree; every other relaunch keeps upstream's endpoint absence proof, adoption, and rebind. Teardown removes both the progress and pi-session records. Fleet snapshot keeps both the parked field and the PR head.
  • Tests (MAIN decision, audit F3): restore upstream's issue Recovery classifier trusts a stale herdr agent registration as 'live', blocking relaunch/spawn after a Pi crew exits to a shell kunchenguid/firstmate#4115 relaunch-over-stale-registration section in tests/fm-control-herdr-smoke.test.sh through the guarded Herdr lab helper (done); keep the fork's stricter "unreadable refused" Herdr verdicts in tests/fm-backend-herdr.test.sh as an explicitly recorded divergence (consolidation on one classifier is later work); keep the 6-to-9 serial shard change and the refusal of raw compound launch commands. The captain requires that no test rigor is lost: test repairs may adapt fixtures to merged contracts or host portability and may raise load-sensitive readiness ceilings, but must not weaken or remove assertions.
  • Do not mutate the primary home, secondmate homes, project copies, live sessions, the shared no-mistakes daemon (v1.64.0), parked validation runs, credentials, or the primary's untracked overload path. Secondmate harness activation (including any Codex switch) is not part of this change; MAIN owns activation and harness changes as a separate step.
  • Push and open the PR only on BenWilcox8/firstmate, never on kunchenguid/firstmate. MAIN owns merge and activation. A branch or PR existing does not mean the live fleet is updated.
  • Use the no-mistakes pipeline for review, tests, documentation, lint, the fork PR, and CI without a deterministic full-suite commands.test override and without bypassing required gates.
  • Deferred follow-ups (MAIN): procevent Lavish doorbell provenance (Atlas c685); away-merge test wording. Upstream ordering stays: a live worker's declared until controls its recheck. Fixtures may adapt to merged contracts without weakening assertions.
  • The private report and preservation evidence live outside the repo in the operator's private data directory (report.md with adopted upstream changes, retained customizations, conflict decisions, tests, unverified limits, a before-and-after file and behavior map, the ancestry proof, and a safe activation plan); they are not part of this diff.

Host limits, not defects: no system python3, lsof, ruby, or Chrome; a stray /tmp/.git; a Codex update prompt blocks the live Codex Herdr test.
Test-step host note: Nix tool paths are in lines 5-6 of ~/firstmate/data/upstream-reconcile-r2/final-evidence/suite-r5/suite-r5.sh. Run tests one at a time under nice -n 10; never run whole-partition or parallel ShellCheck locally. CI runs the full suite. fm-test-run.sh --changed refuses tests/assets/* as upstream does.

What Changed

  • Merges upstream kunchenguid/firstmate main aec043c7 into the fork with a two-parent merge commit (3432404), then merges fork main PRs 44-57 on top with merge commits. This adds upstream behavior: the AGY harness (bin/fm-agy-trust.sh, harness/agy.md), native ultra effort, Codex --disable hooks isolation, mail (bin/fm-mail.sh, bin/fm-mail.py, bin/fm-mail-check.sh), typed dispatch (bin/fm-dispatch-resolve.sh), PR state and reviewer helpers, contributions, the remote Herdr guard, the Claude calm mod, and Herdr stopped-server and drifted-endpoint recovery in bin/backends/herdr.sh.
  • Resolves conflicts by behavior and keeps fork contracts. The fork keeps the full AGENTS.md policy under the .agentsmd-ceiling mechanism, explicit captain merge authority, archive-aware captain answers, Claude account pins, and remoteless pool freshness. A --resume-session relaunch always opens a new endpoint in the recorded worktree. All other relaunches keep upstream's endpoint absence proof, adoption, and rebind. PR 52 pane hooks step aside for upstream stopped-server recovery and refuse a recycled pane. The upstream Pi files and the deleted Rovo verification doc stay as they are.
  • Splits CI lint into four partitions with FM_LINT_JOBS: "1". bin/fm-lint.sh now also accepts 1of4 through 4of4, and each root in the new bin/fm-lint-heavy-roots.list runs alone in its own ShellCheck call. tests/fm-pending-reply.test.sh points its library source directives at /dev/null. The branch also adds or extends tests for the merged contracts, restores the upstream issue Recovery classifier trusts a stale herdr agent registration as 'live', blocking relaunch/spawn after a Pi crew exits to a shell kunchenguid/firstmate#4115 stale-registration section in tests/fm-control-herdr-smoke.test.sh, and adds no-mistakes v1.70.1 capture fixtures.

🤖 Generated with Claude Code

Risk Assessment

✅ Low: The only new change since the last review makes existing assertions stricter. It splits each chained [ ... ] && ... check in tests/fm-local-pane-recovery.test.sh into separate top-level commands, so under set -e a false condition now fails the test. It weakens no assertion and changes no product code.

Testing

I ran each test on its own under nice -n 10, with the Nix host tool paths set. Four token-free tests drove the real Herdr 0.8.2 in isolated, guarded lab sessions (pane lifecycle smoke, restart recovery, worker restore, pane cleanup), and all passed. The fixture tests for the latest review fixes also passed: the recycled-pane relaunch refusal, relaunch adoption of a pane that outlived its stopped server, archiving of the merge-authority and progress records, and the pending-reply test. In /tmp copies, I reverted the guard and, separately, closed the pane before the refusal; the pane-recovery test failed both times. I also ran the real lint CLI in list mode for all partitions. The 1of4..4of4 split covers the same 483 roots as the 2-way split, with no overlap, and bad partition names are refused. With a recording ShellCheck stand-in, each of the 6 heavy roots ran alone in exactly one invocation. With the away record in place, fm-pr-merge.sh refused yolo=off and a missing yolo field, and made no forge call; the yolo=on control merged. The ancestry proof and the AGENTS.md ceiling check (29219/29220) passed. The fixture, fake, and stand-in checks are recorded as untested in live terms. The park/resume live test and the live Codex Herdr test were not run. This change has no UI surface, so there are no screenshots; the evidence is CLI transcripts and test logs.

  • Live validation: ✅ go - 7 of 16 scenarios driven live against the product
Scenario Result Live Evidence
Real Herdr: exit, interrupt, and relaunch of worker panes work end to end, and a stale registration does not block a relaunch into a new pane (restored kunchenguid#4115 section) ✅ pass live tests/fm-control-herdr-smoke.test.sh in a guarded Herdr lab session -> fm-control-herdr-smoke.log
Real Herdr: after a lab server restart, authorized second mates and the primary come back in their own panes, a dormant mate stays down, and there is no double launch or repeat recovery ✅ pass live tests/fm-local-restart-recovery-herdr-e2e.test.sh -> fm-local-restart-recovery-herdr-e2e.log
Real Herdr: after a restart only the working worker is restored in its own slot, and a relaunch into a worktree leased to another task refuses ✅ pass live tests/fm-local-worker-restore-herdr-e2e.test.sh -> fm-local-worker-restore-herdr-e2e.log
Real Herdr: pane cleanup closes ended or husk panes, keeps live, supervisor, and unmanaged panes, and teardown refuses to discard unlanded work ✅ pass live tests/fm-local-pane-cleanup-e2e.test.sh -> fm-local-pane-cleanup-e2e.log
Adversarial: with the home workspace absent, a relaunch refuses a recorded pane id that now belongs to other work, closes nothing, and keeps the task record ⏸️ untested no The prior payload did not establish a live result: the recycled pane-id state was built only with a fake Herdr in tests/fm-local-pane-recovery.test.sh, and the lab tests do not produce that exact stat…
Adversarial: the hardened pane-recovery assertions fail the test when the guard is removed or when the pane is closed before the refusal ⏸️ untested no The prior payload did not establish a live result: the mutation runs were on /tmp copies and test the test itself, not the running product. No live product surface exists for this check.
A Herdr pane that outlived its stopped server is adopted on relaunch and reads as already-stopped on exit ⏸️ untested no The prior payload did not establish a live result: the stopped-server state was built only with a fake Herdr in tests/fm-control-relaunch.test.sh. To drive it live, stop a Herdr lab server while a pan…
Retiring a reassigned-slot task archives its merge-authority and progress records and leaves a live task's record alone ⏸️ untested no The prior payload did not establish a live result: tests/fm-local-cleanup-slot.test.sh used a fixture home and a fake forge, and no real retired task existed to retire. To drive it live, retire a real…
Lint CLI accepts 1of4..4of4, the four partitions cover the full CI inventory with no overlap (same as the 2-way split), and other split names are refused ✅ pass live bin/fm-lint.sh --partition KofN --list-files -> lint-partition-listing.txt
Each heavy lint root runs alone in exactly one ShellCheck invocation, and every root is still checked once ⏸️ untested no The prior payload did not establish a live result: real ShellCheck was replaced by a recording stand-in, because host rules forbid local whole-partition lint. The real four-partition lint runs in CI.
The pending-reply test still passes with its library source directives pointed at /dev/null ⏸️ untested no The prior payload did not establish a live result: tests/fm-pending-reply.test.sh is a unit and fixture test, and the ShellCheck memory gain is measured only in CI.
Adversarial: with an away record confirmed, a merge with yolo=off or no yolo field and no --captain-authorized refuses before any forge call; the yolo=on control merges ⏸️ untested no The prior payload did not establish a live result: bin/fm-pr-merge.sh ran against the fixtures from tests/fm-pr-merge.test.sh with a faked forge. A real merge needs a live GitHub PR and captain author…
Branch keeps both histories: the integration merge has the two pinned parents, and upstream aec043c and every fork main merge (PRs 44-57) are ancestors of HEAD ✅ pass live git ancestry commands -> ancestry-proof.txt
AGENTS.md stays within the tracked token ceiling without raising it ✅ pass live bin/fm-agentsmd-ceiling.sh check -> agentsmd-ceiling.txt (29219 of 29220)
Park and resume a real worker across a Herdr restart (a --resume-session relaunch opens a new endpoint in the recorded worktree) ⏸️ untested no tests/fm-park-resume-live-e2e.test.sh was not run. This test is opt-in because it spends model tokens with real Claude/Codex/Pi credentials. To run it, set FM_PARK_RESUME_LIVE_E2E=1 on an operator-app…
Codex worker lifecycle in a live Herdr lab (Codex hook isolation, idle doorbell) ⏸️ untested no tests/fm-control-herdr-codex-live-e2e.test.sh was not run. Per the host note, a pending Codex update prompt blocks the live Codex Herdr test on this machine. To run it, finish the Codex CLI update, th…
Evidence: Per-test exit summary (including both mutation runs)

Source: Per-test exit summary (including both mutation runs)

tests/fm-local-pane-recovery.test.sh rc=0 secs=48
mutation revert-fix rc=1
mutation close-before-refuse rc=1
tests/fm-control-relaunch.test.sh rc=0 secs=284
tests/fm-local-cleanup-slot.test.sh rc=0 secs=100
tests/fm-pending-reply.test.sh rc=0 secs=120
tests/fm-local-restart-recovery-herdr-e2e.test.sh rc=0 secs=112
tests/fm-local-worker-restore-herdr-e2e.test.sh rc=0 secs=88
tests/fm-local-pane-cleanup-e2e.test.sh rc=0 secs=215
tests/fm-control-herdr-smoke.test.sh rc=0 secs=34
Evidence: Pane recovery test transcript (recycled-pane relaunch refusal)

Source: Pane recovery test transcript (recycled-pane relaunch refusal)

ok - pane-cleanup-on-exit tests are registered through the fork family hook
ok - recovery skips a pane whose fresh spawn holds the task meta lock
closed ended pane=test:w1:p2
ok - recovery closes an ended worker pane while preserving its task record
closed ended pane=test:w1:p2
ok - a duplicated home label resolves to the recorded workspace, never to a guess
ok - automatic cleanup preserves a pane with unreadable process evidence
ok - automatic cleanup preserves a secondmate supervisor pane
ok - watcher poll closes an ended worker pane through its hook
ok - bootstrap cleans restored shells before layout repair
ok - bootstrap skips and reports a layout repair whose plan would close a recorded task pane
ok - bootstrap still heals husks that have no task record in this home
ok - bootstrap runs the existing layout repair when its plan closes no pane
ok - exit treats a missing home workspace as proof that no task pane remains
ok - exit refuses when the home workspace is missing but the recorded pane still exists
ok - exit refuses when the home workspace exists but its crew inventory is unreadable
ok - a sweep reads a failed session inventory once and closes nothing
ok - exit refuses when a split tab hides the recorded worker pane from label resolution
ok - teardown refuses and keeps the records when a split tab hides the recorded worker pane
ok - recovery leaves a split-tab worker pane untouched
ok - exit treats a recorded pane id that now names other work as gone and closes nothing
ok - exit still refuses an unlabeled recorded pane that runs inside this worktree
ok - relaunch with no home workspace refuses a recorded pane id that now names other work
ok - forced teardown does not report a recycled recorded pane as this task leaked pane
ok - teardown stop submits no exit command after an interrupt ends the agent
Evidence: Mutation: guard reverted -> test fails

Source: Mutation: guard reverted -> test fails

ok - pane-cleanup-on-exit tests are registered through the fork family hook
ok - recovery skips a pane whose fresh spawn holds the task meta lock
closed ended pane=test:w1:p2
ok - recovery closes an ended worker pane while preserving its task record
closed ended pane=test:w1:p2
ok - a duplicated home label resolves to the recorded workspace, never to a guess
ok - automatic cleanup preserves a pane with unreadable process evidence
ok - automatic cleanup preserves a secondmate supervisor pane
ok - watcher poll closes an ended worker pane through its hook
ok - bootstrap cleans restored shells before layout repair
ok - bootstrap skips and reports a layout repair whose plan would close a recorded task pane
ok - bootstrap still heals husks that have no task record in this home
ok - bootstrap runs the existing layout repair when its plan closes no pane
ok - exit treats a missing home workspace as proof that no task pane remains
ok - exit refuses when the home workspace is missing but the recorded pane still exists
ok - exit refuses when the home workspace exists but its crew inventory is unreadable
ok - a sweep reads a failed session inventory once and closes nothing
ok - exit refuses when a split tab hides the recorded worker pane from label resolution
ok - teardown refuses and keeps the records when a split tab hides the recorded worker pane
ok - recovery leaves a split-tab worker pane untouched
ok - exit treats a recorded pane id that now names other work as gone and closes nothing
ok - exit still refuses an unlabeled recorded pane that runs inside this worktree
not ok - recycled relaunch output ({"tab_id":"w1:t9","label":"fm-other"}): ●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
●  WATCHER DOWN - SUPERVISION IS OFF
●  1 task(s) in flight, but no watcher has a fresh beacon (last beat: never, grace 300s).
●  Trust the emitted supervision protocol for this harness; do not use shell & for watcher repair.
●  This is a supervision warning only; the guarded operation WILL still run.
●  repair missing watcher supervision according to the session-start block for this harness; do not use shell &.
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
error: task ended's recorded worktree '/tmp/tmp.5VjrNeimy9/worktree' is missing; refusing to relaunch without the local copy its work lives in
Evidence: Mutation: pane closed before refusal -> hardened assertion fails

Source: Mutation: pane closed before refusal -> hardened assertion fails

ok - pane-cleanup-on-exit tests are registered through the fork family hook
ok - recovery skips a pane whose fresh spawn holds the task meta lock
closed ended pane=test:w1:p2
ok - recovery closes an ended worker pane while preserving its task record
closed ended pane=test:w1:p2
ok - a duplicated home label resolves to the recorded workspace, never to a guess
ok - automatic cleanup preserves a pane with unreadable process evidence
ok - automatic cleanup preserves a secondmate supervisor pane
ok - watcher poll closes an ended worker pane through its hook
ok - bootstrap cleans restored shells before layout repair
ok - bootstrap skips and reports a layout repair whose plan would close a recorded task pane
ok - bootstrap still heals husks that have no task record in this home
ok - bootstrap runs the existing layout repair when its plan closes no pane
ok - exit treats a missing home workspace as proof that no task pane remains
ok - exit refuses when the home workspace is missing but the recorded pane still exists
ok - exit refuses when the home workspace exists but its crew inventory is unreadable
ok - a sweep reads a failed session inventory once and closes nothing
ok - exit refuses when a split tab hides the recorded worker pane from label resolution
ok - teardown refuses and keeps the records when a split tab hides the recorded worker pane
ok - recovery leaves a split-tab worker pane untouched
ok - exit treats a recorded pane id that now names other work as gone and closes nothing
ok - exit still refuses an unlabeled recorded pane that runs inside this worktree
Evidence: Live Herdr lab: pane lifecycle smoke (incl. kunchenguid#4115 stale registration)

Source: Live Herdr lab: pane lifecycle smoke (incl. #4115 stale registration)

ok - real herdr: exit closes an agent-free pane and preserves the supervisor and worktree
ok - real herdr: interrupt refuses when herdr's own agent registry reports no agent
ok - real herdr 0.8.2: a gone session reads recoverable while a live pane and a malformed target do not
warning: /tmp/fm-local-pane.Q82trU/home/data/hsmoke/launch-brief.md records no delivery contract line (scaffolded before ship briefs recorded one); launching on the explicit --mode no-mistakes - confirm its definition of done matches
ok - real herdr: a drifted agent-free shell relaunches in a new pane in its worktree and the old pane closes
ok - real herdr: stale lifecycle-hook status does not keep a shell-only pane alive
warning: /tmp/fm-local-pane.Q82trU/home/data/hsmoke/launch-brief.md records no delivery contract line (scaffolded before ship briefs recorded one); launching on the explicit --mode no-mistakes - confirm its definition of done matches
ok - real herdr: a stale registration does not block relaunch into a new pane, and the local copy survives
ok - real herdr: interrupt protects an exact foreground agent process
ok - real herdr: interrupt preserves the live worker, supervisor, and local copy
ok - real herdr: an agent behind an unproven composer fails closed instead of typing an exit command into it
Evidence: Live Herdr lab: restart recovery

Source: Live Herdr lab: restart recovery

ok - live: a primary in a lab pane records its endpoint through the real restart-record path
ok - live: with no restart, the boot unit relaunches nothing
ok - live: after a reboot the authorized second mates run again in their own panes, and the dormant one stays down
ok - live: the primary comes back last, in its recorded pane, resuming its recorded session
ok - live: the pass leaves one wake and a status boundary for each relaunched second mate
ok - live: a second mate relaunched by hand during a restart is detected and not launched twice
ok - live: a primary whose pane is gone comes back in a new tab of its recorded workspace, named in the pass record
ok - live: a finished restart is never recovered twice
ok - live: a restart loop stops at the rate limit with one alert and no launches
ok - live: a primary whose workspace is gone comes back in a new workspace with its label, named in the pass output
Evidence: Live Herdr lab: worker restore after restart

Source: Live Herdr lab: worker restore after restart

ok - five Claude workers run in their own lab panes and worktrees
{"session":{"default":false,"name":"fm-lab-reboot-recovery-1955091-5975","running":false,"session_dir":"~/.config/herdr/sessions/fm-lab-reboot-recovery-1955091-5975","socket_path":"~/.config/herdr/sessions/fm-lab-reboot-recovery-1955091-5975/herdr.sock"},"stopped":true}
ok - the lab restart stopped every worker
check: worker restore after machine reboot: restored 1 (working); skipped: running 1; parked 1; finished 1; captain-waiting 1 (waiting: open decision: scope); slot-reused 1 (reused: its worktree is now recorded for taker in /tmp/fm-local-pane.OmnXWu/home); failed 0
ok - only the working worker came back, in its own slot and worktree; the others stayed down
ok - a relaunch into a worktree leased to another task refuses and leaves that task running
ok - a restart is restored once
Evidence: Live Herdr lab: pane cleanup

Source: Live Herdr lab: pane cleanup

herdr client=0.8.2 protocol=20
already-stopped ended harness=pi backend=herdr endpoint=fm-lab-pane-cleanup-on-2109762-14524:w1:p2 worktree=/tmp/fm-local-pane.qFqeCj/worktree
ok - exit removes a shell pane and frees its slot while work and supervisor remain
already-stopped ended harness=pi backend=herdr endpoint=fm-lab-pane-cleanup-on-2109762-14524:w1:p2 worktree=/tmp/fm-local-pane.qFqeCj/worktree
ok - repeated exit is idempotent after its pane is gone
closed crashed pane=fm-lab-pane-cleanup-on-2109762-14524:w1:p3
ok - recovery skips held relaunch and spawn locks and cleans the pane after release
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
●  WATCHER DOWN - SUPERVISION IS OFF
●  3 task(s) in flight, but no watcher has a fresh beacon (last beat: never, grace 300s).
●  Trust the emitted supervision protocol for this harness; do not use shell & for watcher repair.
●  This is a supervision warning only; the guarded operation WILL still run.
●  repair missing watcher supervision according to the session-start block for this harness; do not use shell &.
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
warning: /tmp/fm-local-pane.qFqeCj/home/data/restart-2109278/launch-brief.md records no delivery contract line (scaffolded before ship briefs recorded one); launching on the explicit --mode no-mistakes - confirm its definition of done matches
relaunched restart-2109278 harness=pi from=pi model=default effort=default backend=herdr endpoint=fm-lab-pane-cleanup-on-2109762-14524:w1:p5 worktree=/tmp/fm-local-pane.qFqeCj/worktree
ok - relaunch replaces the pane in the same slot and worktree
ok - recovery retains a live foreground agent
ok - relaunch refuses a held task-set reservation before it stops the old agent
closed restart-2109278 pane=fm-lab-pane-cleanup-on-2109762-14524:w1:p5
ok - recovery closes a killed agent husk without changing its work
WARNING: watcher still down (same stale episode; last beat: never, grace 300s) - full banner already printed this episode.
warning: /tmp/fm-local-pane.qFqeCj/home/data/restart-2109278/launch-brief.md records no delivery contract line (scaffolded before ship briefs recorded one); launching on the explicit --mode no-mistakes - confirm its definition of done matches
relaunched restart-2109278 harness=pi from=pi model=default effort=default backend=herdr endpoint=fm-lab-pane-cleanup-on-2109762-14524:w1:p6 worktree=/tmp/fm-local-pane.qFqeCj/worktree
closed restart-2109278 pane=fm-lab-pane-cleanup-on-2109762-14524:w1:p6
ok - recovery closes an agent that exits on its own
WARNING: watcher still down (same stale episode; last beat: never, grace 300s) - full banner already printed this episode.
warning: /tmp/fm-local-pane.qFqeCj/home/data/restart-2109278/launch-brief.md records no delivery contract line (scaffolded before ship briefs recorded one); launching on the explicit --mode no-mistakes - confirm its definition of done matches
error: the replacement agent for restart-2109278 did not come up within 2s (endpoint reads 'dead')
error: restart-2109278 was relaunched on pi but no running agent could be confirmed; its work is preserved at /tmp/fm-local-pane.qFqeCj/worktree
ok - a failed replacement leaves no shell pane and preserves the task and work
WARNING: watcher still down (same stale episode; last beat: never, grace 300s) - full banner already printed this episode.
warning: /tmp/fm-local-pane.qFqeCj/home/data/restart-2109278/launch-brief.md records no delivery contract line (scaffolded before ship briefs recorded one); launching on the explicit --mode no-mistakes - confirm its definition of done matches
error: pane cleanup: restart-2109278 cannot reclaim occupied slot t1/s1
error: the replacement agent for restart-2109278 could not be launched on pi
error: restart-2109278's agent was stopped but the replacement did not launch; no agent is running, and its work plus the recorded progress note are preserved at /tmp/fm-local-pane.qFqeCj/worktree
ok - relaunch refuses an occupied slot without closing its occupant
closed occupant pane=fm-lab-pane-cleanup-on-2109762-14524:w1:p8
ok - recovery retains supervisor, secondmate, and unmanaged panes
ok - a local record cannot authorize cleanup in another home workspace
WARNING: watcher still down (same stale episode; last beat: never, grace 300s) - full banner already printed this episode.
teardown retired complete (window fm-lab-pane-cleanup-on-2109762-14524:w1:pC, worktree /tmp/fm-local-pane.qFqeCj/retired-worktree)
Backlog: retired just finished (this home keeps no markdown backlog at /tmp/fm-local-pane.qFqeCj/home/data/backlog.md). Update /tmp/fm-local-pane.qFqeCj/home/data/backlog.md - move retired to Done, keep Done to the 10 most recent, then re-scan Queued and dispatch only work whose blockers are gone and date is due.
ok - teardown removes a landed worker pane and refuses to discard unlanded work
ok - teardown refuses before removing anything when it cannot stop a live agent
WARNING: watcher still down (same stale episode; last beat: never, grace 300s) - full banner already printed this episode.
teardown live complete (window fm-lab-pane-cleanup-on-2109762-14524:w1:pD, worktree /tmp/fm-local-pane.qFqeCj/retired-worktree)
Backlog: live just finished (this home keeps no markdown backlog at /tmp/fm-local-pane.qFqeCj/home/data/backlog.md). Update /tmp/fm-local-pane.qFqeCj/home/data/backlog.md - move live to Done, keep Done to the 10 most recent, then re-scan Queued and dispatch only work whose blockers are gone and date is due.
ok - teardown stops a live agent, then closes its proven-gone pane
{"session":{"default":false,"name":"fm-lab-pane-cleanup-on-2109762-14524","running":false,"session_dir":"~/.config/herdr/sessions/fm-lab-pane-cleanup-on-2109762-14524","socket_path":"~/.config/herdr/sessions/fm-lab-pane-cleanup-on-2109762-14524/herdr.sock"},"stopped":true}
closed restored pane=fm-lab-pane-cleanup-on-2109762-14524:w1:pE
closed unlanded pane=fm-lab-pane-cleanup-on-2109762-14524:w1:pB
ok - recovery closes restored bare shells by task identity after a session restart
Evidence: Relaunch adoption of a pane that outlived its stopped server

Source: Relaunch adoption of a pane that outlived its stopped server

ok - fm-control relaunch: a same-harness relaunch replaces the agent in the same endpoint and worktree
ok - fm-control relaunch: pending composer text refuses before the exit command is typed
ok - fm-control relaunch: an unreadable composer fails safe before the exit command is typed
ok - fm-control relaunch: a linked spawning home preserves committed and unfinished work in the recorded copy
ok - fm-control relaunch: durable task metadata survives replacement launch publication
ok - fm-control relaunch: delivery and concurrent task metadata publication serialize
ok - fm-control relaunch: disabling tracing clears metadata and pane context
ok - fm-control relaunch: progress and the Firstmate-worktree worker identity reach the replacement
ok - fm-control relaunch: a ship task refuses without the progress note its replacement needs
ok - fm-control relaunch: switching harness is one ordinary relaunch, and the old wiring goes with the old agent
ok - fm-control relaunch: a harness switch resets model and effort unless they are named too
ok - fm-control relaunch: a prefixed recorded harness can switch adapters transactionally
ok - fm-control relaunch: a prefixed command requires an explicit replacement harness
ok - fm-control relaunch: a same-harness relaunch keeps the profile axes it was running with
ok - native Ultra relaunch preserves its profile and rejects an unsupported model before stopping
ok - fm-control relaunch: explicit model and effort win over the recorded ones
ok - fm-control relaunch: refuses to relaunch onto an adapter with no verified mechanics
ok - fm-control relaunch: the retired incarnation's global turn-end token is revoked
ok - fm-control relaunch: wiring cleanup failure refuses replacement arming
ok - fm-control-lib: one owner resolves each harness's turn-end registry entry, and refuses a malformed token
ok - fm-control relaunch: a secondmate relaunch re-resolves its durable configured harness pin
ok - fm-control relaunch: invalid configured effort is ignored before stop
ok - fm-control relaunch: an adapter unverified for this task kind refuses before the agent is stopped
ok - fm-control relaunch: explicit secondmate harness resets unnamed profile axes
ok - fm-control relaunch: a ship task keeps its recorded harness instead of re-reading crew config
ok - fm-spawn --relaunch: with no explicit harness it reuses the task's recorded one, never the crew default
ok - fm-promote/fm-spawn --relaunch: the current ship contract supersedes stale scout delivery text
ok - fm-spawn --relaunch: wiring armed under a prefixed harness name is still retired
ok - fm-spawn --relaunch: switching away from muse retires its session binding
ok - fm-spawn --relaunch: switching away from cursor retires its session binding
ok - fm-control relaunch: an unaccountable local copy refuses before the agent is touched
ok - fm-control relaunch: a worker with nothing to work from is never launched
ok - fm-control relaunch: a refusal before the agent is stopped leaves the durable record untouched
ok - fm-control relaunch: checkpoint inspection failures refuse before stopping
ok - fm-control relaunch: a launch failure after the stop keeps the prior record and reports the real state
ok - fm-control relaunch: unpublished rollback keeps concurrent durable metadata
ok - fm-control relaunch: post-publication failure keeps the new durable record
ok - fm-control relaunch: partial stop reconciles actual agent state
ok - fm-control relaunch: failed journal replacement preserves durable phase
ok - fm-spawn relaunch: prepublication abort removes replacement state
ok - fm-control relaunch: the checkpoint records the exact unlanded work it preserved
ok - fm-control relaunch: a secondmate's child work is accounted for and its charter is left alone
ok - fm-control relaunch: a secondmate home that is not this secondmate's is refused
ok - fm-control relaunch: unreadable and untraversable child state fails checkpoint
ok - fm-control relaunch: two control actions on one task serialize instead of interleaving
ok - fm-spawn relaunch: direct entry participates in lifecycle serialization
ok - fm-promote: promotion participates in lifecycle serialization
ok - fm-spawn --relaunch: refuses to launch a second agent into a live endpoint
ok - fm-spawn --relaunch: symlinked records refuse before inspection
ok - fm-spawn --relaunch: keeps its early meta lock continuous
ok - fm-spawn --relaunch: pending closes refuse before replacement begins
ok - fm-spawn --relaunch: every identity axis comes from the record, and a contradicting flag refuses
ok - fm-spawn --relaunch: an unrecorded task is refused
ok - fm-spawn --relaunch: refuses to start a replacement outside the copy holding its work
ok - tmux: a window absent from its session refuses both verbs rather than being assumed gone
ok - tmux: an unfindable session refuses both verbs, so a live agent is never duplicated
ok - tmux: a dead server on this socket refuses both verbs rather than proving absence
ok - reclaim: an unclassifiable endpoint is still refused, so two agents cannot share one
ok - reclaim: a herdr pane that outlived its stopped server is adopted, never orphaned beside a new tab
ok - fm-control exit: a herdr pane that outlived its stopped server is already-stopped, not gone
ok - resume launch: a stopped herdr server cannot hide the task's leftover pane
ok - fm-control resume: a stopped herdr server is re-read before its leftover pane is judged gone
ok - reclaim: a herdr rebind is created in the session the record names, never the ambient one
ok - reclaim: a herdr agent that came back with its server refuses, so one worktree keeps one agent
ok - reclaim: a herdr reclaim rebinds the endpoint and leaves the whole rest of the task alone
ok - reclaim: a herdr secondmate whose endpoint is gone is sent to its own respawn owner
ok - reclaim: a rebind refused from a plain shell reports the real cause, not a fabricated session mismatch
ok - relaunch re-reads the backlog item instead of blindly re-running the transition
ok - relaunch heals an item that drifted out of In flight while the task stayed live
Evidence: Retire archive of merge-authority and progress records

Source: Retire archive of merge-authority and progress records

ok - unsafe PR merge markers refuse before endpoint or pool changes
ok - retirement proves the Atlas task identity before any ticket or node mutation
ok - parent refusal retains the slot and endpoint; a repaired channel permits cleanup
ok - an outcome written before the endpoint stopped reaches the parent before the record retires
ok - cleanup-late-outcome-retry: an undelivered late outcome completes teardown after the slot owner exits and the watcher delivers it
ok - cleanup-late-outcome-retry: an older failed outcome never follows a newer done outcome of the same task
ok - cleanup-late-outcome-retry: a reused task id delivers the older incarnation outcome once, before the replacement outcome
ok - cleanup-late-outcome-retry: a teardown rerun delivers the owed older outcome before the replacement outcome, once each
ok - cleanup-late-outcome-retry: a late report that overlaps another retry delivers its outcome once, without a false warning
ok - cleanup-late-outcome-retry: a relaunch with the same ledger delivers its owed outcome once
ok - a refused Treehouse return keeps the status log, busy state, and record for a rerun
ok - busy generation refusal retains the leased slot and endpoint
ok - finished scout retirement preserves the live slot, record, branch, dirty work, and process
ok - unfinished, unowned, ambiguous, decision-held, and unlanded records refuse retirement
ok - Atlas refusal keeps a keyed blocker and the record; retry closes only the old ticket and backlog
ok - retirement releases and lands a node that no other open ticket holds
ok - landed ship retirement keeps Git refs and archives only its own task artifacts
ok - retirement refuses while another lifecycle action owns the live task
Evidence: Pending-reply test transcript

Source: Pending-reply test transcript

ok - normal correlated reply resolves once (idempotent)
ok - completed turn with no report triggers exactly one recovery
ok - recovery attempts reconcile without reinjection
ok - recovery reply resolves the original expectation
ok - second missed turn escalates once and remains durable
ok - escalations and replies wake; the home's own escalation close stays quiet
ok - failed escalation publication remains retryable and publishes once
ok - legacy escalation closes under the shared default key
ok - legacy escalation cannot close an unrelated default-key decision
ok - foreign correlated blocker cannot impersonate a pending-reply escalation
ok - concurrent resolution closes one keyed escalation exactly once
ok - concurrent escalation yields to a late correlated reply
ok - transport success cannot masquerade as reply success
ok - undelivered records remain immutable across scan paths
ok - delivery confirmation fallback reconciles durably
ok - delivery confirmation serializes with reconciliation
ok - unrelated events and stale correlation ids cannot resolve
ok - restart preserves expectation and exact parent destination
ok - wrong-home reports are detected but do not silently acknowledge
ok - direct unmarked captain input creates no expectation
ok - fm-send marked secondmate path creates pending and embeds corr
ok - status-pointed document resolves the expectation
ok - optional helper report resolves without being required for correctness
ok - backend busy/idle observation covers Pi/Claude paths without conversation scrape
ok - tmux and zellij unknown states use bounded capture fallback
ok - pending replies scope Kimi capture fallback by recorded harness
ok - tick skips terminal records and reuses target observations
ok - correlations are reused only for matching open task records
ok - tick end-to-end: miss -> one recovery -> escalate -> durable
ok - failed transport discards undelivered expectation only
ok - fm_pending_reply_new_id is set -u safe and returns valid id without openssl
ok - a remote repost waits for the reply channel and still fires on a real miss
ok - a mirrored correlated remote reply resolves without any repost
ok - same-basename self-home corr= is restated onto the parent channel and resolves
ok - same-basename reply resolves at the recovery failure boundary
ok - a child-file mate-home sighting is not copied and still escalates
ok - mechanical helper writes the parent channel from verb, corr, and note
ok - remote parent-replies.status is not classified as wrong-home
ok - local parent-replies.status remains wrong-home evidence
ok - an escalated correlation stays retryable only while undelivered
ok - all pending-reply tests passed
Evidence: Lint partition listing (1of4..4of4 coverage and rejection)

Source: Lint partition listing (1of4..4of4 coverage and rejection)

1of2 rc=0 roots=241
2of2 rc=0 roots=242
1of4 rc=0 roots=121
2of4 rc=0 roots=121
3of4 rc=0 roots=120
4of4 rc=0 roots=121
full inventory (CI context) roots=483
4-way union=483 unique=483
2-way union=483 unique=483
4-way and 2-way splits cover the same inventory
reject 5of4 rc=2 : fm-lint.sh: --partition must be 1of2, 2of2, or 1of4 through 4of4, got 5of4.
reject 0of4 rc=2 : fm-lint.sh: --partition must be 1of2, 2of2, or 1of4 through 4of4, got 0of4.
reject 1of3 rc=2 : fm-lint.sh: --partition must be 1of2, 2of2, or 1of4 through 4of4, got 1of3.
reject 3of2 rc=2 : fm-lint.sh: --partition must be 1of2, 2of2, or 1of4 through 4of4, got 3of2.
heavy roots listed: 6
1of4 heavy roots: 3
2of4 heavy roots: 3
3of4 heavy roots: 0
4of4 heavy roots: 0
Evidence: Heavy lint roots run alone per partition

Source: Heavy lint roots run alone per partition

partition 1of4: rc=0 shellcheck-invocations=5 every-listed-root-checked-once=yes
   heavy root bin/fm-procevent-remote-reply.sh -> invocations containing it: 1, invocations where it is alone: 1
   heavy root tests/fm-remote-reply.test.sh -> invocations containing it: 1, invocations where it is alone: 1
   heavy root tests/fm-pending-reply.test.sh -> invocations containing it: 1, invocations where it is alone: 1
partition 2of4: rc=0 shellcheck-invocations=5 every-listed-root-checked-once=yes
   heavy root bin/fm-teardown.sh -> invocations containing it: 1, invocations where it is alone: 1
   heavy root bin/fm-pr-merge.sh -> invocations containing it: 1, invocations where it is alone: 1
   heavy root bin/fm-send.sh -> invocations containing it: 1, invocations where it is alone: 1
partition 3of4: rc=0 shellcheck-invocations=2 every-listed-root-checked-once=yes
partition 4of4: rc=0 shellcheck-invocations=2 every-listed-root-checked-once=yes
Evidence: Away posture alone never authorizes a merge

Source: Away posture alone never authorizes a merge

away record files in state: .afk-contract 
afk-contract status: Usage:
afk-contract status:   fm-afk-contract.sh propose [--words-file <path> | --words <text>]
afk-contract status:       [--expected-return <UTC ISO 8601>] [--spend <n>]
--- case away-yolo-off (away record present, yolo=off, no --captain-authorized): exit 1
stderr: error: merge refused for task task-x1: yolo=off (expected on or --captain-authorized for an explicit captain merge instruction)
forge merge call: none
merge-authority record: none
away record files in state: .afk-contract 
afk-contract status: Usage:
afk-contract status:   fm-afk-contract.sh propose [--words-file <path> | --words <text>]
afk-contract status:       [--expected-return <UTC ISO 8601>] [--spend <n>]
--- case away-yolo-on-control (away record present, yolo=on, no --captain-authorized): exit 0
stderr: ●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
stderr: ●  WATCHER DOWN - SUPERVISION IS OFF
stderr: ●  1 task(s) in flight, but no watcher has a fresh beacon (last beat: never, grace 300s).
stderr: ●  Trust the emitted supervision protocol for this harness; do not use shell & for watcher repair.
stderr: ●  This is a supervision warning only; the guarded operation WILL still run.
stderr: ●  repair missing watcher supervision according to the session-start block for this harness; do not use shell &.
forge merge call: MADE
merge-authority record: written
not ok - away-yolo-on-control: away posture alone authorized a merge
control yolo=on: merged (driver refusal assertion tripped as expected, exit 1)
away record files in state: .afk-contract 
afk-contract status: Usage:
afk-contract status:   fm-afk-contract.sh propose [--words-file <path> | --words <text>]
afk-contract status:       [--expected-return <UTC ISO 8601>] [--spend <n>]
--- case away-yolo-absent (away record present, yolo=<absent>, no --captain-authorized): exit 1
stderr: error: merge refused for task task-x1: yolo=<missing> (expected on or --captain-authorized for an explicit captain merge instruction)
forge merge call: none
merge-authority record: none
ok - away posture alone never authorizes a merge
Evidence: History ancestry proof

Source: History ancestry proof

HEAD=609b3e6ee09faf13cff6eba71653b33fdc271c47
ancestor-of-HEAD: af823b7f569479c0a5057e127e904df55c2ec4b4 Merge pull request #43 from BenWilcox8/fm/watcher-staging-loop-r1
ancestor-of-HEAD: aec043c718ef186a1d28c607fd5852b041e24435 test: fix Claude session-start drain live E2E (#5165)
ancestor-of-HEAD: 712aed80 Merge fork main with Atlas close-out gates and the agent limit
ancestor-of-HEAD: 534a987a Merge fork main with the one-worker lint job and park and resume
ancestor-of-HEAD: c91b03a5 merge: fork main a21721c9 (idle Codex doorbell fix, PR 49)
ancestor-of-HEAD: d5fbb6c2 merge: fork main 5607ed74 (typed delivery provenance, PR 50)
ancestor-of-HEAD: 0089320a merge: fork main b2ef7600 (idle Codex frames without prompt-row cells,
ancestor-of-HEAD: eeb841b9 merge: fork main e1b269ed (cleanup slot gates PR 51, typing provenance
ancestor-of-HEAD: cc9045e0 merge: fork main fe5909bb (pane cleanup PR 52, restart recovery PR 56)
ancestor-of-HEAD: d978c181 merge: fork main d2cda540 (restore working workers after a restart, PR
ancestor-of-HEAD: d2cda540 feat(bin): restore working workers after a fleet restart (#57)
integration merge 34324044 parents: af823b7f569479c0a5057e127e904df55c2ec4b4 aec043c718ef186a1d28c607fd5852b041e24435
non-merge-linear check: merges in range af823b7f..HEAD = 9
Evidence: AGENTS.md ceiling check

Source: AGENTS.md ceiling check

$ bin/fm-agentsmd-ceiling.sh check
AGENTS.md is 29219 estimated tokens, within ceiling 29220 (1 to spare).
rc=0
tracked ceiling: 29220

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 1 issue found → auto-fixed ✅
  • ⚠️ tests/fm-local-pane-recovery.test.sh:273 - The new regression's no-close assertion is not enforced. Line 273 is [ ! -e &#34;$tmp/closed&#34; ] &amp;&amp; cmp -s ... inside a for loop. Under set -e, bash does not exit when a command before the final &amp;&amp; fails, and it does not exit when a loop returns non-zero for that reason. I checked this with bash -c &#39;set -euo pipefail; for x in 1; do [ ! -e /etc/passwd ] &amp;&amp; cmp -s a a; done; echo reached&#39;, which prints reached. Scenario: a later change makes the refused relaunch run pane close (the fixture touches $tmp/closed) before it refuses. The meta cmp is skipped, the loop continues, and the test still prints 'ok - relaunch with no home workspace refuses...'. The test then does not prove that the recycled pane is left alone. Fix: split the line into two separate commands, or use [ ! -e &#34;$tmp/closed&#34; ] || { echo &#39;not ok - ...&#39; &gt;&amp;2; exit 1; } followed by cmp -s ... || { ...; exit 1; }. The same weak pattern exists in older fork lines of this file, outside this diff: :183, :191, :224, :232, :250. It is optional to harden them in the same edit.

🔧 Fix applied.
✅ Re-checked - no issues remain.

✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 7 of 16 scenarios driven live against the product
Scenario Result Live Evidence
Real Herdr: exit, interrupt, and relaunch of worker panes work end to end, and a stale registration does not block a relaunch into a new pane (restored kunchenguid#4115 section) ✅ pass live tests/fm-control-herdr-smoke.test.sh in a guarded Herdr lab session -> fm-control-herdr-smoke.log
Real Herdr: after a lab server restart, authorized second mates and the primary come back in their own panes, a dormant mate stays down, and there is no double launch or repeat recovery ✅ pass live tests/fm-local-restart-recovery-herdr-e2e.test.sh -> fm-local-restart-recovery-herdr-e2e.log
Real Herdr: after a restart only the working worker is restored in its own slot, and a relaunch into a worktree leased to another task refuses ✅ pass live tests/fm-local-worker-restore-herdr-e2e.test.sh -> fm-local-worker-restore-herdr-e2e.log
Real Herdr: pane cleanup closes ended or husk panes, keeps live, supervisor, and unmanaged panes, and teardown refuses to discard unlanded work ✅ pass live tests/fm-local-pane-cleanup-e2e.test.sh -> fm-local-pane-cleanup-e2e.log
Adversarial: with the home workspace absent, a relaunch refuses a recorded pane id that now belongs to other work, closes nothing, and keeps the task record ⏸️ untested no The prior payload did not establish a live result: the recycled pane-id state was built only with a fake Herdr in tests/fm-local-pane-recovery.test.sh, and the lab tests do not produce that exact stat…
Adversarial: the hardened pane-recovery assertions fail the test when the guard is removed or when the pane is closed before the refusal ⏸️ untested no The prior payload did not establish a live result: the mutation runs were on /tmp copies and test the test itself, not the running product. No live product surface exists for this check.
A Herdr pane that outlived its stopped server is adopted on relaunch and reads as already-stopped on exit ⏸️ untested no The prior payload did not establish a live result: the stopped-server state was built only with a fake Herdr in tests/fm-control-relaunch.test.sh. To drive it live, stop a Herdr lab server while a pan…
Retiring a reassigned-slot task archives its merge-authority and progress records and leaves a live task's record alone ⏸️ untested no The prior payload did not establish a live result: tests/fm-local-cleanup-slot.test.sh used a fixture home and a fake forge, and no real retired task existed to retire. To drive it live, retire a real…
Lint CLI accepts 1of4..4of4, the four partitions cover the full CI inventory with no overlap (same as the 2-way split), and other split names are refused ✅ pass live bin/fm-lint.sh --partition KofN --list-files -> lint-partition-listing.txt
Each heavy lint root runs alone in exactly one ShellCheck invocation, and every root is still checked once ⏸️ untested no The prior payload did not establish a live result: real ShellCheck was replaced by a recording stand-in, because host rules forbid local whole-partition lint. The real four-partition lint runs in CI.
The pending-reply test still passes with its library source directives pointed at /dev/null ⏸️ untested no The prior payload did not establish a live result: tests/fm-pending-reply.test.sh is a unit and fixture test, and the ShellCheck memory gain is measured only in CI.
Adversarial: with an away record confirmed, a merge with yolo=off or no yolo field and no --captain-authorized refuses before any forge call; the yolo=on control merges ⏸️ untested no The prior payload did not establish a live result: bin/fm-pr-merge.sh ran against the fixtures from tests/fm-pr-merge.test.sh with a faked forge. A real merge needs a live GitHub PR and captain author…
Branch keeps both histories: the integration merge has the two pinned parents, and upstream aec043c and every fork main merge (PRs 44-57) are ancestors of HEAD ✅ pass live git ancestry commands -> ancestry-proof.txt
AGENTS.md stays within the tracked token ceiling without raising it ✅ pass live bin/fm-agentsmd-ceiling.sh check -> agentsmd-ceiling.txt (29219 of 29220)
Park and resume a real worker across a Herdr restart (a --resume-session relaunch opens a new endpoint in the recorded worktree) ⏸️ untested no tests/fm-park-resume-live-e2e.test.sh was not run. This test is opt-in because it spends model tokens with real Claude/Codex/Pi credentials. To run it, set FM_PARK_RESUME_LIVE_E2E=1 on an operator-app…
Codex worker lifecycle in a live Herdr lab (Codex hook isolation, idle doorbell) ⏸️ untested no tests/fm-control-herdr-codex-live-e2e.test.sh was not run. Per the host note, a pending Codex update prompt blocks the live Codex Herdr test on this machine. To run it, finish the Codex CLI update, th…
  • nice -n 10 bash tests/fm-local-pane-recovery.test.sh (all cases pass, including the new recycled-pane relaunch refusal)
  • Mutation run 1: the absent-workspace guard in bin/fm-local-pane-lib.sh reverted in a /tmp copy; the same test fails (rc=1, not ok - recycled relaunch output)
  • Mutation run 2: a herdr pane close injected before the refusal in a /tmp copy; the same test fails (rc=1) at the hardened no-close assertion
  • nice -n 10 bash tests/fm-control-relaunch.test.sh
  • nice -n 10 bash tests/fm-local-cleanup-slot.test.sh
  • nice -n 10 bash tests/fm-pending-reply.test.sh
  • nice -n 10 bash tests/fm-local-restart-recovery-herdr-e2e.test.sh (live, guarded Herdr lab)
  • nice -n 10 bash tests/fm-local-worker-restore-herdr-e2e.test.sh (live, guarded Herdr lab)
  • nice -n 10 bash tests/fm-local-pane-cleanup-e2e.test.sh (live, guarded Herdr lab)
  • nice -n 10 bash tests/fm-control-herdr-smoke.test.sh (live, guarded Herdr lab)
  • bin/fm-lint.sh --partition {1of2,2of2,1of4..4of4} --list-files, compared with CI=true bin/fm-lint.sh --list-files, plus rejection of 5of4/0of4/1of3/3of2
  • FM_LINT_JOBS=1 bin/fm-lint.sh --partition KofN for each of the four partitions, with a recording ShellCheck stand-in on PATH, to observe the invocation grouping
  • Driver on the fixtures in tests/fm-pr-merge.test.sh: bin/fm-pr-merge.sh with a confirmed away record, run with yolo=off, with yolo absent, and with a yolo=on control
  • git merge-base --is-ancestor for af823b7f, aec043c7, and fork merges 712aed80..d978c181; git rev-list --parents for integration merge 34324044
  • bin/fm-agentsmd-ceiling.sh check
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

NewAiCoder-bot and others added 30 commits September 7, 2026 19:34
…nchenguid#3946)

* fix(bin): derive the away-mode beacon grace from the poll cadence

fm-turnend-guard.sh's away-mode branch required the watcher beacon to be
fresh within the flat FM_GUARD_GRACE default (300s), but the daemon starts a
fresh one-shot watcher only after it finishes handling the previous wake, and
that handling can legitimately outrun a fixed 300s window under load (a slow
registered check, a busy supervisor pane) with the daemon perfectly healthy
throughout. That misread a live, correctly-cycling daemon as down and blocked
the turn.

Add fm_poll_derived_grace, the single owner of the max(300, FM_POLL + 60)
formula, and have the away-mode branch, fm-claude-stop-autoarm.sh, and
fm-watch.sh's own runtime beacon-staleness check all derive their default
grace from it instead of the flat default. A dead daemon pid or a beacon
older than that grace still blocks, so a genuinely lapsed away mode still
alarms; every other check is unchanged.

fm-claude-stop-autoarm.sh computed the derived grace into GRACE but its two
fm-watch-arm.sh invocations called the wrapper bare, so the wrapper fell back
to its own flat 300s default and could reject a healthy long-poll watcher.
Both invocations now pass FM_GUARD_GRACE="$GRACE" through explicitly, and a
new test proves a long FM_POLL with FM_GUARD_GRACE unset reaches
fm-watch-arm.sh with the derived value.

Also drops fm_last_activity_age, added alongside the derivation but never
called anywhere in the tree; fm-inactive-reconcile.sh already owns that
computation.

* no-mistakes(review): Remove dead WATCHER_STALE_GRACE assignment in fm-watch.sh

* no-mistakes(document): Update FM_WATCHER_STALE_GRACE default note for poll-derived grace

---------

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
…henguid#3904)

* fix(procevent): bind a source runner to the session that owns it

A process-event source runner is detached into its own process group so a
persistent source survives the turn that armed it. Nothing bounded that
detachment, so a runner could reparent to init and keep its blocking child -
and every process that child spawned - running with nothing left to reap it.
One such runner outlived its home for about a day; the cost was not the runner
but the exec churn of the poll stubs under it, which stalled every fresh
process launch on the host.

Each runner now starts a small guard beside it, in a separate process group,
that re-reads its home's process-event lease and stops the runner's whole
process group once that lease can no longer be proved fresh. Every ordinary
entry point an owning session runs refreshes the lease, and the watcher's
reconcile cycle keeps it fresh in a live home; nothing a runner spawns can
refresh it, so a source cannot certify its own owner. Scope is the owning state
root and one runner generation, never a script or process name, so a live
source in another home is untouched and a live home simply starts a
replacement runner on its next cycle.

The test scaffolding that starts real runners could not reap them either: the
bearings-board and board-render suites tracked their homes in a shell array
appended to inside a command substitution, so the array was always empty and
every listener they started survived the run. Home registration moves to a
`$$`-keyed registry in tests/lib.sh, which sweeps it from every cleanup path,
now including HUP and QUIT, and the blocking fixture stubs stop themselves at a
bound so an escaped one cannot keep spawning processes indefinitely.

Adds a regression test that reproduces the orphan shape - a reparented listener
with a live descendant tree under it - and proves the whole group and its
process churn stop once its session is gone, that an identical listener in a
home whose session is still there is untouched, and that retirement still
reaches a reparented listener and everything under it.

* no-mistakes(review): Bound source launches and fail closed on guard startup

* no-mistakes(document): Document runner lease and storm containment

* no-mistakes(ci): Fixed the Greptile watchdog finding: failed runner cleanup now retries on each watchdog tick instead of abandoning the orphaned process group. The shared state-root lease behavior remains unchanged because it is an explicitly accepted ownership policy. Verified with bash syntax checks, git diff checks, and the complete fm-procevent test suite

* test(procevent): pin that an unprovable stop is retried, not abandoned

The owner guard used to call stop_runner_pid and exit unconditionally, so a
stop it could not prove - a descendant still finishing uninterruptible work
outlives even the group signal, and an unreadable process identity proves
nothing - left a still-running expired runner with nothing watching it. That is
the best-effort reaping this mechanism exists to remove, and the fix that made
the guard retry landed without a test holding it in place.

The unprovable attempt is injected through the signal the real path actually
reads: `ps` answers exactly one process-group query for the runner with a group
it does not lead, which is how a stop that cannot be proved is reported, and
every other call is the real command. The test also asserts that the injected
attempt happened, so it cannot pass vacuously if the fixture stops arming.

Fails against the exit-after-one-attempt guard, where the runner survives its
expired lease, and passes once the guard retries on its check cadence.

* docs(procevent): scope the no-self-refresh rule to confused-agent grade

The runner-lease documentation asserted as an absolute that nothing a runner
spawns can refresh the lease, so a source cannot certify its own owner.
That overclaims what the inherited FM_PROCEVENT_IN_RUNNER marker actually
enforces.

The marker holds at confused-agent grade: a runner and its ordinary children
inherit it and skip every refresh, which is exactly the accidental case this
boundary exists for.
A source that deliberately strips the marker from its environment can still
refresh, so adversarial-grade unforgeability is explicitly out of scope and
tracked as separate follow-up design work.

This states the real scope in docs/configuration.md, which owns the operating
contract, and corrects the two matching comments in bin/fm-procevent.sh.
The process-event-sources skill keeps its cross-reference and gains one line
in its never-to-be-claimed list so the overclaim is not reintroduced from the
agent-facing side.
The lease mechanism itself is unchanged.

* no-mistakes(review): Fix process-event lease and launch pacing edge cases

* no-mistakes(review): Scope launch pacing and clarify lease boundaries

* no-mistakes(review): Reap leftover groups and use monotonic launch pacing

* no-mistakes(review): Keep guards alive across runner PID reuse

* no-mistakes(review): Use monotonic leases and simplify launch generation identity

* no-mistakes(review): Prevent pacing identity reuse and bound reused-group guards

* no-mistakes(review): Preserve active pacing state on failed registration

* no-mistakes(review): Reap reused runner groups with registration evidence

* no-mistakes(review): Avoid ambiguous group kills and encode pacing identities

* no-mistakes(review): Abort kill escalation after runner identity reuse

* no-mistakes(review): Gate group signals and prune stale pacing state

* no-mistakes(review): Document bounded PID reuse signaling safety

* no-mistakes(review): Align leaderless group ambiguity guidance

* no-mistakes(review): Expire reboot stamps and preserve publication success

* no-mistakes(review): Bind owner leases to physical state roots

* no-mistakes(document): Clarify process-event lease and pacing contracts

* fix(procevent): drop a platform-dependent post-TERM test assertion

CI ran red on two lanes that the local gate could not see.

Lint failed with SC2034 on two reads in cmd_owner_watchdog that
destructure the state-root identity into five fields while using only the
device and inode.
Local changed-file mode suppresses the cross-file codes that need
--external-sources, so the warning cleared the pre-push lint step and
failed CI's full analysis, exactly as bin/fm-lint.sh's header describes.
The unused fields now read into `_`.

The behavior shard failed on this suite's own post-TERM assertion, which
required the stubbed identity source to be consulted more than once.
Whether that happens is platform-dependent: where the runner leader keeps
waiting on its TERM-ignoring source child, the post-TERM check sees a live
leader whose identity no longer matches, and where the leader dies
promptly it sees a leaderless group carrying the same numeric id.
fm_procevent_pid_state reaches that second verdict without consulting
process identity at all, so the identity source is never read twice and
the count assertion fails through no fault of the behavior.

The case now asserts the invariant both forms share: retirement refuses,
and the ambiguous group is not signalled.
Scoping a mutation to this fixture and making the refusal signal instead
confirms the case still fails, so dropping the count does not leave it
passing vacuously.

* no-mistakes(review): Prevent superseded runners recreating stale pacing stamps

* no-mistakes(document): Document pacing and ambiguity boundaries

* fix(procevent): retire under the recorded identity source and state the home-scoped lease

The reused-group case started its runner with the proc-root override in
place, so the runner recorded a ps-derived identity, then retired it
without that override.
Where /proc exists the retirement read identity from a different source
than the one recorded, the guard correctly refused an identity it could
not confirm, and the case failed on Linux while passing on macOS.
It now retires under the same source, and clearing the stub marker first
turns that cleanup into the complementary assertion: once the ambiguity is
gone, retirement reaps the whole group instead of leaving it behind.

The lease prose claimed a runner is bound to the session that owns it,
while the mechanism binds it to the home.
That gap is what makes a replacement session or an inspection command look
like a defect: any activity in the same home refreshes the lease.
The granularity is deliberate, because a persistent source is meant to
outlive the session that armed it, and binding a runner to that session
would stop the sources this mechanism exists to keep running.
A runner whose source is no longer wanted in a live home is stopped by
reconcile when that source is retired, independently of the lease, so the
lease is the backstop for a home that is gone - the torn-down sandbox this
change bounds - and the residual is recorded as a known limit.

* no-mistakes(review): Rate-limit polls and skip superseded runner launches

* fix(procevent): build the claim-only sweep case as a runnerless owned claim

A superseded generation now observes the registration-identity mismatch,
self-retires, and releases its claim, which is the behavior we want: it
clears its own residue rather than leaving a claim with no runner for the
home sweep to find.

The claim-only sweep case was built by deleting a registration out from
under a live runner, which used to leave that runner in place. It now
makes the runner retire itself, so the sweep raced that exit and retired
one source or two depending on which won. The case failed three runs in
four, alternating between a preflight-count failure and `attempted=1`.

It now builds the state it means to test: kill the runner's group so it
cannot run its own cleanup, assert the owned claim survived that kill, and
only then drop the registration. Coverage is unchanged - a runnerless
owned claim must still be swept - and the result no longer depends on
whether the runner had exited yet. Three consecutive runs pass.

The superseded exit also skipped the runner-marker cleanup the normal path
performs. The marker is written before the launch floor is waited on, and
a home sweep counts a marker with no owned claim as a preflight failure,
so exiting without clearing it would make that home refuse to sweep.

`FM_LAVISH_POLL_RETRY_DELAY= ` trips SC1007 under the full analysis CI
runs, though not under the changed-file mode the pre-push gate uses.

* test(procevent): retire a quiet reparented listener instead of racing a storm

Explicit retirement was exercised against the spawn-churning stub, which
made it nondeterministic. Retirement refuses rather than signalling when it
cannot confirm the runner's identity, that identity is read through `ps`,
and the stub's 0.1s spawn loop starves that read often enough that a single
attempt is a race - the suite failed on this case roughly one run in four,
reporting `cannot confirm runner identity; source remains registered`.

The refusal is correct: it is the documented preserve-for-retry contract,
and a separate case already asserts it. So this is a fixture problem, not a
behavior problem.

The storm is still covered where the evidence for it lives. The owner-loss
home keeps the churning stub and still asserts its tick log stops, which is
what proves the churn ended rather than one pid going away. The retirement
home never asserted ticks; it only ever read the descendant pid, so the
spawn loop bought this case nothing while costing it determinism.

It now uses a quiet stub that still reparents and still holds a real
descendant in its process group, so the assertions are unchanged: retiring
the source must reap the reparented listener's whole group and the
descendant under it. Four consecutive runs pass.

* no-mistakes(review): Serialize registration replacement through source child launch

* chore(no-mistakes): require honest test-step scenario marking

The test step recorded scenarios as passing that were only reached through
a stubbed dependency or the executable suite, and its validator refused
them, because `pass` asserts a scenario was verified against the real live
product.

That refusal is correct, so the fix is to mark honestly rather than to
weaken the gate: a scenario driven live stays a pass and cites its live
transcript, while one reached only through a stub or the suite is recorded
as untested with the reason and a pointer to its executable coverage.
Untested scenarios are reported rather than treated as failures, so real
coverage stays visible without claiming verification that did not happen.

The instruction also forbids dropping a scenario to avoid marking it
untested, since that would hide the gap instead of stating it.

* no-mistakes(review): Remove unrelated test scenario policy

* no-mistakes(document): Clarify process-event home lease documentation

* no-mistakes(document): Correct owner guard failure wording
…nguid#4037)

* test(lib): set fixture mtimes through one portable epoch helper

On macOS the visible symptom was ONE red case in the turn-end guard suite. The
actual damage was TWO cases that had quietly stopped testing their subject. The
red one was the harmless half - people read one red case as one broken thing,
and here that intuition is wrong.

`touch -d @<epoch>` is a GNU extension; BSD touch rejects it outright and leaves
the file at its current mtime. So on macOS the three away-mode beacon cases
never aged their beacon at all.

The 400s case exists to pin that 400s is stale under the flat 300s default but
fresh under the poll-derived grace (660s at FM_POLL=600). Deleting the grace it
guards (FM_POLL=60, so max(300,120)=300) and re-running proves what it was
worth on this platform:

  pre-fix input (beacon left at now):  ok     - passes with the feature DELETED
  post-fix input (beacon 400s old):    not ok - expected exit 0, got 2

It was green while measuring nothing, and could not have caught a regression in
the grace it names. Only the 700s case broke loudly.

`touch -t [[CC]YY]MMDDhhmm[.SS]` is POSIX and both platforms accept it, so the
only host-specific step left is formatting the epoch into that stamp, which
date(1) spells two incompatible ways. fm_touch_epoch in tests/lib.sh owns that
probe once and fails loudly rather than leaving an unset timestamp behind - the
failure mode that caused this. Verified on BSD touch/date here and on GNU
coreutils 9.7 in a container.

Three real sites, and one consistency change - not four fixes. The stale
destination lock in tests/fm-remote-backlog-handoff.test.sh was never a defect:
its `uname = Darwin` branch made the `touch -d` line unreachable on macOS, and
BSD touch accepts that space-separated form anyway. `touch -t` takes the date
directly on both platforms, so the branch goes rather than standing as a second
copy of the same platform assumption.

Known limit: the third case (away mode off) is only HALF recovered here. It now
receives the input its name claims, but it is still insensitive after this fix -
its verdict is identical with a 0s and a 400s beacon, because the fixture
records a daemon lock and no watcher lock, and with away mode off the daemon
lock proves nothing. Not fixed here; tracked separately, with the requirement
that any fix be shown to FAIL when the protection is removed.

FULL SUITE ON macOS: 189 scripts, four red, none of them this change. The
turn-end guard and remote-handoff suites are clean. Attribution was established
by running the four failures at the base commit and at this head on an idle
machine, because base-idle against head-under-load moves two variables at once:

  script                  head/loaded  base/idle  head/idle  verdict
  fm-calm-pi-extension    red          red        red        pre-existing
  fm-backlog-atomicity    red          red        red        pre-existing
  fm-procevent            red          red        red        pre-existing
  fm-startup-network      red          green      green      cause unestablished,
                                                             load-sensitive under
                                                             a full run

Reported, not fixed. fm-calm-pi-extension deserves its own note: it FAILS
because Chrome is absent instead of declaring the capability it needs and
standing aside, so its verdict is about the machine rather than its subject -
the same family as the defect above, with the red at least announcing itself.

Neighbouring class, reported not changed: `file_mode()` - a verbatim
`uname = Darwin ? stat -f %Lp : stat -c %a` - is copy-pasted across at least
five test scripts plus a `reread_mode` variant, and epoch-mtime reads are
open-coded as `date -r … || stat -c %Y` in three more; same one-owner shape as
the defect above. `git init` without `-b main` depends on the host's
init.defaultBranch in several scripts (branch-name case, tracked elsewhere).
timeout, sha256sum and sed -i uses are all correctly guarded where checked.

Observed while building the check rather than the fix: the first watcher I
wrote to wait for the suite matched its own command line, so it was waiting on
its own existence and could never fire. Same shape as the cases above -
machinery answering confidently about something other than its subject, by
including itself in the evidence it was meant to judge. The file sentinel it
was replaced with cannot be produced by the observer that reads it.

* fix(review): Pin fixture timestamps to UTC across DST transitions

* fix(document): Clarify shared fixture suite coverage
…ision (kunchenguid#4038)

* fix(pi): preserve native Codex effort and guarded supervision

* test(pi): identify native compatibility guard versions

* no-mistakes(review): share native-main follow rule between build and picker

* no-mistakes(document): document native progress marker and ultra effort owners

* no-mistakes(ci): Failing check "Behavior portable serial 1" was caused by this PR. The shard ran the default-on live guard tests/fm-pi-branch-responsiveness-live-e2e.test.sh, whose idle arm loads .pi/extensions/fm-branch-supervision.ts into a scratch project with a fixed list of copied libs. This PR added `import { registerFirstmateTool } from "./lib/fm-native-contract.ts"` to that extension and updated every other loading fixture's copy list, but missed this guard. Pi 0.85.1 therefore refused to load the extension ("Cannot find module './lib/fm-native-contract.ts'"), never drew its TUI, the test failed with "Pi 0.85.1 never drew its TUI in the idle arm", and the job hit its 20-minute cap. Fix (one line): added fm-native-contract to the lib copy loop in tests/fm-pi-branch-responsiveness-live-e2e.test.sh. Swept all other suites referencing fm-branch-supervision.ts / fm-primary-pi-watch.ts; the remaining ones without the new lib only hash, path-reference, or string-match the files and do not load them into Pi, so no further fixture changes are needed. Verification: reproduced the mechanism against the installed Pi 0.85.1 by building the lab copy with the old lib list (load error as above) and the fixed list (loads cleanly). tmux is not installed on this machine, so the live guard itself gate-skips locally ("skip: live: tmux absent") and could not be run end to end here; CI (which has tmux and Pi) will exercise it. shellcheck is clean on the edited file

---------

Co-authored-by: Talon Stark <talonstark@gmail.com>
…guid#4041)

* fix(herdr): step around a stale client the running server refuses

A remote host can carry a self-updated herdr in ~/.local/bin beside a
package-managed one, and the fixed remote-job PATH resolves ~/.local/bin
first. After the server upgraded to 0.9.0 (protocol 22) the stale 0.8.2
client (protocol 20) was answered with protocol_mismatch on every command,
which the read classifiers folded into `unreadable`: the live remote
secondmate read unknown, every doorbell into it failed, and both the spawn
and relaunch recovery paths refused, so the defect trapped itself.

The adapter's session-scoped CLI wrapper now recognizes that refusal, reads
status per session from each distinct herdr on PATH, adopts the first one
the running server reports compatible, retries once, and keeps it for the
process. The happy path makes no extra call and no other failure reselects.
An endpoint that still reads unreadable names the refused client, both
protocols, and the fix on stderr; the remote state read, fm-crew-state, and
the launch refusal carry that reason, and fm-remote-doctor reports the
selected client and rebinds the launch agent to it.

Regression coverage: fake two-client hosts in the herdr unit suite, the
doctor suite, the crew-state remote arm, and the real host-local control
script in the remote lifecycle e2e; the real-herdr smoke refreshes the
status shape the selection reads.

* no-mistakes(review): Reselect Herdr client after every protocol mismatch

* no-mistakes(review): Remove unrequired Herdr diagnostics and launch-agent rebinding

* no-mistakes(review): Scope cached Herdr clients to their selected session

* no-mistakes(review): Restrict herdr client selection to reactive CLI calls

* no-mistakes(document): Document session-scoped Herdr client reselection

* no-mistakes(document): Clarify Herdr client selection documentation

* no-mistakes(ci): Updated the trusted fm-remote-doctor.sh SHA-256 in bin/fm-remote-entrypoint.sh after the PR changed the doctor, restoring git-unavailable bootstrap authentication. Verified tests/fm-on.test.sh, tests/fm-backend-herdr.test.sh, bin/fm-lint.sh, and git diff --check all pass
* feat(afk): record the away posture and its lifecycle (phase 1)

Away mode becomes a posture of the one supervision session, recorded in
state/.afk-contract by the new bin/fm-afk-contract.sh: the one owner of the
record schema, the mandate-clause grammar and compiler, refusal naming the
missing part, the read-back rendering, the entry announcement (hold-for-return
only, no phone channel), and the archive at return. This release records
clauses and does not execute them; the announcement and return brief say so.

bin/fm-afk-launch.sh gains propose and confirm, confirms the record before any
daemon launch, refuses to launch the daemon on Pi and pi-signed, and archives
the record last on stop. bin/fm-afk-return.sh snapshots supervisor health
before shutdown, renders the return brief (health, mandate, waiting on the
captain, could not fix, handled, cost) from the archived record, the outcome
store, the held set, and the status logs, and shrinks the blocker gate to what
the away session could not fix.

While the record exists the watcher and the daemon never recheck an item held
for the captain. Declared external waits get a four-hour default cadence and
honor `until <UTC ISO 8601>` on the paused line, in both postures, bounded by
FM_PAUSE_UNTIL_MAX_SECS.

The /afk skill, AGENTS.md's layout and away-mode stub, the session-start
digest, and the architecture, Pi branch, configuration, and scripts docs
describe the record. The Pi/Herdr e2e now proves the no-daemon posture on a
real Pi primary; its verification record carries the 2026-09-08 run.

* no-mistakes(review): Fix AFK confirmation, grammar, waits, and return gating

* no-mistakes(review): Harden AFK authority and posture lifecycle

* no-mistakes(review): Preserve AFK history and tighten authority grammar

* refactor(afk): record clause fields with no natural-language parser

By the captain's mandate the away-posture record keeps no static parser
that tries to understand natural language. A mandate clause is now given
as explicit fields (--action, --object, --when, optional --stop) that
bin/fm-afk-contract.sh records verbatim. The structural check asserts
only that the action, object, and precondition fields are present and
that the action is a listed verb; whether a precondition holds is the
supervision session's judgment at execution time in a later phase.

The never-set stays as a forbidden-concept safety scan: fields mentioning
credentials, passwords, logins, legal or financial acceptance, payments,
invoices, one-time codes, or an attended prompt are refused, matched at
token prefixes after punctuation normalization so compound and plural
spellings are caught. The red-check grammar, class-word rejection,
unconditional-word detection, clause-reference resolution, and condition
aliases are removed. --words-file keeps the captain's words verbatim,
trailing newline included.

The skill, docs, launcher help, and tests describe the field form.

* no-mistakes(review): Preserve AFK words and tighten safety refusals

* no-mistakes(review): Preserve clause bytes and honor declared waits

* no-mistakes(review): Harden deny-list and gate unreadable outcomes

* no-mistakes(review): Demote never-set scan and clarify authority

* no-mistakes(review): Gate return on unreadable held and status data

* no-mistakes(review): Validate posture archives and enforce Pi detection

* fix(afk): make the never-set a non-refusing flag and keep return fail-safe

Per the captain's decision the never-set scan is a coarse best-effort
flag, never a refusal and never the gate: a clause naming a listed
concept is still recorded with a flag the read-back, announcement, and
return brief show, and the scan matches listed terms exactly or with a
plain inflection at punctuation-delimited token boundaries, so unrelated
names such as ping-service or tokenize-worker are never flagged and
joined compounds remain a documented miss. Authoritative never-set and
forbidden-action enforcement is the supervision session's judgment at
execution time in phase 4.

A replacement copies the superseded record through a temporary name and
renames it atomically so a failed copy leaves no partial archive, the
record owner gains validate and flags subcommands, and the return keeps
catch-up gated when a superseded archive cannot be read.

* no-mistakes(review): Harden AFK record validation and return reconciliation

* no-mistakes(review): Harden AFK record validation and simplify commands

* no-mistakes(review): Harden mandate validation and retain missing records

* no-mistakes(review): Refuse blank explicit mandate stops

* no-mistakes(review): Recover restored posture epoch before return

* no-mistakes(review): Prevent return brief status symlink reads

* no-mistakes(document): Refresh AFK posture documentation

* no-mistakes(ci): Fixed both CI failures: updated lint telemetry for the new fourth source directive, quoted the hyphenated fixture value, and removed unreachable test cleanup. Verified with tests/fm-lint.test.sh, targeted CI-mode ShellCheck, bin/fm-lint.sh, bash syntax checks, and git diff checks

* no-mistakes(ci): Bound structured pause deadlines by FM_PAUSE_RESURFACE_SECS in watcher and daemon housekeeping, added distinct bounded-horizon reasons, regression coverage for near, passed, and wrong-year deadlines, and updated documentation. Verified targeted behavior tests, full daemon tests, ShellCheck source-following lint, syntax, and diff checks
)

Opt-in IMAP/SMTP mail plane (fm-mail.sh / fm-mail-check.sh). Absent FM_MAIL_* stays off.

Speaking as Kun's firstmate: this is merged. Thank you @feilipu — really appreciate you taking the time on this.
* fix(remote): start the fm-remote Herdr agent through a login shell

Launchd was exec-ing herdr directly, so the Aqua agent inherited a background session without login-keychain access. Start it via /bin/zsh -lc exec so panes keep login env and can refresh OAuth tokens after reboot.

* fix(remote): start fm-remote Herdr via the account login shell

Resolve UserShell from Directory Services and invoke it with separate -l and -c so bash, fish, and zsh all get login-keychain access. Fall back to SHELL, then /bin/zsh, then /bin/sh without failing the render.

* no-mistakes(review): Fix launch-agent shell fallback resolution

* no-mistakes(review): Preserve and escape Directory Services shell paths

* no-mistakes(document): Document login-shell LaunchAgent behavior

* no-mistakes(ci): Updated the trusted fm-remote-doctor SHA-256 identity in bin/fm-remote-entrypoint.sh. Verified with tests/fm-on.test.sh, tests/fm-remote-doctor.test.sh, bash syntax checks, and git diff --check

* no-mistakes(ci): Resolved the login shell exactly once per doctor invocation and threaded it through plist rendering, installed/loaded contract validation, repair reporting, and post-repair checks. Added a regression test proving repeated repair remains healthy and performs no reload when a hypothetical second Directory Services lookup would differ. Updated the trusted doctor hash. Verified doctor, fm-on, remote-entrypoint, lint tests, ShellCheck, syntax, and diff checks

* no-mistakes(ci): Made Darwin shell resolution hermetic with executable injection and a 2-second Directory Services timeout. Updated tests to inject shells by default, isolate dscl-specific cases, parse plists semantically, and verify stalled dscl fallback. Updated the trusted doctor hash. Doctor, fm-on, entrypoint, syntax, hash, and diff checks pass

* no-mistakes(ci): Raised portable serial CI timeout from 20 to 30 minutes, refreshed the specified timing hints, added missing hints, and recomputed shard documentation. Verified coverage, runner behavior tests, workflow lint tests, shell syntax, requested timing maxima, and diff checks
…enguid#3710)

* fix(bearings): keep captain-approved deliveries in Recently Landed

A closed task is never held: tasks-axi clears the held flag when a task
closes and keeps hold-kind and the hold reason as the record of the call
that was made. Recently Landed excluded every Done row whose hold-kind was
captain, so the marker it treated as "closed while still waiting on the
captain" was in fact the proof that the captain had approved the work. Every
merge routed through a captain decision disappeared from the list of what
shipped, including under --all-landed.

The selector now asks whether the closed row delivered something. Recently
Landed is merged PRs, completed scouts, and finished local-only merges, so a
row carrying one of those artifacts belongs there whoever approved it. A
captain question closes with an answer and no artifact of its own, and that
is what still stays out, so an answered question is never rendered as
shipped work.

The same rule was written twice - the bearings projection selects this
home's Done rows and the fleet snapshot selects each secondmate home's Done
rows into the roll-up the same section merges in - which is why one defect
hid deliveries in every home. Both now share bin/fm-landed-lib.sh.

* fix(review): Normalize landed evidence and exclude answered captain questions

* fix(review): Normalize captain delivery evidence across relocated data

* fix(review): Record authoritative delivery provenance with legacy fallback

* fix(review): Harden delivery provenance across forced and pruned completions

* fix(review): Replace premature merge closure with existing release contract

* fix(review): Document provenance-based Recently Landed selection

* fix(document): Align documentation with completion provenance

* fix(lint): Fix targeted ShellCheck warnings

* fix(ci): order the pinned tasks-axi install before its stock-Bash consumers

In `.github/workflows/ci.yml` the pinned tasks-axi install now precedes both
stock-Bash consumers, and the Bearings expectation is updated from 49 to 50
tests.

Verified with macOS Bash 3.2: snapshot 16/16, Bearings 50/50, public-followup
1/1. Full repository lint and all three workflow validations pass, and
`git diff --check` is clean.

* fix(review): Make completion provenance unambiguous

* fix(review): Make completion verdict authoritative over quoted provenance

* fix(review): Preserve retained artifacts through resumed captain closes

* fix(review): Unify completion provenance ordering across writer and reader

* fix(review): Preserve artifacts across failed captain closes

* fix(review): Refresh v1 assertions; provenance authority remains unresolved

* fix(review): Remove unreliable provenance while preserving landed deliveries

* fix(review): Reject stale home summaries visibly

* fix(review): Restore retained deliverable recording

* fix(review): Match landed artifacts and restore retention documentation

* fix(review): Disambiguate captain calls and restore landed artifact matching

* fix(review): Persist retained report and PR artifacts

* fix(review): Preserve staged artifacts before captain answers

* fix(review): Avoid wedging answers on unsupported report paths

* fix(review): Exclude unreleased captain-held pull requests

* fix(review): Exclude held local-only answers from landed

* fix(review): Preserve retained scout reports across snapshot rendering

* fix(document): Align landed lifecycle documentation with release semantics

* fix(review): Enforce landed artifact-kind ownership

* fix(review): Infer canonical task kinds in snapshots

* fix(review): Require captain-hold release before merges

* fix(review): Qualify merge lifecycle regression evidence

* fix(review): Serialize captain holds with merge operations

* fix(review): Document merge cleanup residuals honestly

* fix(test): Replace vacuous Bearings regression with behavioral cases

* fix(document): Align Bearings verification and merge lifecycle documentation

* fix(review): Serialize merges and exclude captain calls from landed

* fix(review): Harden merge identity and landed selection

* fix(document): Clarify landed selector compatibility filtering

* fix(bin): keep merge entrypoints usable on records without an incarnation

The merge identity guard refused any task record with no spawn_gen field.
That field identifies one exact incarnation, so comparing it across the wait
for the merge lock is what catches a task relaunched while the merge was
queued. Requiring it to be present is a different rule, and it refused every
record written before the field existed: a legacy task could no longer be
merged at all, and five behaviour suites refused before reaching the check
they were written to exercise.

The comparison only needs to notice a change. An absent field is now read as
an empty incarnation and compared like any other value, so a record that
gains, loses, or alters one is still refused, while a record that simply
predates the field merges. An ambiguous or unreadable field stays an error,
because a record that cannot name one incarnation cannot be compared. The
missing-record message each entrypoint had before the guard is restored, so
a genuinely absent record still says so in its own words.

The role partition now precedes reading the record. Refusing the supervision
branch is a statement about the actor, not about the task, so it cannot
depend on a record the wrong actor may not have.

A backlog file that does not exist meant "no longer an open captain call".
For a caller that asked to tell absence apart it now means absent, so a board
card whose home carries no backlog stays visible instead of being dropped as
resolved.

Fixture repositories pin their initial branch instead of inheriting
init.defaultBranch, which resolved to main on a developer machine and master
on a runner, so a fixture naming main failed only in CI.

* fix(review): read local-only note from body; surface pending-close failures

* fix(review): keep kindless local-only landings in Recently Landed

* fix(review): bind local-only note scan to the tasks-axi note line

* fix(review): Guard unavailable captain-hold authority records

* fix(document): Document unreadable authority predicate outcome

* fix(bin): read an absent backlog as absence, not an unreadable record

The merge gate refused every task whose home carries no backlog file. A
backlog that does not exist holds no captain call, so nothing can be held and
the merge is safe; only a backlog that exists and cannot be read may hide a
live hold. Those two states were collapsed into one refusal, which stopped
merges in any home that keeps no backlog.

The predicate now reports a missing backlog file as absence, alongside a row
the backlog does not carry. A record that exists but cannot be read still
leaves by the existing cannot-tell path, which both merge entrypoints already
refuse, so the restrictive direction is unchanged.

That leaves no way to reach the separate unavailable-record result, so the
result and the two branches that handled it are removed rather than left
describing an outcome that can no longer occur. The lifecycle documentation
loses the same claim.

Regressions cover both directions in each entrypoint: a home with a task
record and no backlog merges, and a backlog present but unreadable refuses
without reaching the forge.

* fix(review): Fail closed unreadable backend configuration

* fix(tests): pin the bare origin's initial branch in the remote seed fixture

The fixture created its bare origin with no initial branch, so that
repository's HEAD followed init.defaultBranch while the source repository
pushed the branch fm_git_init_commit pins. On a host that still defaults to
master the two disagreed: the bare origin's HEAD named a branch the push never
created, cloning it warned that the remote HEAD referred to a nonexistent ref
and checked out nothing, and the seed assertion for the cloned README failed.

A machine whose default is already main paired the two by accident and hid it,
which is why the fixture passed locally and failed on the runner.

Pinning the bare origin to the same branch removes the dependency on the
ambient default from both sides. Verified under both conditions: with
init.defaultBranch set to master, and set to main, the suite passes 26 of 26.

* fix(review): Fail closed unreadable user backend configuration

* fix(bin): republish the home summary as v1 and record two load-bearing rules

The published home-summary schema had moved to v3, which routed every
secondmate home still emitting the earlier version to the stale branch: their
landed rows, open decisions and holds all came back empty and their state read
as unknown until each home was updated. The payload never justified that. Its
field set, field order, truncations and the landed array construction are
byte-identical to v1, so only which rows the selector places in landed
differs, and a v1 consumer reads that the same way.

Republishing as v1 removes the rollout regression and, with it, the tolerance
machinery that existed only to soften the bump: the stale-schema predicate,
its two collection branches, the flag and its provenance branch, the omitted
surface that can no longer be reached, and the fixtures and assertions that
covered them.

Two rules that a scope review proposed removing are kept, each now carrying
the reason it exists, because both were measured to be load-bearing:

The artifact-kind ownership clause is what keeps an explicit scout that
recorded no report out of Recently Landed. Without it such a row has none of
the three artifacts, satisfies the compatibility fallback and renders as
shipped work with an empty artifact.

The kind fallback is needed because tasks-axi omits the kind metadata
entirely when a title begins with a canonical keyword. Without it a scout
titled "SCOUT ..." reports no kind, its recorded report stops counting as a
delivery, and it drops out of the section this selector exists to repair.

* fix(bin): move the scout guard note onto the rule and drop two dead pieces

The LOAD-BEARING note sat on an unreachable branch. Measured in both
directions: removing that branch together with the kind-is-not-scout guards
lets an explicit reportless scout into Recently Landed and fails
tests/fm-captain-hold-lifecycle.test.sh, while removing the branch alone
leaves that suite passing at 49 assertions. The guards carry the rule, so the
note now sits on them and the unreachable branch is gone. A note pointing a
later reader at the wrong line is the hazard this change corrects elsewhere.

summary_file_has_schema lost its only caller when the stale-schema machinery
was removed, so it goes with it.

* fix(review): Fix legacy report artifacts and canonical keyword boundaries

* fix(review): Update pinned Bearings test count to 56

* fix(document): Clarify landed summary compatibility documentation

* fix(review): Preserve unreadable backend configuration errors

* fix(review): Honor backend resolution errors at existing call sites

* fix(test): Stabilize remote collector tests under host load

* fix(document): Document backend resolution failure contracts

* fix(lint): Suppress intentional deferred probe expansion warnings

* fix(ci): Captain, quoted the two literal test IDs in tests/fm-backlog-atomicity.test.sh to fix SC2100 without changing behavior. Both warnings reproduced before the fix; the targeted fm-lint.sh run now passes with ShellCheck 0.11.0. Bash syntax and git diff --check also pass

* fix(ci): Fixed the resolver’s two configuration-parent checks to return 2 for inaccessible directories while preserving genuine absence. Added two behavioral tests; RED/GREEN and both requested mutation proofs confirmed. All 10 focused checks, targeted lint, syntax, and whitespace checks passed. Broader merge suite stopped after 10 passing cases under host load. Declined portable checks and merge-authority code remain unchanged
…nguid#4090)

* fix(remote): let the Aqua launch agent own the fm-remote Herdr session

A herdr server keeps the macOS audit session of whatever started it, and
only the Aqua login session (gui/<uid>) can read the login keychain
without a prompt. Herdr's SSH remote attach starts the fm-remote server
as its own child when it finds none, wins the socket at boot because sshd
accepts connections before the login session exists, and every claude
pane under that server then gets `security` exit 36, falls back to a stale
plaintext credentials file, and reports "Login expired". launchd's own job
lost the socket on every KeepAlive retry and the doctor still reported the
session ready because it only asked whether any server answered.

- Add bin/fm-remote-herdr-guard.sh, the launch agent's exec target: start
  the server in the foreground when nothing owns the socket, exit 0 when an
  Aqua-born server does, and otherwise stop the foreign server, wait for the
  socket, and exec the server at once.
- Add bin/fm-remote-herdr-owner-lib.sh, the single owner of socket-owner
  discovery (lsof; pgrep cannot see herdr's argv on macOS) and the birth
  markers (SSH_*, XPC_SERVICE_NAME, FM_REMOTE_JOB_ACTIVE, sshd or
  remote-client-bridge ancestry matched on argv[0] and whole arguments).
- Render the agent as the login shell exec'ing the guard with
  KeepAlive={SuccessfulExit=false} and ThrottleInterval=10, check the loaded
  job's successful-exit semaphore, and report a session served outside the
  Aqua login session as fixable so --fix retakes it through launchd; the
  reload waits for an Aqua-born owner rather than any running server.
- Correct the doctor and docs: the launch shell provides environment parity,
  the launchd domain provides keychain access.
- Pin the guard's decision table and the doctor's verdicts against real
  marker-carrying processes, and record the dated audit-session evidence.

* no-mistakes(review): Verify Aqua ownership through launchd domains

* no-mistakes(document): Document macOS lsof ownership requirement
…uid#2881)

* fix(bin): prefer a live no-mistakes run over a terminal one

A worktree can bind to more than one recorded no-mistakes run at once.
The branch-and-code-identity rule in bin/fm-nm-run-lib.sh accepts both an
exact-equal commit and a worktree-is-an-ancestor match, but never stated
which wins when both bind, so the tie fell to whichever candidate the
caller reached first.

Observed on a live fleet: a crashed validation daemon left a FAILED run
at the worktree's own commit while the live run that replaced it
validated a descendant commit on the same branch. Bare `axi status`
answers with the most-recently-touched run - the corpse - and it bound by
the equal-commit rule, so every recomputation reported `failed` for a
task whose real run was healthy. The same label had also read `failed`
earlier while the work was genuinely stalled, so the signal was wrong in
both directions.

State the live-over-terminal policy in the matching rule's own contract,
where the equal-commit and ancestor rules already live, and add
fm_nm_run_status_class as the one classifier that decides liveness from a
recorded status word. fm-crew-state.sh applies it on both selection
paths: the runs listing now scans past a terminal row for a live one, and
a terminal `axi status` answer is provisional until the listing has been
asked whether this worktree also has a live run.

Same-liveness-class candidates keep the listing's newest-first
precedence, and a status word the classifier cannot place keeps the
caller's own ordering rather than displacing a known result, so a
single-run task and a task whose runs are all terminal are unchanged.

Regression coverage reproduces the proven case (terminal run at the
worktree's exact commit plus a live run descending from it) and its
runs-list twin; both fail under the old tie-break. Two companion cases
pin the no-widening half - two terminal rows still resolve newest-first,
and a terminal run with no live sibling keeps its full run-step detail -
and both pass before and after the change.

* no-mistakes(review): accept unfetched live sibling anchored at exact worktree head

* docs(bin): name both ledger reads behind the runs-limit setting

The FM_CREW_STATE_RUNS_LIMIT comment in bin/fm-crew-state.sh still described
the runs ledger as scanned only by the cross-branch fallback, but the
live-over-terminal fix also consults it as the live-sibling probe behind a
terminal axi status answer. Point the comment at docs/configuration.md as the
setting's owner instead of restating a second copy.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…chenguid#4009)

* fix(procevent): make the ordinary stop signal actually stop a runner

The owner guard that shipped in kunchenguid#3904 half-reaps. Against a poll child that
handles the ordinary stop signal and keeps waiting, the guard signals the group,
loses the runner leader to its own signal, then reads that success as a
leaderless group and exits without escalating. It destroys the only proof of
ownership that would have authorised the forced signal, so the survivor becomes
unreachable by retire, reconcile, sweep-home and the guard alike. A guard that
turns a leaking-but-identifiable generation into a permanently unreachable one
is worse than no guard at all.

Two defects, and they hid each other:

- The escalation re-derived ownership from the leader. `runner_group_signal`
  now takes a `proved` mode, passed only by the escalation inside the stop that
  already proved and signalled that exact generation moments earlier. A leader
  dying to our own signal is the ordinary outcome, not fresh ambiguity.
- Every stop held the per-source lock across its wait while the runner's own
  exit cleanup waited unboundedly for that same lock. That circular wait was
  broken only by the forced signal, so the forced signal silently became the
  normal path - and, by keeping the leader alive through the whole window, it
  masked the escalation defect above. The runner's exit cleanup now refuses that
  lock instead of waiting for it, which is what its existing `return 0` already
  said it did.

Fixing the lock alone would have turned every stop of a signal-proof child into
a refusal that leaves it running, so both land together and the tests pin that.

Measured on macOS with a stand-in poll child that traps TERM, INT and HUP:
the guard left it running past 70s and now clears the group within the lease
plus one check; retiring a healthy runner fell from ~2.8s with a forced group
signal every time to ~0.6s on the ordinary signal alone.

Unchanged and stated deliberately: a leader lost to anything other than the
stop's own signal still leaves a group that retire, reconcile, sweep-home and
the guard all refuse, permanently - and that source stops listening without
saying so. Whether such a group may ever be signalled is an open decision and
is not answered here.

* fix(review): Fix proved escalation race and stop regression assertions

* fix(review): Preserve proved escalation through transient identity failures

* fix(review): Simplify proved escalation and correct guard timing documentation

* fix(document): Clarify process-event stop ownership and cleanup limits

* fix(document): Clarify process-event stop ownership and fixture comments

* revert(skills): restore the leaderless-ambiguity limit to the loaded skill

An automatic documentation step in this branch's validation edited
.agents/skills/process-event-sources/SKILL.md, which no instruction in this
change asked it to touch. That file is not documentation about the code: it is
the agent-loaded instruction surface, what an agent reads to know what it is
permitted to do.

The step deleted this line:

  - leaderless PID/PGID-reuse ambiguity preserves the claim without signalling
    or replacement, as owned by the operating contract in
    [`docs/configuration.md`](../../../docs/configuration.md#process-to-event-sources-stateprocevent);

and folded it, with its neighbour, into a generic "registration and ownership
transitions, stop authority, and claim reclamation follow the operating
contract".

That deleted line states a PROHIBITION - that such a group is preserved WITHOUT
SIGNALLING - and it is the exact limit an open captain decision currently rests
on. Folded into a pointer, an agent reading the skill to learn what it may do
would have to chase a second document to discover it may not signal. A
prohibition that requires a second lookup is not a prohibition. The effect was
to weaken, in the instructions themselves, the boundary that keeps one home from
signalling another's process group - while the question of whether that boundary
should move at all is still open.

This is a deliberate revert, not an oversight, and it restores the file exactly
to its pre-branch state. The full statement also survives in
docs/configuration.md; that does not rescue it, because the agent handling a
process-event wake loads the skill and not the documentation.

* revert(procevent): restore the open-question marking beside the escalation

The same automatic documentation step that edited the loaded skill also removed
this from the comment above runner_group_signal:

  A leaderless group nobody in this call ever proved remains refused too, for
  every caller. That untouched refusal is what makes a crashed leader's group
  permanent, and relaxing it is a separate open question, not something this
  path assumes.

and replaced it with a pointer to docs/configuration.md.

This one fails differently from the skill deletion, which is why it is restored
separately. There, a prohibition was moved out of the reader's path, and a
missing prohibition gets violated. Here the prohibition survives in code - the
unproved path still refuses - and what was removed is the fact that the limit is
UNDECIDED. A prohibition that has quietly lost its "this is still open" reads as
settled design, and settled design gets relied on, extended, and eventually
relaxed by someone confident they understand why it is there. That question is
open right now.

The rule this branch's four instances produce, stated once here because this is
the point of decision: an unresolved question must be marked unresolved AT THE
POINT OF DECISION, not only where the contract is documented. A reader who does
not know something is open will treat it as closed, and that default is stronger
than any pointer overcomes.

The pointer added by that step is kept alongside; this restores what it replaced
rather than reverting it.

* docs(verification): restore the measured guard bound and its reason

The document step's rewrite of this record dropped the concrete figure while
keeping the surrounding measurements. What went missing was the bound itself -
lease plus two consecutive failed checks plus the stop's grace, roughly 630
seconds at the shipped 600-second lease and 15-second check - together with the
reason there are two checks rather than one: a single unreadable read must not
be enough to kill a live runner.

The mechanism survived elsewhere and the reason survived in
docs/configuration.md, so nothing was lost from the repository. The concreteness
was, and that is what this restores. A number recorded without why it is that
number is the one a later reader shortens; the reason is the whole safety
argument for the debounce, and the debounce is what stops the reaper killing a
live runner on one bad read.

* fix(document): Replace stale stop-authority summaries with owner pointers

* test(procevent): make the guard-bound case able to fail for its own reason

An automated reviewer observed that this case allowed sixty seconds for a bound
of roughly eight, so it could not go red for the reason it names: it would have
passed a guard that took fifty-five seconds. That is correct, and it is the same
family as the defect the case exists to defend against - a check that is green
because it cannot fail, rather than because the thing it guards is working.

The deadline is now derived from the bound itself - the lease, plus the two
consecutive failed checks the guard debounces on, plus the stop's own ordinary
and forced signal windows - rather than from a flat wall-clock number, and the
shortened lease and check the fixtures run under have a single definition so a
derived deadline cannot silently diverge from the settings the guard is given.
The doubling that remains is a load allowance and is documented as one; widening
it to make a slow guard pass would convert the assertion back into decoration.

Proven by mutation rather than by argument. Against the repaired case:

  correct code                                        ok
  guard debounces on 20 misses instead of 2           not ok - "still holding
                                                      the group after 16s,
                                                      against a documented
                                                      bound of 8s"
  proved escalation removed (the original defect)     not ok - same
  code restored                                       ok

The previous sixty-second version passes every one of those mutations.

The reviewer's other claim, that the guard can survive past the announced bound
when an owner disappears immediately after a check, was measured and does not
hold against what this branch announces. Sweeping the phase deliberately at
0.0, 0.2, 0.4, 0.6 and 0.8 of a check interval gave 7.21s, 7.31s, 6.75s, 6.49s
and 6.31s, worst 7.31s, against the announced lease plus two consecutive failed
checks plus stop grace, which is up to 8s at those settings. The mechanism the
reviewer describes is real and is the announced mechanism; the bound it was
measured against is a phrasing this branch no longer carries.

* revert(scope): return the instruction surfaces to their base state

This delivery is being split. It carries the two proven process fixes alone; the
instruction text travels separately, through a run that removes the
documentation step rather than refusing it at its gate.

Two surfaces are therefore returned to exactly what the base branch has, so this
delivery neither adds to them nor removes from them:

  .agents/skills/process-event-sources/SKILL.md - identical to base again. Three
  bullets an automatic documentation step had folded into a pointer, including
  that leaderless PID/PGID-reuse ambiguity preserves the claim WITHOUT
  SIGNALLING and that there is one identity-matched owner per canonical source
  across homes sharing one store.

  The header comment block of bin/fm-procevent.sh, which is what the script
  prints as its own help. Seven lines were removed from it: that a live owner is
  never displaced, that only a claim whose stale owner and independently absent
  process group prove its whole generation gone is reclaimed, that a crashed
  leader or reused pid whose process group still has members cannot relax
  ownership cleanup, and that reconcile signals only a live identity-matched
  runner group and otherwise keeps the claim without starting a replacement.
  The help output is now byte-identical to base.

Neither removal was requested by any instruction in this change, and both were
made to text that predates it. Returning them is scoping, not a third
restoration: nothing is being added to those files here.

* fix(ci): Captain, live CI revealed a fixture deadlock: it suspended the runner before startup released its lock. Added a public-list synchronization barrier in tests/fm-procevent.test.sh. Forced-delay reproduction detected the deadlock before the fix; all four cases passed afterward. Targeted lint, Bash syntax, and whitespace checks passed. Greptile’s watchdog requirement conflicts with the recorded R2 decision; runtime behavior and documentation remain unchanged. Full CI rerun belongs to the outer executor

* test(procevent): make the post-TERM cases report what they saw when they fail

On the failure path only, these cases now print what they actually saw: the
identity recorded at claim time, the identity readable at that moment, the size
of the signals file, the leader's state and wchan, every live member of the
runner's process group with its own state and wchan, the elapsed time since the
stop began, and what retire said. None of it runs when a case passes.

WHY THIS IS KEPT, stated accurately rather than by its original reason. It was
written to make an unexplained CI failure verifiable. That failure is now
explained - it was a fixture deadlock, diagnosed and repaired in the preceding
commit - so that justification has expired and is not the reason given here.
The reason it stays is smaller and independent of that failure: it is already
written, it is small, it sits in the file whose assertion this change reworked,
and an assertion that could not say why it failed cost most of a morning to
diagnose from the outside. The next failure will not be this one.

WHAT A PASSING RUN WOULD NOT MEAN: a pass is a sample of behaviour already
observed many times, not proof that anything is fixed. Only a failure carrying
the evidence above establishes a cause.

* fix(document): Clarify process-event fixture diagnostic rationale

* fix(ci): Captain, fixed two cleanup races in tests/fm-procevent.test.sh: removed premature child completion and waited for runner exit before retiring the restart fixture. Controlled Linux reproductions demonstrated failure before and success after. The full Linux process-event suite, six focused macOS checks, targeted ShellCheck, Bash syntax, and whitespace checks passed. Runtime behavior, guard debounce, and documentation remain unchanged. CI rerun belongs to the outer executor

* fix(procevent): bound owner-guard cleanup at one check interval, not two

A THIRD WAY, not a capitulation to the reviewer and not a refusal of it.

The automated reviewer's grievance was the LOOSENESS OF THE BOUND, not the
number of observations the guard makes before it acts. It asked for a single
read because that was the only route it could see to an acceptable bound. There
was another route, and this change takes it: the bound is reached and both reads
are kept.

TIGHTENED - the SPACING of the guard's two reads, not their number. The owner
watchdog now sleeps half the configured check interval and still requires two
consecutive failing reads, so the pair completes inside one check interval
instead of costing two. Worst-case detection falls from the lease term plus TWO
check intervals to the lease term plus ONE. At the shipped 600s lease and 15s
interval the stated bound falls from ~635s to ~620s.

PRESERVED - the second read. bin/fm-procevent.sh's two-consecutive-miss rule is
untouched. WHY IT PROTECTS: the guard's inputs are a lease read and a state-root
identity read, and either can fail transiently on a live, healthy home. Acting
on the first failure would let one isolated unreadable read kill a live service.
Requiring a second, independent read is what makes that impossible, and it is a
protection rather than padding. Nothing was traded away to reach the bound.

Both properties are now guarded by their own case, and each was proven by
MUTATION rather than asserted:

  - putting a full interval back between the two reads fails the bound case:
    "still running 17.0s after the last owner activity, against a documented
    bound of 15s";
  - acting on one failed read fails the new debounce case: "one unreadable lease
    read ended a runner whose home was still alive" - while the bound case then
    passes FASTER, 9.9s against 13.1s. The unsafe variant being the quicker one
    is exactly why these are two cases: one elapsed-time case would have
    registered the removal of the protection as an improvement.

MEASURED, sampling the phase between the guard's check clock and the lease clock
across eight runs per variant, on macOS (Darwin 25.5.0). Reaping an orphaned
listener whose home stopped refreshing its lease:

  lease 2s / interval 1s:  4.41-5.29s before,  3.48-4.65s after
  lease 2s / interval 4s:  7.69-8.12s before,  5.94-6.13s after

The 4s configuration is the informative one: the gap is about one check
interval, which is precisely the term that was removed.

A previously unstated term of the bound surfaced while measuring: the lease age
is compared in whole seconds, so a configured lease of N is honoured until that
age reads N+1. It is now part of the documented bound and of the regression's
derivation instead of being absorbed into a fudge factor.

The bound regression derives its deadline from the documented bound instead of a
flat number, and PINS the phase between the guard's check clock and the lease
clock rather than sampling it, because with a sampled phase a guard spending two
intervals passes about half the time on a lucky alignment. Its load slack is
additive and stays under half a check interval, so an extra whole interval
cannot hide inside it. The two flat deadlines that were there before (40s and
20s) and the doubling allowance on the derived one are gone; that looseness was
the reviewer's third complaint.

The stop's own grace is untouched: 2s for the ordinary signal, then 2s for the
forced one. It is a ceiling paid only by a group that outlives the signal it was
sent, not a delay every stop pays - a healthy runner's whole retire measures
0.40-0.66s on this host. The reviewer's literal "lease plus one tick" is
unreachable by any implementation, since signalling a process and giving it any
chance to exit takes non-zero time; detection now meets it and the stop runs
inside its own ceiling, and the contract says so rather than glossing it.

NECESSARY BUT NOT SUFFICIENT, and written BEFORE this head's integration runs
start rather than after they report. On the previous head, "Behavior portable
serial 1" and "Behavior portable serial 4" were both CANCELLED at the job
ceiling, independently of this finding. A new head triggers fresh runs, so those
two lanes MAY complete this time. IF THEY DO, THAT IS NOT EVIDENCE THE CEILING
DEFECT IS FIXED. It is one more sample of a lane that has been cut repeatedly
and sometimes is not; the shard-packing repair for it is open separately. Do not
reread a lucky pass here as a resolution.

Relatedly, and deliberately: the per-script duration hint in bin/fm-test-run.sh
was NOT updated even though the two new cases add ~19s of wall clock.
docs/fm-test-portable-shards.md says those hints are replaced wholesale from CI
timing artifacts of green runs, and that repair is the open request doing it; a
hand-edited estimate here would collide with it and silently repack the shards.
This suite runs in portable serial shard 3, which was green in the last run.

Verification: tests/fm-procevent.test.sh green, plus
tests/fm-captain-hold-lifecycle.test.sh, the test-coverage guard, and
bin/fm-lint.sh. The unrelated "reconcile stops a runner whose registration was
removed" case flaked in 4 of 7 local full runs; an isolated 20-trial
reproduction measured it at 13/20 unclean before this change and 11/20 after, so
it is issue 4080 and is not aggravated here.

* fix(procevent): repair our decimal-interval regression and enforce the timing phase

REPAIRED BEFORE PUBLICATION, AND IT WAS OURS. The half-interval arithmetic added
by the previous commit read a zero-prefixed interval as octal: 010 halved to 4
instead of 5, and 08 was not a number at all, so the owner guard died before
reporting ready and the runner failed closed and never listened. The validator
accepts those values and `[` compares them as decimal, so this broke a
configuration that worked before. Introduced by this delivery, found in review,
repaired here. Forcing base ten before the arithmetic is the whole runtime fix.

Proven by driving it rather than by reading the source: a new case starts a real
listener at 08 and at 010 and observes the guard's actual sleep argument - 4s and
5s. Removing the normalisation turns that case red with "a zero-prefixed decimal
interval (08) prevented the listener from starting".

THE TIMING PHASE IS NOW OBSERVED AND ENFORCED, NOT ASSUMED. The bound case
pinned its phase by CONSTRUCTION, from an assumed startup time, and enforced
nothing. Review was right that this is not enough: once startup reaches about two
seconds the expiry lands in a different part of the interval and the case
silently stops rejecting a two-interval guard while still reporting success. A
bound that cannot fail for the reason it names is the defect this whole delivery
exists to correct, so it must not ship inside the fix for it.

Now the lease is synchronised to the guard's own FIRST observed lease read,
every later real read is recorded, and the case REFUSES unless one recorded read
proves the required phase: it read the synchronised reference, it was still
fresh, and it began late enough that two further full intervals could not finish
before the deadline. An unestablished precondition refuses; it does not proceed
on trust. The derived deadline, the two-read debounce and the additive slack are
unchanged, and the slack invariant is now asserted rather than left to a comment.

Review also found the deadline was only ever checked while the group was still
alive, so a sampler descheduled past it would see the group gone and certify
success. The observed completion time is now checked too.

PROVEN BY MUTATION, each one run against this code:

  - remove the decimal normalisation -> the interval case fails on 08;
  - a full interval between the two reads -> "the guard exceeded its bound:
    group still running 17.1s ... against a documented bound of 15s";
  - a full interval WITH startup forced to ~2.5s, which is exactly the condition
    the old construction pin could not survive -> still red, same message;
  - the same ~2.5s startup with the correct guard -> still passes, 13.0s against
    the 15s bound, so the delay alone does not break the case;
  - phase evidence made unavailable -> "could not establish the required
    pre-expiry guard-read phase", a refusal rather than a pass, even though the
    group stopped quickly;
  - act on one failed read -> the debounce case fails and the bound case passes
    FASTER, 9.5s against 12.7s, which is why these remain separate cases.

Verification: full tests/fm-procevent.test.sh green, and bin/fm-lint.sh clean.

* fix(document): Correct process-event timing and debounce comments
…kunchenguid#3417)

* fix(backlog): honor configured task adapters

* no-mistakes(review): Harden backend purity lint against prefixed Beads calls

* no-mistakes(document): Document configured backend lifecycle transitions

* fix(backlog): preserve markdown exemptions

* no-mistakes(review): Enforce backend purity for explicit lint paths

* no-mistakes(document): Update lifecycle backend documentation

* no-mistakes(lint): Remove redundant backend lint pattern

* fix(backlog): close adapter routing gaps

* no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint

* no-mistakes(document): Document environment-selected backlog adapters

* no-mistakes(lint): Fix empty local variable assignment

* fix(backlog): close quoted path gaps

* no-mistakes(review): Reject partially quoted direct Beads commands

* no-mistakes(document): Align lifecycle documentation with configured adapters

* test(backlog): keep structural cases markdown-only

* fix(backlog): honor configured task adapters

* no-mistakes(review): Harden backend purity lint against prefixed Beads calls

* no-mistakes(document): Document configured backend lifecycle transitions

* fix(backlog): preserve markdown exemptions

* no-mistakes(review): Enforce backend purity for explicit lint paths

* no-mistakes(document): Update lifecycle backend documentation

* no-mistakes(lint): Remove redundant backend lint pattern

* fix(backlog): close adapter routing gaps

* no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint

* no-mistakes(document): Document environment-selected backlog adapters

* no-mistakes(lint): Fix empty local variable assignment

* fix(backlog): close quoted path gaps

* no-mistakes(review): Reject partially quoted direct Beads commands

* no-mistakes(document): Align lifecycle documentation with configured adapters

* test(backlog): keep structural cases markdown-only

* no-mistakes(review): Harden markdown lifecycle routing and close recovery

* fix(lint): catch dollar-quoted beads commands

* no-mistakes(review): Drop unused authorized-data-dir param; canonicalize fm-lint ROOT with pwd -P

* no-mistakes(document): Document fm-lint backend-purity check in header and CONTRIBUTING

* no-mistakes(ci): Two failing checks, one code-caused and one attestation-only. 1. Greptile Review (P1: ANSI-C quoting bypasses lint) — REAL DEFECT, FIXED. The backend-purity normalizer in bin/fm-lint.sh (invokes_bd awk function) stripped $'...' quote delimiters without decoding ANSI-C escapes, so a core script containing $'\x62\x64' close fm-example executes as `bd close` while the lint accepted it. Fix: the normalizer now tracks whether a single-quote context came from $' (ansi flag) and decodes ANSI-C escapes while building the command word: \xHH (1-2 hex), \uHHHH / \UHHHHHHHH (4/8 hex), \NNN (1-3 octal), \cX control characters, simple escapes (a/b/e/E/f/n/r/t/v decode to a placeholder that can never spell bd), NUL (value 0) truncates the word per bash C-string semantics, and unknown escapes drop the backslash per bash. Values outside printable ASCII decode to a placeholder so they can neither falsely match nor collide into `bd`. Regression coverage added to the existing lint-interface test test_rejects_direct_beads_cli_invocations in tests/fm-lint.test.sh: $'\x62\x64', $'\142\144' (octal), and b$'\x64' (split word), each written into a fixture and asserted rejected through the real fm-lint.sh executable. Verified regression property: with the fix stashed, the new hex case passes lint (reproduces the reported bypass); with the fix, it is rejected. Verification: all 29 tests in tests/fm-lint.test.sh pass, and the full CI-parity lint (CI=true bin/fm-lint.sh: ShellCheck 0.11.0 full analysis over the canonical set, backend-purity check, and actionlint workflow validation) exits 0. One iteration was needed because an awk comment containing a literal $'\x62\x64' sample broke shell-level quoting (bash -n / SC1001/SC2026); the comment was reworded without quotes. 2. PR must be raised via no-mistakes — NOT code-caused. The check failed with 'Pipeline attestation head_sha does not match the current PR head': the PR body attestation binds to 82b41c7 while the PR head is 1995cdb because a later pipeline push moved the head. This is exactly the stale-attestation condition the user intent describes; it clears when the outer pipeline re-runs 'git push no-mistakes' and re-binds the attestation to the new head (which now includes this Greptile fix). No code change can or should address it. Files changed: bin/fm-lint.sh (ANSI-C escape decoding in the backend-purity normalizer), tests/fm-lint.test.sh (three encoded-bd rejection cases)

* no-mistakes(review): fix tasks.toml hang, root authorization, lint quoting

* no-mistakes(review): validate tasks config before exemption; fix lint quote gap

* fix(backlog): address the markdown backlog as <data>/backlog.md

Resolving the markdown backlog through a configured `[markdown] path` was
scope this task never asked for. It is absent from main, which addresses
`<data>/backlog.md` everywhere, and it came from an earlier review round
rather than the task brief.

Making it effective on the transition path alone put that path at odds
with every other consumer of the same backlog - fm-captain-hold.sh,
fm-session-start.sh, fm-fleet-snapshot.sh, fm-inbox.sh,
fm-backlog-handoff.sh - which all still address `<data>/backlog.md`. In
fm-captain-hold.sh the split was live: its reads had already moved to the
shared gate while its writes had not, so the two could address different
files.

Address `<data>/backlog.md` from the shared gate, delete the unused
resolver, and drop the two tests that pinned the withdrawn behaviour.

What this task actually changes is unaffected: a configured non-markdown
adapter is still addressed by its own root, without `--file`.

* fix(backlog): honor configured task adapters

* no-mistakes(review): Harden backend purity lint against prefixed Beads calls

* no-mistakes(document): Document configured backend lifecycle transitions

* fix(backlog): preserve markdown exemptions

* no-mistakes(review): Enforce backend purity for explicit lint paths

* no-mistakes(document): Update lifecycle backend documentation

* no-mistakes(lint): Remove redundant backend lint pattern

* fix(backlog): close adapter routing gaps

* no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint

* no-mistakes(document): Document environment-selected backlog adapters

* no-mistakes(lint): Fix empty local variable assignment

* fix(backlog): close quoted path gaps

* no-mistakes(review): Reject partially quoted direct Beads commands

* no-mistakes(document): Align lifecycle documentation with configured adapters

* test(backlog): keep structural cases markdown-only

* fix(lint): catch dollar-quoted beads commands

* no-mistakes(review): Drop unused authorized-data-dir param; canonicalize fm-lint ROOT with pwd -P

* no-mistakes(document): Document fm-lint backend-purity check in header and CONTRIBUTING

* no-mistakes(ci): Two failing checks, one code-caused and one attestation-only. 1. Greptile Review (P1: ANSI-C quoting bypasses lint) — REAL DEFECT, FIXED. The backend-purity normalizer in bin/fm-lint.sh (invokes_bd awk function) stripped $'...' quote delimiters without decoding ANSI-C escapes, so a core script containing $'\x62\x64' close fm-example executes as `bd close` while the lint accepted it. Fix: the normalizer now tracks whether a single-quote context came from $' (ansi flag) and decodes ANSI-C escapes while building the command word: \xHH (1-2 hex), \uHHHH / \UHHHHHHHH (4/8 hex), \NNN (1-3 octal), \cX control characters, simple escapes (a/b/e/E/f/n/r/t/v decode to a placeholder that can never spell bd), NUL (value 0) truncates the word per bash C-string semantics, and unknown escapes drop the backslash per bash. Values outside printable ASCII decode to a placeholder so they can neither falsely match nor collide into `bd`. Regression coverage added to the existing lint-interface test test_rejects_direct_beads_cli_invocations in tests/fm-lint.test.sh: $'\x62\x64', $'\142\144' (octal), and b$'\x64' (split word), each written into a fixture and asserted rejected through the real fm-lint.sh executable. Verified regression property: with the fix stashed, the new hex case passes lint (reproduces the reported bypass); with the fix, it is rejected. Verification: all 29 tests in tests/fm-lint.test.sh pass, and the full CI-parity lint (CI=true bin/fm-lint.sh: ShellCheck 0.11.0 full analysis over the canonical set, backend-purity check, and actionlint workflow validation) exits 0. One iteration was needed because an awk comment containing a literal $'\x62\x64' sample broke shell-level quoting (bash -n / SC1001/SC2026); the comment was reworded without quotes. 2. PR must be raised via no-mistakes — NOT code-caused. The check failed with 'Pipeline attestation head_sha does not match the current PR head': the PR body attestation binds to 82b41c7 while the PR head is 1995cdb because a later pipeline push moved the head. This is exactly the stale-attestation condition the user intent describes; it clears when the outer pipeline re-runs 'git push no-mistakes' and re-binds the attestation to the new head (which now includes this Greptile fix). No code change can or should address it. Files changed: bin/fm-lint.sh (ANSI-C escape decoding in the backend-purity normalizer), tests/fm-lint.test.sh (three encoded-bd rejection cases)

* fix(bin): preserve captain calls during teardown (kunchenguid#3595)

* fix(bin): never close a captain call during cleanup

A scout that held its own work item for the captain, which is what
captain-hold-lifecycle prefers ("hold the work item the question gates"),
was closed by bin/fm-teardown.sh's automatic backlog transition. The
completion gate passed, cleanup ran, and the captain's question moved to
Done with no recorded answer: the one thing the policy says must never
happen. `tasks-axi done` closes a held row silently, and nothing in
teardown asked whether the row was the captain's own call.

bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when
the task is still an open captain call, 1 when it is not, 2 when that
cannot be established. It reads the row through the transition library's
backend-aware probe, so it addresses the same backlog teardown does; the
script's other commands now address the configured data directory the
same way instead of FM_HOME, which also fixes captain holds in a home
with a relocated data directory.

Teardown asks `open` before any destructive step and refuses on 2. On 0
only the close changes: after cleanup and still under the task's own
lock, the row gets one "Deliverable of the finished work" line at the end
of its body and returns to Queued through `tasks-axi reopen`, keeping its
hold, so it lands in Captain's Call instead of reading as work under way.
--force does not lift this: it authorizes discarding unlanded work, never
the captain's question. The deliverable goes into the body because
`tasks-axi update --report` rewrites the title of a row that is not Done.

The crash window reuses the pending-close record teardown already stages:
a `mode=retain` line makes the existing replay record the deliverable and
reopen instead of closing, with the same validator, stale-generation
check, cleanup-incomplete marking, and non-blocking bootstrap lock as an
ordinary close. A retained row the captain answered first simply retires
the record. No parallel record type, recovery command, or second bootstrap
loop is introduced.

Regressions run the real executables: the captain-held scout survives
cleanup queued, held, with its deliverable and on the board, only
`answer` closes it, --force keeps it open, and an ordinary scout still
closes with its report; an interrupted cleanup leaves the row untouched
and the next session start retains it; a relocated backlog keeps the
retention in its one configured file; and a ship row whose hold cannot be
read refuses cleanup before anything destructive.

Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np

* no-mistakes(review): Serialize captain holds and fix backend-aware listing

* no-mistakes(document): Update captain-call retention documentation

* no-mistakes(document): Fix relocated captain-hold backlog diagnostics

* fix(backlog): honor configured task adapters

* no-mistakes(document): Update lifecycle backend documentation

* fix(backlog): close adapter routing gaps

* no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint

* no-mistakes(lint): Fix empty local variable assignment

* fix(backlog): close quoted path gaps

* test(backlog): keep structural cases markdown-only

* no-mistakes(review): Harden markdown lifecycle routing and close recovery

* no-mistakes(review): fix tasks.toml hang, root authorization, lint quoting

* no-mistakes(review): validate tasks config before exemption; fix lint quote gap

* fix(backlog): address the markdown backlog as <data>/backlog.md

Resolving the markdown backlog through a configured `[markdown] path` was
scope this task never asked for. It is absent from main, which addresses
`<data>/backlog.md` everywhere, and it came from an earlier review round
rather than the task brief.

Making it effective on the transition path alone put that path at odds
with every other consumer of the same backlog - fm-captain-hold.sh,
fm-session-start.sh, fm-fleet-snapshot.sh, fm-inbox.sh,
fm-backlog-handoff.sh - which all still address `<data>/backlog.md`. In
fm-captain-hold.sh the split was live: its reads had already moved to the
shared gate while its writes had not, so the two could address different
files.

Address `<data>/backlog.md` from the shared gate, delete the unused
resolver, and drop the two tests that pinned the withdrawn behaviour.

What this task actually changes is unaffected: a configured non-markdown
adapter is still addressed by its own root, without `--file`.

* no-mistakes(review): restore home boundary guard and tighten purity lint

* no-mistakes(review): authorize home boundary for every backlog adapter

* no-mistakes(test): complete tasks-axi stubs in fm-gotmp teardown fixtures

* no-mistakes(document): align backlog transition docs with adapter-neutral addressing

* no-mistakes(review): label adapter data-dir authorization, drop dead row_probe local

* no-mistakes(review): pin markdown backend at relocated-data addressing roots

* no-mistakes(document): point lint-definition mention at fm-lint.sh header

* no-mistakes(document): point mutate comment at adapter addressing owner

* no-mistakes(review): Fix leftover-symlink refusal on non-markdown homes; hoist config check and lint/dedup cleanups

* no-mistakes(document): Align fm-lint purity scope header with bin/backends

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
…guid#4119)

* fix(bin): stop false watcher-down alarms on long Claude turns

A healthy Stop auto-arm rewake or open claim already explains a mid-turn
beacon that has aged past grace, because turn-end will re-arm. Keep the
supervision-off banner for a missing, failed, or exhausted generation.

* no-mistakes(review): Bind Claude rewakes to active recovery generation

* no-mistakes(document): Document Claude long-turn supervision exception

* no-mistakes(lint): Fix empty ShellCheck assignment

* no-mistakes(ci): Fixed all reported CI failures: quoted the hyphenated recovery-delivery value to satisfy ShellCheck SC2100, and updated the session-lock auto-arm fixture to emit the recovery marker and watcher beacon now required for a valid rewake. Verified fm-session-lock-ancestry, fm-test-run, stale-banner, Claude auto-arm, targeted lint, ShellCheck, and workflow lint checks pass
* fix(herdr): classify a gone session's endpoint as recoverable

A task whose Herdr endpoint could not be read was classified `unreadable`,
which blocks recovery by design. The commonest reason that read fails is
that the recorded session's server is not running at all - a host reboot, a
server exit, a session never restored - and that is authoritative absence
for every pane in that session, not an ambiguous answer about one of them.
Tasks in that state had no sanctioned way back.

The recovery-grade read now settles an uninterpretable pane read with the
session server's own `.server.running` state: positively stopped reads
`missing`, while a running server, or a server state that cannot itself be
read, still reads `unreadable`. Resting the verdict on that field rather
than on the `server_not_running` error code is what keeps it working across
Herdr 0.8.x and 0.9.0, since the field is present on both and the code is
not.

Only that one boundary is widened. The husk classifier under it stays
strict, so duplicate prevention, rollback, and teardown - the paths that can
destroy something - keep refusing on exactly the reads they refused on
before.

Separately, a relaunch refused outright when the endpoint's shell had
drifted out of the recorded worktree. An agent's own exit routinely leaves
its shell somewhere else, so that refusal stranded tasks whose work was
sitting untouched on disk. The shell is now told once to return, and only a
shell that will not go refuses; the replacement still never starts outside
the copy holding the work.

Herdr 0.8.x is not installed on this host, so protocol-20 coverage is
structural plus the adapter fixture exercising both response shapes, and is
recorded as such rather than as a live result.

Fixes kunchenguid#4091.

* no-mistakes(review): Restrict drift recovery to Herdr endpoints

* no-mistakes(review): Correct Herdr recovery verification coverage

* no-mistakes(document): Document Herdr endpoint recovery boundaries
kunchenguid#4033)

* fix(bin): keep an escalated undelivered handoff wake retryable

A remote backlog handoff holds its outbox until the backlog receipt and
the receiver wake are both confirmed, and retries the wake under the same
pending-reply correlation on every resume. When that wake's remote
transport was lost, the correlation stayed undelivered in delivery_unknown
and the watcher's next pending-reply tick escalated it. Both the reuse
predicate and the known-undelivered reset refused an escalated record, so
the resume refused to resend the wake forever and every later handoff to
that mate jammed behind the outbox.

Treat an escalated record with no confirmed delivery as the undelivered
correlation it is: fm_pending_reply_corr_reusable accepts it for its own
task and fm_pending_reply_reset_known_undelivered returns it to
awaiting_report for the idempotent remote resend, while a delivered
record is still never reset and a missed-report escalation keeps its
meaning. The published delivery-unknown decision stays open until the
record resolves, so a repeat loss neither re-notifies nor strands it.

Reproduce the deadlock end to end in the remote handoff test (lost wake
transport, watcher escalation, resume) and pin the predicate contract in
the pending-reply suite; the fm-send fixture that pinned the refusal now
uses a genuinely stale delivered escalation.

* no-mistakes(review): Decouple durable outboxes from best-effort wake retries

* no-mistakes(review): Align handoff documentation with durable receipt release policy

* no-mistakes(review): Handle unrecordable wake state as dropped

* no-mistakes(review): Prevent stale wake markers blocking handoffs

* no-mistakes(review): Prevent stale delivered markers suppressing new wakes

* no-mistakes(document): Clarify retry escalation decision lifecycle

* no-mistakes(document): Document pending receiver wake retries
* docs: bound the mandatory captain address to the chat channel

AGENTS.md's opening address rule said "address the user as captain at
least once in every response" and never said what a response is. The
artefact exclusion two lines below governed only the optional nautical
seasoning, not the mandatory address. An agent that reads this file
without being the first mate - a pipeline corrector agent running inside
a copy of this repo - therefore read the obligation as applying
everywhere and the exclusion as applying only to flavour, and opened its
delivery message with "Captain,". That reading was correct.

Patch the existing owner rather than adding a rule elsewhere:

- bound the obligation to chat messages sent to the captain;
- state the artefact exclusion once, explicitly binding every agent that
  reads this file whether or not it is the first mate, and naming commit
  messages, PR and issue descriptions, briefs, code and comments;
- fold the seasoning under the same bound instead of carrying a second,
  narrower copy of the exclusion.

The obligation itself is unchanged: the captain is still addressed in
every chat message.

AGENTS.md goes from 603 to 602 lines: the redundant "never send a
response with zero direct address" clause and the duplicated seasoning
exclusion pay for the new bound.

The two cross-references that paraphrased the unbounded wording
(bin/fm-parent-channel-lib.sh's header and
docs/secondmate-parent-channel.md's problem statement) now match the
owner; neither restates the rule.

* fix(review): Limit address exclusions to artifacts while preserving public replies

* fix(document): Consolidate captain address guidance
…nguid#4131)

* fix(herdr): close persisted-focused tabs when no live client is attached

The teardown active-tab guard treated Herdr's last-focused pointer as a live viewer, so detached sessions could not close panes on that tab.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(herdr): allow detached seeded-tab prune after live-client gate

Projection create still restored the persisted focused tab after a successful prune, so a detached last-focused seeded tab still quarantined the spawn.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(herdr): probe live client after seeded prune only when that tab was focused

The extra title-clear read after every prune shifted canned CLI fixtures and failed projection create.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Tighten Herdr active-tab close guard

* no-mistakes(review): Guard Herdr mutations with fresh target focus

* no-mistakes(document): Document Herdr live-viewer teardown guard

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
…guid#4169)

* fix(bin): escalate decision-owned wakes once as the decision

The away-mode daemon treated a needs-decision: queued payload as an
unknown wake, so suppression markers never committed and the same open
decision re-escalated on every poll.

Classify that payload through the existing signal path so it escalates
once, labelled as the decision, and an unchanged repeat is suppressed
on the same terms as any other signal.

Fixes kunchenguid#4096

* no-mistakes(review): Escalate captain-held decision-owned rows once as the decision

* no-mistakes(review): Self-handle captain-held decision-owned rows instead of escalating them

* no-mistakes(document): Name away daemon as needs-decision payload reader
* test(herdr): pin leftover-shell vs live-idle via agent get

Herdr 0.9.0 already distinguishes a Pi that exits to a surviving pane shell
from a sibling live idle occupant. Pin that pair through agent get and the
recovery classifier so a lagged pane-get status cannot silently reclaim the
leftover shell as alive.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(document): Document Herdr leftover-shell liveness regression

* no-mistakes(ci): Fixed Lint failure SC2034 by replacing the unused wait-loop variable with `_`. Verified with the pinned project lint command, Bash syntax check, and git diff check

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
…nchenguid#4151)

* ci: rebalance the portable parallel lanes on measured runner durations

Both portable parallel lanes are capped at 10 minutes. Lane 1 was cancelled at
that cap on every request raised on 2026-09-10 while lane 2 finished in about
3.5 minutes, so no request could go green.

CONTRACT CLASS: RESTORE.
The workflow already promises two duration-balanced lanes and the shard
documentation already claims a measured wall; this re-establishes both against
what the lanes now cost, and changes no lane count, no cap, and no scope of what
runs. The counter-argument, so nobody has to take that on trust: two pieces here
are genuinely new rather than restored, and either could be argued to make this
a NEW-behavior change. `--list-scheduled` now ranks a parallel lane on measured
durations where it previously handed every parallel script the serial default
weight and returned an alphabetical order; and `--check-coverage` gains three
reported fields. I classify the change RESTORE because both exist only to make
the already-promised property checkable, but they are named here rather than
folded into the restoration.

=== PART 1: THE TOTAL, AND HOW IT WAS OBTAINED ===

This section stands on its own. It establishes what the parallel set costs. It
derives no packing; Part 2 does that, from this number.

THE TOTAL: 828568 ms, about 13 min 49 s of serial work across the 24 scripts.
Lane 1 held 624299 ms of it and lane 2 held 204269 ms, a 3.06:1 split.

HOW IT WAS OBTAINED. The difficulty was that lane 1 had never finished, so its
duration did not exist as a recorded figure anywhere and no timing artifact was
expected for it. It turned out to be recoverable from the real lane without
estimating, by two routes, across six CI runs on 2026-09-10 (34459949083,
34460760299, 34462530836, 34462758357, 34466966385, 34470382458):

  - Run 34462758357's lane-1 job finished its suite 18 s BEFORE the wall and
    uploaded a complete fm-test-timing-portable-parallel-1 artifact carrying all
    11 scripts, FM_TEST_SUMMARY total=11 failed=0 duration_ms=598225. The
    upload step is if: always(), so the cancellation did not suppress it. This
    is one full, untruncated lane-1 measurement.
  - The five other lane-1 jobs were cancelled mid-suite, but each logs every
    script that had already finished as an FM_TEST_END duration_ms= marker.
    Those per-script records are complete measurements of completed scripts;
    only the script in flight at cancellation is lost, and it differs by run.

Lane 2 completed in all six runs, so its scripts come from the six uploaded
fm-test-timing-portable-parallel-2 artifacts.

Every one of the 24 scripts therefore carries at least one untruncated
measurement: 20 of them measured in all six runs, two in three or four runs, and
two (fm-brief, fm-transition-lib, the tail of lane 1) in the single complete run.
Each hint is the SLOWEST value that script reached, so the total is an upper
envelope rather than an average. NO FIGURE IN IT IS DERIVED FROM A TRUNCATED
LANE, and no lower bound was ever extrapolated into a total.

THE ENVIRONMENT, AND WHETHER IT TRANSFERS. Every hint is a serial run of the
real portable parallel lane on a GitHub ubuntu-latest runner, produced by the
lane's own CI job. It transfers because it is not a proxy for the lane; it is
the lane. Nothing in the total came from this machine or from any harness of
mine.

That mattered, and here is what it would have cost. A same-day macOS
cross-check of the same scripts ran 1.7x to 5.0x slower with the ratio varying
per script (fm-test-run 157420 ms against 92944 ms, fm-x-mode 67217 ms against
31870 ms, fm-composer-ghost 10521 ms against 2120 ms). Local timings therefore
do not scale the lane, they REORDER it, so a packing derived from them would
have balanced the wrong thing while looking clean.

WHAT IT REPLACES, which is the root cause. The lanes were packed from the
2026-08-20 concurrent isolation proof: 24 candidates across four LOCAL workers.
That record answers whether the candidates are isolation-safe, not how long a
SERIAL CI lane runs, so it was structurally incapable of representing lane wall
clock even when it was fresh. It was also never refreshed while the set grew
about 3.2x. Both the wrong instrument and the staleness are fixed here: the
hints now come from the lane itself and carry their run ids and date.

=== PART 2: THE SPLIT DERIVED FROM THAT TOTAL ===

Longest-processing-time assignment over those hints gives 414269 ms and
414299 ms, 30 ms apart, against 624299/204269 before.

tests/fm-pi-primary-types.test.sh stays in lane 1 because that is the job which
installs the Pi package, so ci.yml needs no step changes.

=== PART 3: DOES THE MARGIN SURVIVE MACHINE VARIANCE ===

Stated explicitly, because 6.90 min against a 10 min cap is 69% of cap before
any variance is applied, and the cap covers the whole job rather than the suite.

  worst lane, script time                         414299 ms   6.90 min
  job overhead, measured on the real lane             ~18 s   (see below)
  expected healthy job                            ~432300 ms  7.21 min
  x1.29 on the script time, plus overhead         ~552400 ms  9.21 min
  cap                                             600000 ms  10.00 min
  room left after the multiplication                ~47.6 s   7.9% of cap

The 1.29x is the runner variance measured today on the SIBLING SERIAL lane, as
supplied; it is not this lane's own figure. This lane family does have its own,
and it is tighter: the six full lane-2 sums today span 192939 ms to 203451 ms,
a spread of 1.054x. At that figure the worst lane lands near 7.58 min with about
2.4 min of room. I have used the LARGER, borrowed 1.29x for the verdict rather
than the tighter one this lane actually shows, and note that the hints are
already per-script maxima, so 1.29x on top is conservative twice over.

THE MARGIN SURVIVES THE MULTIPLICATION, so this proceeds rather than stopping.
The 18 s overhead is measured, not assumed: in run 34462758357 the lane-1 job
ran 10 min 16 s against a 598.2 s suite, and lane 2 ran 3 min 21 s against a
192.9 s suite, a ~10 s difference that matches lane 1's extra Pi package install.

The cap is unchanged, the lane count is unchanged, and nothing in the serial
lane, its shard count, its guard or its hint table is touched.

=== PART 4: THE RECORDED FACT ===

The workflow comment no longer restates the shard wall as a literal, which is
how "~1 min of serial sum" survived a 10x change without announcing it. It now
points at bin/fm-test-run.sh --check-coverage, which prints parallel_max_ms,
parallel_imbalance_ms and parallel_unhinted derived from the hint table, so the
current number is computed on demand. The shard documentation carries the dated
run ids, which route it was taken by, and the local cross-check that shows why
local numbers are not admissible as hints.

Two regressions pin what rotted: lane membership must be stored
longest-measured-first, and the lanes must be fully hinted and packed within 5%
of each other. Both were run against the old composition and both fail on it
(420030 ms imbalance against a 624299 ms worst lane). The ordering assertion they
replace named a specific script by hand and had itself gone stale.

=== PART 5: NAMED AND LEFT, OUTSIDE THIS REBALANCE ===

tests/fm-captain-hold-lifecycle.test.sh alone is 296481 ms, 36% of the whole
set, so it is the floor of any two-lane split: no repacking can put a lane below
it. After this rebalance the cap is about 1.45x the healthy lane where the
sibling serial lane keeps roughly 2x.

Nothing refuses a stale parallel hint the way PORTABLE_SERIAL_MAX_UNHINTED_PERCENT
bounds the serial lane. parallel_unhinted is reported, not enforced, which is
what let this drift for three weeks unnoticed.

* fix(review): Restrict parallel scheduling hints to portable parallel lanes

* fix(document): Clarify parallel lane scheduling and timing evidence
… PR (kunchenguid#4148)

pr_for_task fell back to scraping the whole status log with tail -1, so
any PR URL a worker ever mentioned in prose - including a scout citing
someone else's PR - became the task's delivered PR in the parent-channel
terminal report. Recorded meta pr= is now the only authoritative source,
the fallback scrape accepts only a preferred terminal line in a mode's
ready-signal shape (done: PR <url> or done: PR <url> checks green), and
a scout never carries pr= at all.
…claims instead of counting a dead drop as started (kunchenguid#4212)

* fix(procevent): stop a dead runner owning a source and reconcile reporting it

The captain answered ten calls on a bearings board, the board accepted
them, and nothing collected them. He had to answer all ten again in chat.
A surface that presents as armed while being a dead drop is worse than one
that visibly fails, because the answers looked recorded.

Two independent defects, reproduced together in an isolated home where
reconcile reports started=1 on every run while ownership never moves and
no runner ever attaches.

1. reconcile counted a launch it never verified. detach_runner is
   fire-and-forget and discards the child's stderr, so a runner that died
   before it could claim was counted exactly like one that is listening.
   Launches are now confirmed - the source observed owned, or its runner
   record moved - before being reported as started; the rest are reported
   as failed= with a non-zero exit. The runner-record clause is what keeps
   a fast-completing source from being reported as a failure when it
   finished between two polls. One bounded window covers a whole cycle's
   launches, so a home full of broken sources costs the same wait as one.

2. A claim whose whole generation is provably gone could be refused
   forever. Reclaiming it ran cleanups over that dead generation's own
   leftovers, and any failure vetoed the claim - permanently, because none
   of those conditions clears on its own. Every one of those leftovers is
   keyed by the dead generation's claim token and a replacement always
   claims a fresh one, so none can collide with what replaces it.
   fm_procevent_claim_capture_reservation_reclaim_locked already said this
   for the reservation record; the staging file and the shape check on the
   registry directory recorded to hold it now take the same rule. Removing
   the claim record itself stays a hard precondition: two owners is the one
   outcome worse than none.

Two smaller repairs to the same "registered is not listening" confusion:

- `list` reported OWNER=none for a source nothing can claim. A reused PID
  whose process group survives reaches that state through the stale branch
  rather than the leaderless one, so it read as an idle source waiting to
  be started - the reassuring answer this surface gave while a board
  collected nothing. It now reports the orphaned state it shares.
- reconcile relaunched into that same unclaimable state on every cycle,
  spawning a runner that could only die on the claim. docs/configuration.md
  already promised it preserves such a claim without starting a
  replacement; the code now does that and reports it as uncertain.

This is NOT a third instance of today's two lock-identity defects
(4e1bf9aa and its replayed predecessor). Those were wrong liveness
predicates: a reused PID read as a live holder, then an exec'd holder read
as dead. Here the predicate is right - the code correctly proves the owner
dead and refuses the claim anyway, on a condition unrelated to liveness.

Regression coverage, each failing on the parent commit for its own reason:
- tests/fm-procevent.test.sh: a source that cannot start is reported as
  failed rather than started; a dead generation whose leftovers cannot be
  tidied no longer keeps owning its source (the parent reports a start
  while nothing ever runs); the existing reused-PID fixture now also
  asserts the orphaned listing and that no doomed relaunch is reported.
- tests/fm-captain-hold-lifecycle.test.sh: a board answer reaches the
  keyed-answer intake through the runner end to end - durable capture, the
  wake, and the closed task carrying the captain's selection. This one
  passes on the parent, because that chain was never what broke.

fm-procevent 100, fm-bearings-board 18, fm-captain-hold-lifecycle 50,
fm-procevent-when 13 and fm-procevent-quota 18 pass; bin/fm-lint.sh and
bin/fm-doc-audience-check.sh clean. tests/fm-extension-binding.test.sh has
two failures identical on the parent commit (EACCES on package install in
this sandbox) and unrelated to this change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gxgshn5jkWJ3GEYWy7vTG

* no-mistakes(review): confirm reconcile launches on durable launch stamps

* no-mistakes(review): announce stranded sources and refuse bad confirm windows

* no-mistakes(review): announce leaderless strands, bound confirm window, fix recovery docs

* no-mistakes(review): announce unconfirmed launches once per episode, qualify start reclaim

* no-mistakes(review): nonce launch-failed keys, refuse bad window at arm

* no-mistakes(review): state only observed launch outcome, shorten episode nonce

* no-mistakes(test): assert launch-failed headline not re-delivered, allow recovery wake

* no-mistakes(document): docs: cover strand and launch-failure wakes in skill trigger and verification record

* no-mistakes(lint): restructure SC2015 chain into explicit if-block

* test(watch-triage): fix two timing-exposed defects the pipeline found

Both surfaced in the no-mistakes test step on this branch, each failing one
full run of tests/fm-watch-triage.test.sh; neither was accepted as a flake to
retry past.

1. The new launch-failed delivery test assumed an already-surfaced key never
   wakes the watcher again. That is false: a fresh watcher legitimately
   re-surfaces any unacknowledged queue row through its downtime-recovery
   path ("check: rearm-resurface"), so the assertion failed whenever a
   re-arm landed between its two checks. The pipeline's own fix tolerated any
   wake lacking the repeated key's headline; this tightens it to exactly one
   tolerated reason, by its exact line, with a failure message that names the
   expectation so a reworded path reads as "the tolerated recovery path
   changed" rather than as a mystery - and so nobody restores the strict
   silence check. The positive assertion (a fresh-suffix key is delivered
   under its own headline) is unchanged.

2. seed_captured_procevent_result retired its source in the gap between the
   runner publishing its wake and releasing its claim, so retire read the
   exiting runner's ownership as uncertain and refused ("cannot confirm
   runner identity"). The fixture and retire path pre-date this branch; the
   confirm window returns reconcile closer to the moment of capture, which
   made the gap easier to hit. The fixture now waits, bounded, for the claim
   release the publish promises, with the reason at the wait.

Verified on this head with tasks-axi on PATH: fm-watch-triage 113/113 with
no skips, fm-procevent 106/106, fm-captain-hold-lifecycle 50/50,
fm-watch-arm 15/15, fm-bearings-board 18/18, fm-procevent-when 13/13,
fm-procevent-quota 18/18; bin/fm-lint.sh and bin/fm-doc-audience-check.sh
exit 0. First attempt, no retries.

* no-mistakes(document): docs: route stranded and launch-failed wakes in skill handling

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…gistration (kunchenguid#4191)

* fix(herdr): verify agent registrations at process level before trusting them

Herdr keeps a Pi registration (`agent get` -> agent=pi, agent_status=idle)
after the Pi process has exited to a plain shell whenever a nested interactive
shell sits under the pane's top shell, which is the crew shape `treehouse get`
leaves behind. The pane classifier trusted that registration alone, so
`fm-control.sh <id> relaunch`, `fm-spawn.sh --relaunch`, and the crew-state
recovery read all treated a shell-only pane as a live agent and refused
recovery for as long as the record lived.

The Herdr adapter now reads `pane process-info` plus the real process table
through a shared harness-process classifier (bin/fm-agent-process-lib.sh,
moved verbatim out of the tmux adapter so both backends mean the same thing by
agent, shell, and other) before a registered agent counts as live. A
registration over a shell-only pane is the new explicit `stale-agent` pane
state, which the recovery-grade read maps to `dead`; husk detection, reclaim,
presentation recovery, and session cleanup keep refusing it, so recovery reuses
the pane and nothing gains close authority. A working record is verified the
same way before the native busy verdict reports busy, so the recovery
classifier never reports a shell-only pane as working. An unreadable process
view reads unknown, trusting neither the registration nor its absence.

Reproduced and measured on Herdr 0.9.0 with Pi 0.85.1 in an isolated lab; the
new default-on live guard tests/fm-herdr-pi-stale-registration-live-e2e.test.sh
exercises the real stale record, tests/fm-control-herdr-smoke.test.sh proves
exit and relaunch through the control plane, and the portable suites pin the
classifier over real processes.

Fixes kunchenguid#4115. Duplicates: kunchenguid#3639, kunchenguid#3487, kunchenguid#2908, kunchenguid#3545.

* no-mistakes(review): settle transient prompt helpers before trusting herdr process state

* no-mistakes(review): drop stray codegraph file; read spaced comm whole in descendant walk

* no-mistakes(review): untrack stray .codegraph/.gitignore

* no-mistakes(review): untrack codegraph file; make spaced-path walk test discriminating

* no-mistakes(review): untrack stray .codegraph/.gitignore

* no-mistakes(review): untrack stray .codegraph/.gitignore re-added by fix round

* no-mistakes(review): untrack stray .codegraph/.gitignore

* no-mistakes(review): untrack codegraph file, drop dead control case, record process-info floor

* no-mistakes(review): refuse stale-agent on fresh herdr spawn preflight

Documented non-goal: fresh-spawn, reclaim, and presentation-recovery auto-recovery for a stale-agent pane is a separate design change, out of scope here, to be proposed upstream as its own issue if wanted.

* no-mistakes(test): Fix herdr flake: don't misread transient empty foreground as unreadable

* no-mistakes(document): Add fm-agent-process-lib.sh to scripts inventory

* no-mistakes(fix): update remote herdr fixture to the real pane process-info shape

The shared remote-secondmate herdr fixture still returned the old flat
process-info body ({"result":{"process":{"name":...}}}). The process-level
liveness classifier added for kunchenguid#4115 requires the real
{"result":{"type":"pane_process_info","process_info":{...foreground_processes}}}
shape and treated the old body as unreadable, so an already-launched remote
endpoint's agent-state read failed and any relaunch attempt against it died
with "remote endpoint state is unreadable; refusing duplicate launch"
instead of reaching the state it was actually exercising
(tests/fm-remote-secondmate-parent-binding.test.sh,
tests/fm-remote-secondmate-lifecycle-e2e.test.sh).

* no-mistakes(review): test: add empty-foreground regression test for herdr flake fix

* no-mistakes(document): docs: register new stale-registration live-e2e test in herdr entry points
* Bind GitHub merges to a live green head and require an away-task grant.

A GitHub merge now re-reads the pull request and passes
--match-head-commit, so a red or moved head cannot land the way GitLab
already refused. While an away record exists, only yolo or a named
grant may merge, so hold-for-return cannot ship an ungated PR.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Harden away merge authorization and grant parsing

* no-mistakes(review): Restrict fallback outcomes to proved GitHub merges

* no-mistakes(document): Refresh merge safety documentation

* no-mistakes(ci): Fixed all three CI failures by updating legacy GitHub merge fixtures for live-head verification/direct gh merges and removing a process-event runner cleanup race. Verified fm-pr-check-security, fm-captain-hold-lifecycle, and fm-watch-triage pass locally; shell syntax and git diff checks also pass

* no-mistakes(document): Document attended red-check exception

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
…m epoch (kunchenguid#4221)

The --claude guard's re-block budget charged the auto-arm ledger epoch, not
the re-block: `budget_account_current_epoch` advanced the session count only
when `state/.claude-autoarm-epoch` named a different generation than the
previous accounting. The epoch advances only inside the auto-arm hook's
generation claim, so a hook kept inert before that claim - a session lock
held by a live harness outside its ancestry, a hook that never fires, or an
identity or write failure ahead of `fm_autoarm_claim_next` - left the ledger
frozen at its last outcome and the count frozen with it. Reproduced in a
fixture: twelve consecutive Stops re-blocked with the count at 0 and the
attended fail-open never fired, leaving only Claude's silent 8-block
override, the blind end the bounded alarm exists to prevent.

The budget now charges a re-block against an epoch the previous re-block
already charged, while still charging each epoch at most once per Stop so
the wait loop's repeated observations of one fresh terminal outcome and the
same invocation's block decision cannot double count. The advancing-epoch
progression is unchanged: three re-blocks, then one attended fail-open for a
verified failure episode, and a frozen epoch now follows the same shape.
Budget exhaustion without a verified failure still blocks, by the existing
contract, and positive watcher recovery still clears the whole episode.

Regression coverage drives the real auto-arm hook against a foreign session
lock holder, asserts the ledger itself stays frozen, and fails before the fix
in both the verified and unverified shapes; the existing unverified budget
test now proves its budget actually ran out.
…enguid#4242)

* feat(herdr): attach a real foreground viewer so the live-client teardown cases can be driven

PR kunchenguid#4131 gated the Herdr active-tab close refusal on a live foreground client
instead of the persisted `.focused` pointer, but only its two detached
scenarios could be validated live. Every pseudo-terminal the runner built
started at a zero-sized window grid, so Herdr registered no foreground client
and `terminal title clear` kept answering `no_foreground_client`, leaving the
four attached-client scenarios untested. That was a harness limit, not a
product one.

Add `fm-herdr-lab.sh viewer start|stop <session>`, backed by
`bin/fm-herdr-lab-viewer.py`. The launcher sets the pty window size on the
master fd BEFORE the fork, so the TUI cannot read the grid until it is already
non-zero, and scrubs the inherited `HERDR_*` variables so Herdr's nested-viewer
refusal does not fire when the helper runs inside one of its own panes. Attach
and detach are both confirmed against the session's own foreground-client
reason rather than assumed from a signal.

The viewer inherits the lab's isolation contract: it attaches only to a session
carrying this lab's ownership tripwire, never to `default`, and it signals only
the processes it recorded, so a client someone else attached is never touched.
Teardown now refuses while an owned viewer is still attached.

Turn the reproduction into the regression with
`tests/fm-herdr-attached-viewer-live-e2e.test.sh`, which drives kunchenguid#4131's
scenarios 3, 4, 5, and 7 live against real Herdr and asserts the close refusal
fires. Scenarios 4 and 5 need a focus change at one exact product boundary, so
a PATH shim performs the real `tab focus` when the close helper issues its
planning `pane get`. Removing either half of the recipe from the launcher makes
the guard fail with the same `no_foreground_client` symptom kunchenguid#4131 reported.

* test(herdr): fail loudly when an attached-viewer fixture cannot be created

The fixture helpers run inside command substitutions, where fail() exits only
the subshell and leaves the script running with empty ids. Return non-zero
instead and carry the message at each call site.

* fix(herdr): stop the viewer launcher's kill timer from raising on an exited child

The SIGALRM escalation called os.kill unguarded, so a viewer that exited
during the grace window turned an ordinary shutdown into a traceback inside
the signal handler.

* docs: list the lab viewer's pty engine in the bin toolbelt

* no-mistakes(review): Harden Herdr viewer ownership and live CI coverage

* no-mistakes(review): Validate viewer startup timeout and process ownership

* no-mistakes(document): Document Herdr viewer safety contracts

* no-mistakes(review): Fix viewer timeout to two seconds

* no-mistakes(review): Cancel timed-out viewers and fix PTY grid

* no-mistakes(review): Serialize viewer transitions and verify process parentage

* no-mistakes(review): Harden viewer ownership locks and deduplicate CI

* no-mistakes(review): Release interrupted locks and preserve viewer escalation

* no-mistakes(review): Remove viewer locks and cancel interrupted launches

* no-mistakes(review): Close viewer launch signal races

* no-mistakes(document): Document attached Herdr viewer regression
…unchenguid#3766)

* fix(bootstrap): allow nonvisual work without Lavish

* no-mistakes(review): Gate scout brief Lavish line on bootstrap version floor
…kunchenguid#3825)

* fix(tests): isolate fixture Git configuration from host preferences

Ignore global and system Git configuration in the shared test library,
which all four fixture helper entry points source. Keep local config,
command-line overrides and explicitly supplied test config usable without
changing the caller's environment or real project signing preferences.

Exercise global and system signing inputs through all four helpers, real
fixture and child commits, explicit signing overrides, unchanged input
files, and signing refusal outside fixture subprocesses.

Verification evidence for issue kunchenguid#3770:
On pristine upstream f09de8a, all 12 reported suites failed and each logged
"No secret key" using a private GIT_CONFIG_GLOBAL containing
commit.gpgsign=true and gpg.format=openpgp, GIT_CONFIG_NOSYSTEM=1, and an
empty private GNUPGHOME (GIT_CONFIG_COUNT and GIT_CONFIG_PARAMETERS unset).
With this change, all 12 pass in the identical environment through
bin/fm-test-run.sh --per-script-timeout-secs 900:
fm-backlog-atomicity, fm-bootstrap-network-parallel, fm-bootstrap,
fm-crew-state, fm-fleet-sync, fm-gate-refuse, fm-grok-harness,
fm-session-start, fm-sessionstart-nudge, fm-tangle-guard, fm-test-run,
and fm-update (all tests/<name>.test.sh).
The new fm-test-fixtures regression failed before the library change and
passes after it. Canonical bin/fm-lint.sh passes.

Additional verification exposed fm-teardown's
herdr-preflight-missing-adapter assertion on both this branch and an
unchanged f09de8a archive with signing neutralized. That pre-existing
failure needs separate disposition; it is not repaired or skipped here.
The separately owned Muse and composer fixture defects remain untouched.

Fixes kunchenguid#3770

* no-mistakes(review): Complete fixture Git isolation and scope config assertions

* no-mistakes(review): Share Git isolation across standalone fixture entry points

* no-mistakes(review): Map git-config helper changes to lib.sh dependents

* no-mistakes(review): Select fixture-isolation regression on runner change; halve config matrix

* no-mistakes(review): Scope fixture-isolation regression selection to the runner alone

* no-mistakes(document): Give fixture Git isolation helper its owning header

* no-mistakes(document): Record fixture Git-isolation coverage in fixtures suite header

* no-mistakes(review): Fix linked-worktree fixtures and remove redundant Git isolation

* no-mistakes(document): Correct stale runner-selection documentation

* no-mistakes(document): Clarify family antecedent in isolation-proof runner evidence
… in auto mode (kunchenguid#4239)

* feat(spawn): add config/claude-permission-mode to launch Claude workers in auto mode

Every Claude worker launched with --dangerously-skip-permissions, and a
captain who refuses bypass mode had no way to select Claude Code's
classifier-reviewed auto mode instead. A new one-token local config,
config/claude-permission-mode, selects the permission flag for every
Claude launch: absent or `bypass` keeps today's launch byte-for-byte,
`auto` swaps in --permission-mode auto, and any other value refuses the
spawn before any endpoint, worktree, or record exists and names the
accepted values.

fm-spawn resolves the file on every spawn and relaunch, threads the flag
through the Claude launch template for crewmates, scouts, and secondmates
alike, and records claude_permission_mode=auto in the task meta only
under auto so the default meta stays unchanged; a relaunch re-resolves
rather than preserving the line. The file is a captain-wide safety
preference, so it joins the inherited local material pushed into
secondmate homes.

The Claude adapter reference records the verified auto launch shape on
Claude Code 2.1.269 and that it never meets the once-per-machine bypass
confirmation dialog; docs/configuration.md owns the schema.

* no-mistakes(review): drop unread claude_permission_mode meta line and its assertions
- fm-control-relaunch gives its Herdr cases a fake process table through
  FM_HERDR_PS_BIN, because the fork's exact recovery proof reads process
  identity and the host table has no fixture processes. The same cases pin
  the native pane path with an empty FM_BACKEND_HERDR_AXI_BIN, because a
  host agent-axi would otherwise drive the fake herdr.
- fm-endpoint-retire's fake tmux shows an empty composer box, because the
  merged exit path types /exit only into a composer proven empty. Its fake
  also follows the staged launch file that a spawn now sources.
- fm-procevent accepts an orphan whose new parent is a child subreaper such
  as systemd --user, provided that parent is outside the test's own tree.
On a loaded machine a 1s checkpoint can end before the watcher's first
secondmate stall tick. The next checkpoint then starts a fresh interval at
idle 0, returns the downtime recovery wake first, and the genuine stall
alert never appears. The same scenarios fail on upstream under equal load.

The observation checkpoints in the foreign-stall and reprovisioned-queue
scenarios now get the same 4s ceiling as the alert checkpoints, and each
asserts the recorded drain position, so a missed observation fails with
its real cause.
Brings fork main 7e4bdf9 (PR 48 one ShellCheck worker, PR 47 park and
resume) into the upstream reconcile. Conflicts resolved by behavior:

- ci.yml: keep upstream's two lint partitions with FM_LINT_JOBS "1",
  pin actions/upload-artifact to v4.6.2, and count 60 Bearings tests.
- fm-spawn.sh: a --resume-session relaunch always opens a new endpoint in
  the recorded worktree; every other relaunch keeps upstream's endpoint
  absence proof, adoption, and rebind. Busy arming keeps the idle resume
  seed in upstream's layout. The resume launch form stays valid with the
  new Codex --disable hooks flag.
- fm-teardown.sh removes both the progress and the pi-session records.
- fleet snapshot keeps the parked field and the PR head.
- herdr backend keeps the endpoint absence recheck and the foreground
  pid reader.
- docs and the Codex reference keep both sides' sections; park and resume
  replace upstream's "resume is not a verb" note.
- AGENTS.md: drop a redundant "to this default" so the PR 47 verb list
  fits the tracked ceiling without raising it.
Upstream now types `. '<launch file>'` instead of the whole launch
command. The park suite's tmux fake only recognized the typed command,
so a correct resume never flipped the fake pane to its harness and read
dead. Read the staged file's command, as the relaunch suite's fake does.
No assertion changed; the suite passes 20/20, as on fork main.
The integration merge kept the fork's smoke file and dropped upstream's
issue kunchenguid#4115 check that a relaunch over a stale Herdr registration still
succeeds. Restore it after the fork's stale-hook case: prove Herdr still
holds the registration, relaunch through fm-spawn, and require the new
harness to start, the endpoint to be reused, and the local copy to stay.
All its Herdr reads go through the guarded lab helper.

The merged drift section reset the recorded harness to upstream's
fixture value (claude), so the fork's later Pi interrupt check failed.
Both relaunch sections now restore the harness they found. The real
Herdr 0.8.2 lab run passes all nine checks.
The concurrent config-push scenario waited at most 2s for the first push
to reach its held send. On a loaded host both this branch and fork main
need 2.2 to 3.8s, so the scenario failed before its real check ran. The
wait ends as soon as the push arrives, so a 10s ceiling weakens nothing.
The pipeline's fix round timed out, so these land between runs, each
decided at the review gate (review-1 and review-2 by MAIN).

- Away posture never grants merge authority: the afk skill, the
  fm-pr-merge.sh header, and the Pi supervision doc now say a merge needs
  explicit captain authority or yolo plus green checks at its live head.
- Teardown refuses on the main Herdr path too when the pane close never
  confirms, keeping the task record unless --force (upstream behavior).
- A live crewmate's declared wait with a future until time still gets its
  one owed report before the declared time holds later sightings.
- Resume on Herdr proves a `missing` endpoint absent with its server
  running, so a stopped server cannot hide the task's leftover pane.
- VISION.md and GROK_BOT.md map to no tests in changed selection.
- Remove unreachable stale-agent arms and the unused
  fm_brief_marked_captain_words; correct stale classifier and promotion
  comments. The watcher's none) arm stays: a status append between the
  caller's check and the class read can reach it.
- Three Herdr test messages now describe the verdicts they assert.
The review round made a live crewmate's declared future wait report
once before its declared time. Upstream deliberately lets a live
worker's declared until time control its recheck, and its own watcher
test (test_live_paused_until_controls_recheck_time) asserts no wake
before that time. The accepted intent keeps new upstream behavior unless
a captain requirement conflicts, and the fork's one-report rule predates
this upstream feature without a captain order behind it, so the merged
ordering stands.
Keep both Codex composer verification entries in runtime-backends.md.
… PR 53)

Keep the PR 53 composer change. The merged tree still needs it: without it,
a bright footer-like row under an idle Codex frame reads empty.

Seams ledger: keep the budgets re-measured against aec043c. Add the
codex-idle-animation row at its merged delta of 248 lines. Set the teardown
budget to its measured delta of 361 lines after the document step.

Idle-frame test: upstream kunchenguid#4532 reads braille-only rows as composer furniture.
Thus a Claude glyph over a dim placeholder and braille-only rows is a ghost
(empty). Each frame now declares its Claude-glyph verdict. The frames with
prompt-row cells still assert pending. No assertion is removed.
fm-merge-local now checks the captain-hold record before a merge, which
needs FM_HOME/data. The test pointed FM_HOME at the repo root, which has
no data/ in a fresh CI checkout, so the yolo-on merge was refused. Each
case now has its own home with data/. No assertion changed.

The contributions fake forge rewrote its clock with a truncating
redirect while parallel reads could see an empty file. It now writes a
temporary file and renames it. No assertion changed.
With the fork's content, one Lint 1 shard grew past the memory of a
standard GitHub runner. CI stopped the job twice, and a local run in a
15 GB memory-capped scope stopped at 14.9 GB in the same way.

bin/fm-lint.sh now accepts 1of4 through 4of4 as well as 1of2 and 2of2.
Each root goes to the least-loaded partition, and the lowest number
wins ties, so the two-partition split is unchanged. CI runs four lint
partitions with one ShellCheck worker. A new test proves the four
partitions are complete, disjoint, deterministic, and full analysis.

Upstream: generic lint memory fix, upstreamable
… PR 55)

bin/fm-teardown.sh validates both the merge-authority record and the
pr-poll merge marker before cleanup. bin/fm-control.sh keeps the merged
empty-composer check and then records the exit-command provenance, so
a refused exit writes no record. The seams ledger keeps every row, with
budgets re-measured against upstream base aec043c.

Upstream's decision fold settles a ship or scout's decisions at its
terminal line. The fork-only retire repair still refuses a decision that
was never resolved by key, so it folds without the task kind.
Merge fork main with a merge commit. Do not rebase this branch.

- fm-spawn: keep the upstream line format and add the PR 52 abort hook.
- Herdr control smoke test: use the PR 52 lab fixture. Exit closes the
  pane, so each case gets a fresh pane. A relaunch now opens a new pane,
  so the upstream relaunch cases check that the old pane is gone and
  that the new pane is in the worktree. The lab server starts with the
  inert harness on PATH and a shell that reads no profile, so a new pane
  cannot start an installed harness.
- Unforced Herdr worker close: the PR 52 early refusal now also calls
  endpoint_close_refusal, like the main Herdr close. The refusal names
  --force and keeps the records while the slot is still leased.
- fm-teardown: the slot-return comment now names the early Herdr worker
  close refusal.
- seams.tsv: keep both sets of rows, and measure every budget again
  against aec043c.
The PR 52 pane hooks read the Herdr inventory before upstream's
stopped-server absence proof ran. A stopped server answers no inventory
read, so a pane that outlived its server could not be adopted. The
exit and relaunch hooks now step aside when the recorded server is
stopped, and the relaunch hook also steps aside when no home workspace
exists, because upstream's rebind recreates it.

A relaunch test readiness loop waits up to 30s instead of 2s under
load. No assertion changed.
Four CI lint partitions still ran partition 1 out of runner memory.
Shard memory is set by which heavy roots share one source-following
ShellCheck process, not by the root count.

bin/fm-lint.sh now reads bin/fm-lint-heavy-roots.list. In each shard,
the other roots run in one invocation as before, then each listed root
runs alone, one at a time. The list holds roots whose single-file,
memory-capped measurement reached about 4 GB or more: fm-teardown.sh
7.0 GB, fm-send.sh 7.0 GB, fm-procevent-remote-reply.sh 6.7 GB,
fm-remote-reply.test.sh 6.4 GB, fm-pending-reply.test.sh over 4 GB
(stopped before it finished), and fm-pr-merge.sh 3.6 GB. A new test
proves each listed root runs alone and every canonical root runs once.

Upstream: generic lint memory fix, upstreamable
… 57)

Clean merge. Seams budgets re-measured against upstream base aec043c.
…k graph

tests/fm-pending-reply.test.sh alone needed more ShellCheck memory than
a CI runner has: the fork's copy was killed at 14.65 GB under a 14 GB
cap, and upstream's own copy peaks at 12.69 GB. Its source directives
for bin/fm-pending-reply-lib.sh now point at /dev/null, and the file
peaks at 6.36 GB. The library stays its own lint root (4.85 GB), so it
is still checked. Directive comments mark the test's stub overrides,
which the library calls, and FM_FROMFIRST_MARK, which the library sets.
No test behavior changed.

Upstream: generic lint memory fix, upstream candidate
@BenWilcox8
BenWilcox8 merged commit 2812e3c into main Sep 25, 2026
21 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.