Skip to content

feat(bin): merge upstream Firstmate through 1b82b77a - #37

Merged
ashkonmousavi merged 86 commits into
mainfrom
fm/firstmate-upstream-catchup-0927
Sep 28, 2026
Merged

ashkonmousavi merged 86 commits into
mainfrom
fm/firstmate-upstream-catchup-0927

Conversation

@ashkonmousavi

@ashkonmousavi ashkonmousavi commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

Intent

Update firstmate and all its dependencies from Kun's repo only (2026-09-27 22:39 and 22:51), before a Pi firstmate takes over this home.
Upstream Firstmate (kunchenguid/firstmate) has 32 commits after this fork's last catch-up at 72b63ee, including quieter no-change supervision (ade7133), a memory and swap thrashing guard (6839202) and restored primary rewakes (1b82b77). Merge upstream into this fork.
Fetch upstream and merge upstream/main at its current head into a branch from this fork's main as a merge commit, with no rebase of published history, force-push, or reset. Keep every fork-unique feature with no upstream equivalent: off-table crew spawn refusal and missed-reply closes, exit and relaunch for quota-exhausted agents, the journey-walk skill, the cross-harness push skill, the prep gate and nav-prep install, Claude worker settings merge, ordered crew dispatch profiles, the Claude permission mode, board fixes, and the advisor line for eligible Claude workers. Where upstream fixed the same defect differently, take upstream's version and drop ours, naming the replaced fork commit. Carry any fork-only AGENTS.md text moved by 49a218b into the matching new skill or keep it in AGENTS.md. The full test suite passes; bin/fm-session-start.sh runs clean in an isolated scratch home. Complete the no-mistakes delivery with green checks.

What Changed

  • Merges upstream kunchenguid/firstmate main through 1b82b77a as a merge commit (no rebase or force-push), with the fork's main merged back in on top. The upstream changes are across bin/, .pi/extensions/, docs and tests. They include quieter no-change supervision outcomes (ade71339), /quiet as a statement once attended supervision is ready (8c5493a0), and a new Jev guard framework with a memory RSS/swap thrashing guard (bin/fm-jev-mem-guard.{py,sh}, docs/jev-guards.md, 6839202a). They also restore primary rewakes after attended main-only closes (1b82b77a), add the config/keep-ai-trailers opt-out and the shared bin/fm-path-lib.sh, and fix watcher, remote-worker, Herdr composer, contribution-poll and dispatch-resolver behavior.
  • Moves situational AGENTS.md sections into on-demand skills, following upstream 49a218bb. The new skills are agent-skill-trigger-index, away-quiet-supervision, operational-home-layout, scout-completion, session-start-recovery, ship-landing and validation-supervision. Fork-only wording from those sections moves word for word into the matching skills. Doc and skill references that pointed at the moved sections now point to the skills, and the fork's Codex quiet-mode refusal is kept in the quiet check, the header, the docs and the quiet skill.
  • Widens the Stop-hook retry wait in tests/fm-supervision-host.test.sh to the named timeout budgets, so the 30s+2s successor-confirmation path no longer flakes. Adds or extends tests for the upstream features, including fm-jev-mem-guard, fm-fork-free-helpers, fm-supervision-host-attended-live-e2e, fm-lint, fm-crew-state and fm-pi-branch-extension.

Merge conflict resolutions

The upstream merge resolved the following 50 conflicts against fork head a3f9be78.
combined means both sides were carried forward; upstream or fork identifies the chosen implementation for that conflicting file.

Path Resolution Reason
.agents/skills/afk/SKILL.md combined Carry fork away guidance into upstream skill.
.agents/skills/captain-hold-lifecycle/SKILL.md combined Keep keyed hold rules with upstream lifecycle.
.agents/skills/harness-adapters/references/harness/cursor.md upstream Use updated Cursor adapter contract.
.agents/skills/quiet/SKILL.md combined Keep Codex refusal with upstream quiet guidance.
.pi/extensions/lib/fm-branch-dispatch.ts upstream Use upstream branch dispatch behavior.
AGENTS.md combined Keep fork-only rules in AGENTS or moved skills.
README.md combined Keep fork overview and update quiet guidance.
bin/backends/herdr.sh combined Keep per-home Herdr safety with upstream backend changes.
bin/fm-afk-launch.sh combined Keep Codex quiet refusal and upstream away flow.
bin/fm-afk-return.sh combined Keep fork return handling and upstream rewakes.
bin/fm-branch-outcome.sh upstream Use upstream branch outcome handling.
bin/fm-branch-prompt.sh combined Preserve fork branch policy and upstream prompts.
bin/fm-branch-report.sh upstream Use upstream report handling.
bin/fm-config-inherit-lib.sh combined Keep fork worker settings and upstream inheritance.
bin/fm-git-strip-ai-trailers.sh upstream Use upstream trailer helper and opt-out.
bin/fm-host-mirror.sh upstream Use upstream host mirror helper.
bin/fm-lab-home.sh upstream Use upstream disposable-home helper.
bin/fm-spawn.sh combined Keep fork dispatch and off-table guard with upstream spawn.
bin/fm-supervise-daemon.sh fork Retain fork supervision behavior.
bin/fm-supervision-host.sh upstream Use upstream host lifecycle behavior.
bin/fm-teardown.sh combined Keep fork retirement safety and upstream lease guards.
bin/fm-test-run.sh combined Keep fork test invocation with upstream runner.
bin/fm-wake-drain.sh upstream Use upstream wake drain and primary rewakes.
bin/fm-watch-arm.sh upstream Use upstream watcher arming.
docs/calm.md upstream Use upstream Calm contract.
docs/captain-hold-lifecycle.md fork Retain fork hold documentation.
docs/configuration.md combined Document fork options and upstream config.
docs/herdr-backend.md combined Keep per-home Herdr rules with upstream changes.
docs/pi-supervision-branch.md combined Keep fork branch policy and upstream roles.
docs/remote-secondmates.md upstream Use upstream remote secondmate contract.
docs/scripts.md combined Keep fork script coverage and upstream additions.
docs/sessionstart-nudge.md combined Keep fork harness notes and upstream routing.
docs/supervision-host.md upstream Use upstream host contract.
docs/supervision-protocols/supervision-host.md upstream Use upstream host prompt protocol.
docs/turnend-guard.md combined Keep fork guard rules and upstream wake behavior.
docs/watcher-continuity.md upstream Use upstream continuity contract.
tests/fm-afk-launch.test.sh combined Cover upstream away flow and fork Codex refusal.
tests/fm-afk-return.test.sh combined Cover fork return and upstream rewakes.
tests/fm-backend-herdr-workspace-per-home-e2e.test.sh fork Retain fork per-home Herdr regression.
tests/fm-backend-herdr.test.sh combined Cover fork Herdr safety and upstream backend.
tests/fm-branch-supervision.test.sh combined Cover fork branch rules and upstream behavior.
tests/fm-daemon.test.sh fork Retain fork daemon regression.
tests/fm-gate-refuse.test.sh upstream Use upstream gate refusal tests.
tests/fm-git-strip-ai-trailers.test.sh upstream Use upstream trailer opt-out tests.
tests/fm-remote-job.test.sh upstream Use upstream remote job tests.
tests/fm-remote-secondmate-lifecycle-e2e.test.sh combined Cover fork lifecycle and upstream remote flow.
tests/fm-spawn-dispatch-profile.test.sh combined Keep fork ordered profiles with upstream spawn tests.
tests/fm-supervision-host.test.sh upstream Use upstream host tests; later fix wait budget.
tests/fm-watch-arm.test.sh upstream Use upstream arming tests.
tests/fm-watcher-lock.test.sh combined Cover fork lock behavior and upstream watcher.

Fork main moved while this branch was in preparation.
Merge commit 838a8972 carried its Codex supervision repair (e599911e) through these four additional conflicts:

Path Resolution Reason
AGENTS.md combined Keep upstream skill move plus fork Codex supervision facts.
bin/fm-watch.sh combined Keep bounded watcher cleanup and fork Codex checkpoint.
docs/turnend-guard.md combined Document both watcher and Codex turn-end rules.
docs/watcher-continuity.md combined Document both upstream continuity and fork checkpoint.

Fork work retained: c9a45c02 (off-table dispatch and missed-reply close), 9a3aacc5 (quota exit and relaunch), cdb1b34f (journey walk), a3f9be78 (cross-harness push), and the earlier 5c0addef catch-up features (prep gate, Claude worker settings and permission mode, ordered dispatch profiles, board repairs, and eligible-worker advisor line).
No fork commit was dropped wholesale or replaced by one upstream commit; individual conflicting files chosen from upstream are recorded above.
No manual state migration was identified in this catch-up; Firstmate will sync this home after merge.

Risk Assessment

⚠️ Medium: This merge spans shared supervision paths. Review fixes restored the fork-only AGENTS.md wording and Codex quiet-mode contract. The final published head passed all 19 PR checks, including both full source-aware lint partitions. Live-home cutover remains Firstmate-owned after merge.

Testing

The final published head 1046735f also repairs CI lint in bin/fm-teardown.sh. The unchanged file reproduced ShellCheck's memory failure under the existing 12 GiB address-space guard; removing four lines of unreachable duplicate sourcing passed the same guarded root at 7,152,892 KiB peak RSS, with four focused teardown/backend tests and bash -n passing. Earlier validation re-drove key fleet paths live in disposable lab homes. All 19 PR checks passed on the published head. With the old 25s wait, the new retry regression fails the same way the reported flake did. With the named 67s budget it passed twice (35-36s each, above the old limit), and the original write-failure test passed once. The loop was stopped early to report, so repeat coverage is limited to those runs, and the adjacent Stop-hook batch was not run. The following ran live in disposable lab homes, which were then torn down: bin/fm-session-start.sh in a scratch home (clean, exit 0), a real Claude primary on the lab tmux socket running session start (lock acquired, silent bootstrap, empty wake queue), off-table crew spawn refusal, and the Claude permission-mode and worker-settings refusals. All passed. There is no UI surface; the evidence is CLI transcripts and tmux pane captures. The transient runner scripts and the orphaned test fixture processes and temp directory from the stopped loop were removed, and the worktree is clean.

  • Live validation: ✅ go - 7 of 8 scenarios driven live against the product
Scenario Result Live Evidence
Stop-hook retry after a failed at-turn downtime write still notifies main (new regression) within the named budget ✅ pass live r2-retry-loop.log: 2/2 ok at 36s and 35s, above the old 25s wait; each exits 2 with the auto-arm FAILED notice and a committed failed outcome
Adversarial: the old 25s wait fails on the retry path (proves the regression reproduces the flake) ✅ pass live r2-retry-regression-oldwait.log: not ok - hook retry write failure: the Stop hook did not finish with HOOK_RETRY_POLLS=250
Original failed-downtime-write test keeps its assertions and passes ✅ pass live r2-retry-loop.log iter 1: ok in 4s
bin/fm-session-start.sh runs clean in an isolated scratch home ✅ pass live r2-session-start-scratch-home.txt: lock acquired, BOOTSTRAP silent, no queued wakes, exit=0
Real Claude primary in a lab home runs session start and takes the lock ✅ pass live r2-lab-claude-primary-session-start.txt: the primary reports LOCK acquired, BOOTSTRAP silent, WAKE QUEUE empty; state/.lock holds the primary pid
Adversarial: off-table crew spawn is refused with no state written ✅ pass live r2-offtable-spawn-refusal.txt: both off-profile spawns exit 1 with the named-override refusal; no state/offtable-z1.* written
Adversarial: invalid Claude permission mode or worker settings refuse the spawn ✅ pass live r2-claude-config-refusals.txt: both exit 1 with a precise error; no state/scout-z1.* written
Adjacent Claude Stop-hook host+hook cases (pass-through, at-turn close, successor close during main turn) still pass with the refactored helper ⏸️ untested no This batch was queued after the retry loop and did not start because the loop was stopped early to deliver this result. It can be driven by running tests/fm-supervision-host.test.sh or CI's supervisio…
Evidence: Retry regression fails with the old 25s wait (reproduces the flake)

Source: Retry regression fails with the old 25s wait (reproduces the flake)

== HEAD 4a3542e9: pre-fix wait (HOOK_RETRY_POLLS=250, the old 25s) on the new retry regression
not ok - hook retry write failure: the Stop hook did not finish
exit=1
Evidence: Retry and original write-failure Stop-hook tests with the named budget

Source: Retry and original write-failure Stop-hook tests with the named budget

== HEAD 4a3542e9: new retry test + original write-fails test, HOOK_RETRY_POLLS=670
-- iter 1
ok - host+hook: failed at-turn downtime write notifies main after the Stop hook's retry
elapsed test_claude_stop_hook_retry_notifies_when_at_turn_downtime_write_fails: 36s
ok - host+hook: failed at-turn downtime write notifies main despite a healthy successor
elapsed test_claude_stop_hook_notifies_when_at_turn_downtime_write_fails: 4s
exit=0
-- iter 2
ok - host+hook: failed at-turn downtime write notifies main after the Stop hook's retry
elapsed test_claude_stop_hook_retry_notifies_when_at_turn_downtime_write_fails: 35s
(loop stopped early by the tester after the results above; later iterations and the r2-stop-hook-cases batch were not run)
Evidence: fm-session-start.sh in a scratch lab home at 4a3542e

Source: fm-session-start.sh in a scratch lab home at 4a3542e9

$ FM_HOME=<lab> bin/fm-session-start.sh   (commit 4a3542e9)

================================================================================
SESSION START - <lab>
================================================================================

LOCK
--------------------------------------------------------------------------------
lock acquired: harness pid 3440603

BOOTSTRAP
--------------------------------------------------------------------------------
(silent - all good)

WAKE QUEUE
--------------------------------------------------------------------------------
(no queued wakes)
================================================================================
SUPERVISION OPERATING INSTRUCTIONS - primary harness: claude
================================================================================
Current state:
- Lock: held by this session; this session owns normal supervision unless away mode says otherwise.
- Away/quiet mode: inactive.
- X mode: inactive; use the default watcher cadence.
- Ordinary wake: the Stop-owned auto-arm (bin/fm-claude-stop-autoarm.sh) already owns watcher continuity; drain and handle the wake, and do not arm another cycle yourself.

Mode: Claude Stop-hook-owned supervision.

When this session owns supervision and away mode is not active:
1. Drain first with `bin/fm-wake-drain.sh`.
   After handling all emitted wakes and reconciling open decisions and unread status lines, run the exact `--ack-through` command printed as `WAKE_ACK_REQUIRED`; until then the work remains durable for idempotent re-handling after interruption.
2. Routine watcher arm and re-arm are owned by the Stop `asyncRewake` hook (`bin/fm-claude-stop-autoarm.sh`), never by you.
   Every turn end while supervision is needed launches or attaches one home-scoped watcher cycle with no model command and no model tokens.
   An actionable close wakes you through the hook's exit-2 rewake, delivered as a `Stop hook feedback` message.
3. On a `Stop hook feedback` wake (`signal:`, `stale:`, `check:`, or `heartbeat`), run `bin/fm-wake-drain.sh` first and handle the wake.
   Do not run `bin/fm-watch-arm.sh` after an ordinary wake; the next turn end re-arms automatically when supervision is still needed.
   Do not invent a wake from an attach-status line alone; drain and act only on real wake records, the drain's `OPEN DECISIONS` and `UNREAD STATUS` entries, or a real watcher reason line.
4. On the one `Stop hook feedback` automatic-mechanism failure notice (`firstmate watcher auto-arm FAILED ...`), drain, inspect the automatic mechanism failure, and do not turn the notice into a repeating manual-arm loop.
5. If the Stop hook does not claim the home or reports an exhausted failure, inspect its registration and watcher startup path before ending blind.
   Keep the Stop-owned automatic mechanism as the only Claude arm owner.
6. Treat `watcher: started ...` and `watcher: attached ...` inside automatic arm output as proof that one live cycle exists.
   On attach, the arm follows verified identity-matched successors instead of exiting when the first cycle ends.
7. The durable wake queue preserves actionable events between a rewake and the next Stop-launched arm, while the bounded turn-end guard prevents a blind Stop when recovery did not start.
   No PreToolUse hook denies fleet commands based on watcher status.
   [`watcher-continuity.md`](../watcher-continuity.md) owns the exact session-lock recovery boundary.
8. The turn-end guard (`bin/fm-turnend-guard.sh --claude`) remains the final backstop.
   It requires the PID-strict live-watcher and fresh-beacon predicate at the Stop boundary, except for the Claude-specific foreign-live-owner safe exit owned by [`turnend-guard.md`](../turnend-guard.md#guard-predicates); that document also owns the distinct model-aware mid-turn pull-guard rules.
   Otherwise, it allows the stop when a watcher is healthy or an open auto-arm generation claim owns recovery, while fresh failure epochs advance the bounded one-time attended fail-open progression described there.
9. Waiting on the hook-owned cycle is silent: do not send idle progress while the watcher is parked.

The watcher itself remains `bin/fm-watch.sh`, and `bin/fm-watch-arm.sh` remains the verified arm wrapper that the Stop hook foregrounds unless this home opts into the [supervision host](../supervision-host.md).
Re-arm attaches to an existing healthy cycle when one is already present and follows its verified successor chain.
See [`watcher-continuity.md`](../watcher-continuity.md) for the arm-layer successor and clean-close failure contract and the Claude ownership model.


================================================================================
READ-ONCE CONTRACT
================================================================================
Everything below is printed in full for this session start: every state/*.meta,
a compact data/backlog.md listing, a bounded tail of every state/*.status,
data/projects.md, data/secondmates.md, data/captain.md, data/captain-shared.md,
and data/learnings.md.
Do NOT re-read any of them after reading this digest, and do NOT bulk-read
data/backlog.md or state/*.status: re-reading everything defeats the entire
point of this command.

Go to a source directly only when:
  - this digest flagged it ABSENT (then rebuild or create it per AGENTS.md),
  - its contents looked unparseable or corrupt,
  - an individual full status log is needed for older wake-event history, or a
    status line was capped and its tail matters (each task's full log path is
    printed with its tail),
  - a full task body is needed (bin/fm-tasks-axi.sh show <id> --full, or data/backlog.md),
  - the backlog listing disclosed omitted queued items and this turn needs them,
  - the NETWORK CHECKS section reported its checks still IN PROGRESS and this
    turn needs their verdict (bin/fm-startup-network.sh report),
  - or a STARTUP TRUNCATED banner named the stage that would have printed it, in
    which case that stage's sources were never emitted and must be reconciled.

================================================================================
FLEET STATE
================================================================================

data/backlog.md
--------------------------------------------------------------------------------
ABSENT

Work under way (state/*.meta)
--------------------------------------------------------------------------------
(none)

Orphan status logs (state/*.status without matching .meta)
--------------------------------------------------------------------------------
(none)

AFK
--------------------------------------------------------------------------------
absent

================================================================================
NETWORK CHECKS
================================================================================
IN PROGRESS - the deferred network checks have not finished yet.
NOT yet confirmed: GitHub authentication, dead-secondmate relaunch, secondmate convergence, pending handoff delivery, project clone refresh with its drift reporting, and inactive terminal-outcome reconciliation.
Started 1s ago, bounded at 120s.
Only a FAILED or otherwise actionable result arrives as a `check: startup-network` wake; a clean success stays silent.
The durable result is readable on demand with ~/.no-mistakes/worktrees/befb827bbae4/01M3KPWBK1K6340ZCX5T29093S/bin/fm-startup-network.sh report; until it finishes, treat none of it as confirmed.

================================================================================
CONTEXT
================================================================================

data/projects.md
--------------------------------------------------------------------------------
ABSENT

data/secondmates.md
--------------------------------------------------------------------------------
ABSENT

data/captain.md
--------------------------------------------------------------------------------
ABSENT

data/captain-shared.md (shared, main-authoritative, read-only in secondmate homes)
--------------------------------------------------------------------------------
ABSENT

data/learnings.md
--------------------------------------------------------------------------------
ABSENT

================================================================================
NEXT STEP
================================================================================
Follow the supervision operating instructions block above for harness 'claude'.
This script never starts supervision itself.

The digest above is complete for this session start. The READ-ONCE CONTRACT
section near the top of it governs what may still be read from disk.
exit=0
Evidence: Real Claude primary running session start in a lab home

Source: Real Claude primary running session start in a lab home

# commit 4a3542e9: real Claude primary on lab tmux socket, FM_HOME=<lab>
 ▐▛███▛█   Claude Code v2.1.283
▝▜██████▀  Opus 5.5 with medium effort · Claude Max
 ▝▝   ▝▝   ~/.no-mistakes/worktrees/befb827bbae4/01M3KPWBK1K6340ZCX5T29093S · /rc
❯ This is a disposable lab home. Run bin/fm-session-start.sh, then reply with just its LOCK line, whether BOOTSTRAP was silent, and the WAKE QUEUE line. Do not start any other work.
  Ran 1 shell command
● Captain, here's what the session start reported:
  - LOCK: lock acquired: harness pid 3528046
  - BOOTSTRAP: silent. It printed "(silent - all good)".
  - WAKE QUEUE: (no queued wakes)
✻ Worked for 9s · done 4:25 AM
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
❯ 
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
  Lab home setup verification · WORKTREE:01M3KPWBK1K6340ZCX5T... · opus/medium · default · 6% · 5h8%/7d99% · $0.41
  ⏵⏵ auto mode on (shift+tab to cycle) · ← 2 agents
## <lab>/state after the turn:
. .. .inactive-outcome-reconcile .lock .lock-session .session-start-complete .startup-network.report .startup-network.status .startup-network.timings .trace-context-effective .wake-queue home-summary.json terminal-outcomes 
## state/.lock pid: 3528046
Evidence: Off-table crew spawn refusal

Source: Off-table crew spawn refusal

# commit 4a3542e9; config/crew-dispatch.json: rules[0]=grok/grok-4/high, default=codex/gpt-5/high
$ bin/fm-spawn.sh offtable-z1 <lab>/projects/demo --scout --harness claude --model opus --effort high
fm-gate-refuse: gate agent lifecycle permitted only against lab home <lab>
error: spawn refused - claude/opus/high does not match crew-dispatch rule default: no candidate profile of that rule has this harness, model, and effort; record the captain's explicit override at <lab>/data/offtable-z1/dispatch-override or pass one of its profiles
exit=1
$ bin/fm-spawn.sh offtable-z1 <lab>/projects/demo --scout --dispatch-rule 0 --harness codex --model gpt-5 --effort low
fm-gate-refuse: gate agent lifecycle permitted only against lab home <lab>
error: spawn refused - codex/gpt-5/low does not match crew-dispatch rule 0: no candidate profile of that rule has this harness, model, and effort; record the captain's explicit override at <lab>/data/offtable-z1/dispatch-override or pass one of its profiles
exit=1
## state/ entries for offtable-z1:
(none written)
Evidence: Claude permission-mode and worker-settings refusals

Source: Claude permission-mode and worker-settings refusals

# commit 4a3542e9
# config/claude-permission-mode = 'yolo' (invalid)
$ bin/fm-spawn.sh scout-z1 <lab>/projects/demo --scout --harness claude
error: config/claude-permission-mode holds 'yolo'; accepted values are: bypass (--dangerously-skip-permissions, the default when the file is absent), auto (--permission-mode auto)
exit=1
# config/claude-permission-mode = auto, config/claude-worker-settings.json = '[1,2]'
$ bin/fm-spawn.sh scout-z1 <lab>/projects/demo --scout --harness claude
error: config/claude-worker-settings.json must be a readable regular file holding one JSON object of Claude Code settings
exit=1
## state/ entries for scout-z1:
(none written)
- Outcome: 🔧 1 issue found → auto-fixed ✅ across 2 runs (1h27m58s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - 1 info
  • ⚠️ bin/fm-afk-launch.sh:309 - The merge took upstream's quiet-mode text but kept the fork's Codex rule that enter refuses quiet mode (line 399). That leaves the script's runtime output and its header contract wrong for Codex. On a Codex home with config/supervision-host, the attended host never counts as ready, because FM_HOST_MIRROR_VERIFIED='claude cursor'. quiet-check then prints "...quiet mode enters through the quiet daemon instead." and exits 1, and the following quiet enter refuses because Codex runs no daemon. Other places that still describe the old Codex behavior: bin/fm-afk-launch.sh:17-24, where the header now says only Pi/pi-signed refuse start and 'every other harness still runs the daemon' (the fork's Codex sentence about the legacy-daemon handoff to the foreground checkpoint was dropped, but the code at 335 and 788-800 still does that); bin/fm-afk-launch.sh:38-40, where the QUIET MODE header says quiet 'enters through the daemon' when the host is unready; bin/fm-afk-launch.sh:324, a comment saying 'no longer launched on Pi' although the case at 335 also covers codex; docs/supervision-host.md:137, which says '/quiet names the missing part and enters the quiet daemon'; and .agents/skills/quiet/SKILL.md:47-48, which dropped the fork's instruction to tell the captain that quiet mode is unavailable on Codex, that /afk is the away posture there, and to stop. Fix: make the quiet-check line Codex-aware and restore the fork's Codex sentences in the upstream-structured header, comment, doc and skill.
  • ⚠️ .agents/skills/agent-skill-trigger-index/SKILL.md:15 - Intent: "Carry any fork-only AGENTS.md text moved by 49a218b into the matching new skill or keep it in AGENTS.md." Some fork-only AGENTS.md lines were shortened instead of carried. (1) agent-skill-trigger-index:15 research-first-decisions drops the triggers 'before recording such a selection as decided', 'integrating a library, SDK, API, CLI, framework, or service', and 'its Context7 version check is mandatory whether or not a candidate is being chosen'. (2) Line 16 wayfinding drops 'the queue looks fully gated' and 'the ready frontier lists only umbrellas' (now just 'empty'). This index says these skills load 'only at their precise triggers', so the shorter lines now narrow when they load, although each skill's own frontmatter still lists the full triggers. (3) operational-home-layout/SKILL.md:22 .env drops '(presence-gates bin/fm-dispatch-resolve.sh)', which both upstream and the fork had. (4) away-quiet-supervision/SKILL.md:16 drops 'Codex inside its foreground checkpoint loop'. (5) operational-home-layout/SKILL.md:25 claude-worker-settings drops its 'switch off add-on servers' example. Restoring the fork wording verbatim is the smallest remedy. It is flagged ask-user because it concerns what the intent requires to be carried.
  • ℹ️ tests/fm-branch-supervision.test.sh:430 - The merge commit changes upstream's 47h age sample to 47h30m, with a comment about a clock tick crossing the 48-hour boundary. recordedAgo in bin/fm-branch-outcome.sh:140-142 floors the age, so a 1-second tick cannot move 47h to 48h. The edit is harmless, but its stated reason is wrong and it differs from upstream for no need; consider reverting to upstream's line.

🔧 Fix applied.
3 issues (2 warnings, 1 info) still open:

  • ⚠️ .agents/skills/agent-skill-trigger-index/SKILL.md:15 - Intent: "Carry any fork-only AGENTS.md text moved by 49a218b into the matching new skill or keep it in AGENTS.md." Some fork-only AGENTS.md lines were shortened instead of carried. (1) agent-skill-trigger-index:15 research-first-decisions drops the triggers 'before recording such a selection as decided', 'integrating a library, SDK, API, CLI, framework, or service', and 'its Context7 version check is mandatory whether or not a candidate is being chosen'. (2) Line 16 wayfinding drops 'the queue looks fully gated' and 'the ready frontier lists only umbrellas' (now just 'empty'). This index says these skills load 'only at their precise triggers', so the shorter lines now narrow when they load, although each skill's own frontmatter still lists the full triggers. (3) operational-home-layout/SKILL.md:22 .env drops '(presence-gates bin/fm-dispatch-resolve.sh)', which both upstream and the fork had. (4) away-quiet-supervision/SKILL.md:16 drops 'Codex inside its foreground checkpoint loop'. (5) operational-home-layout/SKILL.md:25 claude-worker-settings drops its 'switch off add-on servers' example. Restoring the fork wording verbatim is the smallest remedy. It is flagged ask-user because it concerns what the intent requires to be carried.
  • ℹ️ tests/fm-branch-supervision.test.sh:430 - The merge commit changes upstream's 47h age sample to 47h30m, with a comment about a clock tick crossing the 48-hour boundary. recordedAgo in bin/fm-branch-outcome.sh:140-142 floors the age, so a 1-second tick cannot move 47h to 48h. The edit is harmless, but its stated reason is wrong and it differs from upstream for no need; consider reverting to upstream's line.
  • ⚠️ .agents/skills/agent-skill-trigger-index/SKILL.md:15 - Still present after the round 1 fix; it was left unselected and is still waiting on a decision. The intent says: "Carry any fork-only AGENTS.md text moved by 49a218b into the matching new skill or keep it in AGENTS.md." The merge resolution shortened several fork-only AGENTS.md lines instead of carrying them over. (1) agent-skill-trigger-index/SKILL.md:15 research-first-decisions drops 'before recording such a selection as decided', 'integrating a library, SDK, API, CLI, framework, or service', and 'its Context7 version check is mandatory whether or not a candidate is being chosen' (fork AGENTS.md:620). (2) Line 16 wayfinding drops 'such as a stage, a release, a migration, or a campaign', 'the queue looks fully gated', and 'the ready frontier lists only umbrellas or nothing' (fork AGENTS.md:621). This index says skills load 'only at their precise triggers', so the shorter lines narrow when these skills load. (3) operational-home-layout/SKILL.md:22 .env drops '(presence-gates bin/fm-dispatch-resolve.sh)' (fork AGENTS.md:71). (4) away-quiet-supervision/SKILL.md:16 drops 'Codex inside its foreground checkpoint loop' (fork AGENTS.md:505). (5) operational-home-layout/SKILL.md:25 claude-worker-settings drops 'e.g. to switch off add-on servers workers never use' (fork AGENTS.md:74). The smallest remedy is to restore the fork wording verbatim. This is marked ask-user because it turns on how much the carry requirement covers.

🔧 Fix applied.
1 info still open:

  • ℹ️ tests/fm-branch-supervision.test.sh:430 - The merge commit changes upstream's 47h age sample to 47h30m, with a comment about a clock tick crossing the 48-hour boundary. recordedAgo in bin/fm-branch-outcome.sh:140-142 floors the age, so a 1-second tick cannot move 47h to 48h. The edit is harmless, but its stated reason is wrong and it differs from upstream for no need; consider reverting to upstream's line.
🔧 **Test** - 1 issue found → auto-fixed ✅
  • ⚠️ tests/fm-supervision-host.test.sh:1140 - test_claude_stop_hook_notifies_when_at_turn_downtime_write_fails flakes with 'hook write failure: the Stop hook did not finish'. It failed once in the targeted batch, once in two solo reruns, and intermittently in focused loops. The test file, bin/fm-supervision-host.sh and bin/fm-claude-stop-autoarm.sh are byte-identical to upstream 1b82b77, and upstream's untouched tree fails the same assertion. So this was inherited from upstream, not introduced by the merge. With a longer wait, both trees always reach the correct outcome (exit 2 plus the failure notice to main). The Stop hook's exit time is bimodal on both trees: 0-5s or 34-57s, with nothing in between, while the test waits only 25s (wait_until 250). The ~32s slow path matches the handling-successor confirmation budget in start_handling_successor (FM_ARM_CONFIRM_TIMEOUT default 30, plus 2s) in bin/fm-claude-stop-autoarm.sh:339-367. That suggests the successor arm sometimes fails to confirm after the failed downtime write and the hook waits out the full budget. CI runs this file and the intent requires green checks, so this can flake a required shard. The fix is to either bound the wait on that budget or set FM_ARM_CONFIRM_TIMEOUT low in this case, ideally upstream, rather than silently widening the wait. Evidence: flake-branch-timing.log, flake-upstream-timing.log, flake-upstream-1b82b77a.log, flake-branch-diag.log.
  • Live validation: ✅ go - 7 of 12 scenarios driven live against the product
Scenario Result Live Evidence
Operator runs bin/fm-session-start.sh in an isolated scratch home and gets a clean digest ✅ pass live session-start-scratch-home.txt: exit=0, lock acquired, BOOTSTRAP '(silent - all good)', no queued wakes; the fm-startup-network.sh report said '(silent - no problems found)'
A real Claude primary in a lab home runs session start and reports lock, bootstrap and wake queue ✅ pass live lab-claude-primary-session-start.txt: the Opus primary on the fm-lab socket reported 'lock acquired', bootstrap silent, '(no queued wakes)'
Upstream memory RSS/swap thrashing guard (6839202) reports host state and fails --check above threshold ✅ pass live jev-mem-guard-live.txt: OK status and top RSS list, structured JSON, and --check with a 1%/2% threshold returned CRITICAL with exit=1
Fork off-table crew spawn refusal: a scout spawn with a harness/model/effort not in crew-dispatch.json is refused before any task state exists ✅ pass live offtable-spawn-refusal.txt: both the default-rule and --dispatch-rule 0 attempts exit=1 with 'does not match crew-dispatch rule', and no state/offtable-z1.* was written; dispatch-select shows 'selecti…
Fork Claude permission mode and worker-settings merge: invalid config refuses spawn; valid config reaches the real worker launch ✅ pass live claude-config-refusals.txt: 'yolo' and '[1,2]' both refuse with exit=1 and no state. claude-scout-launch-argv.txt: the real worker ran `claude --permission-mode auto ... --settings {deniedMcpServers..…
Fork advisor line: an eligible Claude worker (Sonnet) gets the Opus advisor section; a Fable relaunch drops it ✅ pass live claude-relaunch-fable-no-advisor.txt: before, launch-brief had '# Advisor tool'; after fm-control.sh relaunch --model claude-fable-5-1 the meta read model=claude-fable-5-1 and the brief had 0 adviso…
Fork exit on a just-exited Claude worker is idempotent and preserves the endpoint ✅ pass live claude-exit-idempotent.txt: first exit 'stopped', second exit 'already-stopped', both exit=0, and the pane survived with a 'claude --resume' hint
Fork exit/relaunch of a quota-exhausted agent (9a3aacc) ⏸️ untested no A real quota exhaustion on a live Claude, Codex or Grok account cannot be triggered on demand from this sandbox. Driving it live needs an account that is actually at its usage limit. Only the idle and…
Fork missed-reply close settlement (second half of c9a45c0) ⏸️ untested no Driving it live needs a live secondmate home with an outstanding pending-reply expectation that is then missed. A secondmate provisioning run was out of scope for this time-boxed lab. The fm-send-reso…
Upstream quiet no-change supervision (ade7133) and restored primary rewakes (1b82b77) through the supervision host ⏸️ untested no A real Claude primary with config/supervision-host and a live scout in flight was started from the gate worktree. At turn end no host or watcher armed, because fm_primary_scope_matches deliberately ke…
Merge shape: upstream/main head merged into fork main as a merge commit with no rebase, and fork feature commits kept ⏸️ untested no The prior payload did not establish a live result: this was a git-history inspection (merge-topology.txt showed df2e83a with parents a3f9be7 and 1b82b77, and the fork feature commits as ancestors o…
Full test suite passes ⏸️ untested no The rules for this test phase forbid running the complete repository suite, and remote CI owns broad regression. Only a targeted 21-suite subset was run (see targeted-suites.log); the one flake found…
  • bin/fm-lab-home.sh create $LAB then FM_HOME=$LAB bin/fm-session-start.sh in a scratch lab home, followed by bin/fm-startup-network.sh report
  • Real claude primary started on the lab's private tmux socket (tmux -L fm-lab new-session ... -e FM_HOME=$LAB claude) from the gate worktree and prompted to run bin/fm-session-start.sh
  • bin/fm-jev-mem-guard.sh --help, bare run, --json, and --check --warn-mem-pct 1 --crit-mem-pct 2 on the live host
  • bin/fm-dispatch-select.sh &lt;lab&gt;/config/crew-dispatch.json 0, plus two off-table bin/fm-spawn.sh ... --scout attempts (default rule and --dispatch-rule 0)
  • bin/fm-spawn.sh refusals with config/claude-permission-mode=yolo and with config/claude-worker-settings.json=[1,2]
  • Real Claude scout spawn with claude-permission-mode=auto and a worker settings object, with the launched argv and child MCP processes inspected
  • bin/fm-control.sh scout-z1 relaunch --model claude-fable-5-1 --note ..., checking the launch-brief advisor section before and after
  • bin/fm-control.sh scout-z1 exit called twice (stopped, then already-stopped), then bin/fm-teardown.sh scout-z1 --force
  • bin/fm-test-run.sh --per-script-timeout-secs 1500 on 21 targeted suites: fm-jev-mem-guard, fm-supervision-host, fm-afk-return, fm-branch-supervision, fm-supervision-instructions, fm-claude-stop-autoarm, fm-control-relaunch, fm-control, fm-spawn-dispatch-profile, fm-dispatch-resolve, fm-prep-install, fm-brief, fm-bearings-board, fm-captain-hold-lifecycle, fm-spawn-compact-adviser-disable, fm-send-resolve-key, fm-pending-reply-retire-notice, fm-codex-session, fm-session-start, fm-watcher-lock, fm-afk-launch
  • bash tests/fm-supervision-host.test.sh rerun twice alone, plus an 8-iteration focused loop of test_claude_stop_hook_notifies_when_at_turn_downtime_write_fails on this branch and on a git archive 1b82b77a tree (temporary drivers, since removed)
  • A second lab with config/supervision-host and a live Claude scout in flight, a real Claude primary turn end, and the fm_primary_scope_matches scope check
  • git log/git merge-base --is-ancestor checks of the merge parents and of fork and upstream feature commits; git diff e599911e HEAD on the fork skill and feature files

🔧 Fix applied.
✅ Re-checked - no issues remain.

  • Live validation: ✅ go - 7 of 8 scenarios driven live against the product
Scenario Result Live Evidence
Stop-hook retry after a failed at-turn downtime write still notifies main (new regression) within the named budget ✅ pass live r2-retry-loop.log: 2/2 ok at 36s and 35s, above the old 25s wait; each exits 2 with the auto-arm FAILED notice and a committed failed outcome
Adversarial: the old 25s wait fails on the retry path (proves the regression reproduces the flake) ✅ pass live r2-retry-regression-oldwait.log: not ok - hook retry write failure: the Stop hook did not finish with HOOK_RETRY_POLLS=250
Original failed-downtime-write test keeps its assertions and passes ✅ pass live r2-retry-loop.log iter 1: ok in 4s
bin/fm-session-start.sh runs clean in an isolated scratch home ✅ pass live r2-session-start-scratch-home.txt: lock acquired, BOOTSTRAP silent, no queued wakes, exit=0
Real Claude primary in a lab home runs session start and takes the lock ✅ pass live r2-lab-claude-primary-session-start.txt: the primary reports LOCK acquired, BOOTSTRAP silent, WAKE QUEUE empty; state/.lock holds the primary pid
Adversarial: off-table crew spawn is refused with no state written ✅ pass live r2-offtable-spawn-refusal.txt: both off-profile spawns exit 1 with the named-override refusal; no state/offtable-z1.* written
Adversarial: invalid Claude permission mode or worker settings refuse the spawn ✅ pass live r2-claude-config-refusals.txt: both exit 1 with a precise error; no state/scout-z1.* written
Adjacent Claude Stop-hook host+hook cases (pass-through, at-turn close, successor close during main turn) still pass with the refactored helper ⏸️ untested no This batch was queued after the retry loop and did not start because the loop was stopped early to deliver this result. It can be driven by running tests/fm-supervision-host.test.sh or CI's supervisio…
  • git diff --stat b4c245d2 HEAD (only tests/fm-supervision-host.test.sh changed since the round-1 live validation)
  • Focused runner (the test file's definitions plus selected calls), with HOOK_RETRY_POLLS forced to 250 (the old 25s wait): test_claude_stop_hook_retry_notifies_when_at_turn_downtime_write_fails, which gave not ok - hook retry write failure: the Stop hook did not finish
  • Focused runner at HEAD (HOOK_RETRY_POLLS=670): test_claude_stop_hook_retry_notifies_when_at_turn_downtime_write_fails passed twice (36s, 35s) and test_claude_stop_hook_notifies_when_at_turn_downtime_write_fails passed once (4s); the loop was stopped early after that
  • bin/fm-lab-home.sh create $LAB then FM_HOME=$LAB bin/fm-session-start.sh at 4a3542e9 (lock acquired, BOOTSTRAP silent, no queued wakes, exit 0)
  • Real Claude primary on the lab tmux socket (tmux -L fm-lab new-session ... -e FM_HOME=$LAB2 claude), driven to run bin/fm-session-start.sh, then tmux -L fm-lab kill-server and rm -rf $LAB2
  • bin/fm-spawn.sh offtable-z1 ... --harness claude --model opus --effort high and --dispatch-rule 0 --harness codex --model gpt-5 --effort low against a lab config/crew-dispatch.json (both refused, no state written)
  • bin/fm-spawn.sh scout-z1 ... --harness claude with an invalid config/claude-permission-mode, then with a non-object config/claude-worker-settings.json (both refused, no state written)
⚠️ **Document** - 1 info
  • ℹ️ bin/fm-merge-outcome-lib.sh:66 - The runtime merge-landed reminder still ends with 'see AGENTS.md section 7'. The merge moved the post-merge verification rule into the ship-landing skill, and section 7 now only points there, so the reminder still leads the reader to the right rule, one hop later. Changing the string is a behavior change: the pinned MERGE_LANDED_REMINDER test fixture would have to change with it, so this documentation phase left both alone. Follow-up: point the reminder and its fixture at ship-landing together.
✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

tiago-peixoto and others added 30 commits September 24, 2026 15:22
…ning (kunchenguid#5566)

* fix(bin): report a Lavish source armed only after its listener is running

Registration alone was treated as ready, so arm could succeed before anything was collecting from the board.

* no-mistakes(review): Guard Lavish arm launches, keep retire refusals, report live prior listener

* no-mistakes(review): Keep polling through window before reporting a still-live prior listener

* test: wait for a capture's claim to drop before the next arm

The result is stored before the runner exits, so a re-arm in that gap was meeting a live claim.

* no-mistakes(document): Record Lavish arm readiness evidence in verification doc

* no-mistakes(ci): Both failures were caused by this PR, and both are fixed with test-only edits. Lint 2 (ShellCheck SC2034): this branch removed the only use of `reply_id` (a `start "$reply_id"` call) from tests/fm-procevent.test.sh, which left the assignment at line 1450 unused. I deleted that assignment. It was the only `reply_id` in the file. ShellCheck is now clean on both test files. Behavior portable serial 4: the failing test was tests/fm-bearings-board.test.sh, in the check "registration consumed its answer before the any-origin binding existed". I reproduced it locally: the hold was still `state: queued` when the test checked it. - What must hold: the test's check that the hold is closed must run after the listener has captured the answer. - Why it broke: the test used a stand-in adapter that ran `fm-procevent.sh start` in the foreground after `arm`, so capture finished before build returned. On this branch, `arm` starts the listener itself in the background, so the real listener captures the answer and closes the hold a moment after build returns. - Fix: removed the now-redundant stand-in adapter, the copied runtime directory, and its extra environment variables. The test now runs the real build through the existing `run_board` helper and waits up to about 10s for the hold to reach `state: done`. The checks that follow are unchanged: `Resolution mode: answered` and the any-origin binding. - Other tests: this was the only test in the file that stood in for the adapter this way. The shard's other pure-contract-unit test (tests/fm-trace-context-lib.test.sh) passed unchanged. Verification: - tests/fm-bearings-board.test.sh passed 3 times in a row via bin/fm-test-run.sh, all 18 checks, about 53s per run. - tests/fm-procevent.test.sh was not rerun, because the lint fix only removed an unused assignment
…elivered (kunchenguid#5599)

* fix(bin): acknowledge a delivered unknown-wake escalation

The same unrecognized wake was escalated again after it had already been handled, because delivery never recorded that identity.

* no-mistakes(review): Scope unknown-wake acknowledgements to one away session

* no-mistakes(review): Clear delivered digest when unknown-wake ack write fails

* no-mistakes(review): Limit unknown-wake suppression to acknowledged lines

* no-mistakes(document): List unknown-wake ack file among away-session artifacts
…ess wait (kunchenguid#5587)

* fix(bin): keep a stated default retraction from cancelling a keyless wait

A resolved line that names the shared default decision bucket was closing the keyless live wait that only prints as that same key. Keyless self-retraction still closes the keyless wait.

* no-mistakes(review): Keep declared waits standing past foreign-key resolved lines

* no-mistakes(review): Bound declared-wait read and share one decision-key parser

* no-mistakes(document): Document supervisors' key-aware declared-wait read
…unchenguid#5544)

* fix(bin): terminate a remote job worker that lost ownership when it receives TERM

A serving worker whose lock directory is gone can no longer quarantine shutdown, and resuming service publishes a false ready heartbeat. Exit after stopping only that worker's own command tree, without removing a replacement owner's lock.

* no-mistakes(review): Check worker lock ownership before publishing shutdown quarantine

* no-mistakes(document): Correct worker shutdown comment on replacement-owned lock

* fix(bin): keep an ousted remote job worker off the replacement quarantine

Shutdown can lose the lock after the first ownership check and before it
writes or clears quarantine. Bind both operations to the directory object
this process still owns so a replacement's quarantine stays untouched.

* no-mistakes(review): Make ousted-worker shutdown test reliably reach quarantine clear

* no-mistakes(document): Reattach worker_shutdown doc comment to its function

* no-mistakes(ci): Fixed the failing check (Behavior portable serial 7) with a test-only change to the stall test in tests/fm-remote-job.test.sh. Product code is unchanged; no other test changed. Cause: after the decoy dies, both workers run the same check-exists, read, delete sequence on the job records. On the CI runner the replacement deleted a record between the ousted worker's check and its read. The ousted worker exited 125, and because the file runs under set -e the unguarded `wait` ended the test with 125. The exit trap then killed the replacement, which produced the "Killed" line. Reproduction: a temporary 0.3 s delay between the check and the read, applied to the ousted worker only, made the committed test fail exactly as in CI (exit 125 and the "Killed" line). The new test passed with the same delay. The delay is reverted, along with a similar debug hook that the timed-out attempt had left in bin/fm-remote-job-worker.sh. Test changes: - The replacement is frozen (and confirmed stopped) before the decoy is killed and resumed only after the ousted worker exits, so only one worker touches the job records at a time. - The ousted worker is stopped only once its quarantine exists and its lane is reaped, which places it inside its stop loop. - Every fixed poll loop is now a wait on a named condition with a 30 s deadline and an explicit failure message. Exit detection also handles zombies. - The exit trap kills and waits for the decoy and both workers on every path. - A non-zero exit from the ousted worker now fails with its exit code and stderr instead of silently ending the file. The test still proves that the resumed ousted worker exits 0 and leaves the replacement's lock, quarantine contents and quarantine inode unchanged. Verification: the full test file passed four times on its own and three times under nice -n 10 with four busy-loop CPU hogs; bin/fm-lint.sh passes. Changes are not committed

* no-mistakes(ci): I fixed the failing check (Behavior portable serial 7) by changing only the stall test in tests/fm-remote-job.test.sh. Product code is unchanged. **What failed:** "an ousted worker in shutdown leaves the replacement quarantine untouched" failed on CI with the ousted worker exiting 125 ("could not stop the active command tree"). **Why:** during shutdown, the worker retries the still-running decoy command group a fixed 100 times, 0.01 s apart, then gives up and exits 125. The test tried to freeze the worker partway through those retries by sending SIGSTOP from outside. On a slow runner the retries ran out before the stop arrived, so the worker had already given up. The invariant is that the test must hold the ousted worker inside that retry loop until the replacement owns the lock. That was the only place the test depended on timing. The other waits already watch for a named state change with a 30 s deadline. **Fix:** - The ousted worker now starts with a small `sleep` wrapper at the front of its PATH, and the SIGSTOP race is gone. - The wrapper only holds a `sleep` called directly by that worker's own process (it checks its parent pid against a hold file) while its quarantine file exists. - The only such `sleep` is the first retry in the shutdown stop loop, so the worker waits there as long as needed. - The wrapper writes a marker when it starts holding. The test waits for that marker, then hands the lock to the replacement, freezes the replacement, and kills the decoy. - The test releases the worker by deleting the hold file. Deleting the whole temp directory also releases it, so a failed run cannot leave the wrapper looping. - A process leak: the test overwrites the job's command-group record with the decoy, so no worker ever stopped the job's real command. `fm-hold-job.sh` and its `sleep 30` stayed running for up to 30 s after the test. The test now records that group before overwriting it and kills it at the end of the test and in the exit cleanup. - The test still asserts the same things: the ousted worker exits 0, and the replacement's lock, quarantine contents and quarantine inode are unchanged. **Verification:** - The full file passed twice on its own, twice under `nice -n 10` with six busy-loop CPU hogs, and twice more after the leak fix. - `pgrep` found no leftover processes afterwards. - With the worker from just before the fix commit (cf45cb6^), the test still fails with "the ousted worker wrote or cleared the replacement quarantine during shutdown", so it still proves the fix. - `bin/fm-lint.sh` passes. - I did not reproduce the CI failure locally. The cause comes from the fixed retry limit and the CI error message. The changes are not committed

* no-mistakes(ci): I changed only the stall test ("an ousted worker in shutdown leaves the replacement quarantine untouched") in tests/fm-remote-job.test.sh. Product code is unchanged, and so is every other test. **Invariant:** the pid written to the job's group record must be a process-group leader whose group dies when that one process is killed. Otherwise the worker's bounded stop loop never sees the group die, gives up, and exits 125 ("could not stop the active command tree") before it reaches the lost-ownership exit. The decoy is the only place in this test that depends on this. **Fix:** - The decoy used to be `set -m; sleep 30 &`. It now starts as `perl -MPOSIX=setsid -e 'setsid() >= 0 or exit 1; exec @argv' sleep 30 &`, which gets its own session and group without shell job control. tests/fm-procevent.test.sh already uses the same idiom. - The test now waits, with the file's usual 30 s deadline and a named failure, until `ps -o pgid=` of the decoy equals its pid before writing it into the group record. This way the worker can never read the record before `setsid` has run. - The existing steps are unchanged: the test kills the decoy, reaps it with `wait` before releasing the hold file, and the exit trap still kills and reaps the decoy and both workers. - The assertions are unchanged: the ousted worker exits 0, and the replacement's lock pid, quarantine text and quarantine inode stay the same. **Cleanup:** I reverted a debug `printf` hook that the timed-out previous attempt had left in bin/fm-remote-job-worker.sh, and deleted its untracked `.tmp-repro/` directory. Neither was committed. **Verification:** - The full tests/fm-remote-job.test.sh passed twice normally and once under `setsid -w` with stdin from /dev/null (no controlling terminal). - `bin/fm-lint.sh` passes. - No leftover `sleep 30` processes afterwards. **Not reproduced:** I could not reproduce the CI failure locally. On this host `set -m` made the decoy its own group leader even without a controlling terminal, so the cause on the runner is not confirmed. The change removes the test's reliance on shell job control, as the user asked. Changes are not committed

* no-mistakes(ci): I changed only the stall test ("an ousted worker in shutdown leaves the replacement quarantine untouched") in tests/fm-remote-job.test.sh. Product code is unchanged, and so is every other test. **Invariant:** the group record the ousted worker checks in its stop loop must stay the job's own command group, and the test must stop that group before it releases the hold. Otherwise the bounded retry keeps seeing a live group, gives up, and exits 125 ("could not stop the active command tree") before it reaches the lost-ownership exit. The test overwrote this record in one place (the decoy) and stopped the group in one place (killing the decoy); both are changed. **Fix:** - I removed the setsid decoy and the overwrite of `.claim/group`. The record keeps the job's real command group, which the test still saves as `STALL_JOB_GROUP`. - The two-line `group_start` stays. It is still needed: without it the worker kills the real group on its first pass, before the replacement takes over, so the hold would never matter. - The `sleep` wrapper that holds the worker at its first stop-loop retry is unchanged. - After the replacement owns the lock, its quarantine is planted and it is frozen, the test runs `kill -KILL -- -$STALL_JOB_GROUP`. It then waits, with the file's usual 30 s deadline and a named failure, until `kill -0` on the group fails. Only then does it remove the hold file. The worker therefore always sees its own command already stopped and never races its retry budget. - The exit trap still kills the saved command group if the test fails. It can't `wait` on that group because the group is not a child of the test shell. The decoy variable and its cleanup entry are gone. - The assertions are unchanged: the ousted worker exits 0, and the replacement's lock pid, quarantine text and quarantine inode stay the same. **Verification:** - The full tests/fm-remote-job.test.sh passed twice normally. - It passed once under `setsid -w` with stdin from /dev/null (no controlling terminal). - It passed once under `nice -n 10` with six busy-loop CPU hogs. - With the worker from before the fix (cf45cb6^), the test still fails with "the ousted worker wrote or cleared the replacement quarantine during shutdown", so it still proves the fix. - No `fm-hold-job` or `sleep 30` processes were left afterwards. - `bin/fm-lint.sh` passes. **Not reproduced:** I couldn't reproduce the CI failure locally; the decoy version also passed on this host. So I can't confirm why the decoy group stayed alive on the runner. The new wait turns any leftover live group into a clear named failure instead of an exit 125. The changes are not committed

* fix(bin): keep a dead command group dead on bash 5.2

A bare return inside the liveness check drops the failing kill status when the check runs in a conditional, so shutdown keeps treating a stopped group as alive and exits 125.

* no-mistakes(review): Use bash 3.2 fd syntax and fix trap return comments
…henguid#5589)

* docs: make configuration settings easier to find and understand

* no-mistakes(review): Restore dropped qualifiers and fix misplaced config doc labels

* no-mistakes(review): Restore three dropped qualifiers in configuration reference
)

* fix: bound worker edits of project AGENTS.md/CLAUDE.md to factual corrections

These files are loaded into every agent session of a project, so additions
should be a deliberate human choice rather than automated task output. The
ship brief's project-memory section and AGENTS.md section 6 previously invited
workers to record durable knowledge, which let project AGENTS.md files accrete
detail the codebase or README already carries. Workers now edit only to fix
factually wrong content - including content their own change made wrong - and
fm-ensure-agents-md.sh runs only alongside such a correction. Stow no longer
routes project-memory additions through ship tasks, and the generated skeleton
no longer invites discovery-driven additions.

* no-mistakes(review): Stop running fm-ensure-agents-md.sh on memory-file corrections

* no-mistakes(document): Clarify manual project-memory initialization and remove duplicate guidance
…enguid#5635)

* fix(bin): let gate agents drive lifecycle against marked lab homes

Part 2 of the kunchenguid#5615 split. A no-mistakes gate agent runs inside a
checkout carrying the fleet-captain identity, so fm-gate-refuse-lib
refuses fleet mutation on the gate signal. That refusal was absolute,
which kept gate validation from ever exercising the real lifecycle.

Stamp a disposable lab FM_HOME with a .fm-lab-home marker file that
only bin/fm-lab-home.sh writes, and only onto a fresh empty dir, so
no call path can mark a populated real home. fm_refuse_if_gate_agent
then permits lifecycle only when FM_HOME carries the marker and is
driven through its stock layout - any FM_*_OVERRIDE relocation stays
refused so part of the "lab" cannot be split back onto the real fleet.
The threat model is a confused agent touching the real fleet, not
deliberate forgery, so the marker is a plain token file rather than a
bound record. FM_GATE_REFUSE_BYPASS is unchanged: it still serves the
test harness, which cannot mark hundreds of temp homes.

Teardown's slot-ownership scan compared state-dir paths textually
while fm_firstmate_root_home canonicalizes, so a lab home under a
symlinked TMPDIR scanned its own record twice and self-collided;
compare file identity (-ef) instead.

* no-mistakes(review): Refuse unlistable lab homes and hardlinked slot records

* no-mistakes(review): Mint lab markers only on verified-empty fresh dirs

* no-mistakes(document): Clarify lab-home gate documentation and comment contracts

* no-mistakes(document): Clarify lab-home gate documentation and remove stale claims

* no-mistakes(document): Clarify gate lab-home documentation and boundary wording
…ad of refusing every re-arm (kunchenguid#5594)

* fix(bin): replace a watcher whose beacon stalls past a hard bound instead of refusing every re-arm

A fleet watcher that is alive but whose liveness beacon has gone stale could
never be replaced: every re-arm was refused because the lock holder was a live
pid, and the holder was never evicted because it was not dead. Add
FM_WATCHER_STALL_BOUND (default 3x the stale grace): below it the refusal is
unchanged; at or past it the arm re-verifies the holder against the lock's
recorded identity, sends TERM, waits boundedly, and takes the lock the normal
way, ledgering a stalled-holder-replaced row. A holder that survives TERM keeps
the old refusal.

Fixes kunchenguid#4400

* no-mistakes(test): poll for replacement message to fix watcher-lock test flake

* no-mistakes(document): document FM_WATCHER_STALL_BOUND in config inventory
…n can keep them (kunchenguid#5563)

* fix(pi): hide queued Firstmate notifications under Calm only when the session can keep them

Calm now keeps authenticated Firstmate operational inputs out of Pi's queued-message
listing, but only after proving the live session exposes every member needed to keep
them across Escape. A session missing any of them keeps stock rows and Escape and shows
one generic warning. Escape and the dequeue key return only captain-authored messages to
the editor and re-queue hidden notifications in order; after an abort that kept any in
Pi's agent queue, the adapter starts the delivery turn itself because Pi 0.87.1 does not
continue an aborted run. Compaction-held notifications stay with Pi's compaction flush and
never start or announce a turn.

Fixes kunchenguid#1588

* docs(calm): record Pi 0.87.1 queued-row retention verification

* no-mistakes(review): Deliver kept Calm notifications after tree-navigation aborts too

* no-mistakes(review): Defer Calm notification turn until tree navigation finishes

* no-mistakes(lint): Silence SC2016 for literal JavaScript in queue-retention e2e test
…uid#5548)

* fix(bin): refuse teardown when a required source disappears

A missing sibling was sourced after cleanup had started, so Bash 3.2
exited 0 from the EXIT trap and Bash 5 continued and reported success.

* no-mistakes(review): Remove unused FM_TEST_ONLY hook from teardown tests

* no-mistakes(review): Check task backend sources before any teardown cleanup

* test(gotmp): give teardown fixtures every tmux adapter sibling

Teardown now refuses when a sibling the recorded backend's adapter sources
is missing, so the fake bin must carry fm-session-lock-lib.sh,
fm-agent-process-lib.sh and fm-gemini-lib.sh.
Restructure the supervision host doc's prose into shorter sections, lists,
and tables without changing documented behavior. Every original heading,
anchor, identifier, number, quoted string, and link target is preserved.
* docs: make herdr-backend easier to read

Restructure the Herdr backend doc's prose into shorter sections, lists, numbered procedures, and tables without changing documented behavior. Every original heading, anchor, fenced code block, link target, and documented fact is kept.

* no-mistakes(document): Restore composer-proof reason and complete Herdr topic table
Restructure the prose into sections, lists, and tables without changing
documented behavior. Every original heading and anchor, inline-code span,
link target, number, and quoted string is kept, and each sentence sits on
its own line. Adds a topic navigation table and short subsections under
the existing headings.
* docs: make watcher-continuity easier to read

Restructure the prose into sections, lists, and tables without changing documented behavior. Every original heading, anchor, identifier, link target, and number is kept.

* no-mistakes(review): Fix actor and supervision-host scope in watcher-continuity doc

* no-mistakes(review): Make readiness TERM and retry conditional on unready successor
* docs: make sessionstart-nudge easier to read

Restructure the prose into sections, lists, and tables without changing documented behavior. Every original heading, inline-code span, link target, number, and fact is preserved, and a harness-to-tier table now sits near the top.

* no-mistakes(review): Drop helm glossary line and dedupe exit-code lead-in
* docs: make captain-hold-lifecycle easier to read

Restructure the captain-hold lifecycle prose into sections, lists, and tables without changing documented behavior. Every original heading, anchor, identifier, number, quoted string, and link target is kept.

* no-mistakes(review): Fix verification record subjects and grouping headings

* no-mistakes(review): Clarify task-body read-back cases belong to the suite
* docs: make remote-secondmates easier to read

Restructure the remote second mates prose into sections, lists, numbered procedures, and tables without changing documented behavior. Every original heading, anchor, fenced code block, identifier, link target, and qualifier is preserved.

* no-mistakes(review): Merge remote-home table cell into one sentence

* no-mistakes(review): Tighten readiness lead-in, restore causal link, fix dangling reference
…nguid#5554)

* fix(bin): bound the away digest and log why a delivery failed

The away daemon joined every buffered escalation into one unbounded
digest. A start-up catch-all span can exceed what one transport argument
carries (tmux rejects the send-keys command; Linux refuses to exec any
argument above 131,071 bytes, which is how herdr receives it), so the
initial send failed on every housekeeping pass and was logged as an
unconfirmed Enter with text possibly in the composer.

escalate_flush now builds the injected digest under a fixed byte budget:
each event is cut at a UTF-8 boundary with an omitted-bytes marker, the
joined events stop with a "+K more event(s)" tail, and a bounded digest
names a state/.subsuper-digests/ file that keeps every buffered event
verbatim. The buffer itself is untouched, so the return catch-up stays
complete.

The tmux submit core and the herdr literal send now replay the
transport's stderr on failure, and inject_msg logs the failing stage
(initial send versus Enter confirmation) with the byte count and that
stderr. The wedge alarm line and marker carry the last failure reason.

Fixes kunchenguid#4382

* no-mistakes(review): Drop digest pruning; label send-failed as send-or-Enter stage

* no-mistakes(review): Keep digest full text once submit ran; reuse on retry

* no-mistakes(lint): Count digest files with find instead of ls

---------

Co-authored-by: firstmate-oss <firstmate@kunchenguid.local>
…henguid#5638)

* feat(tests): add FM_TEST_SEAM launch seam and gate lab-primary recipe

Part 1 of the kunchenguid#5615 split: the pieces that let the no-mistakes pipeline
live-validate firstmate changes, without the gate-refusal rescoping.

- bin/fm-afk-launch.sh: FM_TEST_HARNESS pins the detected harness only
  alongside the FM_TEST_SEAM=1 marker test suites set, so a leaked variable
  in a real primary's environment stays inert and unknown tokens fall
  through to real detection.
- tests/lib.sh: export FM_TEST_SEAM=1 for every suite.
- .no-mistakes.yaml: per-harness recipe for running a real fixture primary
  from a gate run - a plain mktemp lab FM_HOME on a private tmux socket,
  with FM_GATE_REFUSE_BYPASS=1 scoped to it and NO_MISTAKES_GATE scrubbed.
- tests/fm-wake-queue.test.sh: stop the owned watcher fixture with KILL and
  clear its lifecycle state so the next leg starts clean; TERM could leave
  bash waiting in a child on some runners.
- tests/fm-remote-secondmate-lifecycle-e2e.test.sh: wait for the liveness
  lock holder's post-acquire marker instead of the lock dir, which is
  published before the claim finishes.

* no-mistakes(review): Scrub lab home overrides and require FM_TEST_SEAM separately

* no-mistakes(document): Clarify test seam and disposable lab bypass documentation

* no-mistakes(document): Clarify lab isolation and test-seam documentation

* no-mistakes(ci): Fixed the CI failure: test cleanup killed the remote worker child but left its supervisor able to restart it during fixture removal. Cleanup now stops the worker tree. The lifecycle test passed locally; ShellCheck and diff checks passed
… home is gone (kunchenguid#5552)

* fix(bin): refuse watchers from disposable checkouts and exit when the home is gone

Fixes kunchenguid#321
Fixes kunchenguid#4760

A watcher armed from a disposable no-mistakes validation checkout under
.no-mistakes/worktrees/ outlived the validation step and kept writing the
real home's state, and a running watcher never noticed when its home,
state directory, or code root disappeared. The arm now refuses from such
a checkout with the typed failure line, the watcher checks once per poll
that its home, state directory (or its own lock holder record), and bin
directory still exist and exits with a logged reason scoped to itself,
and the shared test helpers reap every watcher a suite armed for a
temporary home through the home-scoped stop.

* no-mistakes(lint): fix SC1007 by assigning empty string in watch-arm test

* no-mistakes(ci): Found and fixed a genuine, reproducible hang introduced by this branch's test-watcher reaper, which is what killed both CI checks (serial-2 cancelled at the 30-min cap; Lint 2 exit 143 = the suite's own TERM-trap code). Root cause: test_drain_asserts_watcher_liveness (tests/fm-wake-queue.test.sh) fabricates a .watch.lock whose pid is the test runner's own $$ with the runner's real identity, to make the drain believe a live watcher exists. The new make_case tracking registers that state dir for reaping, so at fm_test_cleanup the new fm_test_reap_watchers drives fm-watch-arm.sh --stop; its identity check matches (the fixture recorded the runner's identity) and it kill -TERMs the test runner. tests/lib.sh:231 is `trap 'fm_test_cleanup; exit 143' TERM`, so the TERM re-enters cleanup -> reap -> kills $$ again -> infinite loop until the runner cap. I reproduced this locally: the suite ran all tests then looped forever in cleanup spawning fm-watch-arm.sh --stop against a lock naming its own PID. Fix (tests/lib.sh, +5 lines): in fm_test_reap_watchers, skip any tracked lock whose pid equals our own $$ before driving --stop. This is the single shared reap boundary; seven $$-self-lock fixtures across four test files are all covered by the one guard, and real armed watchers (pid != $$) are still reaped. Invariant: the test reaper must only signal real armed watcher processes, never the test runner itself. Verified locally: tests/fm-wake-queue.test.sh -> EXIT 0 (63 ok, no hang); tests/fm-watch-arm.test.sh -> EXIT 0 (21 ok, including test_reaper_stops_a_tracked_watcher, confirming the guard does not over-skip). Lint 2's exit 143 was the same shard/cap signature; a fresh CI run on this new commit will re-evaluate it

---------

Co-authored-by: firstmate-oss <firstmate@kunchenguid.local>
…m as silence (kunchenguid#5588)

* fix(bin): surface an unrecognized status prefix instead of dropping it

A parked or holding declaration, and a verb whose correlation token did not parse, never became an event, so the supervisor still saw the earlier line.

* no-mistakes(review): Require verb-shaped unrecognized status prefixes, add continuation tests

* no-mistakes(document): Document unrecognized status prefix escalation in afk skill

* no-mistakes(ci): I reproduced the "Behavior portable serial 6" failure locally and fixed it by changing the test data in one test. No product code changed. **What failed:** `tests/fm-session-start.test.sh`, in `test_orphan_status_logs_are_printed`, with "matched status log was printed 2 times". **Why:** the test writes status lines with made-up prefixes, `matched: surfaced once` and `orphan: step N`. The test only uses them as placeholder text. It checks that the session-start digest prints each task's status tail exactly once. This PR (kunchenguid#4763) deliberately makes an unrecognized one-word lowercase prefix a status event. So those lines now surface as captain-relevant events, and the wake queue's STATUS OUTCOME BACKSTOP section prints them a second time. The code under review is behaving as the issue asks. Only the test's placeholder data had become meaningful. **Rule the test depends on:** its status lines must not be captain-relevant, so the digest is the only place they are printed. Both lines in this test broke that rule. The orphan line would have failed the same count check right after the matched line did. **Fix:** in that test only, I switched both lines to the recognized, non-captain verb `working:`: `working: surfaced once` and `working: orphan step 1..6`. I updated the matching assertions and counts to use the new text. What the test checks is unchanged: orphan logs are labelled, the tail is bounded, the log path is printed, and each tail appears once. **Verification:** before the fix, the test failed locally the same way as in CI. After it, `bash tests/fm-session-start.test.sh` reports "all assertions passed

* no-mistakes(review): Detect unrecognized prefixes on unstamped lines; share verb list

---------

Co-authored-by: Kun's firstmate <kunchenguid+firstmate@users.noreply.github.com>
…henguid#5658)

Fixes kunchenguid#5295

Session start now reports a remote inheritance failure using the
push's own error line instead of the first unchanged item that
happened to print before it, and the shared captain preferences
header check now names the first required phrase it did not find,
on both the local and remote inheritance paths.
…yloads (kunchenguid#5657)

* fix(bin): stand down the Claude Stop auto-arm on pi-code-delivered payloads

pi-code loads the tracked Claude settings but has no asyncRewake, so it
awaits every Stop hook; without a stand-down the auto-arm runs
synchronously inside Pi's turn end and holds it open for the declared
multi-hour timeout. Stand down when the payload's transcript_path
contains a /.pi/ path component, the same discriminator the closed-but-
unmerged fix in kunchenguid#3352 used, with an explicit string-type check on the
jq filter.

Fixes kunchenguid#3343

* no-mistakes(document): document pi-code stand-down in harness integrations reference
…kunchenguid#5659)

* fix(bin): match whole multi-word project names in the registry lookup

bin/fm-project-mode.sh matched a registered project name against only the
first whitespace-delimited token of a registry row, so a name containing a
space never matched, silently defaulting the project to no-mistakes off
instead of its declared posture.

The lookup now matches the whole registered name against the raw line text,
so a name is compared literally (never as a regex) and a name that is a
leading prefix of another registered name still resolves to its own row.

* no-mistakes(document): docs already accurate for multiword registry name match

* chore: drop accidental empty err file

Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
…d#5546)

* fix(bin): classify the stdin program of `bash -s` with operands in the arm policy

With -s, sh/bash/zsh read the program from stdin even when operands follow;
the operands are only positional parameters. The arm policy treated the first
operand as a script path, so heredoc and here-string payloads were never
classified and a hidden bin/fm-watch.sh execution was allowed.

A protected path in the operand position still fails closed as before.

Fixes kunchenguid#1489

* no-mistakes(document): Clarify stdin shell operand documentation

* no-mistakes(ci): Captain, fixed `shellInvocation` so `bash -- -s` treats `-s` as a script name, and updated R21 to test the exact command. The targeted policy suite, lint, documentation check, and diff check pass. Both hosted workflows show `action_required` before any jobs ran; that external approval state remains unresolved

* fix(bin): keep main's handling of words after a leading `--`

Revert the pipeline CI-step change that made the first word after a leading
`--` always a script. It turned forms that main denies today into allow
(for example `bash -- -c 'bin/fm-watch.sh'`), which is outside kunchenguid#1489 and
loosens a fail-closed policy. `--` after `-s` still ends option parsing.
…nchenguid#5695)

* fix(bin): strip AI co-author trailers from fleet-launched commits

Cursor and other non-Claude runtimes append the trailer after the typed
message. A per-task commit-msg hook removes it and leaves human co-authors
and the author identity untouched.

* no-mistakes(review): Export pane hooksPath override and drop generated-with stripping

* no-mistakes(ci): This PR caused all three CI failures, and the fix is test-only: 4 test files change, no product code. **Cause.** `fm-spawn.sh` now installs the AI-trailer strip hooks for every spawn, secondmates included. The installer refuses a worktree that is not a git repository, and the PR deliberately keeps that fail-closed rule because real secondmate homes are firstmate clones. Four test fixtures still gave secondmates a plain directory as their home, so each spawn failed with "not a git worktree ... could not install the AI-trailer strip hooks": - serial 5: `tests/fm-backlog-atomicity.test.sh` ("secondmate spawn failed"). - serial 8: `tests/fm-secondmate-harness.test.sh` ("split: no meta written"). - Herdr: `tests/fm-backend-herdr-launcher-workspace-e2e.test.sh` and `tests/fm-backend-herdr-workspace-per-home-e2e.test.sh`. **Rule that must hold.** Every home a test spawns as a secondmate must be a git worktree. I checked the other places in the changed area: the only secondmate spawns in these tests are the ones listed. The earlier rounds already fixed the other fixtures (`fm-secondmate-liveness`, `fm-secondmate-safety`) the same way. **Fix.** - Each of those four secondmate homes now gets the same `.gitignore` plus `git init -q -b main` that the liveness and safety tests already use. - The two Herdr tests clean up with their own plain `rm -rf "$TMP_ROOT"`, not the shared `tests/lib.sh` helper. Because the installer leaves each `state/<id>.git-hooks` directory read-only, that cleanup printed "Permission denied" and left the directories behind. Both cleanups now restore the owner's write bit on every directory before removing (`find ... -exec chmod u+rwx`), which is what `fm_test_remove_tree` in `tests/lib.sh` does. **Verification.** - `tests/fm-secondmate-harness.test.sh` passes. - `tests/fm-backlog-atomicity.test.sh` passes (99 ok, exit 0). - shellcheck is clean on all four files. - I could not run the two real-Herdr tests locally: the Herdr lab on this host refuses to start because it needs exactly one running default session, and I did not change the host's Herdr state to get around that. Instead I checked their two changed steps directly: the installer succeeds on a home set up the new way, and the new cleanup removes the read-only hooks directory completely. Those two tests will only be proven on CI
…id#5683)

* fix(bin): treat Pi's dollar-first cost footer as furniture

An idle Pi status row opening with $0.000 was read as a dead-shell prompt, so exit and relaunch refused on an empty composer.

* test: wait for the draining holder to exec sleep before reading its identity

The procevent drain fixture read fm_pid_identity immediately after
backgrounding setsid sleep, racing the child's exec chain. Mid-exec the
cmdline can read empty, failing the fixture on a loaded CI runner.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…5534)

* fix(bin): refuse a merge when a required check never reported

fm-pr-merge.sh built its GitHub refusals only from checks present in
statusCheckRollup, so a required check that never ran was simply absent and
the merge proceeded on the subset that reported, contradicting its own
"every required check green" claim.

The GitHub verify now reads the base branch's required contexts from the
forge itself - the classic branch protection summary on
GET repos/{o}/{r}/branches/{b} and the active ruleset rules on
GET repos/{o}/{r}/rules/branches/{b} - and refuses when a required context
has no entry in the same rollup, at the same head, that the merge is bound
to. Absence reads as unknown, never green. The required-set read joins the
existing refusal list, so a draft, a red check, and an unreported required
check are all reported together.

Could not read vs nothing required: both endpoints need only repository
read access. The admin-only GET .../branches/{b}/protection endpoint is
deliberately not used: it answers a non-admin token with the same 404 an
unprotected branch gets (observed live on kunchenguid/firstmate main with
this token), which would read a missing permission as "nothing required".
Any failed or malformed read of either source (auth, missing fine-grained
permission, rate limit, network, 404, unexpected shape) refuses the merge
with a line naming the unreadable source. The one exception is GitHub's
plan-gated 403 on the rules endpoint ("Upgrade to GitHub Pro or make this
repository public"), which already means "this repository has no branch
rules" for the merge-queue reader; that check moves into one shared helper
and the classic summary still decides for such a repository.

Attended waiver: --allow-missing <check-name> is the twin of --allow-red and
follows the same design and recording path: once, separate name argument,
waives only that exact unreported required check, still requires every
other required check reported and every check green, never waives an
unreadable required set, refused while the away-posture record exists, and
refused on GitLab. Merge-state BLOCKED policy is unchanged.

How this differs from the withdrawn kunchenguid#5353 (read from its diff):
- kunchenguid#5353 read the admin-only branches/{b}/protection endpoint and treated
  its 404 as "no required checks", so for any non-admin token the required
  set silently read as empty; this change reads the read-access branch
  summary and treats every failure as unreadable.
- kunchenguid#5353 ignored rulesets; this change also reads required_status_checks
  rules from the effective branch rules.
- kunchenguid#5353 made separate per-head REST reads of statuses and check-runs capped
  at per_page=100 with no pagination; this change checks presence in the
  same statusCheckRollup view the red-check gate already reads at the
  verified head.
- kunchenguid#5353 stopped at the first unreadable read; this change reports it as one
  refusal among all the others.
- kunchenguid#5353 also claimed kunchenguid#5345 (lock stealing) and changed 39 files, most
  unrelated; this change is kunchenguid#5344 only.

Live proof, read-only (a gh wrapper refused every merge and mutating call):
- cli/cli#14474 (trunk requires 3 classic build contexts, none ran):
  refused, naming build (macos-latest), build (ubuntu-latest),
  build (windows-latest); with --allow-missing "build (macos-latest)" it
  still refused, naming the other two.
- cli/cli#13665 (ran build (ubuntu-24.04-firewall) instead): refused,
  naming build (ubuntu-latest).
- hashicorp/terraform#39262 (ruleset-required checks absent): refused,
  naming Code Consistency Checks, End-to-end Tests, Race Tests, Unit Tests.
- cli/cli#14485 (all required reported and green): verified; the wrapper
  blocked the merge call and the pull request read back open.

Fixes kunchenguid#5344

* fix(review): Preserve required-check producers and aggregate independent read failures

* fix(document): Clarify required-check verification and waiver documentation

* fix(bin): match an app-bound required commit status by name

The producer-identity check resolved an app-bound required context only
against check runs, so a required context that the required app reports as
a commit status could never match and always read as "has not reported".
A commit status carries no app id to compare, so an app-bound requirement
that arrives as a status now matches by name, as before producer binding;
check runs keep requiring the configured producer app.

Live, read-only: hashicorp/terraform#39262 requires license/cla from
integration 865473, reported green as a commit status by the CLA app. The
previous head refused it as unreported; this head no longer does, while
still naming the four required check runs that never ran there.

Refs kunchenguid#5344

* fix(document): Clarify accepted commit-status producer verification limitation

---------

Co-authored-by: firstmate-oss <firstmate-oss@kunchenguid.local>
…uid#5696)

* fix(bin): never offer a persistent secondmate for teardown

The return brief's "Landed, cleanup due" scan listed every state/*.meta
record carrying a pr= and a merge-notified marker without regard to kind, so
a secondmate record holding a relayed child's merged PR put the mate itself
up for "bin/fm-teardown.sh <mate>" cleanup. A secondmate is a persistent
worker, never landed work.

- bin/fm-afk-return.sh: skip kind=secondmate in the landed-cleanup scan.
- bin/fm-pr-check.sh: refuse to record pr= or arm a merge watch on a
  kind=secondmate record before any side effect; a PR reported on its routed
  status channel belongs to a task in the mate's own home, which arms its
  own watch.
- bin/fm-watch.sh: a merged result from a poll already armed on a secondmate
  retires the poll silently - no merge outcome, marker, or wake.

* no-mistakes(document): Document secondmate merge-watch and return-brief exclusions

* no-mistakes(ci): CI failed in an unchanged watcher-shutdown test whose three-second wait was sensitive to runner load. Increased the wait for both state- and home-deletion cases without changing watcher behavior. The full fm-watch-arm suite passed locally; syntax and diff checks passed
…unchenguid#5702)

* fix(bin): refuse unknown dash-leading args in public-posting fm-x scripts

fm-x-reply.sh collected any unrecognized argument into the positional
pool and took the first one as the reply text, so an invocation like
"fm-x-reply.sh <id> --followup --final <text>" posted the literal
string "--final" to X and silently dropped the real text.

Make argument parsing strict in every script that can post publicly:
an unknown dash-leading argument, a dash-leading request_id/task id, a
dash-leading option value, or a surplus positional now exits 2 with a
usage error before any config load, outbox write, or network call. Reply
text starting with '-' is still accepted via --text-file or stdin, and
--help is honored wherever it appears instead of becoming text (a --help
forwarded through fm-x-followup.sh would have counted as a posted
follow-up and mutated the link).

fm-x-link.sh and the fm-public-followup scripts already refuse unknown
arguments; fm-x-poll.sh takes none.

* no-mistakes(review): Refuse surplus follow-up text sources; drop post-ID help branches

* no-mistakes(document): Clarify reply and follow-up argument usage

* no-mistakes(review): Refuse dash-leading --text-file operands in fm-x-reply

* no-mistakes(document): Correct follow-up argument parsing comment

* no-mistakes(document): Document dismiss argument rejection in script header
tiago-peixoto and others added 29 commits September 26, 2026 15:24
…nchenguid#4806)

* fix: stop quarantining ordinary shared-captain source updates

* no-mistakes(document): Rewrap remote inherit header so usage prints fully
* docs: make calm easier to read

Restructure the Calm mode prose into sections, lists, and tables without changing documented behavior. Every original heading, anchor, fenced code block, inline-code span, link target, number, and quoted string is preserved.

* docs: restore reload case in calm override lead-in

The Built-in tool override collisions lead-in covers a session that reloads with Calm already on, as the original text did.
* docs: make turnend-guard easier to read

Restructure the turn-end guard doc's prose into shorter sections, lists, and tables without changing documented behavior. Every original heading, anchor, inline identifier, link target, and number is kept.

* no-mistakes(review): Restore legacy-only scope on TERM retirement sentence

* no-mistakes(review): Name Cursor park behavior in live e2e test line
…henguid#5872)

* docs: move situational AGENTS.md sections into on-demand skills

Backpass memory optimization: shrink the always-loaded AGENTS.md by moving
situational contracts (home layout, session-start recovery, validation and
landing supervision, scout completion, away/quiet supervision, Relay
ownership) into agent-only skills loaded at their triggers, with a trigger
index skill.

* docs: classify the new on-demand skills' documentation audience

Register the seven new agent-only skills as agent-runtime docs and fix a
link in validation-supervision that kept its AGENTS.md-relative path.

* docs: close load-timing gaps found by the live regression check

- load validation-supervision whenever an ask-user finding is decided or
  answered, so forbid --yes and process-every-return reach the worker
- keep the mid-task captain-ask rule, the unconfirmed network-checks rule,
  and the worker account pin rule inline in AGENTS.md
- fix cross-references that still pointed at moved AGENTS.md sections
…guid#5879)

* fix: route second-mate signal wakes by their new status span

A second mate's status log is a shared channel carrying many independently
keyed decisions, so judging its signal rows by every decision still open in
the whole log pinned each routine update to main behind any unrelated
parked hold. scopeForUnreadWake (the one owner for Pi and the attended
supervision host) now judges a second-mate signal row by the lines presented
since the last drain, bounded by the existing status-presentation cursor:
a decision, blocked, resolution, or captain-held line, or a line declaring
the key of a still-open decision, keeps the whole row on main, and any
cursor problem falls back to the whole log. Keys are read only at the
status parser's declared positions, with readable time stamps stripped as
bin/fm-classify-lib.sh does. Single-task crewmate and stale routing are
unchanged, and stale and signal rows for one mate keep independent verdicts.

The supervision branch now treats a second mate's done and merged lines as
relayed child outcomes, and fm-teardown refuses the branch actor second-mate
retirement through the existing role-partition helper in both postures.

* no-mistakes(review): Route second-mate resolutions to main only when closing open decision

* no-mistakes(review): Guard bare-verb fold lines and fold resolution spans incrementally

* no-mistakes(document): Clarify second-mate wake routing and retirement documentation

* no-mistakes(ci): The CI failure was a timing-sensitive watcher teardown test, not the PR’s signal-scope code. Extended the bounded wait for both state-directory and home removal. The watcher suite passes locally, and git diff --check is clean
…extension (kunchenguid#5882)

register-extension took the extension lifecycle lock and then the source
lock, while reconcile republishing an unhandled extension result holds the
source lock and reaches the lifecycle lock through the extension host's
process-event path. Both waits are unbounded and both owners stay alive, so
the two could wait on each other forever and freeze the home's monitoring
cycle.

register-extension now takes the source lock first, matching every other
path that holds both. The lifecycle lock still spans binding resolution
through registration publication, so binding retirement stays serialized.

A new lifecycle-order section in the extension-binding suite, run in the
default aggregate, holds a re-registration inside binding resolution while
reconcile republishes that source's unhandled result and requires both to
finish within a bound.

Fixes kunchenguid#5866
* fix(bin): run the repository's own hooks when git -c carries the per-task hooksPath

The per-task hooks wrappers cleared only the GIT_CONFIG_COUNT override before
looking up the repository's own hooks directory. When core.hooksPath reached git
through git -c (GIT_CONFIG_PARAMETERS), directly or inherited by a child
process, the lookup found the wrapper directory again and exited 0, so the
repository's real hook - such as a pre-push publish guard - never ran and the
push succeeded.

The lookup now ignores GIT_CONFIG_PARAMETERS too, so only the repository's
config files decide its hooks directory, and a failed lookup exits nonzero
instead of skipping the hook. AI-trailer stripping is unchanged.

Fixes kunchenguid#5871

* no-mistakes(document): Document git hook chaining and lookup failure behavior
…unchenguid#5744)

* feat(bin): add an optional never-send list to typed dispatch resolution

* Added config/dispatch-never-send, an optional local list of literal
  values and re: regular expressions checked against every string of
  the resolver request before it is sent to typesafe.ai
* A match, an unreadable list, or an empty or invalid pattern now stops
  the request and falls back to the off path, so firstmate dispatches
  through its existing intake; the one stderr diagnostic names at most
  the list line number and never the value
* No list, or a list with no match, leaves resolution unchanged

* no-mistakes(review): Match never-send literals across whitespace, drop regex mode

* no-mistakes(review): Inherit the never-send list into secondmate homes

* no-mistakes(ci): ci-2 (Lint 2), caused by this PR and now fixed. The rule that broke: the test script must not share shell variables with a library it sources. The new secondmate-inheritance test in tests/fm-dispatch-resolve.test.sh sourced bin/fm-config-inherit-lib.sh inside a `( ... )` subshell. That lib assigns `out`, so ShellCheck flagged every later `$out` in the test with SC2031. The subshell was the only place this PR sources that lib. The fix runs the propagation in a child shell instead (`bash -c '. "$1" && propagate_inheritable_config "$2" "$3"' _ lib from to`), so the test shell never sources the lib. `bin/fm-lint.sh tests/fm-dispatch-resolve.test.sh` now exits 0, and `bash tests/fm-dispatch-resolve.test.sh` passes. That includes the check that an inherited list is enforced in a secondmate home. ci-1 (Behavior portable serial 8), not caused by this PR, so no code change for it. The only failing test is tests/fm-remote-secondmate-relaunch.test.sh, which fails with "not ok - could not arm the PR poll fixture for the relaunch-ordering test". This PR doesn't touch that test or the code it runs. I reproduced the same failure locally on the base commit ea7c7f7. Main's own CI run on ea7c7f7 (run 36212602588) fails only this job, with the same message. The recent main runs before it also concluded failure
…hrashing guard (kunchenguid#5903)

* feat(jev): add the guard framework and the memory RSS/swap thrashing guard

A Jev guard is a bounded read-only host diagnostic that turns one class of
resource pressure into a machine-readable audit record and a one-line verdict.
This lands the framework contract (docs/jev-guards.md) with one representative
family (Pattern 46, memory RSS and swap thrashing) as the pilot: thin bash
wrapper, stdlib-only python engine, and a behavioral test through the CLI.

* no-mistakes(review): fix jev mem guard fail-open unknown and contract

* no-mistakes(review): fix mem guard rss pairing, memfree, tests, docs

* no-mistakes(review): fix inverted UNKNOWN branch in fail-forcing test

* no-mistakes(review): register docs/jev-guards.md in audience inventory

* no-mistakes(review): simplify wrapper, drop dead top_n, single-source thresholds

* no-mistakes(review): assert exact exit code in fail-forcing test leg

* no-mistakes(review): tolerate any stdout encoding in text output

* no-mistakes(document): Fix guard contract dash style and output wording
…enguid#5900)

* fix(bin): stop slow GitHub reads from starving and waking the contributions poll

The poll's fixed budget cut the tail URL's reads short, and any read killed at the bound was recorded as a forge failure, so routine GitHub slowness produced an observation-unavailable wake. A configured FM_CONTRIBUTIONS_BUDGET now rides the generated check shim and is cut down to the watcher's per-check bound with a margin, a read killed at the bound or the deadline is budget refusal that leaves records untouched, and the poll moves to the next URL that still has a full observation reserve instead of ending.

* fix(bin): report the bound when a signal death leaks through fm_run_timed

fm_run_external_timeout trusted the wrapper's recorded status over the runner's own bound verdict, so a read killed by the bound's TERM could surface as 143 and callers such as the contributions poll classified a budget-bound read as a forge failure. When the runner reports the bound, a signal-death status is the bound's own TERM racing the wrapper's bookkeeping and is reported as 124; a natural exit still passes through. The replace-shell probe also moves to the repo's portable BASHPID fallback so the suite runs on the macOS system bash.

* no-mistakes(review): Rotate contribution polling and verify generated budget behavior

* no-mistakes(review): Stabilize contribution rotation across successful observation refreshes

* no-mistakes(review): Exclude settled contributions from live observation rotation

* no-mistakes(document): Document contribution poll rotation and observation reserves

* no-mistakes(ci): Captain, fixed timeout handling with a one-line change preserving natural exit 137. Full session-start and contributions suites, targeted regressions, and lint passed. The timeout suite still fails on the known Bash 3.2 BASHPID issue, left unchanged as instructed. Logs retained in .no-mistakes/ci-evidence/. Remote CI was not rerun

* no-mistakes(ci): Fixed the fixture’s lock-acquisition race with a one-line bounded wait. Forced contention reproduced the CI error before the fix and passed afterward; the ordinary held-lock case, lint, syntax, and diff checks passed. Local Bash 3.2 failures remain: “cleanup lock bound 08 gave up before the marker lock freed” (also reproduced without the fix) and “TERM did not stop a watcher blocked inside a poll”. Remote CI was not rerun
…railers (kunchenguid#5859)

* feat: add keep AI trailers setting

* no-mistakes(document): Drop duplicate keep-ai-trailers line, update Cursor attribution note

* no-mistakes(document): Honor keep-ai-trailers for Devin worker attribution

* no-mistakes(document): Qualify Devin attribution note with keep-ai-trailers flag

* no-mistakes(review): Inherit keep-ai-trailers into secondmate homes
…sh-command popup cannot hide the composer (kunchenguid#5876)

* Fix Herdr composer reads blinded by the slash-command popup

Capture the full visible viewport for every Herdr composer state and content read instead of a bounded tail.
Claude Code renders its slash-command popup between the composer and the pane bottom, which pushes the composer outside a tail window.
The pre-Enter payload proof then read an empty composer, judged a typed command unsent, and cleared it without pressing Enter.
The proof-lines value now bounds only the clear cost, not the capture size.
Growing the window adds rows above the composer only, so bottom-most shape selection and prior verdicts are unchanged.
The suffix refusal is kept, and the shared inbox pending-line read stays a bounded tail.
Portable regressions cover the popup-below-composer layout, and the live submit-confirmation guard gains a third exit scenario.
The runtime-backends verification record documents the Herdr 0.9.0 and Claude Code 2.1.283 run.

Closes kunchenguid#5533

* no-mistakes(document): Clarify composer capture bound ownership

* no-mistakes(review): Drop unrequired bracketed-paste Enter fallback from live guard

* no-mistakes(review): Latch trust prompt, check idle composer arm first
…configurable bound (kunchenguid#5917)

Each per-task endpoint read in the session-start fleet digest runs in its own crash-isolated child under the existing timeout helper with a configurable FM_SESSION_START_ENDPOINT_TIMEOUT bound (default 10s).

A hung or killed read becomes that task's endpoint error while the digest continues, and the wrapper banners any abnormal child exit naming its status.
…nchenguid#5886)

* Keep lab tmux sockets on short private paths

* no-mistakes(review): Fix Linux stat checks and fail-closed lab tmux teardown

* no-mistakes(document): Update lab helper documentation for isolated tmux sockets
* Allow silent task-level no-change outcomes

* no-mistakes(review): Exclude silent outcomes from captain-return handoffs

* fix: look up supervision receipts by exact sequence

* no-mistakes(review): Suppress silent notes in away-return brief

* no-mistakes(review): Clarify visible notes; remove unused mode

* no-mistakes(review): Clarify silent outcomes and avoid false drain promises

* no-mistakes(document): Clarify silent supervision outcome documentation
…chenguid#5928)

* feat: make /quiet a statement where the attended supervision host runs

On a home that opted into the supervision host, quiet mode is what the
attended host already does, so /quiet now enters nothing there instead of
launching the quiet daemon and writing a record that would park a present
captain's main.

- bin/fm-afk-launch.sh quiet-check says quiet mode needs nothing where the
  attended host runs, or that the session is paused while its
  broken-session latch holds; a quiet enter refuses there before writing
  anything.
- Where the home opted in but the attended host lacks a part (engine,
  tools, verified mirror writer, identifiable main session, valid mirror),
  quiet-check names it and quiet mode falls back to the daemon.
- Under a live away record on that home, quiet-check and a quiet enter
  refuse and name the record, so the return runs first, whatever
  state/.afk says.
- A quiet enter records mode: quiet in the posture record, so start and
  start-native launch the quiet daemon without FM_AFK_MODE, and the away
  refusal wording fires only for away.
- bin/fm-host-mirror.sh check validates the dialog mirror read-only and
  exits 1 on a missing, unreadable, or invalid mirror.
- The quiet and afk skills and the supervision-host docs describe the new
  behavior; homes without the opt-in and Pi homes keep the daemon path.

* no-mistakes(review): Archive the quiet record when a quiet daemon start fails

* no-mistakes(document): Clarify quiet-mode documentation and remove stale duplicates
…uid#5884)

* fix(bin): grant Claude workers their task-channel dirs via --add-dir

Since Claude Code 2.1.257, a file-tool read (Read/Glob/Grep, and an
Edit's mandatory prior read) of a path outside the working directories
parks --permission-mode auto panes on a one-time interactive question,
and a "Block" answer lands permissions.blockReadsOutsideWorkingDirectories
in user settings, refusing the same reads even under bypass. Firstmate
launches Claude with no --add-dir, so a secondmate's parent-home steering
inbox and a ship or scout worker's launch record, steering inbox, brief
dir, and code-root .agents/skills were all outside: workers wedged on
the question the first time they read a steer.

Every Claude launch, spawn and relaunch, in both permission modes, now
grants exactly the task's channel directories: state/<id>.inbox for a
secondmate (in the parent home), or state/operational-inbox,
state/<id>.inbox, data/<id>, and the code root's .agents/skills for a
ship or scout. Paths resolve to real paths and lazily created channel
dirs are made before launch so the grant never names a not-yet-existing
directory; the whole state/ is deliberately never granted.

The grant keeps the bypass-mode launch argv changed on purpose: it also
protects bypass workers against a machine-recorded Block answer.

* no-mistakes(document): Consolidate Claude launch guidance in configuration reference

* no-mistakes(document): Clarify Claude permission documentation reference
…nguid#4819)

* fix(supervision): prevent idle recovery loops without stranding wakes

* no-mistakes(review): Remove unused wake-append rollback helper
…id#5889)

* fix(bin): stop the remote-job worker busy-polling an idle queue

The serving loop slept 50ms between passes and re-ran state preparation
(chmod on every queue directory), the heartbeat publish, and the stale sweep
on every pass. It now blocks on a worker.wake FIFO that staging,
cancellation, and lane exit nudge, keeps a short fast-poll window after
activity, refreshes the heartbeat at most once a second, and runs the sweep
(which re-applies the queue directories' 0700 modes) at startup and then on
a bounded interval. Lane-owned records are no longer re-read every pass.

Measured with a fork/execve-interposing counter on a --serve worker in a
disposable HOME and queue, bash 3.2, 20-second windows (the counter slows
the old loop to about 5 passes a second, so real-host rates were higher):
  idle worker             146 forks/s, 61 execs/s -> 4.8 forks/s, 3.1 execs/s
  one running long job    232 forks/s, 100 execs/s -> 15 forks/s, 11 execs/s
Stage-to-result latency for a no-op job, idle and back to back, stayed at
about 0.8-1.2s in both versions (dominated by job execution, not pickup).

* perf(bin): drop per-cycle forks from watcher, drain, and lock helpers

The watcher, drain, inactive-reconcile scan, and branch-outcome reads forked
small external commands on every cycle where bash can do the same work.

- fm-wake-lib.sh gains fm_dirname_to, fm_basename_to, and
  fm_epoch_seconds_to, exact stand-ins for $(dirname --), $(basename --),
  and $(date +%s); the clock uses printf %(%s)T on bash 4.2+ and still forks
  date exactly once on stock macOS bash 3.2.
- fm_lock_abs_path, fm_wake_signal_seen_path, fm_path_age, the watcher's
  age_of and wedge timer, and the recovery-marker line count use them or
  plain reads instead of dirname/basename/tr/date/wc.
- window_to_task reads a meta file once instead of two
  grep | tail -1 | cut -d= -f2- pipelines per file per call.
- fm-classify-lib.sh reads uname -s once at source time instead of in every
  status stat helper.
- Libraries sourced every cycle derive their own directory without forking
  dirname, including the backend adapter siblings a subshell re-sources on
  each probe.

tests/fm-fork-free-helpers.test.sh pins each replacement against the command
it replaces on edge-case inputs, under every available bash and both the C
and a UTF-8 locale; CI's stock macOS bash lane runs it under /bin/bash 3.2.

Measured with a fork/execve-interposing counter in a disposable home, one
tmux crew task, FM_POLL=1 (forks and execs per watcher cycle, per run
otherwise):
  watcher cycle      bash 5.3  299/138 -> 199/66   bash 3.2  341/146 -> 224/80
  drain              bash 5.3  492/238 -> 430/200  bash 3.2  567/250 -> 491/212
  inactive scan      bash 5.3   27/14  ->  17/4    bash 3.2   37/14  ->  17/4
  branch-outcome     bash 5.3   40/21  ->  35/16   bash 3.2   48/24  ->  38/19

* test: note the interpreter-expanded version probe for shellcheck

* no-mistakes(review): Fix worker idle bounds and fork-free contributions snapshot

* no-mistakes(review): Coalesce buffered worker wake nudges into one wake

* no-mistakes(review): Coalesce wake nudges via pending marker so publishers never block

* no-mistakes(review): Claim wake nudges atomically via noclobber pending marker

* no-mistakes(review): Release abandoned wake claims only after a 30-second bound

* no-mistakes(review): Drop worker wake FIFO; load path helpers side-effect free

* no-mistakes(document): Document remote worker polling and preemption cadence

* no-mistakes(ci): Fixed both failing CI shards: isolated remote and teardown test fixtures now include fm-path-lib.sh, which fm-wake-lib.sh requires. The three affected tests, fm-lint.sh, and git diff --check pass locally
…uid#5941)

* fix: keep a successor watcher and remote-reply listeners across the gaps that dropped them

A main-only supervision pass-through exited without leaving a watcher, and each remote-reply poll released its claim until the next cycle, so short-lived listeners stayed down.

* no-mistakes(document): Clarify listener and supervision continuity documentation

* no-mistakes(ci): Fixed the three Greptile findings: failed ingestion leaves one durable capture, failed reads exit instead of relistening, and the disposable-checkout guard rejects state paths outside the marked lab. Added behavioral tests; the remote-reply and watcher-lock suites, shell syntax checks, and git diff checks passed

* no-mistakes(ci): Fixed a race in the wake-queue interruption test: it now waits for the drain to own the lock and enter handling before signaling it. The wake-queue suite, shell syntax check, and diff check pass
…enguid#5925)

* fix: date replayed branch outcomes and ask main to check current state first

A captain outcome main never acknowledged is presented again, which after a
harness or posture switch, or the first drain after the upgrade whose earlier
presenter never advanced the read cursor, can be days after its situation
settled. The replay read as fresh news, so a PR since merged looked ready.

bin/fm-branch-outcome.sh now adds a "recordedAgo" age (minutes, hours, then
days) to present and unprocessed rows, one owner of that wording for both
presenters. The drain's BRANCH OUTCOMES captain lines and the Pi branch's
processing request name that age and ask main to check the task's current
state first; an outcome already settled needs only the acknowledgement, with
nothing relayed to the captain. Nothing is adopted as processed, so a fresh
home's first outcome is still presented until acknowledged.

* no-mistakes(review): Absent processed marker reads 0; never adopt read cursor

* no-mistakes(review): Require recordedAgo in Pi requests; report undated rows to main

* no-mistakes(review): Keep recordedAgo on captain rows only in present output

* no-mistakes(document): Correct cutover documentation and retire stale migration guidance

* fix: keep settled branch outcomes out of main's reply to the captain

A live Pi primary that took over a host-drain home received the carried-over
outcomes dated and check-first, but its processing reply still told the
captain about an outcome whose decision had since been answered. The request
also claimed every outcome was already shown as an anchor entry in this
transcript, which is false for an outcome carried over from before a restart
or a switch of primary.

The Pi processing request now says each outcome was recorded earlier and may
already have been seen or handled, and that a settled outcome gets no
captain-facing mention at all in the reply or any recap, not even that it is
settled. The drain's BRANCH OUTCOMES header and the supervision docs state the
same rule, and the tests check both delivered texts.

* fix: scope main's outcome reply to what is still open

Telling main what not to say about a settled outcome was not enough: in two
live Pi trials the processing reply still told the captain that an answered
decision was settled. Main now sorts the outcomes by current state first, and
its reply to the captain covers only the still-open ones, written as if the
settled ones had never been listed. With that framing three live Pi trials
kept the settled outcome out of the reply and relayed the open one each time.

The drain's BRANCH OUTCOMES header and the supervision docs use the same
framing, and the tests check both delivered texts.

* no-mistakes(review): Clarify that main acknowledges every presented captain outcome

* no-mistakes(document): Clarify outcome cursor ownership across Pi and host

* no-mistakes(ci): Fixed Pi replay by batching unprocessed captain outcomes oldest-first and acknowledging only through each batch. Verified a marker-less backlog over 1 MiB replays through all batches. A real-drain regression confirms an older keyed decision remains under OPEN DECISIONS after a newer branch row is acknowledged; the check-first instruction now names those decisions. Relevant targeted tests and branch-supervision tests passed; the full host suite timed out

* no-mistakes(ci): Fixed the host drain’s check-first wording in bin/fm-wake-drain.sh; the CI fixture now passes. The full host suite passed the affected fixtures but timed out later. Syntax and diff checks passed

* no-mistakes(ci): Fixed ci-1: abbreviated Pi outcome summaries now stay within 1,024 characters and include a row-specific full-outcome lookup command. The delivered instruction requires reading the full outcome before acting, relaying, or acknowledging it. The new extension-driver regression failed before the fix and passes now; the Pi and supervision-host suites pass

* no-mistakes(ci): Corrected the batching sentence in docs/pi-supervision-branch.md. The cancelled CI check needs no code fix; its clean rerun passed. The Pi branch extension suite and git diff check passed
…guid#5961)

* fix: wake an idle Claude primary for attended main-only hand-backs

An attended main-only pass-through confirmed a handling handoff for the
successor it leaves running, which flipped the recovery marker to handling.
The Claude Stop hook only rewakes main while that marker reads downtime, so
the close reached no one and an idle primary slept with wakes queued.

The pass-through now leaves the marker at downtime, and a close that turns
main-only at its turn hands the consumed handoff back to downtime before it
reaches main. Regression tests drive the real Stop hook around the real host
on both paths and for the successor's own later close, and a new opt-in live
guard proves it against an idle interactive Claude primary with a pre-fix
negative control.

* no-mistakes(review): Assert live lab Stop-hook registration via parsed settings JSON

* no-mistakes(document): Correct supervision hand-back documentation

* no-mistakes(ci): Fixed the failed downtime-write path so the Stop hook notifies main instead of silently dropping the close. Corrected the live guard’s tracked-hook check and added the requested at-turn main-only scenario. The new regression failed before the fix and passed after it; the host suite, syntax checks, and diff check pass. The credentialed live guard was not run in this CI phase because it writes outside the worktree

* no-mistakes(ci): Fixed the Stop hook’s retry ordering: a crashed host gets its bounded retry before a non-crash hand-back failure is reported. The Stop-hook suite passes, including the crash regression. The host suite passed the failed-marker-write regression but timed out before completing; syntax and diff checks pass
…n/fm-teardown.sh. No lint limits, roots or other files changed. Cause: the CI log shows only one failing root, bin/fm-teardown.sh, which ended with reason=memory at about 8.39 GiB RSS ("shellcheck: out of memory"). Every other root ended reason=ok. The merge brought in upstream growth in fm-wake-lib.sh, fm-classify-lib.sh and fm-lease-lib.sh, plus fm-path-lib.sh, and that pushed teardown's full-analysis ShellCheck past the address-space limit. Measured locally, peak RSS was 8.25 GiB at base and 8.55 GiB at HEAD. Invariant: fm-teardown.sh should load each of its sourced libraries once, from its top-level source block, and should not carry a source path that can never run. teardown_herdr_require_prerequisites had a second `. "$SCRIPT_DIR/fm-wake-lib.sh"` behind `if ! declare -F fm_lock_try_acquire`. That branch could never run. fm_lock_try_acquire is defined only in fm-wake-lib.sh, which is sourced unconditionally at the top of the script (line 386). All three callers of teardown_herdr_preflight_target (lines 3125, 3333, 3447) run after that line. ShellCheck still followed this dead source and analysed the whole wake-lib closure a second time. I removed the dead branch, as the removal-first rule asks. The existing check that refuses teardown when the lock functions are missing still stands, so the path still fails closed. Source-following itself was not narrowed or redirected, as the fm-lint.sh:880 comment forbids. The same wake-lib follows inside fm-public-followup-lib.sh and fm-pending-reply-lib.sh were left alone, because those libraries are part of other files' lint runs too and the user limited this round to teardown. Verification: - I ran the real guard, `FM_LINT_REQUIRE_BOUNDS=1 FM_LINT_JOBS=1 bin/fm-lint.sh --telemetry ... bin/fm-teardown.sh`, with its default 12 GiB cap. - The original file reproduces the CI failure: reason=memory, rc=251, rss 8386716 KiB, "shellcheck: out of memory". - The fixed file passes: reason=ok, rc=0, no findings, rss 7152892 KiB. That is about 1.2 GiB below the limit where CI failed, and lower than base. - These tests all pass with no failures: tests/fm-teardown.test.sh, tests/fm-teardown-endpoint-safety.test.sh, tests/fm-remote-secondmate-parent-binding.test.sh and tests/fm-backend.test.sh. - fm-teardown.test.sh includes the Herdr preflight cases (lock contention, preflight refusal, forced secondmate Herdr child preflight), so the edited function is exercised. - `bash -n` is clean. - I did not add a test, because a lint-memory test would be new machinery and a test that reads source text is forbidden. The before/after guard runs above are the regression evidence. - I did not run a full local `--partition 1of2`. Only teardown failed in the CI log, and that root was checked directly under the same guard
@ashkonmousavi
ashkonmousavi merged commit 85228ca into main Sep 28, 2026
20 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.