Skip to content

feat(bin): merge upstream firstmate main and adapt fork supervision contracts - #84

Merged
andrewesweet merged 50 commits into
mainfrom
fm/fm-upstream-sync-0929
Sep 30, 2026
Merged

andrewesweet merged 50 commits into
mainfrom
fm/fm-upstream-sync-0929

Conversation

@andrewesweet

@andrewesweet andrewesweet commented Sep 30, 2026 •

Copy link
Copy Markdown
Owner

Intent

Merge the upstream kunchenguid/firstmate main branch, 41 commits from 30ef650 through 1f2c954, into this fork's main with a true merge commit rather than a rebase or squash. Conflicts are resolved so that upstream's version wins where the fork had merely carried upstream code and the fork's own additions are re-applied on top, with the fork's behaviour kept wherever the two genuinely differ: the Claude Code supervision-branch mod and its mutual exclusion with the supervision host, the Pi supervision-branch extension and its shared modules, typed dispatch resolution, the per-task spend ledger, the per-host no-mistakes validation cap, pooled slot leases, the compact adviser in auto mode on the function-hooks launch path, the three-part brief contract with its published-intent rules, and the fork's intake and discipline skills. Upstream's new behaviour, including the supervision host enabled by default for Claude primaries, AGENTS.md sections moved into on-demand skills, the memory thrashing guard, the AI-trailer opt-out, never-send values withheld from dispatch resolver requests, and bounded merge retries, comes in as upstream ships it, adapted only where the fork's contracts require it. The merged tree is validated with the full regression.

What Changed

  • Merged 41 upstream commits into the fork: the supervision host is now enabled by default for Claude primaries, watcher arms and remote reply listeners survive slow cycles, fm-pr-merge.sh retries a bounded number of times while GitHub reports mergeable: UNKNOWN and accepts a task's next PR after its bound PR merges, pending-reply grace is measured from turn completion rather than delivery, contribution polling converges every open owner onto a terminal state, fm-git-strip-ai-trailers.sh honours a config/keep-ai-trailers opt-out, and the dispatch resolver withholds never-send values from its requests.
  • Added upstream's new surfaces: the Jev guard framework with the memory RSS/swap thrashing guard (bin/fm-jev-mem-guard.{sh,py}, docs/jev-guards.md), the disposable live supervision lab builder (bin/fm-live-lab.sh), bin/fm-path-lib.sh, Claude Code Calm supervision notes (.claude/mods/firstmate-calm/lib/fm-branch-notes.ts), and the situational AGENTS.md sections relocated into on-demand skills under .agents/skills/.
  • Kept the fork's own behaviour on top of the merge: the Claude and Pi supervision-branch mods with their shared eligibility fold and span-based second-mate wake routing, typed dispatch resolution, the spend ledger, the per-host validation cap, pooled slot leases, the auto-mode compact adviser on the function-hooks launch path, and the three-part brief contract - with quiet/away supervision, teardown, spawn, watch-arm and host tests extended to cover the merged behaviour (about 13.2k lines added across 182 files, roughly half in tests/).

Risk Assessment

⚠️ Medium: A 41-commit upstream merge touching 178 files is large and broad by nature, but the fork-adaptation seams I traced (silent-verdict rule, host default-on gate and its mod exclusion, quiet-vs-away split, never-send check, attribution and compact-adviser launch paths) are internally consistent and the four fixer commits land correctly, so it is safe to merge with the full regression as the remaining gate.

Testing

Drove the span-routing change through the real product surfaces: the Pi dispatch module and shared eligibility fold under the merged test, the bin/fm-branch-dispatch.mjs host entry against a disposable lab FM_HOME whose presentation cursor was written by the real bin/fm-classify-lib.sh, and the mod plugin and agent-prompt generators. The round-1 no-go scenario now passes and is a true regression — it fails at the parent commit and passes at HEAD — while the adversarial case, a new keyed decision inside the same unread span, correctly refuses the branch and leaves the wake to main. Parity between bash, the Pi extension, the mod and the lib holds, the committed agent prompt regenerates byte-identically, and the two longer suites completed clean: the whole Pi branch-extension file at 57 passing cases and the spawn dispatch-profile suite that consumes the new fixtures pane-log hunk. No UI surface is involved, so evidence is CLI transcripts and suite output rather than screenshots. Worktree left clean.

  • Live validation: ✅ go - 9 of 9 scenarios driven live against the product
Scenario Result Live Evidence
A second mate's routine status update behind an unrelated parked decision reaches the supervision branch on both the Pi and attended-host paths ✅ pass live bash tests/fm-pi-branch-extension.test.sh filtered to test_branch_dispatch_routes_secondmate_signal_by_new_span — "ok - second-mate signal rows route by their new span while crewmate and stale routi…
The same routing decision through the host's own command-line entry on a real FM_HOME whose presentation cursor was written by bin/fm-classify-lib.sh ✅ pass live secondmate-span-host-routing.txt — FM_HOME=$LAB node bin/fm-branch-dispatch.mjs scope prints status=safe rows=1 tasks=mate for the routine span
Adversarial: a new keyed decision inside the same unread span keeps the wake on main instead of handing it to the branch ✅ pass live secondmate-span-host-routing.txt — after appending needs-decision [key=new-call], the same host entry prints status=unsafe with no eligible rows
The fix is a real regression fix, not a test rewrite: the merged test fails before it and passes after ✅ pass live Same filtered case with lib/fm-branch-eligibility.ts restored from a338057 fails with "expected branch, got pi=true host=false"; restoring HEAD makes it pass
Bash truth, the Pi extension, the Claude mod, and the shared lib still agree on every eligibility fixture, including through the host entry script ✅ pass live bash tests/fm-branch-eligibility.test.sh — all six cases ok, including "host entry: bin/fm-branch-dispatch.mjs scope prints the shared fold's verdict on every fixture in each mode the host uses"
The Claude Code supervision-branch mod still loads the changed eligibility core and ships an agent prompt regenerated from bin/fm-branch-prompt.sh ✅ pass live bash tests/fm-branch-claude-mod.test.sh (mod surface, lib resolution, eligibility wrapper load, agents/fm-branch.md currency) plus bash bin/fm-branch-agent-md.sh --check clean
Task-scoped wake classification unaffected elsewhere in the fold (open captain decision, symlinked log, unresolvable project, torn read) ✅ pass live bash tests/fm-branch-scope.test.sh — four cases ok
No sibling regression across the whole supervision-branch dispatch, delivery, and mirror surface that shares the changed fold ✅ pass live bash tests/fm-pi-branch-extension.test.sh full file — exit 0, 57 cases ok, zero failures (pi-branch-extension-full.log)
The fixtures.sh FM_FAKE_PANE_LOG pane-export capture still drives the spawn dispatch-profile launch assertions ✅ pass live bash tests/fm-spawn-dispatch-profile.test.sh — exit 0, "# all fm-spawn-dispatch-profile tests passed"
Evidence: Live host-entry routing on a disposable lab home: routine second-mate span reaches the branch, a new keyed decision in the span does not

Source: Live host-entry routing on a disposable lab home: routine second-mate span reaches the branch, a new keyed decision in the span does not

--- host entry: fm-branch-dispatch.mjs scope (routine span behind unrelated parked hold) --- status=safe corrupted=0 rows=1 tasks=mate unscoped=0 --- adversarial: append a NEW keyed decision to the same span --- status=unsafe corrupted=0 rows= tasks= unscoped=0

== lab home: $LAB (disposable) ==
--- presented span (drained) ---
needs-decision [at=1790000000] [key=old-hold]: deferred captain call
--- new unread span ---
done [at=1790000100]: sample-a PR merged
--- host entry: fm-branch-dispatch.mjs scope (routine span behind unrelated parked hold) ---
status=safe
corrupted=0
rows=1
tasks=mate
unscoped=0

--- adversarial: append a NEW keyed decision to the same span ---
status=unsafe
corrupted=0
rows=
tasks=
unscoped=0
Evidence: Full Pi branch-extension suite: 57 cases ok, zero failures

Source: Full Pi branch-extension suite: 57 cases ok, zero failures

ok - fm_branch_outcomes hides through ToolExecutionComponent while Calm-off and HTML export stay stock
ok - the installed Pi still bounds the picker's list and ranks its search
ok - branch owns accepted wakes with a stable prefix and deterministic verdict-driven delivery
ok - a captain outcome reaches main's model as one typed, sequence-keyed processing request while routine notes stay plain
ok - a captain-classified wake passes to main with cover rows, the guard file, and its sweep
ok - the pi classification record captures the evidence text it judged beside its byte range, with the cap named
ok - the Pi classifier default resolves the branch's model (pin, else main's session model) before any completion call
ok - the explicit classifier config wins over the pin and records the model used
ok - the unresolvable classifier name falls back once to the branch's model and records the model used
ok - the gated shadow trial appends joined records and the away posture skips classification
ok - requested and unsolicited healthy outcomes keep distinct delivery and event ownership
ok - captain outcomes are exact and exactly once across crash, reload, busy main, compaction, and an unrelated assistant response
ok - a captain outcome opens one sequence-keyed processing turn, survives empty and unrelated answers, is re-presented at run end and session start, and closes only on its acknowledgement
ok - Pi: an abbreviated processing request points to the full outcome and instructs main to read it first
ok - Pi: a marker-less >1 MiB backlog is presented oldest first in bounded requests and continues after batch acknowledgement
ok - an unprocessed captain row the store lists without its age is reported to main, never formatted undated, and stays unprocessed until the store is healthy
ok - scopeForUnreadWake excludes every main-only class without vetoing eligible task-local rows, and writes the eligible snapshot
ok - second-mate signal rows route by their new span while crewmate and stale routing stay unchanged
ok - branch prompt_cache_key is stable per home across sessions and distinct between homes
ok - branch default-on eligibility (task-scoped, heartbeat, legacy flag ignored) binds and a broken branch rejects to watcher fallback
ok - under the away-posture record the wake carries the verbatim read-back tail, claims every row, opens no processing turn, cancels a pending request, and presents the accumulated rows after archive
ok - Pi branch: an unchanged held outcome reaches the captain once across cadences, and a later red check on the task still does
ok - an accepted away-only wake rejects after archive, while a drained task-local wake stays a quiet no-op
ok - a claimed heartbeat row on a non-heartbeat away wake lifts task scoping for the fleet report
ok - a heartbeat review survives a check row arriving before its drain
ok - fm_branch_report refuses a task the wake did not name, fleet included, while a heartbeat is unscoped
ok - pre-drain eligibility re-check excludes a newly main-owned row without deferring eligible work
ok - a co-present needs-decision row neither vetoes nor falsely settles routine branch delivery
ok - a settled branch turn without a durable outcome falls back and releases its grant for main replay
ok - provider-error latches cool down, re-probe once with backoff, and recover through a durable report
ok - report-less error-free turns latch the branch, recover through probes, and double the cooldown
ok - selection changes preserve in-flight transcript ownership and reset provider-error streaks
ok - a stale main claim returns the durable wake to watcher delivery
ok - pre-drain eligibility re-check no-ops an already-drained wake
ok - dialog mirror filters tool and operational traffic, lands before wakes, and keeps a durable cursor
ok - the dialog mirror re-anchors for each session's new branch conversation and stays incremental within it
ok - every main session start begins a new branch conversation while one session keeps its own
ok - the current pin state binds every branch build, and clearing it returns the branch to main's model
ok - unpinned branches follow main model changes live while pinned branches stay fixed
ok - supervision-model command persists the captain's pick and rebinds the live branch
ok - supervision-model opens a bounded searchable list, follow main first, and pins the branch alone
ok - branch model picker keeps follow main first and filters the eligible catalog
ok - the effort pin binds every branch build, and clearing it returns the branch to main's effort
ok - unpinned branches follow main effort changes live while pinned branches stay fixed
ok - an extension-registered provider resolves in the isolated branch runtime
ok - supervision-model runs an effort picker after the model picker and persists both independently
ok - an unusable model pin rejects to watcher fallback and an unparseable one is treated as no pin
ok - replacement activation cleans old branch leases and retries failed cleanup
ok - branch activates on a cold start once the lock is acquired, never before
ok - queued wakes and mirrors stop mutating branch state after lock ownership is lost
ok - stale reports, shells, mirrors, cursors, leases, and prompts perform no side effects
ok - a Pi session that does not own the lock accepts nothing and mutates no branch state
ok - an extension rebind re-mirrors undelivered dialog instead of dropping it
ok - outcome delivery keeps the event loop running and interleaved reports stay ordered and exactly once
ok - a session replaced mid-delivery cancels cleanly and the stored outcome still arrives exactly once
ok - a failing store script surfaces to the branch, its outcome is neither lost nor delivered twice, and fm_branch_processed renders the unified wordings byte-exactly
ok - a failed cursor write re-delivers a routine note exactly once more while a captain outcome stays deduplicated
Evidence: Regression proof: the span case at the parent commit
Error: unrelated open hold plus a routine merged line: expected branch, got pi=true host=false
Node.js v24.21.0: expected exit 0, got 1
Evidence: Spawn dispatch-profile suite driving the new fixtures pane-log capture
ok - active crew-dispatch profile does not block secondmate launches
ok - claude launches with function hooks keep the compact adviser in auto
ok - a function-hooks claude launch clears a kill switch the pane already carried
# all fm-spawn-dispatch-profile tests passed
- Outcome: 🔧 4 issues found → auto-fixed ✅ across 2 runs (1h11m46s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

🔧 **Rebase** - 1 issue found → auto-fixed ✅
  • ⚠️ docs/calm.md - merge conflict rebasing onto origin/main

🔧 Fix applied.
✅ Re-checked - no issues remain.

⚠️ **Review** - medium risk

✅ No issues found.

🔧 **Test** - 4 issues found → auto-fixed ✅
  • 🚨 lib/fm-branch-eligibility-core.ts:435 - Merge dropped upstream's second-mate span routing (fix: route second-mate signal wakes by presented status span kunchenguid/firstmate#5879). Upstream ships the fold inside .pi/extensions/lib/fm-branch-dispatch.ts (readPresentationCursor / secondmates set / spanRule, upstream tip 1f2c954 lines 313-500); the fork keeps its shared-module canon and the rule was never ported into lib/fm-branch-eligibility-core.ts (or its bash twin bin/fm-classify-lib.sh, which tests/fm-branch-eligibility.test.sh pins). The merged test tests/fm-pi-branch-extension.test.sh::test_branch_dispatch_routes_secondmate_signal_by_new_span fails: with an unrelated still-open needs-decision hold in a second mate's log and only 'done: sample-a PR merged' in the presented span, the attended-host path still routes to main (pi=true host=false), so every routine second-mate update stays pinned to main behind any parked hold. Fix in the shared module plus bash parity, not in the test.
  • ⚠️ tests/fm-procevent.test.sh:3831 - tests/fm-procevent.test.sh:3831 requires the orphaned listener's ppid to be exactly 1. On this host the user session runs systemd --user (pid 552), a PR_SET_CHILD_SUBREAPER process, so orphans reparent there and the reproduction precondition can never hold; the reaping behaviour itself is therefore unverifiable locally. The assertion predates this change (upstream b84e0e3), so it is not a merge regression. Decide whether to make the precondition subreaper-aware (accept any reparenting away from the test's own session) or leave it as a Linux-init-only guard that CI hosts satisfy.
  • ℹ️ tests/fixtures.sh:157 - Applied two test-harness repairs the merge left behind: tests/fixtures.sh gained upstream's FM_FAKE_PANE_LOG send-keys text-line capture (upstream added it in its own tests/fixtures.sh; the merge imported the tests that need it but not the fixture hunk), and tests/fm-spawn-dispatch-profile.test.sh's model/effort launch assert now uses the existing claude_worker_add_dirs helper so it matches the merged launch shape with the new task-channel --add-dir grants. Also regenerated .claude/mods/fm-branch-mod/agents/fm-branch.md through bin/fm-branch-agent-md.sh, which the last review round's prompt change left stale.
  • 🚨 live validation verdict: no-go (13 of 14 scenarios were driven live against the product); failed: A second mate's routine status update behind an unrelated parked decision reaches the supervision branch on both the Pi and attended-host paths
  • Live validation: ❌ no-go - 13 of 14 scenarios driven live against the product
Scenario Result Live Evidence
A fresh Claude primary home runs the supervision host with no config file, a cursor home does not, and a mod home writing off gets no host protocol in its block ✅ pass live host-default-on.txt: gate(claude) RUNS, gate(cursor) does NOT run, 'off' yields 0 host protocol lines, explicit line opts cursor in; tests/fm-supervision-host.test.sh and `tests/fm-supervision-instr…
An idle real Claude primary is woken for every close the attended host hands back to main, across four consecutive hand-offs ✅ pass live FM_SUPERVISION_HOST_ATTENDED_LIVE_E2E=1 tests/fm-supervision-host-attended-live-e2e.test.sh against installed Claude Code 2.1.285; live-attended-host.log
A worker commit through the spawn-installed hook loses its AI co-author trailer and keeps a human one ✅ pass live ai-trailer-strip.txt plus tests/fm-git-strip-ai-trailers.test.sh
A second mate's routine status update behind an unrelated parked decision reaches the supervision branch on both the Pi and attended-host paths ❌ fail live pi-span-failure.txt: expected branch, got pi=true host=false in tests/fm-pi-branch-extension.test.sh
fm-pr-merge retries a bounded number of times while GitHub reports mergeable UNKNOWN and accepts a task's next PR after its bound PR merges ✅ pass live bin/fm-test-run.sh tests/fm-pr-merge.test.sh exit=0 driving the real script against a stubbed forge
Never-send config values are withheld from dispatch resolver requests and every resolved call records one outcome line ✅ pass live bin/fm-test-run.sh tests/fm-dispatch-resolve.test.sh exit=0
A Claude worker launch carries attribution off, the task-channel --add-dir grants, the compact-adviser kill switch (auto on the function-hooks path), and the requested model and effort ✅ pass live bin/fm-test-run.sh tests/fm-spawn-dispatch-profile.test.sh exit=0 after the stale assert and fixture were repaired; staged launch transcript in spawn-dispatch-after-fix.log
The memory thrashing guard reads this host's real RSS and swap and exits non-zero under fail-forcing thresholds ✅ pass live bin/fm-test-run.sh tests/fm-jev-mem-guard.test.sh exit=0 with the guard's real host report in its output
A fresh primary's state dir is created before the session-start scope check and the digest completes under a lock refusal ✅ pass live bin/fm-test-run.sh tests/fm-session-start.test.sh exit=0
A pending reply's grace is measured from turn completion rather than delivery ✅ pass live bin/fm-test-run.sh tests/fm-pending-reply.test.sh exit=0
Fork contracts survive the merge: per-task spend ledger lines, pooled slot leases with teardown release, and the three-part brief contract ✅ pass live bin/fm-test-run.sh tests/fm-spend-ledger-append.test.sh tests/fm-spawn-slot-lease.test.sh tests/fm-brief.test.sh exit=0
The shipped mod agent definition matches its generator and the mod plugin loads ✅ pass live bin/fm-branch-agent-md.sh --check clean after regeneration; bin/fm-test-run.sh tests/fm-branch-claude-mod.test.sh exit=0
Teardown retires task-keyed watcher markers and orphan journals, watcher arms and the Claude Stop auto-arm survive slow cycles, and the new lab builder holds its contract ✅ pass live bin/fm-test-run.sh tests/fm-teardown.test.sh tests/fm-watch-arm.test.sh tests/fm-claude-stop-autoarm.test.sh tests/fm-host-mirror.test.sh tests/fm-live-lab.test.sh tests/fm-timeout-lib.test.sh exit=…
An orphaned procevent listener is reaped rather than left to storm relaunches ⏸️ untested no This host's user session runs systemd --user (pid 552) as a PR_SET_CHILD_SUBREAPER process, so an orphan never reaches ppid 1 and the test's reproduction precondition is unreachable here. Drive it on…
  • bin/fm-test-run.sh tests/fm-supervision-host.test.sh tests/fm-session-start.test.sh tests/fm-git-strip-ai-trailers.test.sh tests/fm-dispatch-resolve.test.sh tests/fm-pr-merge.test.sh tests/fm-pending-reply.test.sh tests/fm-spawn-dispatch-profile.test.sh tests/fm-spend-ledger-append.test.sh tests/fm-spawn-slot-lease.test.sh tests/fm-brief.test.sh tests/fm-branch-claude-mod.test.sh tests/fm-pi-branch-extension.test.sh tests/fm-backend-orca.test.sh tests/fm-branch-report-sequence.test.sh tests/fm-precompact-skills.test.sh tests/fm-ensure-agents-md.test.sh tests/fm-jev-mem-guard.test.sh (17 scripts, 3 failures at the time)
  • bin/fm-test-run.sh tests/fm-supervision-instructions.test.sh tests/fm-claude-stop-autoarm.test.sh tests/fm-watch-arm.test.sh tests/fm-procevent.test.sh tests/fm-teardown.test.sh tests/fm-host-mirror.test.sh tests/fm-live-lab.test.sh tests/fm-timeout-lib.test.sh (8 scripts, only the procevent subreaper precondition failed)
  • bin/fm-test-run.sh tests/fm-secondmate-safety.test.sh tests/fm-branch-eligibility.test.sh
  • FM_SUPERVISION_HOST_ATTENDED_LIVE_E2E=1 tests/fm-supervision-host-attended-live-e2e.test.sh (real installed Claude Code 2.1.285, disposable lab home)
  • manual host-gate drive in a throwaway home: bash bin/fm-supervision-engine-lib.sh enabled <lab>/config claude|cursor with the file absent, holding off, and holding claude, each cross-checked against the rendered block from bin/fm-supervision-instructions.sh --harness claude
  • manual commit-attribution drive: bin/fm-git-strip-ai-trailers.sh install <hooks> <repo> then a real git commit through core.hooksPath carrying one AI and one human Co-Authored-By trailer
  • bin/fm-branch-agent-md.sh then bin/fm-branch-agent-md.sh --check and bin/fm-test-run.sh tests/fm-branch-claude-mod.test.sh
  • bin/fm-test-run.sh tests/fm-spawn-dispatch-profile.test.sh (re-driven twice after the fixture and assert repairs)

🔧 Fix applied.
✅ Re-checked - no issues remain.

  • Live validation: ✅ go - 9 of 9 scenarios driven live against the product
Scenario Result Live Evidence
A second mate's routine status update behind an unrelated parked decision reaches the supervision branch on both the Pi and attended-host paths ✅ pass live bash tests/fm-pi-branch-extension.test.sh filtered to test_branch_dispatch_routes_secondmate_signal_by_new_span — "ok - second-mate signal rows route by their new span while crewmate and stale routi…
The same routing decision through the host's own command-line entry on a real FM_HOME whose presentation cursor was written by bin/fm-classify-lib.sh ✅ pass live secondmate-span-host-routing.txt — FM_HOME=$LAB node bin/fm-branch-dispatch.mjs scope prints status=safe rows=1 tasks=mate for the routine span
Adversarial: a new keyed decision inside the same unread span keeps the wake on main instead of handing it to the branch ✅ pass live secondmate-span-host-routing.txt — after appending needs-decision [key=new-call], the same host entry prints status=unsafe with no eligible rows
The fix is a real regression fix, not a test rewrite: the merged test fails before it and passes after ✅ pass live Same filtered case with lib/fm-branch-eligibility.ts restored from a338057 fails with "expected branch, got pi=true host=false"; restoring HEAD makes it pass
Bash truth, the Pi extension, the Claude mod, and the shared lib still agree on every eligibility fixture, including through the host entry script ✅ pass live bash tests/fm-branch-eligibility.test.sh — all six cases ok, including "host entry: bin/fm-branch-dispatch.mjs scope prints the shared fold's verdict on every fixture in each mode the host uses"
The Claude Code supervision-branch mod still loads the changed eligibility core and ships an agent prompt regenerated from bin/fm-branch-prompt.sh ✅ pass live bash tests/fm-branch-claude-mod.test.sh (mod surface, lib resolution, eligibility wrapper load, agents/fm-branch.md currency) plus bash bin/fm-branch-agent-md.sh --check clean
Task-scoped wake classification unaffected elsewhere in the fold (open captain decision, symlinked log, unresolvable project, torn read) ✅ pass live bash tests/fm-branch-scope.test.sh — four cases ok
No sibling regression across the whole supervision-branch dispatch, delivery, and mirror surface that shares the changed fold ✅ pass live bash tests/fm-pi-branch-extension.test.sh full file — exit 0, 57 cases ok, zero failures (pi-branch-extension-full.log)
The fixtures.sh FM_FAKE_PANE_LOG pane-export capture still drives the spawn dispatch-profile launch assertions ✅ pass live bash tests/fm-spawn-dispatch-profile.test.sh — exit 0, "# all fm-spawn-dispatch-profile tests passed"
  • bash tests/fm-pi-branch-extension.test.sh filtered to test_branch_dispatch_routes_secondmate_signal_by_new_span (passes at HEAD)
  • same filtered case with lib/fm-branch-eligibility.ts reverted to git show a338057b: (fails: expected branch, got pi=true host=false), then restored
  • bash tests/fm-pi-branch-extension.test.sh full file (exit 0, 57 cases ok, zero failures)
  • bash tests/fm-branch-eligibility.test.sh (bash / Pi extension / mod / lib parity, including the bin/fm-branch-dispatch.mjs host entry)
  • bash tests/fm-branch-scope.test.sh
  • bash tests/fm-branch-claude-mod.test.sh (mod plugin surface, eligibility wrapper load, agents/fm-branch.md currency)
  • bash tests/fm-spawn-dispatch-profile.test.sh (exit 0, all cases passed — exercises the new tests/fixtures.sh FM_FAKE_PANE_LOG capture)
  • live: disposable lab home via bin/fm-lab-home.sh create, presentation cursor committed through bin/fm-classify-lib.sh status_commit_presentation_snapshot, then FM_HOME=$LAB node bin/fm-branch-dispatch.mjs scope on a routine span and on a span carrying a new keyed decision
  • bash bin/fm-branch-agent-md.sh --check
⚠️ **Document** - 1 info
  • ℹ️ docs/claude-supervision-branch.md:25 - Judgment call recorded as documentation, not code: commit 0f51e6d ports the second-mate span rule into the shared eligibility core, and the Pi path picks it up through lib/fm-branch-eligibility.ts's readPresentedSpanOnDisk binding, but the Claude mod's scan (.claude/mods/fm-branch-mod/lib/fm-branch-scope.ts:211 ScopeScanDeps) supplies no readPresentedSpan, so on a mod home every second-mate signal row falls back to the whole-log rule and is classified by its queue marking alone. The core documents that fallback as designed (StatusSpan's comment: 'A consumer that cannot answer omits the binding'), and the blocker is real - statusFileIdentity needs the status file's device, inode, and birth stamp, while the hooks stat seam answers only isLink, size, mtimeMs. I documented the divergence in docs/claude-supervision-branch.md rather than leaving 'the same eligibility as the Pi branch' unqualified. Whether the mod should gain an identity-capable seam so a second mate's decision-carrying span stays on main there too is a behaviour decision for a later change, and tests/fm-branch-eligibility.test.sh carries no presentation-cursor fixture, so the four-fold equivalence test does not currently exercise the difference.
✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

@andrewesweet
andrewesweet force-pushed the fm/fm-upstream-sync-0929 branch from 4665fc4 to d23deaf Compare September 30, 2026 05:33
tmchow and others added 29 commits September 30, 2026 06:38
* docs: make calm easier to read

Restructure the Calm mode prose into sections, lists, and tables without changing documented behavior. Every original heading, anchor, fenced code block, inline-code span, link target, number, and quoted string is preserved.

* docs: restore reload case in calm override lead-in

The Built-in tool override collisions lead-in covers a session that reloads with Calm already on, as the original text did.
* docs: make turnend-guard easier to read

Restructure the turn-end guard doc's prose into shorter sections, lists, and tables without changing documented behavior. Every original heading, anchor, inline identifier, link target, and number is kept.

* no-mistakes(review): Restore legacy-only scope on TERM retirement sentence

* no-mistakes(review): Name Cursor park behavior in live e2e test line
…henguid#5872)

* docs: move situational AGENTS.md sections into on-demand skills

Backpass memory optimization: shrink the always-loaded AGENTS.md by moving
situational contracts (home layout, session-start recovery, validation and
landing supervision, scout completion, away/quiet supervision, Relay
ownership) into agent-only skills loaded at their triggers, with a trigger
index skill.

* docs: classify the new on-demand skills' documentation audience

Register the seven new agent-only skills as agent-runtime docs and fix a
link in validation-supervision that kept its AGENTS.md-relative path.

* docs: close load-timing gaps found by the live regression check

- load validation-supervision whenever an ask-user finding is decided or
  answered, so forbid --yes and process-every-return reach the worker
- keep the mid-task captain-ask rule, the unconfirmed network-checks rule,
  and the worker account pin rule inline in AGENTS.md
- fix cross-references that still pointed at moved AGENTS.md sections
…guid#5879)

* fix: route second-mate signal wakes by their new status span

A second mate's status log is a shared channel carrying many independently
keyed decisions, so judging its signal rows by every decision still open in
the whole log pinned each routine update to main behind any unrelated
parked hold. scopeForUnreadWake (the one owner for Pi and the attended
supervision host) now judges a second-mate signal row by the lines presented
since the last drain, bounded by the existing status-presentation cursor:
a decision, blocked, resolution, or captain-held line, or a line declaring
the key of a still-open decision, keeps the whole row on main, and any
cursor problem falls back to the whole log. Keys are read only at the
status parser's declared positions, with readable time stamps stripped as
bin/fm-classify-lib.sh does. Single-task crewmate and stale routing are
unchanged, and stale and signal rows for one mate keep independent verdicts.

The supervision branch now treats a second mate's done and merged lines as
relayed child outcomes, and fm-teardown refuses the branch actor second-mate
retirement through the existing role-partition helper in both postures.

* no-mistakes(review): Route second-mate resolutions to main only when closing open decision

* no-mistakes(review): Guard bare-verb fold lines and fold resolution spans incrementally

* no-mistakes(document): Clarify second-mate wake routing and retirement documentation

* no-mistakes(ci): The CI failure was a timing-sensitive watcher teardown test, not the PR’s signal-scope code. Extended the bounded wait for both state-directory and home removal. The watcher suite passes locally, and git diff --check is clean
…extension (kunchenguid#5882)

register-extension took the extension lifecycle lock and then the source
lock, while reconcile republishing an unhandled extension result holds the
source lock and reaches the lifecycle lock through the extension host's
process-event path. Both waits are unbounded and both owners stay alive, so
the two could wait on each other forever and freeze the home's monitoring
cycle.

register-extension now takes the source lock first, matching every other
path that holds both. The lifecycle lock still spans binding resolution
through registration publication, so binding retirement stays serialized.

A new lifecycle-order section in the extension-binding suite, run in the
default aggregate, holds a re-registration inside binding resolution while
reconcile republishes that source's unhandled result and requires both to
finish within a bound.

Fixes kunchenguid#5866
* fix(bin): run the repository's own hooks when git -c carries the per-task hooksPath

The per-task hooks wrappers cleared only the GIT_CONFIG_COUNT override before
looking up the repository's own hooks directory. When core.hooksPath reached git
through git -c (GIT_CONFIG_PARAMETERS), directly or inherited by a child
process, the lookup found the wrapper directory again and exited 0, so the
repository's real hook - such as a pre-push publish guard - never ran and the
push succeeded.

The lookup now ignores GIT_CONFIG_PARAMETERS too, so only the repository's
config files decide its hooks directory, and a failed lookup exits nonzero
instead of skipping the hook. AI-trailer stripping is unchanged.

Fixes kunchenguid#5871

* no-mistakes(document): Document git hook chaining and lookup failure behavior
…unchenguid#5744)

* feat(bin): add an optional never-send list to typed dispatch resolution

* Added config/dispatch-never-send, an optional local list of literal
  values and re: regular expressions checked against every string of
  the resolver request before it is sent to typesafe.ai
* A match, an unreadable list, or an empty or invalid pattern now stops
  the request and falls back to the off path, so firstmate dispatches
  through its existing intake; the one stderr diagnostic names at most
  the list line number and never the value
* No list, or a list with no match, leaves resolution unchanged

* no-mistakes(review): Match never-send literals across whitespace, drop regex mode

* no-mistakes(review): Inherit the never-send list into secondmate homes

* no-mistakes(ci): ci-2 (Lint 2), caused by this PR and now fixed. The rule that broke: the test script must not share shell variables with a library it sources. The new secondmate-inheritance test in tests/fm-dispatch-resolve.test.sh sourced bin/fm-config-inherit-lib.sh inside a `( ... )` subshell. That lib assigns `out`, so ShellCheck flagged every later `$out` in the test with SC2031. The subshell was the only place this PR sources that lib. The fix runs the propagation in a child shell instead (`bash -c '. "$1" && propagate_inheritable_config "$2" "$3"' _ lib from to`), so the test shell never sources the lib. `bin/fm-lint.sh tests/fm-dispatch-resolve.test.sh` now exits 0, and `bash tests/fm-dispatch-resolve.test.sh` passes. That includes the check that an inherited list is enforced in a secondmate home. ci-1 (Behavior portable serial 8), not caused by this PR, so no code change for it. The only failing test is tests/fm-remote-secondmate-relaunch.test.sh, which fails with "not ok - could not arm the PR poll fixture for the relaunch-ordering test". This PR doesn't touch that test or the code it runs. I reproduced the same failure locally on the base commit ea7c7f7. Main's own CI run on ea7c7f7 (run 36212602588) fails only this job, with the same message. The recent main runs before it also concluded failure
…hrashing guard (kunchenguid#5903)

* feat(jev): add the guard framework and the memory RSS/swap thrashing guard

A Jev guard is a bounded read-only host diagnostic that turns one class of
resource pressure into a machine-readable audit record and a one-line verdict.
This lands the framework contract (docs/jev-guards.md) with one representative
family (Pattern 46, memory RSS and swap thrashing) as the pilot: thin bash
wrapper, stdlib-only python engine, and a behavioral test through the CLI.

* no-mistakes(review): fix jev mem guard fail-open unknown and contract

* no-mistakes(review): fix mem guard rss pairing, memfree, tests, docs

* no-mistakes(review): fix inverted UNKNOWN branch in fail-forcing test

* no-mistakes(review): register docs/jev-guards.md in audience inventory

* no-mistakes(review): simplify wrapper, drop dead top_n, single-source thresholds

* no-mistakes(review): assert exact exit code in fail-forcing test leg

* no-mistakes(review): tolerate any stdout encoding in text output

* no-mistakes(document): Fix guard contract dash style and output wording
…enguid#5900)

* fix(bin): stop slow GitHub reads from starving and waking the contributions poll

The poll's fixed budget cut the tail URL's reads short, and any read killed at the bound was recorded as a forge failure, so routine GitHub slowness produced an observation-unavailable wake. A configured FM_CONTRIBUTIONS_BUDGET now rides the generated check shim and is cut down to the watcher's per-check bound with a margin, a read killed at the bound or the deadline is budget refusal that leaves records untouched, and the poll moves to the next URL that still has a full observation reserve instead of ending.

* fix(bin): report the bound when a signal death leaks through fm_run_timed

fm_run_external_timeout trusted the wrapper's recorded status over the runner's own bound verdict, so a read killed by the bound's TERM could surface as 143 and callers such as the contributions poll classified a budget-bound read as a forge failure. When the runner reports the bound, a signal-death status is the bound's own TERM racing the wrapper's bookkeeping and is reported as 124; a natural exit still passes through. The replace-shell probe also moves to the repo's portable BASHPID fallback so the suite runs on the macOS system bash.

* no-mistakes(review): Rotate contribution polling and verify generated budget behavior

* no-mistakes(review): Stabilize contribution rotation across successful observation refreshes

* no-mistakes(review): Exclude settled contributions from live observation rotation

* no-mistakes(document): Document contribution poll rotation and observation reserves

* no-mistakes(ci): Captain, fixed timeout handling with a one-line change preserving natural exit 137. Full session-start and contributions suites, targeted regressions, and lint passed. The timeout suite still fails on the known Bash 3.2 BASHPID issue, left unchanged as instructed. Logs retained in .no-mistakes/ci-evidence/. Remote CI was not rerun

* no-mistakes(ci): Fixed the fixture’s lock-acquisition race with a one-line bounded wait. Forced contention reproduced the CI error before the fix and passed afterward; the ordinary held-lock case, lint, syntax, and diff checks passed. Local Bash 3.2 failures remain: “cleanup lock bound 08 gave up before the marker lock freed” (also reproduced without the fix) and “TERM did not stop a watcher blocked inside a poll”. Remote CI was not rerun
…railers (kunchenguid#5859)

* feat: add keep AI trailers setting

* no-mistakes(document): Drop duplicate keep-ai-trailers line, update Cursor attribution note

* no-mistakes(document): Honor keep-ai-trailers for Devin worker attribution

* no-mistakes(document): Qualify Devin attribution note with keep-ai-trailers flag

* no-mistakes(review): Inherit keep-ai-trailers into secondmate homes
…sh-command popup cannot hide the composer (kunchenguid#5876)

* Fix Herdr composer reads blinded by the slash-command popup

Capture the full visible viewport for every Herdr composer state and content read instead of a bounded tail.
Claude Code renders its slash-command popup between the composer and the pane bottom, which pushes the composer outside a tail window.
The pre-Enter payload proof then read an empty composer, judged a typed command unsent, and cleared it without pressing Enter.
The proof-lines value now bounds only the clear cost, not the capture size.
Growing the window adds rows above the composer only, so bottom-most shape selection and prior verdicts are unchanged.
The suffix refusal is kept, and the shared inbox pending-line read stays a bounded tail.
Portable regressions cover the popup-below-composer layout, and the live submit-confirmation guard gains a third exit scenario.
The runtime-backends verification record documents the Herdr 0.9.0 and Claude Code 2.1.283 run.

Closes kunchenguid#5533

* no-mistakes(document): Clarify composer capture bound ownership

* no-mistakes(review): Drop unrequired bracketed-paste Enter fallback from live guard

* no-mistakes(review): Latch trust prompt, check idle composer arm first
…configurable bound (kunchenguid#5917)

Each per-task endpoint read in the session-start fleet digest runs in its own crash-isolated child under the existing timeout helper with a configurable FM_SESSION_START_ENDPOINT_TIMEOUT bound (default 10s).

A hung or killed read becomes that task's endpoint error while the digest continues, and the wrapper banners any abnormal child exit naming its status.
…nchenguid#5886)

* Keep lab tmux sockets on short private paths

* no-mistakes(review): Fix Linux stat checks and fail-closed lab tmux teardown

* no-mistakes(document): Update lab helper documentation for isolated tmux sockets
* Allow silent task-level no-change outcomes

* no-mistakes(review): Exclude silent outcomes from captain-return handoffs

* fix: look up supervision receipts by exact sequence

* no-mistakes(review): Suppress silent notes in away-return brief

* no-mistakes(review): Clarify visible notes; remove unused mode

* no-mistakes(review): Clarify silent outcomes and avoid false drain promises

* no-mistakes(document): Clarify silent supervision outcome documentation
…chenguid#5928)

* feat: make /quiet a statement where the attended supervision host runs

On a home that opted into the supervision host, quiet mode is what the
attended host already does, so /quiet now enters nothing there instead of
launching the quiet daemon and writing a record that would park a present
captain's main.

- bin/fm-afk-launch.sh quiet-check says quiet mode needs nothing where the
  attended host runs, or that the session is paused while its
  broken-session latch holds; a quiet enter refuses there before writing
  anything.
- Where the home opted in but the attended host lacks a part (engine,
  tools, verified mirror writer, identifiable main session, valid mirror),
  quiet-check names it and quiet mode falls back to the daemon.
- Under a live away record on that home, quiet-check and a quiet enter
  refuse and name the record, so the return runs first, whatever
  state/.afk says.
- A quiet enter records mode: quiet in the posture record, so start and
  start-native launch the quiet daemon without FM_AFK_MODE, and the away
  refusal wording fires only for away.
- bin/fm-host-mirror.sh check validates the dialog mirror read-only and
  exits 1 on a missing, unreadable, or invalid mirror.
- The quiet and afk skills and the supervision-host docs describe the new
  behavior; homes without the opt-in and Pi homes keep the daemon path.

* no-mistakes(review): Archive the quiet record when a quiet daemon start fails

* no-mistakes(document): Clarify quiet-mode documentation and remove stale duplicates
…uid#5884)

* fix(bin): grant Claude workers their task-channel dirs via --add-dir

Since Claude Code 2.1.257, a file-tool read (Read/Glob/Grep, and an
Edit's mandatory prior read) of a path outside the working directories
parks --permission-mode auto panes on a one-time interactive question,
and a "Block" answer lands permissions.blockReadsOutsideWorkingDirectories
in user settings, refusing the same reads even under bypass. Firstmate
launches Claude with no --add-dir, so a secondmate's parent-home steering
inbox and a ship or scout worker's launch record, steering inbox, brief
dir, and code-root .agents/skills were all outside: workers wedged on
the question the first time they read a steer.

Every Claude launch, spawn and relaunch, in both permission modes, now
grants exactly the task's channel directories: state/<id>.inbox for a
secondmate (in the parent home), or state/operational-inbox,
state/<id>.inbox, data/<id>, and the code root's .agents/skills for a
ship or scout. Paths resolve to real paths and lazily created channel
dirs are made before launch so the grant never names a not-yet-existing
directory; the whole state/ is deliberately never granted.

The grant keeps the bypass-mode launch argv changed on purpose: it also
protects bypass workers against a machine-recorded Block answer.

* no-mistakes(document): Consolidate Claude launch guidance in configuration reference

* no-mistakes(document): Clarify Claude permission documentation reference
…nguid#4819)

* fix(supervision): prevent idle recovery loops without stranding wakes

* no-mistakes(review): Remove unused wake-append rollback helper
…id#5889)

* fix(bin): stop the remote-job worker busy-polling an idle queue

The serving loop slept 50ms between passes and re-ran state preparation
(chmod on every queue directory), the heartbeat publish, and the stale sweep
on every pass. It now blocks on a worker.wake FIFO that staging,
cancellation, and lane exit nudge, keeps a short fast-poll window after
activity, refreshes the heartbeat at most once a second, and runs the sweep
(which re-applies the queue directories' 0700 modes) at startup and then on
a bounded interval. Lane-owned records are no longer re-read every pass.

Measured with a fork/execve-interposing counter on a --serve worker in a
disposable HOME and queue, bash 3.2, 20-second windows (the counter slows
the old loop to about 5 passes a second, so real-host rates were higher):
  idle worker             146 forks/s, 61 execs/s -> 4.8 forks/s, 3.1 execs/s
  one running long job    232 forks/s, 100 execs/s -> 15 forks/s, 11 execs/s
Stage-to-result latency for a no-op job, idle and back to back, stayed at
about 0.8-1.2s in both versions (dominated by job execution, not pickup).

* perf(bin): drop per-cycle forks from watcher, drain, and lock helpers

The watcher, drain, inactive-reconcile scan, and branch-outcome reads forked
small external commands on every cycle where bash can do the same work.

- fm-wake-lib.sh gains fm_dirname_to, fm_basename_to, and
  fm_epoch_seconds_to, exact stand-ins for $(dirname --), $(basename --),
  and $(date +%s); the clock uses printf %(%s)T on bash 4.2+ and still forks
  date exactly once on stock macOS bash 3.2.
- fm_lock_abs_path, fm_wake_signal_seen_path, fm_path_age, the watcher's
  age_of and wedge timer, and the recovery-marker line count use them or
  plain reads instead of dirname/basename/tr/date/wc.
- window_to_task reads a meta file once instead of two
  grep | tail -1 | cut -d= -f2- pipelines per file per call.
- fm-classify-lib.sh reads uname -s once at source time instead of in every
  status stat helper.
- Libraries sourced every cycle derive their own directory without forking
  dirname, including the backend adapter siblings a subshell re-sources on
  each probe.

tests/fm-fork-free-helpers.test.sh pins each replacement against the command
it replaces on edge-case inputs, under every available bash and both the C
and a UTF-8 locale; CI's stock macOS bash lane runs it under /bin/bash 3.2.

Measured with a fork/execve-interposing counter in a disposable home, one
tmux crew task, FM_POLL=1 (forks and execs per watcher cycle, per run
otherwise):
  watcher cycle      bash 5.3  299/138 -> 199/66   bash 3.2  341/146 -> 224/80
  drain              bash 5.3  492/238 -> 430/200  bash 3.2  567/250 -> 491/212
  inactive scan      bash 5.3   27/14  ->  17/4    bash 3.2   37/14  ->  17/4
  branch-outcome     bash 5.3   40/21  ->  35/16   bash 3.2   48/24  ->  38/19

* test: note the interpreter-expanded version probe for shellcheck

* no-mistakes(review): Fix worker idle bounds and fork-free contributions snapshot

* no-mistakes(review): Coalesce buffered worker wake nudges into one wake

* no-mistakes(review): Coalesce wake nudges via pending marker so publishers never block

* no-mistakes(review): Claim wake nudges atomically via noclobber pending marker

* no-mistakes(review): Release abandoned wake claims only after a 30-second bound

* no-mistakes(review): Drop worker wake FIFO; load path helpers side-effect free

* no-mistakes(document): Document remote worker polling and preemption cadence

* no-mistakes(ci): Fixed both failing CI shards: isolated remote and teardown test fixtures now include fm-path-lib.sh, which fm-wake-lib.sh requires. The three affected tests, fm-lint.sh, and git diff --check pass locally
…uid#5941)

* fix: keep a successor watcher and remote-reply listeners across the gaps that dropped them

A main-only supervision pass-through exited without leaving a watcher, and each remote-reply poll released its claim until the next cycle, so short-lived listeners stayed down.

* no-mistakes(document): Clarify listener and supervision continuity documentation

* no-mistakes(ci): Fixed the three Greptile findings: failed ingestion leaves one durable capture, failed reads exit instead of relistening, and the disposable-checkout guard rejects state paths outside the marked lab. Added behavioral tests; the remote-reply and watcher-lock suites, shell syntax checks, and git diff checks passed

* no-mistakes(ci): Fixed a race in the wake-queue interruption test: it now waits for the drain to own the lock and enter handling before signaling it. The wake-queue suite, shell syntax check, and diff check pass
…enguid#5925)

* fix: date replayed branch outcomes and ask main to check current state first

A captain outcome main never acknowledged is presented again, which after a
harness or posture switch, or the first drain after the upgrade whose earlier
presenter never advanced the read cursor, can be days after its situation
settled. The replay read as fresh news, so a PR since merged looked ready.

bin/fm-branch-outcome.sh now adds a "recordedAgo" age (minutes, hours, then
days) to present and unprocessed rows, one owner of that wording for both
presenters. The drain's BRANCH OUTCOMES captain lines and the Pi branch's
processing request name that age and ask main to check the task's current
state first; an outcome already settled needs only the acknowledgement, with
nothing relayed to the captain. Nothing is adopted as processed, so a fresh
home's first outcome is still presented until acknowledged.

* no-mistakes(review): Absent processed marker reads 0; never adopt read cursor

* no-mistakes(review): Require recordedAgo in Pi requests; report undated rows to main

* no-mistakes(review): Keep recordedAgo on captain rows only in present output

* no-mistakes(document): Correct cutover documentation and retire stale migration guidance

* fix: keep settled branch outcomes out of main's reply to the captain

A live Pi primary that took over a host-drain home received the carried-over
outcomes dated and check-first, but its processing reply still told the
captain about an outcome whose decision had since been answered. The request
also claimed every outcome was already shown as an anchor entry in this
transcript, which is false for an outcome carried over from before a restart
or a switch of primary.

The Pi processing request now says each outcome was recorded earlier and may
already have been seen or handled, and that a settled outcome gets no
captain-facing mention at all in the reply or any recap, not even that it is
settled. The drain's BRANCH OUTCOMES header and the supervision docs state the
same rule, and the tests check both delivered texts.

* fix: scope main's outcome reply to what is still open

Telling main what not to say about a settled outcome was not enough: in two
live Pi trials the processing reply still told the captain that an answered
decision was settled. Main now sorts the outcomes by current state first, and
its reply to the captain covers only the still-open ones, written as if the
settled ones had never been listed. With that framing three live Pi trials
kept the settled outcome out of the reply and relayed the open one each time.

The drain's BRANCH OUTCOMES header and the supervision docs use the same
framing, and the tests check both delivered texts.

* no-mistakes(review): Clarify that main acknowledges every presented captain outcome

* no-mistakes(document): Clarify outcome cursor ownership across Pi and host

* no-mistakes(ci): Fixed Pi replay by batching unprocessed captain outcomes oldest-first and acknowledging only through each batch. Verified a marker-less backlog over 1 MiB replays through all batches. A real-drain regression confirms an older keyed decision remains under OPEN DECISIONS after a newer branch row is acknowledged; the check-first instruction now names those decisions. Relevant targeted tests and branch-supervision tests passed; the full host suite timed out

* no-mistakes(ci): Fixed the host drain’s check-first wording in bin/fm-wake-drain.sh; the CI fixture now passes. The full host suite passed the affected fixtures but timed out later. Syntax and diff checks passed

* no-mistakes(ci): Fixed ci-1: abbreviated Pi outcome summaries now stay within 1,024 characters and include a row-specific full-outcome lookup command. The delivered instruction requires reading the full outcome before acting, relaying, or acknowledging it. The new extension-driver regression failed before the fix and passes now; the Pi and supervision-host suites pass

* no-mistakes(ci): Corrected the batching sentence in docs/pi-supervision-branch.md. The cancelled CI check needs no code fix; its clean rerun passed. The Pi branch extension suite and git diff check passed
…guid#5961)

* fix: wake an idle Claude primary for attended main-only hand-backs

An attended main-only pass-through confirmed a handling handoff for the
successor it leaves running, which flipped the recovery marker to handling.
The Claude Stop hook only rewakes main while that marker reads downtime, so
the close reached no one and an idle primary slept with wakes queued.

The pass-through now leaves the marker at downtime, and a close that turns
main-only at its turn hands the consumed handoff back to downtime before it
reaches main. Regression tests drive the real Stop hook around the real host
on both paths and for the successor's own later close, and a new opt-in live
guard proves it against an idle interactive Claude primary with a pre-fix
negative control.

* no-mistakes(review): Assert live lab Stop-hook registration via parsed settings JSON

* no-mistakes(document): Correct supervision hand-back documentation

* no-mistakes(ci): Fixed the failed downtime-write path so the Stop hook notifies main instead of silently dropping the close. Corrected the live guard’s tracked-hook check and added the requested at-turn main-only scenario. The new regression failed before the fix and passed after it; the host suite, syntax checks, and diff check pass. The credentialed live guard was not run in this CI phase because it writes outside the worktree

* no-mistakes(ci): Fixed the Stop hook’s retry ordering: a crashed host gets its bounded retry before a non-crash hand-back failure is reported. The Stop-hook suite passes, including the crash regression. The host suite passed the failed-marker-write regression but timed out before completing; syntax and diff checks pass
kunchenguid#5916)

* Fix worker launches to enter recorded worktrees

* no-mistakes(review): placeholder

* no-mistakes(document): Update agent-control.md worktree-refusal note to match new universal cd+assert

* no-mistakes(review): Add regression tests for Orca spawn/relaunch worktree carve-outs

* Fix PR relaunch and prelaunch cwd verification

* no-mistakes(document): Fix docs/agent-control.md: worktree cwd check is pre-launch, not post-launch

* fix: slim worktree launch change onto upstream main
…ardown (kunchenguid#5997)

* WIP: retire task-keyed watcher markers and orphan journals at teardown

Re-applies old PR kunchenguid#5584 on current main: teardown retires the
turn-ended .seen-* signature and an orphaned Herdr presentation
journal whose workspace is already gone, and the wake-drain rotates
its own dead scratch files. Not yet validated through no-mistakes.

* no-mistakes(ci): Fixed the Greptile P1 finding in bin/backends/herdr.sh. fm_backend_herdr_projection_token_workspace_gone used `! ... jq -e ... 2>&1`, which swallowed a jq runtime error (thrown when a non-object workspace entry, e.g. a number before a live token-bearing workspace, hits `.label`) and flipped it to a "gone" verdict, causing teardown to delete a still-live v1 presentation journal. Invariant: a workspace-query error/ambiguity must never be read as token absence; only a cleanly-parsed list with no token-bearing label is "gone". Replaced the body with a single jq verdict (unknown/present/gone): a non-array list or any non-object/non-string-label entry yields "unknown", jq errors/empty output fall through `|| return 1` to unknown, and only "gone" returns 0. Sibling fm_backend_herdr_projection_endpoint_matches_journal already fails safe on jq error (empty match -> journal kept), so it needed no change, matching the author's scoping. Added test_teardown_retains_v1_journal_when_workspace_query_ambiguous driving real teardown with a malformed workspace-list entry, proving the journal is kept and no workspace close occurs. Verified the old logic returns GONE on that input (test fails before, passes after); full tests/fm-teardown.test.sh suite passes (exit 0) and shellcheck is clean. Marker-naming finding left untouched per explicit out-of-scope instruction
…unchenguid#6002)

* fix(tests): disable Claude Code's auto-updater during live harness runs

fm_live_gate let a live run proceed without ever setting
DISABLE_AUTOUPDATER, so a live Claude test could let the real updater
repoint ~/.local/bin/claude into a temporary directory and stop every
Claude process on the machine from starting. Export
DISABLE_AUTOUPDATER=1 on every path where the gate lets a live run
proceed, and assert the export in tests/fm-live-gate.test.sh, including
that it reaches a child process the same way a real harness pane would
inherit it.

* no-mistakes(ci): Greptile flagged that the PR's DISABLE_AUTOUPDATER inheritance test only checked a `bash -c` direct child, not the fm-spawn.sh launch path. The user chose to fix it with a regression on that path. In tests/fm-live-gate.test.sh I replaced the generic child test with test_disable_autoupdater_reaches_the_claude_pane_on_the_fm_spawn_launch_path: it drives the real fm-spawn claude launch through the spawn fixtures, captures the exact staged launch command, and runs it as a synthetic pane whose only `claude` is a stub recording the inherited DISABLE_AUTOUPDATER, asserting it saw 1. Switched the file to source fixtures.sh (pulls in lib.sh, guarded) for the spawn helpers. Verified it is a real guard: the stub records `1` when the ambient var is set and `unset` when absent, so it fails if fm-spawn ever scrubbed the variable (e.g. env -i or -u). This confirms fm-spawn's launch construction never references the name and passes it through via ordinary ambient inheritance with no allowlist. Full suite passes (12 tests ok), shellcheck clean. Note for the outer executor: I embedded the daemon caveat as a code comment in the test, but the finding also asks the PR body to state that a backend daemon already running before the gate exported the variable does not inherit it and fully covering that would need launcher support - that forge-side PR-body sentence is outside this CI phase's scope

* no-mistakes(ci): Fixed Greptile finding ci-1. Root cause: fm-spawn.sh handed its launch command to an already-running backend daemon that never inherited the test process's exported DISABLE_AUTOUPDATER, so ambient inheritance dropped it and Claude's auto-updater could still run. Fix (bin/fm-spawn.sh): when DISABLE_AUTOUPDATER is set in the spawn's own environment, embed `export DISABLE_AUTOUPDATER=<value>;` into the LAUNCH command text (same idiom as the adjacent COMPACT_ADVISER_DISABLE export), so it survives a daemon-built pane, the env -i allowlist path, and relaunch alike; gated on presence so ordinary spawns are unchanged. Added regression test test_disable_autoupdater_survives_a_daemon_pane_that_never_inherited_it in tests/fm-live-gate.test.sh: stages a real claude launch with DISABLE_AUTOUPDATER set in the spawn env, then runs that exact command in a synthetic pane with `env -u DISABLE_AUTOUPDATER` and asserts the claude stub still recorded autoupdater=1. Verified the test fails (autoupdater=unset) without the fix and passes with it; the round-1 ambient test stays green either way. Full suite passes (13 ok); test file and isolated snippet shellcheck-clean (full fm-spawn.sh shellcheck kept getting terminated by the memory-constrained host, not by findings). Forge-side note for the outer executor: the PR-body caveat that fully covering the daemon case would need launcher support no longer applies to the Claude launch path and should be corrected
…cevent record (kunchenguid#6010)

* fix(bin): ring the inbox doorbell only for a newly published procevent result

publish_result rewrote a worker's captured Lavish round idempotently on
every reconcile, unconditionally moved an already-acknowledged inbox
record back out of handled/, and rang the doorbell every time - so an
already-processed round rang the owning worker on every cycle. Snapshot
the existing active and handled records before the idempotent write and
ring, or move anything, only when the write actually created a fresh
record; re-delivery of a still-open round is left to the inbox's own
re-ring ladder.

* no-mistakes(document): docs: reflect worker-board doorbell rings only on fresh inbox record

* no-mistakes(ci): Fixed Greptile finding ci-1 in tests/fm-procevent.test.sh (test-only change). The redelivery regression previously moved the delivered note into handled/ before any repeated reconciles, so it only proved an acknowledged note stays quiet and would still pass if an unchanged active note rang every cycle. Per the user's instruction, I inserted (before the mv into handled/) five repeated `pe reconcile` runs with the note still in the active inbox and asserted the ring log holds exactly one line and 001.msg remains active; the existing acknowledged-note assertion after the move is kept unchanged. No product code changed. bash -n confirms syntax is valid; the block mirrors the already-passing post-move reconcile/ring-count assertion directly below it
…unchenguid#6032)

* fix(bin): make the Claude Stop auto-arm refuse arguments before arming

A model running bin/fm-claude-stop-autoarm.sh --help mid-turn armed a real
supervision-host park owned by its short-lived tool process, leaving
supervision down once that process exited. The Stop hook passes no
arguments, so -h/--help now prints usage and any other argument is refused
before anything is sourced, read, or armed.

* no-mistakes(document): Clarify Claude Stop hook documentation for manual invocations

* no-mistakes(ci): Updated the argument-run regression test to compare checksums of state files as well as entry names. The Stop auto-arm test suite passes, and git diff --check is clean

* docs: restore the bin/ toolbelt intro's manual-use clause

The document step dropped "interactive entrypoints work by hand too" from
docs/scripts.md, which still holds for most bin/ scripts.

* no-mistakes(document): Clarify Claude auto-arm manual-use guidance
…uid#6039)

* feat(calm): show supervision sailboat and anchor notes on Claude Code

The Calm mod follows a bounded display tail copy of the outcome store,
which bin/fm-branch-outcome.sh append now refreshes, and the supervision
host's latch, and appends one dim transcript line per visible routine
outcome, captain outcome, and latch change, replaying unread and
unprocessed outcomes at session start. It shows them whenever the mod is
active, regardless of config/calm, and never marks anything read.

* fix(calm): show each supervision note once per session on Claude Code

Claude Code 2.1.283 stores ui.log lines in the session and restores them
on --continue, so the mod records how far each session has followed the
outcome store and a resume replays only newer outcomes. It also checks
file existence before reads so absent files do not log debug errors.
The live guard gains the supervision-notes scenario and the dated
2.1.283 record documents the observed behavior.

* docs: name the Claude supervision note row as the engine draws it

* no-mistakes(review): Seed outcome tail on present and anchor first tail on markers

* no-mistakes(review): Seed outcome tail at session start; replay against start markers

* no-mistakes(review): Bound outcome tail by bytes; reread recently changed files

* no-mistakes(review): Skip store validation when outcome tail already exists

* no-mistakes(document): Clarify bounded Claude supervision note replay

* no-mistakes(ci): Fixed seed-tail to validate only a bounded suffix of complete store rows and write it through the existing byte- and row-limited tail writer. Added a regression test with malformed history outside that window and updated the script header. Outcome tests and shellcheck passed; the full session-start suite timed out after 240 seconds
…enguid#6033)

* fix(bin): read a quiet-mode record as a present captain, never hold-for-return

Daemon-backed quiet mode writes the away-posture record marked mode: quiet,
but the entry announcement, read-back, and session-start digest rendered it
as "hold-for-return only", and the spend cap and PR merge gate treated it as
away. A present captain's requested actions could then be held for a return
that was not coming.

bin/fm-afk-contract.sh now owns which posture a record is (the mode
subcommand, fm_afk_contract_mode, fm_afk_contract_away_present). A quiet
record announces, reads back, and appears in the digest as a present captain
holding nothing; merges under it stay attended and it binds no spend cap. An
away record is unchanged, an /afk entry over quiet mode rewrites the record
as away, and a quiet entry never turns a standing away record quiet.

* no-mistakes(document): Clarify quiet-mode authority and remove stale away guidance

* no-mistakes(ci): The CI failure came from a race in the supervision-host test: its restart fixture could observe a watcher left by the preceding cycle. The test now retires that watcher and waits for the fixture arm to report its own started cycle. The focused test passed three times; the full suite was attempted but stopped at a separate intermittent test failure

* no-mistakes(ci): Fixed daemon refresh mode selection so an unset-mode refresh follows the posture record: /afk over a running quiet daemon now changes state/.afk to away, while a plain quiet refresh stays quiet. Added script-level regression coverage for start and start-native and corrected a quiet-refresh fixture. The launch test suite, syntax checks, and diff check passed

* no-mistakes(ci): Herdr was blocked before tests ran by a GitHub HTTP 500 downloading pinned Treehouse; no code change was warranted for that check. Fixed the Lint 1 ShellCheck warning in tests/fm-afk-launch.test.sh by annotating the intentional background PID capture. The focused test suite, ShellCheck, syntax check, and diff check passed
kunchenguid and others added 18 commits September 30, 2026 06:44
…merge (kunchenguid#6053)

* fix(bin): accept a task's next PR once fm-pr-merge confirms the bound one merged

require_recorded_pr_identity now checks fm_pr_poll_merge_already_notified for
the recorded pr= before refusing a different URL, so a task's later PR is
accepted once its earlier PR's merge is confirmed, while it keeps refusing
while the bound PR is still unmerged.

* no-mistakes(document): docs(fm-pr-merge): note next-PR accepted after bound PR merges
…6064)

* fix(bin): read a live quiet record as a present captain at the host and watcher

A quiet record left without its daemon (a quiet start that never ran or was
interrupted) was read as away by the supervision host, so it parked a present
captain's main and held captain outcomes for a return that never comes, and
the watcher and daemon silenced captain-held rechecks on record presence.

The host's posture checks, the watcher's and daemon's captain-held silencing,
and the host's outcome path (branch report, drain BRANCH OUTCOMES, relocated
branch authority, the owners' away wake note, and the Codex checkpoint bound)
now ask the record owner's away-or-quiet reading, so only an away record is
away. A live away record keeps today's behavior.

* no-mistakes(document): Correct quiet-record documentation and supervision guidance

* no-mistakes(document): Clarify quiet-record posture and captain-held rechecks

* no-mistakes(document): Clarify quiet-record posture in documentation
…kunchenguid#6043)

* fix(bin): name an in-window engine latch in the return brief and drop the false handling GAP line

The away return brief said nothing had failed after the supervision host
latched on engine errors during the window, and printed a GAP: watcher
downtime line whenever a wake was merely being handled or queued at return.

The failures section now reads the host ledger and latch record and names
the latch time, the window's engine-error count, and whether the session
is still paused or recovered. An open recovery episode is reported as
information, and as a gap only when a queued episode outlived the return
grace or the marker cannot be read.

* no-mistakes(review): Fix latch trip time, drop marker-age grace, bound error count

* no-mistakes(review): Report paused latch without ledger trip row; bound errors

* no-mistakes(review): Never report a failed probe's latch row as trip time

* no-mistakes(review): Only a retained trip row marks a pre-window latch

* no-mistakes(document): Clarify return-brief latch and watcher-gap documentation

* no-mistakes(ci): Fixed Lint 1 by marking the shared cooldown constant as used by sourcing scripts. The repository lint command and diff check pass; the return test run was stopped by a 180-second timeout after its completed cases passed

* no-mistakes(ci): Fixed the return brief so the trip time and error count come from the same initial latch row, and ledger rows before the current session’s lock boundary cannot affect its latch report. Added real-script regressions for both findings. The return test suite, repository lint, and diff check pass

* no-mistakes(ci): Fixed the return brief’s restart cutoff so it retains in-window failures, prints one line per initial-trip row, and omits zero-error count wording. Added real-script restart regressions. The return test suite, ShellCheck, and diff check pass

* no-mistakes(ci): Fixed the return brief so a recorded trip followed by recovery stays recovered, while a later pause with a lost trip append gets a separate “trip time unavailable” line. Added a real-script regression that failed before the fix. The return test suite, ShellCheck, syntax checks, and diff check pass

* no-mistakes(ci): Fixed the false second latch during recovery. A real-script regression failed before the fix and passes now; the lost-second-trip test still passes. The return test suite, ShellCheck, syntax checks, and diff check pass
* fix(calm): name the Claude Code Calm plugin fm so supervision notes read "fm: "

Claude Code labels every mod transcript line with the plugin name, so the
notes rendered as "firstmate-calm: ⚓ ...". Rename the plugin to fm, update
the live guard to assert the fm: label, and document the one-time replay for
sessions resumed across the rename.

* no-mistakes(document): Clarify Calm plugin rename in documentation
…#6037)

* feat(bin): add fm-live-lab.sh, a one-command live supervision lab builder

* fix(bin): exact lab windows, per-lab task ids, self-safe teardown

* fix(bin): target lab windows by id, stop lab descendants, add readiness tests

* fix(bin): keep Claude's auto-updater off in live labs; list fm-live-lab.sh

* fix(bin): start the lab tmux server without user config

* no-mistakes(review): Scope lab teardown to its store, root, and task ids

* no-mistakes(review): Record selected user stores at up for check and down

* no-mistakes(document): Clarify live lab documentation and remove stale narratives

* no-mistakes(ci): Fixed the CI failure by checking for an existing lab root before looking up the harness executable. The affected behavioral test and shell syntax check pass; the refusal also works with Claude absent from PATH

* no-mistakes(ci): Fixed all four Greptile findings: teardown signals only recorded lab processes and their descendants; the worker gate is in its granted task directory and its path is exposed; readiness uses current crew state; and mate and worker IDs use 12 nonce hex digits. The CLI behavior tests pass, as do shell syntax, ShellCheck, and diff checks. The Claude no-host path is unchanged

* no-mistakes(ci): Fixed the CI test’s dependence on an installed Claude binary by supplying a test-local stub. The full fm-live-lab test, shell syntax check, and diff check pass

* no-mistakes(ci): Fixed all three selected findings in bin/fm-live-lab.sh: down waits for recorded processes and escalates before cleanup, PID roots are checked against recorded start times, and Claude primary trust is rechecked after mate/worker readiness. Added behavioral tests in tests/fm-live-lab.test.sh. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass

* no-mistakes(ci): Fixed the pre-primary settle wait, worker gate instructions, unused retry variable, and teardown PID revalidation in bin/fm-live-lab.sh. Added behavioral tests in tests/fm-live-lab.test.sh. Both requested commands pass: tests/fm-live-lab.test.sh and bin/fm-lint.sh

* no-mistakes(ci): Fixed teardown to track pre-kill lab processes by PID and start time, including children orphaned when a root exits. Up now rejects an empty pane PID before calling ps. Added regression tests and a Linux-safe worker fixture. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass

* no-mistakes(ci): Fixed teardown tracking for children spawned during shutdown and made the worker fixture verify its exact window with a Linux-available shell. Both requested checks pass. The lab test takes about 66 seconds locally, so the under-one-minute target remains unmet

* no-mistakes(ci): Fixed ci-2 and ci-4 in bin/fm-live-lab.sh and tests/fm-live-lab.test.sh. Teardown now tracks identity-checked members of captured lab process groups, including children orphaned during shutdown, without signaling the caller’s group or unrelated processes. Lint passed, and the lab test passed four times

* no-mistakes(ci): Fixed teardown so an observed-empty process group is permanently dropped, preventing a reused group ID from signalling unrelated work. Added a ps-shim regression test. The lab test, lint, and diff checks pass

* no-mistakes(ci): Fixed ci-1 in bin/fm-live-lab.sh and tests/fm-live-lab.test.sh. The TERM-born-child fixture now waits until its handler is installed before calling down. Down sends SIGKILL to identity-valid survivors on every pass from pass 20 onward and includes survivor process details if it must refuse cleanup. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass locally; Linux CI remains to be verified

* no-mistakes(ci): Fixed down’s teardown wait to require two empty identity-checked scans separated by 0.5 seconds, and removed the unused test loop variable without changing the TERM-born-child test. The lab test, lint, and diff check pass locally
…unchenguid#6103)

* fix(bin): keep slow watcher cycles and preempted reply polls from breaking supervision

- fm_pending_reply_tick selects the records it has work for in one awk pass,
  so settled records cost no lock or fork and the walk no longer grows with
  the never-pruned store.
- An attached arm keeps following a live, identity-matched holder whose beacon
  went stale until the lock changes or the shared stall bound
  (fm_watcher_stall_bound), then fails with a typed stalled-holder line so the
  retry replaces the holder.
- The remote-reply adapter reports the job worker's preemption (exit 76) as a
  closed window, so the listener keeps its claim and polls again instead of
  being relaunched every watcher cycle.

* no-mistakes(document): Clarify watcher grace and attached-arm documentation
…geable is UNKNOWN (kunchenguid#6110)

* fix(bin): retry a bounded number of times when GitHub mergeable is UNKNOWN

Fixes kunchenguid#6020

bin/fm-pr-merge.sh refused a GitHub merge whenever the pull request's
mergeable field was not literally MERGEABLE. GitHub reports UNKNOWN for
a short while after a push or a base-branch change while it recomputes
mergeability, so a green, conflict-free pull request was refused as if
it could not be merged.

github_verify_mergeable now returns a distinct status when mergeable is
the only failing condition and reads UNKNOWN. The caller retries up to
5 times, 3 seconds apart (overridable in tests), re-reading and
re-checking every live condition on each attempt. Once the bound is
spent it reports mergeability as still being computed rather than
unmergeable, with the same nonzero exit as before. Every other refusal
(closed, draft, conflicting, red or missing checks, away authority,
queue protection) is unchanged and never retried.

* no-mistakes(ci): I fixed both review findings the way you asked. The full suite (`bash tests/fm-pr-merge.test.sh`) ran to completion. Its last lines showed all `ok`, and any failure would have stopped the run early. I watched the output through `tail`, so I didn't see the new test's own `ok` line directly. **ci-2 (`bin/fm-pr-merge.sh`), retry delay not validated.** What must hold: the retry wait is always a short, valid `sleep` argument, so a bad `FM_PR_GITHUB_MERGEABLE_RETRY_DELAY` can never trip `set -e` or hold the task lock for a long time. The retry loop is the only place that reads this variable. The script now reads the value once before the loop and accepts only whole numbers from 0 to 10. Anything else (empty, `abc`, `-1`, `1.5`, `11`, a huge number, leading spaces) falls back to 3. I ran those values through the check by hand and each came out as expected. The loop now sleeps on that checked value. **ci-1 (`tests/fm-pr-merge.test.sh`), no test for a check changing between UNKNOWN reads.** What must hold: every retry re-checks all live conditions, not just mergeable. The fake `gh pr view` in the test can now take an optional second word on each line of the mergeable sequence, which sets the first check's result. The new test `test_github_mergeable_unknown_retry_rechecks_checks` feeds `UNKNOWN`, then `UNKNOWN FAILURE`. It asserts: - exit code 1 after exactly 2 reads, - the refusal names `check 'ci' is not green`, - the message does not say mergeability is still being computed, - `pr merge` was never called. If a later change made the retry look only at mergeable, the loop would read UNKNOWN 5 times, end with the "still being computed" message, and this test would fail. I didn't run it against a deliberately broken script to confirm that. `bash -n` passes. `shellcheck` reports only the existing info-level notes about files it can't follow. Only `bin/fm-pr-merge.sh` and `tests/fm-pr-merge.test.sh` changed
kunchenguid#6112)

* fix(bin): converge every open owner onto a known terminal contribution

settle_final only cleared a stale error on retry, so an owner whose saved
row still said open kept projecting a merged or closed pull request as
open after another owner's row had already recorded the terminal
observation. Copy the known terminal observation to every owner whose
saved row is not itself terminal, keeping that owner's own pending and
notified state, and clear its error.

* no-mistakes(review): Carry terminal checked_at when converging existing owner rows

* no-mistakes(ci): I fixed Greptile finding ci-2 as you asked, with a change to tests/fm-contributions.test.sh only. The rule it enforces: when a retry converges an owner onto a URL that is already merged or closed, that owner gets the terminal owner's whole observation, not just its state. The same weak check appeared twice in test_interrupted_multi_owner_poll_settles_every_owner, so I fixed both: - **Open owner (line 784):** the check now also requires `.observation == $terminal[0].records[0].observation`. The existing checks for error, checked_at, pending and notified are unchanged. - **Errored owner (just below):** it only checked state and error before. It now reads the terminal owner's file and makes the same full-observation comparison. Adding the comparison alone would not have caught anything. The test fixtures gave both owners identical observations apart from `state`, so copying only the state would still have passed. In both cases I also set the terminal owner's observation head to HEAD_B, so the two observations now really differ. Verification: - The focused test passes against the current bin/fm-contributions.sh. - I temporarily changed `settle_final` so it copied only the state. The test then failed, reporting the owner still on the old head (HEAD_A). I restored the file afterwards, and `git status` shows only the test file modified. - The full tests/fm-contributions.test.sh suite exits 0. No product code changed. The other CI finding (ci-1, "Behavior portable serial 9") was left alone because you chose to ignore it
…nguid#6124)

* feat: run the supervision host by default on a Claude primary

An absent config/supervision-host on a Claude primary now reads as on with
the default engine, and a file holding `off` opts any home out. Cursor,
OpenCode, omp, Grok, and Codex stay file-gated, with `off` read as disabled
there too. Every reader asks fm_supervision_host_enabled instead of testing
the file, and non-bash readers query it through the lib's `enabled` entry.
A primary's `off` is not inherited by secondmates: each home keeps its own
supervision posture.

* test: pin the watcher-path posture in fixtures that assume no supervision host

Fixtures that drive the watcher arm or assert a non-host drain now write
an explicit off file, and fixtures that copy the Stop auto-arm or the
supervision instructions carry the engine lib they now source. The two
drain suites also stop reading the code root's config.

* fix: name the opt-out when an off home passes an attended wake to main

A host parked when the home writes off now logs that the home does not run
the supervision host, rather than claiming it has no engine.

* no-mistakes(document): Clarify Claude supervision defaults and historical evidence

* no-mistakes(ci): Fixed process leaks in the two added host tests. Each case now stops its recorded watcher and host/arm processes; fake hook sessions exit through session.stop. The full host suite passed before the final cleanup refinement, and both affected cases, bash syntax, ShellCheck, and diff checks passed afterward. CI runtime still needs confirmation
…start scope check (kunchenguid#6125)

* fix(bin): create the state dir on a fresh primary before the session-start scope check

fm_primary_scope_matches required an already-existing state directory, so
bin/fm-sessionstart-run.sh stood down on a fresh clone before anything could
create it. Split out fm_primary_root_matches so the run wrapper can confirm
primary-home identity first, create the gitignored state dir when it is
missing, and only then run the unchanged scope check.

* no-mistakes(document): Document session-start state dir creation on fresh clones

* no-mistakes(ci): I fixed the Greptile P1 the way you asked. When a fresh primary can't create `state/`, the run wrapper no longer stands down silently. **Invariant:** when an otherwise eligible fresh primary cannot create `state/`, startup must never fail silently. This path has only one site: the mkdir in `bin/fm-sessionstart-run.sh`. Other hooks and the nudge wrapper never create `state/`, so they have no equivalent failure. **What changed:** - **Run wrapper** (`bin/fm-sessionstart-run.sh`): it captures mkdir's error and prints one line to stderr before standing down as before (exit 0, or 3 for the Pi prerequisite). The line looks like `fm-sessionstart-run: startup could not create the state directory <path>: <reason>`. - **Test** (`tests/fm-sessionstart-nudge.test.sh`): the new case `test_run_reports_a_state_dir_it_cannot_create` uses a fresh primary with no `state/` and a read-only (0500) root. It checks four things: exit 0, no digest on stdout, no state dir created, and exactly one stderr line ending in "Permission denied". It fails without the fix and passes with it. - **Docs** (`docs/sessionstart-nudge.md`): I added one sentence describing the stderr line and one describing what the new test proves. **Verification:** I ran `tests/fm-sessionstart-nudge.test.sh`, and every test passes. `bin/fm-lint.sh` on the changed scripts (pinned ShellCheck 0.11.0) and `tests/fm-documentation-audiences.test.sh` also pass. As you asked, the wrapper still stands down with the ineligible-checkout status afterwards. It does not report this as a failed eligible startup, which is what the bot suggested
…ery (kunchenguid#6126)

* fix(bin): measure pending-reply grace from turn completion, not delivery

Fixes kunchenguid#6057

The pending-reply guard demanded a repost ("REPOST REQUIRED: previous
marked request had no correlated parent report") while the second
mate's correlated reply was already on its way.
fm_pending_reply_send_recovery measured its grace window from delivery
instead of from the request turn's completion, so any turn longer than
the grace fired the demand the moment the turn ended, before the reply
could have landed. The missed-report escalation had the same gap: it
fired the instant the recovery turn's completion was observed, with no
grace at all.

Both now measure grace from the relevant turn's completion (request
turn for the recovery repost, recovery turn for the escalation), and
both take one fresh, uncached read of the parent status file
immediately before firing, accepting a correlated line regardless of
its verb. Transport-failure escalations stay immediate, and the
one-repost limit is unchanged.

* no-mistakes(review): Document grace window as measured from turn completion

* no-mistakes(ci): Both Greptile findings were real and caused by this PR, so I fixed them. The full `tests/fm-pending-reply.test.sh` suite passes. **ci-1 (a reply could be overwritten by a repost).** The rule that must hold: a recovery send is recorded only if the record is still unresolved, checked under the same per-correlation lock that resolution uses. The escalation path already did this (`_fm_pending_reply_maybe_escalate_locked` reads fresh and publishes under one lock). The recovery path did not: `fm_pending_reply_send_recovery` did its fresh read through `fm_pending_reply_try_resolve`, which let go of the lock before the send was recorded. A reply landing in that gap could be overwritten, and the repost would go out anyway. Now `send_recovery` takes the lock once and, while holding it, re-checks that the phase is still `awaiting_report`, runs the fresh uncached read, and records the send (sender pid and identity, attempt time, phase `recovery_sending`). It releases the lock before actually sending, so the lock is not held during the send. It uses the same lock helpers the other lock wrappers use. Grace timing, the one-repost limit and the escalation path are unchanged. **ci-2 (the test would pass even without the fix).** In `test_recovery_fresh_status_read_resolves_before_firing`, the reply is still appended to the status file, but the stored file signature is then set to the file's new signature. That stands in for a same-size rewrite that the signature cache cannot see. The test first checks that a normal cached read misses the reply, then that the fresh read before sending catches it. I also added the same check for the fresh read before escalation, which the review said was uncovered. The test now sets its own send hook, so it no longer depends on one left over from an earlier test (that leftover had made failures exit silently). **Checks:** - I removed the fresh-read bypass at each site in turn and reran the suite. With it gone from recovery, the test fails with "recovery must not fire once a correlated reply has landed". With it gone from escalation, it fails with "the fresh pre-escalation read should have resolved the record, got escalated". With both in place, all tests pass. - Shellcheck with `-x` timed out locally. Without `-x` and ignoring SC1091, the only warnings are SC2034 on the existing `maybe_escalate` lock wrapper, which is not part of this change. The new code adds no warnings. Changes are in `bin/fm-pending-reply-lib.sh` and `tests/fm-pending-reply.test.sh`. Nothing is committed yet; a plain commit message such as "fix(bin): record the pending-reply recovery send under the fresh-read lock" fits the instruction

* no-mistakes(ci): ci-1 was real and caused by this PR. The same bug was also in the escalation path, so both are fixed. The full tests/fm-pending-reply.test.sh suite passes. The rule that must hold: a recovery repost or an escalation goes out only if the record's phase, read after the fresh-read resolve, is still what it was before. The resolver writes phase=resolved first and only then writes the other resolution fields. If one of those later writes fails, it returns an error even though the record is already resolved. Places this rule applies, both fixed: - Recovery (fm_pending_reply_send_recovery): the fresh-read resolve now runs first, and the phase is re-read right after it, whatever it returned. The send is recorded and made only if the phase is still exactly awaiting_report. This replaces the earlier phase check rather than adding a second one. - Escalation (_fm_pending_reply_maybe_escalate_locked): same bug. After a failed resolve it went on to publish the blocked line and set phase=escalated. One added line after the resolve call returns 1 without publishing if the phase has changed. Test: added test_partial_resolve_write_blocks_firing. It forces a failure on the resolved_epoch write after a correlated reply has landed. It checks that the recovery send hook is never called, that no escalation line is published, and that the phase stays resolved. The forced failure runs in a subshell so it can't affect later tests. Checks: - With the recovery fix reverted, the new test fails with "recovery must not fire after a partial resolve". - With the escalation fix reverted, it fails with "partial resolve should block escalation, got escalated". - With both fixes in, every test passes. - Shellcheck was run with SC1091 excluded and without -x, not through the repo's lint script. The only new message is one SC2329 info on the test's override function; other test overrides in the same file already get that same info, unsuppressed. Changed files: bin/fm-pending-reply-lib.sh and tests/fm-pending-reply.test.sh. Nothing is committed. Suggested plain commit message: "fix(bin): recheck pending-reply phase after the fresh read before sending
@andrewesweet
andrewesweet force-pushed the fm/fm-upstream-sync-0929 branch from d23deaf to 3adcb16 Compare September 30, 2026 07:13
andrewesweet and others added 3 commits September 30, 2026 08:55
…es; both fixed and verified locally. Lint 1 (SC2329, tests/fm-session-start.test.sh): the merge left two copies of the same block. Lines 561-603 re-defined make_fake_herdr_deadly_read and make_fake_herdr_hanging_read, and lines 1667-1826 re-defined test_endpoint_read_death_is_isolated_and_reported, test_endpoint_read_hang_is_bounded_and_reported, test_perl_timeout_fallback_reports_signal_death_nonzero and test_abnormal_digest_death_banners_and_exits_zero. Invariant: each helper and case is defined exactly once, and the definition that runs is the one the suite asserts against. The later (upstream) copies shadowed the earlier ones and dropped the fork's per-read bound contract (fixed 10s wording instead of FM_SESSION_START_ENDPOINT_TIMEOUT) plus the /proc and perl portability skips, while the surviving-only-once test_endpoint_bound_rejects_padded_zero still pins the configurable bound that bin/fm-session-start.sh:392 and docs/configuration.md:2400 document. Removed the shadowing duplicate blocks (removal-first), keeping the fork copies that match bin and docs. Swept every tracked shell file for duplicate function and duplicate test-registration names; no other site. Behavior portable serial 3 (tests/fm-spawn-orca-worktree.test.sh): "not ok - an Orca-backed spawn should succeed", because bin/fm-spawn.sh refuses a brief with no "## Published intent" subsection. Invariant: every brief fixture spawned through the fork's three-part brief contract carries Captain's intent, Published intent, and Firstmate spec. Both briefs in that upstream-shaped test lacked the middle section; added it to both (the second case refuses earlier on the classifier, but the same invariant holds there). Swept the other test briefs: only fm-public-followup, fm-dispatch-resolve, fm-fleet-ledger and fm-dod-lib omit it, and those exercise the missing-subsection path itself. Verification: tests/fm-spawn-orca-worktree.test.sh passes (2/2), tests/fm-session-start.test.sh passes ("all assertions passed", exit 0), shellcheck -x clean on both files
…eal, flaky test defect in tests/fm-backend-herdr.test.sh. The CI log line hidden in the truncated section is `not ok - container_ensure did not start the herdr server (missing: 'HERDR_SESSION=fmtest^_server')`, and the dumped log shows two invocations' appends spliced together: `HERDR_SESSION=fmtestHERDR_SESSION=fmtest^_status^_--json^_--session^_fmtest^_server` followed by an orphan `^_--session^_fmtest`. Root cause: bin/backends/herdr.sh:1677 backgrounds the server launch (`fm_backend_herdr_cli "$session" server >/dev/null 2>&1 &`) and then polls `status --json` in the foreground, so two fake-herdr processes append to $FM_HERDR_LOG at the same time. Each fake wrote its line with three separate printfs inside a `{ ... } >> "$LOG"` group, i.e. three write() calls per line, so the concurrent writes interleaved and neither the server line nor the poll line survived intact. Deterministic in CI under load, rare locally. Invariant: every fake-herdr invocation appends exactly one complete line to the log, atomically, even when a backgrounded call runs concurrently with a foreground one. Sites where that invariant must hold: all three herdr fake binaries in this file that log the unit-separated form - make_herdr_fakebin (line ~55), make_herdr_statefake (line ~189), make_herdr_eventfake (line ~5522). Fixed all three with the same small edit: build the line into one variable, then a single `printf '%s\n' "$fm_log_line" >> "$LOG"`, which is one O_APPEND write. Log format is byte-identical, so every assert_contains/assert_not_contains consumer reads exactly what it did before. Swept the other FM_HERDR_LOG consumer (tests/fm-send-strict.test.sh:72) - already a single printf, no change needed. The tmux/cmux/zellij/orca fakes keep the multi-printf form; none of those backends background a CLI call while another runs, so the race is not reachable there and no edit was made. Verification: `shellcheck -x tests/fm-backend-herdr.test.sh` clean; tests/fm-backend-herdr.test.sh run three times back to back, each exit=0 with 229 ok assertions and no `not ok`
Portable serial shard 5 was cancelled at the 30-minute CI job cap with every
test green, while the other eight shards finished. The packing estimate said
all nine shards carried 732s, so the imbalance was invisible: several merged
scripts had outgrown their hints, and the 26 scripts the merge added had no
hint at all and were packed on PORTABLE_SERIAL_DEFAULT_WEIGHT_MS.

Weighed against a local serial run of the whole lane, shard 5's real load was
813s against that 732s estimate, driven by tests/fm-contributions.test.sh at
134885 ms against a 35676 ms hint, tests/fm-live-lab.test.sh at 74859 ms on
the 27000 ms default, tests/fm-bootstrap.test.sh at 56609 ms against 46634,
and tests/fm-branch-eligibility.test.sh at 5036 ms against 950.

Refreshed portable_serial_weight_hints through its owner bin/fm-test-run.sh,
the documented duration-balancing mechanism: every measured script keeps the
slower of its CI-derived and locally measured sample, so no CI evidence is
lowered. Across the 36 scripts with both samples the local run totalled 1.11x
the CI hints, so the two are comparable in scale. The existing
longest-processing-time packing then repacks membership on its own: all nine
shards now model 714320-714354 ms, and --check-coverage reports
serial_unhinted=0 rather than 26. No timeout-minutes change, no skipped test,
and the partition stays a complete disjoint cover.

Verification: bin/fm-test-run.sh --check-coverage reports
serial=223 serial_shards=9 serial_unhinted=0; tests/fm-test-run.test.sh
passes every shard assertion, including the union, disjointness, coverage
guard and balance cases. Its two `not ok` lines are
tests/fm-session-lock-ancestry.test.sh failing "the pty-host was not
reparented to init after the daemon ended", which reproduces identically when
that script is run alone on this WSL2 host before and after this change, so it
is environmental and untouched here. tests/fm-ci-workflow.test.sh and
tests/fm-documentation-audiences.test.sh pass; shellcheck -x is clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@andrewesweet
andrewesweet merged commit 2bec13f into main Sep 30, 2026
19 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.