Skip to content

fix(pi): route decision-owned wake batches to main - #3776

Merged
kunchenguid merged 17 commits into
mainfrom
fm/fm-needs-decision-clean-r1
Sep 5, 2026
Merged

kunchenguid merged 17 commits into
mainfrom
fm/fm-needs-decision-clean-r1

Conversation

@kunchenguid

@kunchenguid kunchenguid commented Sep 5, 2026 •

Copy link
Copy Markdown
Owner

Intent

Every needs-decision event must skip the supervision branch and come straight to the main session - not only no-mistakes ask-user findings, but every needs-decision: product calls, destructive choices, captain holds, and second-mate escalations.

The accepted routing rule this change implements: when ONE COALESCED signal or stale trigger batch contains any needs-decision row, deliver the ENTIRE batch to main. Do not split that same batch between the supervision branch and a later main wake. This whole-batch routing is the replacement for an earlier per-row approach that split a mixed batch, reserving routine rows for the supervision branch while sending only the needs-decision rows to main; that earlier approach is superseded and must not be reintroduced.

Ordinary task-local heartbeat, signal, and stale routing stays unchanged, and the implementation stays narrowly limited to the needs-decision signal path.

What Changed

  • Classify needs-decision, captain-held, and second-mate escalation signals as main-owned events.
  • Route an entire coalesced signal or stale trigger batch to main when it contains a decision-owned task, including task aliases and unresolved open decisions.
  • Preserve independent routing for unrelated task-local rows and heartbeat reviews, with expanded regression coverage and documentation.

Risk Assessment

✅ Low: The routing changes are narrowly scoped, preserve heartbeat and routine task behavior, and the reviewed signal, stale, alias, cache, and symlink paths satisfy the stated invariants without a substantiated defect.

Testing

The successful baseline was supplemented with focused watcher, branch-dispatch, and Pi-extension tests covering needs-decision, captain-held, second-mate escalation, stale aliases, mixed batches, routine routing, and heartbeat independence. An executable integration artifact confirms a mixed batch is rejected by supervision and delivered intact to main; one unrelated stock-renderer compatibility case skipped because installed Pi 0.80.10 predates 0.84.4.

Evidence: Whole-batch needs-decision routing through the Pi watcher interface

Source: Whole-batch needs-decision routing through the Pi watcher interface

{
  "scenario": "coalesced signal for task-a (needs decision) plus task-b (routine)",
  "queuedRows": [
    "needs-decision: task-a.status",
    "signal: task-b.status"
  ],
  "supervisionBranchOffer": {
    "eligible": false,
    "message": "signal: task-a.status task-b.status",
    "projects": [
      "~/.no-mistakes/worktrees/016d88035d58/01M1R51AGSZ8VZ3NY8XG9T9QXD/.test-evidence-tmp/home/projects/approved"
    ]
  },
  "mainSessionWake": "⁣FIRSTMATE_OP: v1 watcher: FIRSTMATE WATCHER WAKE: signal: task-a.status task-b.status\n\nRun bin/fm-wake-drain.sh first and handle the queued wake. Watcher continuity is extension-owned.",
  "observedWholeBatch": true,
  "result": "PASS"
}

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 1 issue found → auto-fixed (9) ✅
  • 🚨 .pi/extensions/lib/fm-branch-dispatch.ts:202 - Intent requires “every needs-decision” and any stale trigger batch containing one to reach main, but stale classification only recognizes a final captain-held line. With a mapped task whose status ends needs-decision: choose destructive cleanup and a queued stale row, this check does not populate needsDecisionKeys; scope.eligible remains true and offerWakeToBranch can send the wake to supervision. The same bypass occurs when an earlier durable decision remains open behind a later unrelated status line. Classify stale rows using the durable open-decision semantics, while retaining captain-held detection, at this shared boundary.

🔧 Fix: Route stale open decisions directly to main
2 issues (1 error, 1 warning) still open:

  • 🚨 bin/fm-classify-lib.sh:1657 - Intent explicitly requires second-mate escalations to reach main, but fm_pending_reply_maybe_escalate writes them as blocked [key=pending-reply-…] (bin/fm-pending-reply-lib.sh:1239), while the changed classifier sets the main-only marker only when the verb equals needs-decision. The resulting signal row retains a signal: payload, remains branch-eligible, and offerWakeToBranch can deliver it to supervision. Extend the narrowly scoped decision signal classification to recognize this concrete second-mate escalation class and route its whole coalesced batch to main.
  • ⚠️ .pi/extensions/lib/fm-branch-dispatch.ts:151 - The new stale-path fold hardcodes resolved and captain-held, despite the authoritative status fold supporting FM_CLASSIFY_RESOLVE_VERB and FM_CLASSIFY_CAPTAIN_HELD_VERB. For example, with a configured resolution verb, needs-decision [key=x] followed by that resolution remains falsely open here and routes an ordinary stale wake to main; a configured captain-hold verb can instead be missed and offered to the branch. Use the configured verbs so stale routing agrees with the repository's durable fold semantics.

🔧 Fix: Honor configured verbs in stale decision routing
2 errors still open:

  • 🚨 bin/fm-classify-lib.sh:1657 - Intent requires “every needs-decision” including “second-mate escalations” to reach main. However, fm_pending_reply_maybe_escalate writes these as blocked [key=pending-reply-…] (bin/fm-pending-reply-lib.sh:1239), while the changed classifier marks only the needs-decision verb. The resulting signal retains a signal: payload, remains branch-eligible, and can be delivered to supervision. Extend the decision-owned signal classification to recognize this concrete escalation class and route its coalesced batch wholly to main.
  • 🚨 .pi/extensions/lib/fm-branch-dispatch.ts:155 - The new stale decision fold hardcodes pending-reply- as reserved instead of honoring FM_CLASSIFY_RESERVED_KEY_PREFIXES, unlike the authoritative shell fold. For example, with FM_CLASSIFY_RESERVED_KEY_PREFIXES=secret-, needs-decision [key=pending-reply-x]: choose cleanup is a valid open decision in the authoritative fold, but this code skips it and can offer its stale wake to supervision. Make this fold use the configured reserved-prefix semantics so every supported open decision routes to main and ordinary stale routing remains consistent.

🔧 Fix: Route second-mate escalations and configured decisions to main
1 error still open:

  • 🚨 .pi/extensions/lib/fm-branch-dispatch.ts:245 - Intent requires captain holds to reach main, but status parsing removes only empty strings, not whitespace-only lines. For captain-held [key=route]: awaiting captain\n \n, hasOpenNeedsDecision treats the hold as closing the decision and the final-line check examines " ", so the stale row remains branch-eligible. This contradicts the repository’s last_status_line semantics, which ignore whitespace-only lines. Filter status lines by non-whitespace content before selecting the current declaration.

🔧 Fix: Ignore trailing whitespace after captain holds
1 warning still open:

  • ⚠️ .pi/extensions/lib/fm-branch-dispatch.ts:244 - Each queued stale row rereads and refolds its task's entire append-only status log synchronously. Multiple unread reminders for one task therefore multiply work on Pi's UI thread; for example, 50 stale rows for a 20 MB status log parse roughly 1 GB per eligibility check without changing the result. Cache the decision classification per status path within scopeForUnreadWake so each task is read and folded at most once per scan.

🔧 Fix: Cache stale decision classification per status file
1 error still open:

  • 🚨 .pi/extensions/fm-primary-pi-watch.ts:622 - Intent requires ordinary signal routing to remain unchanged and whole-batch main routing only when that coalesced batch contains a needs-decision row. This key-only cross-reference also matches older unread rows. If task-a's needs-decision row remains queued and a later routine signal batch for task-a arrives, scope.needsDecisionKeys still contains task-a.status, so the later routine batch is incorrectly forced to main without error. Distinguishing queue history from the triggering batch may require a batch identity or equivalent protocol change, so the remedy needs user authorization.

🔧 Fix: Document unread decision precedence for later task wakes
1 warning still open:

  • ⚠️ .pi/extensions/lib/fm-branch-dispatch.ts:253 - The per-scan cache still rereads each unread stale row's entire append-only status log on every scopeForUnreadWake call. Concretely, one unread stale decision for a task with a 20 MB status log makes every later routine signal or heartbeat synchronously read and fold that 20 MB file on Pi's UI thread, even when the file is unchanged. Cache classifications across scans using file identity/size/mtime invalidation, preserving fail-closed behavior when metadata or reads fail.

🔧 Fix: Cache unchanged stale decisions across scope scans
2 errors still open:

  • 🚨 .pi/extensions/fm-primary-pi-watch.ts:624 - The documented precedence says an unread decision forces every later signal or stale trigger for the same task to main, but this compares raw queue keys rather than task identity. For task-a, an unread decision signal is keyed task-a.status; after the status is resolved, a later stale trigger is keyed fm-window. The stale fold no longer marks the resolved status, the keys do not match, and the stale batch is offered to supervision. Preserve the approved same-task precedence across status-file and window aliases.
  • 🚨 .pi/extensions/lib/fm-branch-dispatch.ts:157 - The new stale classifier uses statSync and readFileSync, both of which follow task.status symlinks. A mapped stale row can therefore make the Pi extension synchronously read an arbitrary symlink target and let that external content control main-versus-branch routing. This violates the repository's authoritative status-fold boundary, which explicitly rejects status-file symlinks in status_open_decisions. Reject symlinks with lstatSync before versioning or reading, failing the scope closed.

🔧 Fix: Resolve decision aliases and reject symlinked statuses
1 error still open:

  • 🚨 bin/fm-watch.sh:1771 - Intent requires every captain hold to reach main, but signal payloads are marked main-only only when FM_SIGNAL_NEEDS_DECISION_FILES contains the file. status_is_captain_relevant excludes captain-held, so a captain-held append surfaced through the documented no-verb fallback (for example after the crew ends without positive working evidence) is queued with an ordinary signal: payload. scopeForUnreadWake then treats that signal row as branch-eligible and the captain hold can enter supervision. Extend the main-only marker to surfaced captain-held signal files and cover that executable path.

🔧 Fix: Route surfaced captain-held signals directly to main
✅ Re-checked - no issues remain.

✅ **Test** - passed

✅ No issues found.

  • bin/fm-test-run.sh --changed --exclude-family real-herdr-gated
  • Baseline: bin/fm-test-run.sh --changed --exclude-family real-herdr-gated
  • bin/fm-test-run.sh tests/fm-watch-triage.test.sh tests/fm-pi-branch-extension.test.sh tests/fm-pi-watch-extension.test.sh
  • Manual executable Pi integration: queued one needs-decision signal and one routine signal in the same coalesced batch, captured the supervision offer and main-session wake, and verified whole-batch main routing
  • Verified testing left the git worktree clean and removed .test-evidence-tmp
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

kunchenguid and others added 15 commits September 4, 2026 23:32
Skip the supervision branch for every needs-decision status append, the
same way a check-kind wake already skips it. A coalesced signal/stale
trigger batch containing any needs-decision row is delivered wholly to
main, not split between the branch and a later main wake - the whole
batch, including any co-present routine rows for a different task,
travels together. Heartbeat and unread-status scans stay independent.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE
Add a regression using two distinct files (not the same status file
twice) in one coalesced trigger so a some-vs-every regression on the
file-list cross-reference cannot hide behind a degenerate same-key
case, and a heartbeat/needs-decision co-presence test proving a
needs-decision row neither vetoes nor rides along with an otherwise
eligible heartbeat scan.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE
…pervision and wake main, while unrelated unread rows and heartbeats remain independent. Added routing regressions and updated documentation. Pi watcher tests, strict TypeScript checks, lint, and diff checks pass. The no-mistakes attestation failure was pipeline-state related, not a source defect
… SC2031 diagnostics where background PIDs are captured immediately in the same shell. Verified with `CI=true bin/fm-lint.sh` and `git diff --check`. The no-mistakes attestation failure is pipeline-state-related (`test` was skipped), not a source defect
@greptile-apps

greptile-apps Bot commented Sep 5, 2026 •

Copy link
Copy Markdown

Confidence Score: 5/5

The PR appears safe to merge.

No blocking failure remains.

Reviews (3): Last reviewed commit: "no-mistakes(ci): Fixed the CI regression..." | Re-trigger Greptile

Comment thread bin/fm-classify-lib.sh
…crew working evidence is positive, ensuring the watcher delivers their main-only marker. Updated the executable regression test to cover this case. Verified with the full fm-watch-triage suite, bash syntax checks, and git diff checks. Shellcheck reported only pre-existing test harness warnings (SC1091/SC2034)
…retain their established non-actionable stale classification while the signal-routing side-band still surfaces them main-only. Verified with tests/fm-daemon.test.sh, tests/fm-watch-triage.test.sh, bash syntax checks, and git diff --check
@kunchenguid
kunchenguid merged commit d6660d7 into main Sep 5, 2026
16 checks passed
@kunchenguid
kunchenguid deleted the fm/fm-needs-decision-clean-r1 branch September 5, 2026 21:28
RooseveltAdvisors added a commit to RooseveltAdvisors/firstmate that referenced this pull request Sep 5, 2026
Brings in upstream d6660d7 (kunchenguid#3776 pi decision-owned wake batches) and
a29cdce (kunchenguid#3778 reduce local ShellCheck source-analysis cost) so this
branch stops conflicting with the default branch.

Two conflicts, both resolved by keeping the upstream change and layering
this branch's addition on top:

- CONTRIBUTING.md: kept kunchenguid#3778's "invoke it with no arguments" wording and
  its "file-set selection, and analysis flags" sentence, and re-applied
  this branch's backend-purity clause inside the lint-definition list.
- bin/fm-lint.sh: kept kunchenguid#3778's LOCAL_NOX_EXCLUDE definition and its header
  sentence about dropping --external-sources locally, and re-applied this
  branch's backend-purity note plus the pwd -P resolution of SELF_DIR and
  ROOT. LOCAL_NOX_EXCLUDE is consumed at the local ShellCheck call site, so
  dropping it would have broken kunchenguid#3778.

Verified: the diff against origin/main for every file those two upstream
commits touched contains only this branch's additions, no upstream
deletions. CI=true bin/fm-lint.sh exits 0 (ShellCheck 0.11.0 full
analysis, backend-purity, actionlint 1.7.12), and
bin/fm-test-run.sh tests/fm-lint.test.sh tests/fm-backlog-atomicity.test.sh
exits 0 with 0 failures.

Additive merge commit: no rebase, no squash, no history rewritten.
zeeshaanahmad added a commit to zeeshaanahmad/firstmate that referenced this pull request Sep 6, 2026
…able serial 1" (specifically tests/fm-afk-inject-e2e.test.sh Scenario E: "the digest was cut but carries no truncation marker"). Root cause: this PR merges upstream batch 12, including commit d6660d7 (kunchenguid#3776, "route decision-owned wake batches to main (Pi)"). That commit changed bin/fm-watch.sh so a decision-owned wake-queue row's payload is deliberately marked "needs-decision:$files" instead of "signal:$files", solely so Pi's branch dispatcher (fm-branch-dispatch.ts) excludes that row from what it may claim (per docs/pi-supervision-branch.md). This fork's own bin/fm-supervise-daemon.sh (handle_wake) independently drains the same durable wake queue and classifies each row by pattern-matching the payload's literal prefix (signal:/stale:/check:/heartbeat) — it had no case for "needs-decision:", so any decision-owned row fell through to classify_unknown, producing a short "unknown wake: needs-decision: <raw file paths>" digest instead of running the real per-status note text through classify_signal. Scenario E creates 12 simultaneous decision-owned statuses; the real distilled text is long and must arrive truncated with a marker, but the misclassified path produced a short, unmarked message, failing the assertion. Confirmed via `gh run view --log` (the actual CI failure line) and via code tracing (fm-watch.sh:1875-1890, fm-supervise-daemon.sh handle_wake/handle_durable_wakes) — a genuine merge-consequence regression, not flakiness or a stale/infra check; none of the touched files appear elsewhere in this PR's diff. Fix (bin/fm-supervise-daemon.sh, handle_wake): added a `needs-decision:*` case mirroring `signal:*` — strips the prefix and routes to classify_signal, setting kind=signal — so main still classifies these rows exactly like ordinary signal rows; only Pi's branch-claim exclusion (the actual intent of the upstream change) is unaffected. No other files changed. Verification (completed): `bash -n` passed. Full local run of tests/fm-afk-inject-e2e.test.sh now passes all 5 scenarios (A-E), including Scenario E which previously failed. Full local run of tests/fm-daemon.test.sh (the unit suite covering handle_wake and related daemon logic) passed 126/126 assertions with 0 failures, confirming no regression
lytv pushed a commit to lytv/mymate that referenced this pull request Sep 8, 2026
* fix(pi): route needs-decision wakes and mixed batches wholly to main

Skip the supervision branch for every needs-decision status append, the
same way a check-kind wake already skips it. A coalesced signal/stale
trigger batch containing any needs-decision row is delivered wholly to
main, not split between the branch and a later main wake - the whole
batch, including any co-present routine rows for a different task,
travels together. Heartbeat and unread-status scans stay independent.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE

* test(pi): cover distinct-file mixed batches and heartbeat independence

Add a regression using two distinct files (not the same status file
twice) in one coalesced trigger so a some-vs-every regression on the
file-list cross-reference cannot hide behind a degenerate same-key
case, and a heartbeat/needs-decision co-presence test proving a
needs-decision row neither vetoes nor rides along with an otherwise
eligible heartbeat scan.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE

* no-mistakes(document): Clarify needs-decision and heartbeat routing

* no-mistakes(ci): Fixed captain-held stale reminders so they bypass supervision and wake main, while unrelated unread rows and heartbeats remain independent. Added routing regressions and updated documentation. Pi watcher tests, strict TypeScript checks, lint, and diff checks pass. The no-mistakes attestation failure was pipeline-state related, not a source defect

* no-mistakes(ci): Fixed CI lint by narrowly suppressing false-positive SC2031 diagnostics where background PIDs are captured immediately in the same shell. Verified with `CI=true bin/fm-lint.sh` and `git diff --check`. The no-mistakes attestation failure is pipeline-state-related (`test` was skipped), not a source defect

* no-mistakes(review): Route stale open decisions directly to main

* no-mistakes(review): Honor configured verbs in stale decision routing

* no-mistakes(review): Route second-mate escalations and configured decisions to main

* no-mistakes(review): Ignore trailing whitespace after captain holds

* no-mistakes(review): Cache stale decision classification per status file

* no-mistakes(review): Document unread decision precedence for later task wakes

* no-mistakes(review): Cache unchanged stale decisions across scope scans

* no-mistakes(review): Resolve decision aliases and reject symlinked statuses

* no-mistakes(review): Route surfaced captain-held signals directly to main

* no-mistakes(document): Document decision-owned main routing

* no-mistakes(ci): Fixed captain-held spans to remain actionable while crew working evidence is positive, ensuring the watcher delivers their main-only marker. Updated the executable regression test to cover this case. Verified with the full fm-watch-triage suite, bash syntax checks, and git diff checks. Shellcheck reported only pre-existing test harness warnings (SC1091/SC2034)

* no-mistakes(ci): Fixed the CI regression: captain-held transfers now retain their established non-actionable stale classification while the signal-routing side-band still surfaces them main-only. Verified with tests/fm-daemon.test.sh, tests/fm-watch-triage.test.sh, bash syntax checks, and git diff --check

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
dnth added a commit to dnth/firstmate that referenced this pull request Sep 9, 2026
* fix(bin): derive watcher beacon staleness grace from poll cadence

Port of upstream kunchenguid/firstmate kunchenguid#3946, adapted: the fork has no
away-mode daemon predicate, so only the shared derivation and its two
in-scope readers change.

fm-claude-stop-autoarm.sh required the watcher beacon fresh within the
flat FM_GUARD_GRACE default (300s), but a healthy watcher touches the
beacon once per cycle and can legitimately age up to FM_POLL seconds
between touches, so a long-poll home reads stale mid-wait and the hook
misjudges a live cycle as down.

fm-wake-lib.sh gains fm_poll_derived_grace, the single owner of the
max(300, FM_POLL + 60) formula. fm-watch.sh's own pre-acquisition
beacon-staleness check and the auto-arm hook both derive their default
grace from it, and the hook now exports its resolved FM_GUARD_GRACE to
the fm-watch-arm.sh invocations so arm and check agree on one value.

* fix(bin): count only actor-presentable wake rows and retire unusable rows

Port of upstream kunchenguid/firstmate kunchenguid#3950, adapted to the fork's
SCOPED drain and OMP branch-grant machinery.

A branch grant already excluded its rows from a main drain, but
fm-guard.sh still counted every queued row as pending for the calling
actor, so main was told to run a drain that could only print nothing for
as long as the branch held the grant. And a queue row that lost the five
appended fields or its numeric sequence was counted as queued while it
could never be claimed, presented, or named by an --ack-through cutoff,
wedging the queue permanently.

fm-wake-lib.sh gains the shared per-actor helpers:
fm_wake_grant_rows_valid, fm_wake_branch_owner_matches,
fm_wake_branch_grant_live, and fm_wake_actor_pending_count - the last
reports a pending row when a queue exists but cannot be counted, so an
unreadable queue still raises the alarm. fm-wake-drain.sh and
fm-wake-grant.sh now delegate their grant row-list and owner-record
reads to the library. A main drain retires structurally unusable rows
under the queue lock and reports them verbatim (a branch drain never
does), and a main drain left with only branch-held rows names the
holder instead of exiting silently. fm-guard.sh warns only for rows the
calling actor can itself present or retire, and gives main a distinct
held-by-branch advisory when that is the whole non-empty queue.

tests/fm-wake-queue.test.sh covers the branch-held warning boundary end
to end, the uncountable-queue alarm, and main-only retirement.

* fix(bin): gate secondmate wake-loop stall alerts on real queue no-progress

Port of upstream kunchenguid/firstmate kunchenguid#3943, which lands the detector's
final form on this fork (the fork had no foreign-queue stall check at
all, so the whole feature arrives already fixed).

The watcher now reads the oldest structurally valid actionable row in
every endpoint-recorded local secondmate home's durable wake queue, and
times the interval since that position last changed instead of the row's
age. A queue that is draining is not stalled, so a moved oldest row -
drain progress, or a queue reprovisioned under the same task id that
restarts its sequence anywhere - ends the no-progress episode and starts
a fresh observation interval, recorded by
state/.secondmate-wake-progress-<task> as the same epoch-sequence row
identity the stall receipts use. Declared external-wait pause rows are
not actionable evidence, and a mate provably inside an active turn (an
exact busy verdict bounded by FM_BUSY_TURN_MAX_SECS) never escalates.
One keyed check notification covers each no-progress episode across
watcher and handling crashes; the foreign queue is only ever read.

FM_SECONDMATE_WAKE_STALL_SECS defaults to 180 and is the backstop behind
the active-turn gate, not a substitute for it.

Fork adaptation: the parent watcher keeps its secondmate home-beacon
rule, so the ported tests give the fixture mate a fresh
.last-watcher-beat under its (stubbed or real) clock to isolate the
stall detector; assertions are unchanged from upstream.

tests/fm-wake-queue.test.sh covers progress tracking, once-per-episode
alerting, declared-pause exclusion, reprovisioned-generation reset,
active-turn deferral, and marker symlink rejection.

* fix(bin): never type doorbells into dead panes; recover them instead

Port of upstream kunchenguid#3823, adapted: the fork's task-inbox ring reserves
return code 6 for a positively dead or missing endpoint (codes 3-5 are
OMP-native and Hermes-lock outcomes here). The doorbell line itself is
now a shell no-op (": ...") with quoted, printable-only paths, so even a
liveness race that loses to an exiting agent runs nothing in a bare
shell. The watcher escalates a dead/missing endpoint's unhandled record
once through the ordinary stale wake - over stale busy-state evidence -
instead of walking the re-ring ladder, and fm-send reports the skipped
doorbell as recovery, not a re-ring. Remote secondmate send inherits the
rc-6 recovery notice through fm-send without a new branch. Secondmate
restart also resolves all arrived persist answers before any timeout
decision and rechecks at the decision boundary, so a reply landing at
the deadline wins over the fallback nudge.

Regression coverage: doorbell shell-noop across real shells, control-byte
path rejection, dead/missing ring skip, watcher escalate-once-without-
ringing, dead-pane override of stale busy state, and the
resolution/timeout boundary race. Fake tmux inventories across the suite
now list recorded windows and answer pane_current_command so the
recovery-grade endpoint check exercises real paths.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(omp): prove the doorbell turn instead of trusting triggerTurn

The task-inbox doorbell claimed `.delivered` the moment sendMessage
returned, but OMP can downgrade `triggerTurn` to append-only delivery on
an idle, non-streaming session: the steer was accepted and reported
received while no turn ever ran, leaving an idle worker stranded with an
unhandled durable record.

The extension now renames an accepted request to `.awaiting-turn` when no
turn is open and settles it to `.delivered` only on a `turn_start` or
`agent_start` proving the turn actually began. A steer that lands while a
turn is open is still delivered immediately, since it joins that turn.
When the bound passes with no turn, the extension re-drives the same
instruction through `sendUserMessage` - a user prompt an idle session
cannot defer - and a failed re-drive reports `.failed` rather than
stranding silently. Requests left `.awaiting-turn` by a dead generation
re-queue as `.pending` on the next activation. Installs on runtimes
without the event or user-prompt surface keep the prior accept-only
semantics.

On the shell side, `fm_omp_task_doorbell_request_existing` reports an
awaiting-turn request as in-flight, so fm-send honestly reports a
downgraded doorbell as queued (exit 4 semantics) rather than received
while the extension's own recovery still wakes the worker.

Regression coverage drives the downgrade end to end: an accepted
sendMessage with no turn event stays unproven, the bounded grace re-drives
the identical instruction through sendUserMessage exactly once, a firing
turn_start settles without any re-drive, an open turn delivers the steer
immediately, a failing re-drive reports failed, and a dead generation's
unsettled proof re-enters delivery on activate. The suite's tmux fake also
now lists the recorded task window so the recovery-grade endpoint check
exercises the live path instead of classifying the fixture missing.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(bin): bound detached procevent runners to their owning home's lease

A runner is detached into its own process group so it survives the turn
that started it, and on its own that lets a runner outlive its whole
home: once reparented to init, nothing bounds its lifetime, so a blocking
source child - and everything it spawns - can poll forever after the home
that owns the work is gone. This is the upstream kunchenguid#3904 port, adapted to
the fork's simpler procevent (no extension subsystem or capture
reservations).

Every claim now records the canonical physical state root with its
device, inode, owner, and mode, and every claimed runner spawns a
detached owner guard in a separate process group. The guard re-reads the
recorded root's .owner-lease marker - refreshed by register, handled,
reconcile, sweep, and an attached start's keepalive - and after two
consecutive checks cannot prove both the root identity and a fresh lease,
it signals the runner's whole process group. Runner descendants inherit
FM_PROCEVENT_IN_RUNNER so they cannot keep their own lease alive; a
confused agent that deliberately strips the marker is out of scope.

Reconcile no longer signals a leaderless surviving group: once the leader
is gone, PID/PGID reuse makes the group's provenance unprovable, so the
claim stays owned and uncertain instead of risking a foreign group or a
second poller on one canonical source. Retirement and reconcile take
their signal through an identity-and-leadership-gated group path, and a
launch floor paces relaunches per registration generation so a
crash-looping source cannot spin.

Fork deviations from upstream: the state-root check requires ownership
and canonicality but not a private mode (operational homes use
group-readable state roots), the serialized launch-floor section releases
the source lock at the stamp write because this caller holds no further
locked work, central storage keeps its distinct commit point, and the
usage window is widened to keep the durability-boundary line visible.

Regression coverage: an orphaned runner's group is stopped by its guard
after the home's lease expires, a crashed leader's ambiguous group is
preserved rather than signalled or replaced, and the registration
replacement test now waits for the old generation's source to actually
start so the launch-floor supersede check cannot eat the race.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(omp): latch a provider-broken branch with a bounded cooldown probe

A settled branch prompt whose last assistant message carries stopReason
"error" is a provider failure OMP reports on the message instead of
rejecting prompt(), so the old code never latched: every later wake kept
paying for a predictably broken branch turn before its settlement could
reject to main. Port of upstream kunchenguid#3497, adapted to the fork's extension
shape (generation-only guards, offer.accept(promise) settlement).

Two consecutive settled provider errors now latch the branch off and
merge a captain-facing health note. While latched, main keeps every wake
except one recovery probe after each cooldown, which starts at five
minutes and doubles to a one-hour ceiling; a probe that ends in a durable
fm_branch_report clears the latch with a second note, and a failed probe
reschedules by the current cooldown. A settled prompt that produced no
durable outcome now rejects its settlement instead of silently claiming
its granted wake rows.

The branch record keeps its SessionManager so the settled-prompt scan can
read only the entries the prompt appended. The strict typecheck also
caught two stale errors from the earlier doorbell work: the doorbell
options builder returned a Required<> literal missing turnGraceMs, and
its `on` property signature could not accept the runtime's overloaded
event subscription - the latter is now a method declaration so the
parameter stays bivariant.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(omp): start a fresh supervision branch conversation per main session

Port of upstream kunchenguid#3600. A branch conversation that outlives its main
session keeps reasoning from an older thread's accumulated memory, which
competes with today's generated prompt and current fleet rules. The OMP
extension kept one persistent conversation keyed on a durable pointer,
so /new, /resume, /fork, and reload all reopened the stale thread.

The branch record is now scoped to the extension generation: createBranch
reopens only a conversation it built under the current generation, and
both session_start and session_switch advance the generation, drop the
live handle (beginDispose + dispose), and re-anchor the mirror so the
fresh branch receives the new main session's dialog from its start. The
durable outcome store, not the conversation, carries unacknowledged
outcomes across the boundary. The on-disk pointer remains a record for
operators and the effort picker's model lookup, not a reopen source.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(omp): give the supervision branch main's live model registry

Port of upstream kunchenguid#3871 by equivalent mechanism. A provider an extension
registered into main's runtime at run time is invisible to a freshly
discovered registry, so pinning or following such a model built a branch
that could not resolve its own model. OMP's createAgentSession already
takes a modelRegistry option, so instead of upstream's provider-by-provider
config copy into an isolated ModelRuntime, the branch session now receives
main's live registry directly - read-only sharing: the branch never
installs, converts, or overwrites credentials or registrations. The
resolve and picker paths already read that same registry, so the build,
the pin resolution, and the picker now agree on one model catalog.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(bin/omp): route decision-owned signal and stale wake batches to main

Port of upstream kunchenguid#3776. A needs-decision status append, a captain-held
status transfer surfaced through the no-verb signal path, or a blocked
pending-reply escalation is owned by the captain and must stay on main.
The previous logic only forced the check-kind class to main, so a
coalesced signal batch that also contained routine rows could end up split:
the decision row to the branch, the rest to main, which loses the
single-actor ordering the per-actor acknowledgement contract depends on.

The bash side now:

- `status_span_first_actionable_record` returns a side-band flag when its
  span carries a needs-decision, a captain-held declaration, or a pending-
  reply escalation, without changing the event text it already returns.
- `signal_files_actionable` surfaces those files (even when the captain-held
  line is otherwise non-actionable) and populates
  `FM_SIGNAL_NEEDS_DECISION_FILES`.
- The watcher appends the row payload as `needs-decision:$files` for exactly
  the flagged files; other files in the same batch keep the ordinary reason.

The OMP dispatch side now:

- `scopeForUnreadWake` parses each mapped task's `.status` log to detect an
  open needs-decision or a current captain-held declaration, and excludes
  those rows from `eligibleSeqs` while populating `needsDecisionKeys`.
- `offerWakeToBranch` in `fm-primary-omp.ts` cross-references the current
  signal or stale trigger keys against every unread decision-owned row by
  task identity: any overlap forces the entire coalesced batch to main.
- Heartbeat handling remains independent, and a decision-owned row in one
  task does not affect eligibility for other tasks.

Also fixed a pre-existing full-lint `SC2031` info note in
`tests/fm-afk-launch.test.sh` by disabling it file-wide, so the canonical
lint run is green again.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(core/omp/pi): settle watcher delivery on runtime acceptance, consume on turn start

Port of upstream kunchenguid#3513. A follow-up queued while main is streaming joins
that run as a user message without raising before_agent_start, so waiting
for the model to consume the wake stalled the successor chain: no successor
started, later wakes were not delivered, and the turn-end guard re-armed by
hand.

The shared core now:

- Treats sendFollowUp resolution as delivery; consumption is tracked only
  to let a session replacement replay a wake the runtime accepted but had
  not yet read.
- Keeps accepted-but-unconsumed wakes in a per-generation unconsumedWakes
  map keyed by the pending token.
- Consumes a wake when the runtime starts a turn carrying the exact wake
  text: before_agent_start for an idle main, message_start for a streaming
  main.
- Does not mark a pending row delivered or finish cleanup until the wake is
  consumed or the branch owns it.
- Skips over unconsumed rows when picking the next pending record, so the
  successor chain keeps moving while waiting for a streaming main to read
  the prior wake.
- Defers a verified successor's failure close when the pipeline is still
  delivering the wake it was started for; the retry runs once that delivery
  settles.
- Adds armRetired so a child the core itself killed does not earn a
  deferred retry.

The Pi and OMP adapters both gain a message_start handler that extracts the
user message text and passes it to watch.acknowledgeWake.

Documentation in docs/omp-supervision-branch.md now states the delivery
versus consumption boundary for main fall-back follow-ups.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* no-mistakes(review): Make supervision provider latch recoverable after transient failures

* no-mistakes(review): Correlate doorbells and latch rejected supervision prompts

* no-mistakes(review): Correlate doorbell turn proofs and recover prompt failures

* no-mistakes(review): Correlate synchronous turn proofs to each doorbell dispatch

* no-mistakes(review): Correlate turn proof and redrive uncorrelated doorbells

* no-mistakes(test): Fix procevent sweep claims and OMP turn handler coexistence

* no-mistakes(test): Preserve claim-only sources for reliable home sweeps

* no-mistakes(test): Disabled unreliable turn observation for generated OMP workers

* no-mistakes(document): Updated stale branch-session documentation

* no-mistakes(review): Enforce state-root identity before ownership and stale-claim decisions

* no-mistakes(review): Requeue unreadable awaiting-turn requests for redelivery

* no-mistakes(document): Correct stale supervision and process-event verification docs

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
BenWilcox8 pushed a commit to BenWilcox8/firstmate that referenced this pull request Sep 12, 2026
* fix(pi): route needs-decision wakes and mixed batches wholly to main

Skip the supervision branch for every needs-decision status append, the
same way a check-kind wake already skips it. A coalesced signal/stale
trigger batch containing any needs-decision row is delivered wholly to
main, not split between the branch and a later main wake - the whole
batch, including any co-present routine rows for a different task,
travels together. Heartbeat and unread-status scans stay independent.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE

* test(pi): cover distinct-file mixed batches and heartbeat independence

Add a regression using two distinct files (not the same status file
twice) in one coalesced trigger so a some-vs-every regression on the
file-list cross-reference cannot hide behind a degenerate same-key
case, and a heartbeat/needs-decision co-presence test proving a
needs-decision row neither vetoes nor rides along with an otherwise
eligible heartbeat scan.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE

* no-mistakes(document): Clarify needs-decision and heartbeat routing

* no-mistakes(ci): Fixed captain-held stale reminders so they bypass supervision and wake main, while unrelated unread rows and heartbeats remain independent. Added routing regressions and updated documentation. Pi watcher tests, strict TypeScript checks, lint, and diff checks pass. The no-mistakes attestation failure was pipeline-state related, not a source defect

* no-mistakes(ci): Fixed CI lint by narrowly suppressing false-positive SC2031 diagnostics where background PIDs are captured immediately in the same shell. Verified with `CI=true bin/fm-lint.sh` and `git diff --check`. The no-mistakes attestation failure is pipeline-state-related (`test` was skipped), not a source defect

* no-mistakes(review): Route stale open decisions directly to main

* no-mistakes(review): Honor configured verbs in stale decision routing

* no-mistakes(review): Route second-mate escalations and configured decisions to main

* no-mistakes(review): Ignore trailing whitespace after captain holds

* no-mistakes(review): Cache stale decision classification per status file

* no-mistakes(review): Document unread decision precedence for later task wakes

* no-mistakes(review): Cache unchanged stale decisions across scope scans

* no-mistakes(review): Resolve decision aliases and reject symlinked statuses

* no-mistakes(review): Route surfaced captain-held signals directly to main

* no-mistakes(document): Document decision-owned main routing

* no-mistakes(ci): Fixed captain-held spans to remain actionable while crew working evidence is positive, ensuring the watcher delivers their main-only marker. Updated the executable regression test to cover this case. Verified with the full fm-watch-triage suite, bash syntax checks, and git diff checks. Shellcheck reported only pre-existing test harness warnings (SC1091/SC2034)

* no-mistakes(ci): Fixed the CI regression: captain-held transfers now retain their established non-actionable stale classification while the signal-routing side-band still surfaces them main-only. Verified with tests/fm-daemon.test.sh, tests/fm-watch-triage.test.sh, bash syntax checks, and git diff --check

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
friesentius pushed a commit to friesentius/firstmate that referenced this pull request Sep 21, 2026
* fix(pi): route needs-decision wakes and mixed batches wholly to main

Skip the supervision branch for every needs-decision status append, the
same way a check-kind wake already skips it. A coalesced signal/stale
trigger batch containing any needs-decision row is delivered wholly to
main, not split between the branch and a later main wake - the whole
batch, including any co-present routine rows for a different task,
travels together. Heartbeat and unread-status scans stay independent.

Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE

* test(pi): cover distinct-file mixed batches and heartbeat independence

Add a regression using two distinct files (not the same status file
twice) in one coalesced trigger so a some-vs-every regression on the
file-list cross-reference cannot hide behind a degenerate same-key
case, and a heartbeat/needs-decision co-presence test proving a
needs-decision row neither vetoes nor rides along with an otherwise
eligible heartbeat scan.

Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE

* no-mistakes(document): Clarify needs-decision and heartbeat routing

* no-mistakes(ci): Fixed captain-held stale reminders so they bypass supervision and wake main, while unrelated unread rows and heartbeats remain independent. Added routing regressions and updated documentation. Pi watcher tests, strict TypeScript checks, lint, and diff checks pass. The no-mistakes attestation failure was pipeline-state related, not a source defect

* no-mistakes(ci): Fixed CI lint by narrowly suppressing false-positive SC2031 diagnostics where background PIDs are captured immediately in the same shell. Verified with `CI=true bin/fm-lint.sh` and `git diff --check`. The no-mistakes attestation failure is pipeline-state-related (`test` was skipped), not a source defect

* no-mistakes(review): Route stale open decisions directly to main

* no-mistakes(review): Honor configured verbs in stale decision routing

* no-mistakes(review): Route second-mate escalations and configured decisions to main

* no-mistakes(review): Ignore trailing whitespace after captain holds

* no-mistakes(review): Cache stale decision classification per status file

* no-mistakes(review): Document unread decision precedence for later task wakes

* no-mistakes(review): Cache unchanged stale decisions across scope scans

* no-mistakes(review): Resolve decision aliases and reject symlinked statuses

* no-mistakes(review): Route surfaced captain-held signals directly to main

* no-mistakes(document): Document decision-owned main routing

* no-mistakes(ci): Fixed captain-held spans to remain actionable while crew working evidence is positive, ensuring the watcher delivers their main-only marker. Updated the executable regression test to cover this case. Verified with the full fm-watch-triage suite, bash syntax checks, and git diff checks. Shellcheck reported only pre-existing test harness warnings (SC1091/SC2034)

* no-mistakes(ci): Fixed the CI regression: captain-held transfers now retain their established non-actionable stale classification while the signal-routing side-band still surfaces them main-only. Verified with tests/fm-daemon.test.sh, tests/fm-watch-triage.test.sh, bash syntax checks, and git diff --check

---------
friesentius pushed a commit to friesentius/firstmate that referenced this pull request Sep 21, 2026
* fix(pi): route needs-decision wakes and mixed batches wholly to main

Skip the supervision branch for every needs-decision status append, the
same way a check-kind wake already skips it. A coalesced signal/stale
trigger batch containing any needs-decision row is delivered wholly to
main, not split between the branch and a later main wake - the whole
batch, including any co-present routine rows for a different task,
travels together. Heartbeat and unread-status scans stay independent.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE

* test(pi): cover distinct-file mixed batches and heartbeat independence

Add a regression using two distinct files (not the same status file
twice) in one coalesced trigger so a some-vs-every regression on the
file-list cross-reference cannot hide behind a degenerate same-key
case, and a heartbeat/needs-decision co-presence test proving a
needs-decision row neither vetoes nor rides along with an otherwise
eligible heartbeat scan.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE

* no-mistakes(document): Clarify needs-decision and heartbeat routing

* no-mistakes(ci): Fixed captain-held stale reminders so they bypass supervision and wake main, while unrelated unread rows and heartbeats remain independent. Added routing regressions and updated documentation. Pi watcher tests, strict TypeScript checks, lint, and diff checks pass. The no-mistakes attestation failure was pipeline-state related, not a source defect

* no-mistakes(ci): Fixed CI lint by narrowly suppressing false-positive SC2031 diagnostics where background PIDs are captured immediately in the same shell. Verified with `CI=true bin/fm-lint.sh` and `git diff --check`. The no-mistakes attestation failure is pipeline-state-related (`test` was skipped), not a source defect

* no-mistakes(review): Route stale open decisions directly to main

* no-mistakes(review): Honor configured verbs in stale decision routing

* no-mistakes(review): Route second-mate escalations and configured decisions to main

* no-mistakes(review): Ignore trailing whitespace after captain holds

* no-mistakes(review): Cache stale decision classification per status file

* no-mistakes(review): Document unread decision precedence for later task wakes

* no-mistakes(review): Cache unchanged stale decisions across scope scans

* no-mistakes(review): Resolve decision aliases and reject symlinked statuses

* no-mistakes(review): Route surfaced captain-held signals directly to main

* no-mistakes(document): Document decision-owned main routing

* no-mistakes(ci): Fixed captain-held spans to remain actionable while crew working evidence is positive, ensuring the watcher delivers their main-only marker. Updated the executable regression test to cover this case. Verified with the full fm-watch-triage suite, bash syntax checks, and git diff checks. Shellcheck reported only pre-existing test harness warnings (SC1091/SC2034)

* no-mistakes(ci): Fixed the CI regression: captain-held transfers now retain their established non-actionable stale classification while the signal-routing side-band still surfaces them main-only. Verified with tests/fm-daemon.test.sh, tests/fm-watch-triage.test.sh, bash syntax checks, and git diff --check

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
jordanhindo pushed a commit to jordanhindo/firstmate that referenced this pull request Sep 22, 2026
* fix(pi): route needs-decision wakes and mixed batches wholly to main

Skip the supervision branch for every needs-decision status append, the
same way a check-kind wake already skips it. A coalesced signal/stale
trigger batch containing any needs-decision row is delivered wholly to
main, not split between the branch and a later main wake - the whole
batch, including any co-present routine rows for a different task,
travels together. Heartbeat and unread-status scans stay independent.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE

* test(pi): cover distinct-file mixed batches and heartbeat independence

Add a regression using two distinct files (not the same status file
twice) in one coalesced trigger so a some-vs-every regression on the
file-list cross-reference cannot hide behind a degenerate same-key
case, and a heartbeat/needs-decision co-presence test proving a
needs-decision row neither vetoes nor rides along with an otherwise
eligible heartbeat scan.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE

* no-mistakes(document): Clarify needs-decision and heartbeat routing

* no-mistakes(ci): Fixed captain-held stale reminders so they bypass supervision and wake main, while unrelated unread rows and heartbeats remain independent. Added routing regressions and updated documentation. Pi watcher tests, strict TypeScript checks, lint, and diff checks pass. The no-mistakes attestation failure was pipeline-state related, not a source defect

* no-mistakes(ci): Fixed CI lint by narrowly suppressing false-positive SC2031 diagnostics where background PIDs are captured immediately in the same shell. Verified with `CI=true bin/fm-lint.sh` and `git diff --check`. The no-mistakes attestation failure is pipeline-state-related (`test` was skipped), not a source defect

* no-mistakes(review): Route stale open decisions directly to main

* no-mistakes(review): Honor configured verbs in stale decision routing

* no-mistakes(review): Route second-mate escalations and configured decisions to main

* no-mistakes(review): Ignore trailing whitespace after captain holds

* no-mistakes(review): Cache stale decision classification per status file

* no-mistakes(review): Document unread decision precedence for later task wakes

* no-mistakes(review): Cache unchanged stale decisions across scope scans

* no-mistakes(review): Resolve decision aliases and reject symlinked statuses

* no-mistakes(review): Route surfaced captain-held signals directly to main

* no-mistakes(document): Document decision-owned main routing

* no-mistakes(ci): Fixed captain-held spans to remain actionable while crew working evidence is positive, ensuring the watcher delivers their main-only marker. Updated the executable regression test to cover this case. Verified with the full fm-watch-triage suite, bash syntax checks, and git diff checks. Shellcheck reported only pre-existing test harness warnings (SC1091/SC2034)

* no-mistakes(ci): Fixed the CI regression: captain-held transfers now retain their established non-actionable stale classification while the signal-routing side-band still surfaces them main-only. Verified with tests/fm-daemon.test.sh, tests/fm-watch-triage.test.sh, bash syntax checks, and git diff --check

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
NewAiCoder added a commit to NewAiCoder/firstmate that referenced this pull request Sep 26, 2026
…rt (kunchenguid#28)

* fix(bin): preserve subshell lock ownership on Bash 3.2 (#3789)

* fix: distinguish subshell wake-lock owners on stock Bash

Restore distinct process ownership for issue #3743 using the existing PID helper, consistently across lock publication, reclaim, release, role checks, and bounded handoff.

The existing wake-queue regression fails on pristine upstream Bash 3.2 with rc=13. The complete suite now passes on Bash 3.2.57 and Bash 5.3.15, with added coverage for ownership when BASHPID is unset. Canonical lint and stock-Bash syntax checks pass.

* no-mistakes(document): Correct lock grace-period documentation

* no-mistakes(ci): Captain, fixed all 14 SC2031 false positives with nine ShellCheck source-boundary annotations across three tests. Full CI-mode lint and the complete wake-queue suite on stock Bash 3.2 passed. Runtime behavior is unchanged

* fix(bin): resolve captain holds and legacy teardowns on non-markdown backends (#3782)

* fix(bin): close legacy records on the Beads backend honestly

Two pre-Beads reads blocked honest closure of leftover records:

1. fm-captain-hold.sh complete/verify resolved attested legacy hold ids
   only against the live backend and the pre-collapse derived identity, so
   a home whose holds fm-hold-migration rehomed under fm- ids failed with
   an empty-name absence message (the resolve failure was swallowed by the
   command substitution feeding verify_hold_durable). Resolution now falls
   back, on the Beads backend only, to the legacy id under the configured
   beads prefix and to the row whose notes carry the exact marker line
   'migrated from data/backlog.md id <legacy id>'; every refusal names the
   id it could not resolve, and the markdown path is unchanged.

2. fm-teardown.sh refused any record without spawn_gen forever. A record
   that predates the field can now be torn down with an explicit
   --legacy-record flag once the recovery-grade endpoint classifier
   confirms the recorded endpoint dead or agent-less; the accepted
   incarnation is stamped into the record right before its close marker
   binds to it and named in the teardown line. Refusals leave the record
   byte-identical, the unlanded-work refusal is not relaxed, and a corrupt
   (multi-valued) spawn_gen is never accepted.

The companion repair this branch carries (follow-up commit) is the
backend-gated --file and markdown-file requirement in the mutate path and
lifecycle gates: fm_backlog_mutate passed --file and required the markdown
backlog file regardless of the resolved backend, and the transition gate
plus row probe required that file before any backend work, so a home on a
non-markdown backend could neither gate, probe, nor close its rows.

Behavior tests: self-contained beads fixtures over a scratch bd graph
(self-skipping on markdown-only tasks-axi installs), legacy meta fixtures
for every teardown gate, and the relocated markdown backlog coverage stays
green.

* no-mistakes(review): fix(review): report migrated-hold scan refusals and guard legacy spawn_gen stamp against newline-less records

* fix(backlog): address the configured backend for lifecycle writes

Completes the fm-backlog-transition-lib repair the first commit's message
claims: on this base fm_backlog_mutate passed --file and required the
markdown backlog file regardless of the resolved backend, and
fm_backlog_transition_applies plus fm_backlog_row_probe required that file
before any backend work, so a home on a non-markdown backend could neither
gate, probe, nor close its backlog rows. All three now gate the markdown
file on the resolved tasks-axi backend: markdown keeps exactly its explicit
<data>/backlog.md behavior, non-markdown homes address the backend their
own configuration selects with no markdown file requirement.
fm_backlog_row_show and fm_backlog_row_list already gated correctly and
are unchanged. docs/configuration.md owns the contract line.

Also extends the same backend gate to fm-captain-hold.sh's own mutation
wrapper - hold/add/update/answer/done append the markdown --file only when
the resolved backend is markdown, so a captain call on a Beads home reaches
the Beads store end to end - and applies the review round's two direct
remedies there: the [beads] graph path resolves against the backlog root
when relative (never the process CWD), and a failed bd graph read reports
bd's own trimmed stderr reason in the refusal.

Coverage: tests/fm-backlog-atomicity.test.sh gains a stub-driven Beads
completion case proving the transition gate applies, the row probe reads,
and done runs without any markdown file or --file override; the relocated
markdown backlog test stays green.

* no-mistakes(review): Document root-tasks.toml-only beads settings for migrated-hold resolution

* test(gotmp): stub fm_tasks_axi_backend so the fixture matches the backend-aware transition lib

The legacy-records change made fm-backlog-transition-lib.sh resolve the
configured backend via fm_tasks_axi_backend before the markdown-only skip.
The gotmp fixture's fm-tasks-axi-lib stub lacked that function, so the
markdown check fell through and teardown hit the incompatible-backend
error with unbound FM_TASKS_AXI_MIN under set -u. Stub the backend as
markdown and define the floor, restoring the intended no-backlog skip.

* fix(teardown): roll the legacy stamp back when the close marker fails

A legacy-record teardown stamps its accepted incarnation into the record
right before the close marker binds to it; when that marker write then
fails, the stamp survived, so a retried teardown sailed past the
dead-or-agent-less endpoint gate the stamp now proved unnecessary. The
failed marker write now truncates the record back to its exact pre-stamp
bytes (verified by size), restoring the byte-identical-refusal invariant;
when the rollback itself fails the operator is told to re-run with
--legacy-record after reconciling the endpoint.

Also completes the recorded review decision's coverage wording: the
beads stub test now drives the answer close end to end (update and done
through the gated wrapper), asserting no markdown file override reaches
either verb.

* fix(review): harden the legacy stamp rollback and resolve derived migrated ids

The legacy-record stamp rollback now uses perl (already in the teardown
curated PATH; truncate is not, and is absent on stock macOS), routes every
failure branch inside the stamp block through the same size-verified
rollback so the byte-identical-refusal invariant holds on those paths too,
and gains behavior coverage: an unrecordable close (an invalid pr= link)
fails the teardown, leaves the record byte-identical, keeps the backlog
row in flight, and a flag-less retry still refuses.

Migrated-hold resolution now probes the derived pre-collapse identity
(<origin>-decision-<entry>) alongside the raw entry - fm-hold-migration
recorded the DERIVED id in every migrated row's marker note - in both the
prefix and the migration-note forms, with the ambiguity refusal naming
every identity tried, plus behavior coverage for a bare decision key
resolved through its derived identity's marker.

Also aligns fm-backlog-transition-lib.sh's header ADDRESSING/SCOPE
paragraphs with the backend-gated contract, drops an unreachable FORCE
validity guard the parser rewrite left behind, and switches the new stub
fixture to the portable sed -i.bak idiom.

* no-mistakes(review): Name the configured backend in teardown's backlog reminder

* no-mistakes(review): Scan migration markers before the prefix guess

* no-mistakes(review): Document marker-first resolution and cover the prefix branch

* no-mistakes(document): Record prefix-attestation audit and marker-line forms

* no-mistakes(ci): Fixed the Greptile P1 on bin/fm-teardown.sh: a failed rollback of the synthetic legacy stamp let a retry bypass the dead-or-agent-less endpoint gate. Root cause: teardown minted `spawn_gen=legacy-<ts>-<pid>` into the task record before the close marker bound to it. When the close-marker write failed AND the rollback also failed, the record retained that token. On the next invocation `fm_backlog_meta_spawn_gen` succeeded, so `TEARDOWN_LEGACY_PENDING` stayed 0 and the endpoint gate was skipped entirely — even with `--legacy-record`. The script's own error text told the operator to "re-run teardown with --legacy-record", advice the code could not honor. Fix (bin/fm-teardown.sh): - A `legacy-*` spawn_gen is now recognized as a stamp this teardown path minted, never one a spawn published (fm-spawn.sh publishes `s<epoch>.<pid>.<random>`). Such a record still reads as the legacy record it is: it re-enters the endpoint gate, and a flag-less retry refuses naming `--legacy-record`. - Acceptance reuses the retained token instead of minting a second one; the append block is skipped when the record already carries it, so no duplicate spawn_gen is written. - The rollback attempt and its "could not be rolled back" message are guarded to runs that actually appended a stamp, so a run that appended nothing never claims a rollback it did not perform. - Usage header documents the retained-stamp rule. Test (tests/fm-teardown.test.sh): added `test_retained_legacy_stamp_still_faces_the_endpoint_gate`, an end-to-end reproduction — a `perl` stub that fails only the rollback's `truncate` (delegating every other perl call to the real interpreter) leaves the stamp behind, then the retry must still hit the gate, must not stamp a second incarnation, must not close the backlog row, and the flag-less retry must refuse. Verification: the new test fails against the pre-fix script on exactly the reported defect ("the retry skipped the dead-or-agent-less endpoint gate") and passes after. Full tests/fm-teardown.test.sh 80 ok / 0 failures / rc=0; tests/fm-backlog-atomicity.test.sh 80 ok / 0 failures / rc=0; bin/fm-lint.sh (pinned ShellCheck 0.11.0 + actionlint 1.7.12) clean

* fix: reduce local ShellCheck source-analysis cost (#3778)

* fix(lint): drop source following on the local changed-file gate

The local lint step was inlining library closures through --external-sources
and peaking above 8 GB on a single root. Keep full analysis in CI, on main,
and without a merge-base; exclude the four cross-file codes from the local
pass so those findings still land in CI.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Run local ShellCheck per root and document measurements

* no-mistakes(review): Correct local source-following telemetry

* no-mistakes(document): Clarify context-sensitive lint documentation

---------

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(pi): route decision-owned wake batches to main (#3776)

* fix(pi): route needs-decision wakes and mixed batches wholly to main

Skip the supervision branch for every needs-decision status append, the
same way a check-kind wake already skips it. A coalesced signal/stale
trigger batch containing any needs-decision row is delivered wholly to
main, not split between the branch and a later main wake - the whole
batch, including any co-present routine rows for a different task,
travels together. Heartbeat and unread-status scans stay independent.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE

* test(pi): cover distinct-file mixed batches and heartbeat independence

Add a regression using two distinct files (not the same status file
twice) in one coalesced trigger so a some-vs-every regression on the
file-list cross-reference cannot hide behind a degenerate same-key
case, and a heartbeat/needs-decision co-presence test proving a
needs-decision row neither vetoes nor rides along with an otherwise
eligible heartbeat scan.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE

* no-mistakes(document): Clarify needs-decision and heartbeat routing

* no-mistakes(ci): Fixed captain-held stale reminders so they bypass supervision and wake main, while unrelated unread rows and heartbeats remain independent. Added routing regressions and updated documentation. Pi watcher tests, strict TypeScript checks, lint, and diff checks pass. The no-mistakes attestation failure was pipeline-state related, not a source defect

* no-mistakes(ci): Fixed CI lint by narrowly suppressing false-positive SC2031 diagnostics where background PIDs are captured immediately in the same shell. Verified with `CI=true bin/fm-lint.sh` and `git diff --check`. The no-mistakes attestation failure is pipeline-state-related (`test` was skipped), not a source defect

* no-mistakes(review): Route stale open decisions directly to main

* no-mistakes(review): Honor configured verbs in stale decision routing

* no-mistakes(review): Route second-mate escalations and configured decisions to main

* no-mistakes(review): Ignore trailing whitespace after captain holds

* no-mistakes(review): Cache stale decision classification per status file

* no-mistakes(review): Document unread decision precedence for later task wakes

* no-mistakes(review): Cache unchanged stale decisions across scope scans

* no-mistakes(review): Resolve decision aliases and reject symlinked statuses

* no-mistakes(review): Route surfaced captain-held signals directly to main

* no-mistakes(document): Document decision-owned main routing

* no-mistakes(ci): Fixed captain-held spans to remain actionable while crew working evidence is positive, ensuring the watcher delivers their main-only marker. Updated the executable regression test to cover this case. Verified with the full fm-watch-triage suite, bash syntax checks, and git diff checks. Shellcheck reported only pre-existing test harness warnings (SC1091/SC2034)

* no-mistakes(ci): Fixed the CI regression: captain-held transfers now retain their established non-actionable stale classification while the signal-routing side-band still surfaces them main-only. Verified with tests/fm-daemon.test.sh, tests/fm-watch-triage.test.sh, bash syntax checks, and git diff --check

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* feat(bin): add opt-in worker launch environment allowlist (#3802)

* feat(spawn): add an opt-in worker environment allowlist

Honor a home-local launch-env-allowlist at the shared worker command
boundary and inherit it into secondmate homes. Preserve the existing
launch behavior when the file is absent. Keep the operational environment
and explicit launch assignments, and account for filtered Muse credentials.

Refs https://github.com/kunchenguid/firstmate/issues/3742

Verification:
- Red on origin/main 1820316b66ac2c68e244dd04a02512859ee8c1f4:
  the new enabled-allowlist regression observed synthetic-unrelated in
  the worker; the absent-file control passed.
- Green: fm-test-run.sh on fm-spawn-dispatch-profile, fm-muse-harness,
  and fm-trace-context-spawn; all three passed without skips.
- Synthetic emitted-command probes ran through sh, stock Bash, and zsh.
- Canonical lint, documentation audience checks, and stock Bash syntax
  checks passed.

* no-mistakes(review): Reject inaccessible launch environment configuration

* no-mistakes(review): Preserve inherited allowlists on source inspection errors

* no-mistakes(document): Clarify worker environment grants and inheritance documentation

* no-mistakes(lint): Fix inheritance test ShellCheck source boundary

* feat(bin): add rovo crewmate/scout adapter with home-path file access (#3575)

* feat(bin): add rovo as a verified crewmate/scout worker harness

Wire the Atlassian Rovo CLI (202609.1.2) into the TUI-under-tmux/herdr
adapter contract: detection with marker-precedence ordering, one-shot
positional launch with --startup-receipt readiness polling instead of
composer scraping, model/effort flags, a screen-scrape busy fallback
scoped like grok's, and crew/scout-only lifecycle control that refuses
secondmate launches. Ships with a portable regression suite, a live PTY
guard against the real binary, a per-harness reference doc, and a dated
verification record covering the silent OAuth refresh, the interrupt-ack
divergence from the originating scout report, and the still-open
composer-ghost and tmux/herdr pane-liveness gaps.

* no-mistakes(review): revert rovo launch to positional brief, drop startup-receipt

* no-mistakes(test): rewire rovo adapter to kimi-style launch-then-send shape

* no-mistakes(document): add rovo to stale worker-harness enumerations in docs

* docs(verification): close the rovo herdr-liveness gap with live isolated-lab evidence

Placement, launch-then-send, and busy/idle rendering are now verified live
in an isolated non-default Herdr lab session (bin/fm-herdr-lab.sh), driven
directly through fm-spawn.sh's/fm-backend.sh's own shared primitives since
the cross-session launcher-identity guard refuses this task's own ambient
Herdr identity for a full fm-spawn.sh run.

fm_backend_agent_state reported dead for a live, responding rovo pane at
every point checked, because herdr's own agent-integration registry has no
rovo entry (herdr integration status), so herdr agent get returns
agent_not_found regardless of whether rovo is actually running. This is
recorded as a Herdr-side integration gap rather than a firstmate bug, left
unpatched to avoid a false-positive alive verdict for other idle shells.

Updates docs/verification/rovo.md's backend-liveness section and its two
cross-references (docs/verification/runtime-backends.md, docs/configuration.md)
accordingly.

* fix(bin): close rovo's failed-spawn leak and busy-scrape false idle

Greptile P1s on PR #3575: a failed rovo readiness/submission/delivery gate
exited without tearing down the just-created endpoint, leaving the launched
--yolo rovo process running as an orphaned agent outside task control.
Separately, the busy classifier's rendered-tail fallback returned definitive
idle whenever the "Rovo is thinking" marker scrolled out of the last 12
nonblank lines of a long turn, which could make supervision wrongly conclude
a still-working worker had gone idle.

fm-spawn.sh: rovo_spawn_fail now calls rovo_endpoint_cleanup, which kills the
created endpoint (tmux/herdr/zellij/cmux) via the same generic fm_backend_kill
dispatch fm-spawn.sh's own orca-abort path already uses; orca's worktree and
terminal remain owned by the separate ORCA_ABORT_CLEANUP trap.

fm-busy-lib.sh: the rovo classifier arm now reports "unknown rovo-regex"
instead of "idle rovo-regex" when the marker is absent, matching how muse and
cursor already express "can't tell" for their own fallbacks. The positive
busy match is unchanged.

Extends tests/fm-rovo-harness.test.sh: the readiness and delivery failure
tests now assert the endpoint is torn down (and the success test asserts it
is not), and a new test drives the busy marker out of the tail window to
confirm the verdict is unknown, never idle. bin/fm-lint.sh is clean on both
changed files.

* test(rovo): align spawn fixture with the launch-brief validation contract

Upstream main now requires a brief's ## Captain's intent and
## Firstmate spec subsections (or a nonempty legacy # Task body)
before spawn, and rewrites ship+no-mistakes briefs into
launch-brief.md. Update the rovo harness fixture and pointer
assertions to match, mirroring the kimi harness fixture.

* no-mistakes(review): align rovo.md delivery-gate note with live herdr evidence

* fix: prevent stale supervision wake loops (#3672)

* fix(bin): stop the supervision branch's stale-ack and ghost-report loops

Clean-slate implementation of the four authorized recommendations from the
supervision-ghost-retrigger analysis (items 1, 2, 3, and 7), in their minimal
form, superseding PR #3604:

- fm_branch_report refuses a task the wake being handled never named. The
  extension fixes the reportable task set from the eligible rows before each
  prompt (signal and stale rows resolve to their tasks, a heartbeat allows any
  task with a live record, fleet is always allowed), so a report typed from
  memory about a task whose records teardown already removed is never stored
  or delivered.
- An acknowledgement that consumes nothing says "nothing was acknowledged
  through N" and prints the exact --ack-through / --recovery-generation
  command for the current presented wake, instead of "re-run the drain",
  which re-fed the same stale acknowledgement in a loop.
- bin/fm-guard.sh no longer tells the branch actor to drain queued wakes
  while it is handling them; it names the granted rows instead.
- Teardown removes state/.<task>.branch-outcome-index for ordinary tasks and
  descendants; the index rebuild and the append-side index write both skip a
  task with neither a live record nor a status log, so the branch's report of
  a teardown it just performed is stored without recreating the index.

No new locking, no spawn-generation binding, and no retired-task refusal: the
branch can still report the outcome of a task it just tore down, and the
teardown test now proves that path end to end.

* fix(bin): narrow the branch report scope and guard silence to the minimal form

Apply the four review decisions on the clean-slate branch:

- A signal or stale prompt may report only the tasks its own rows resolve
  to; fleet is refused there too. A heartbeat review is not scoped by task
  at all, so the extension no longer tracks live task records and refuses
  nothing by task id during a fleet review.
- The outcome-index rebuild no longer skips retired tasks; the append-side
  skip alone keeps a torn-down task's index from being recreated.
- bin/fm-guard.sh keeps the queued-wakes warning silent for the branch actor
  instead of printing a replacement note.

* no-mistakes(document): Align supervision docs with scoped wake handling

* fix(bin): grant rovo the per-task home paths its standard crewmate flow needs

rovo confines every file-tool operation to its worktree by default, and its
bash tool independently refuses the same external paths regardless of any
grant (confirmed live), so a rovo worker could not read its own brief or
steering messages or write its status/report - all of which live in the
firstmate home outside the worktree - without hand-feeding it. Grant
toolPermissions.allowedExternalPaths for exactly the task's brief directory,
steering inbox, and status file at launch time via --config-override,
merged with agent.efficiencyLevel into one JSON object since that flag is
single-value and silently discards a second occurrence.

Extends the live PTY guard to prove, against the real binary, that the
grant lets rovo read an external brief and append to an external status
file, and that the same flow is blocked without the grant.

* no-mistakes(document): align rovo reference Effort row with merged single --config-override

* no-mistakes(document): document rovo file-access grant in harness reference

---------

Co-authored-by: PUNEET PATWARI <ppatwari@atlassian.com>
Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>

* fix(bin): prevent false pipeline blocks after drive timeouts (#3813)

* fix(bin): read a crew's pipeline-death claim against the live run

A crew's no-mistakes drive call blocks until the next gate or outcome,
routinely far longer than its harness lets one command live, so the call
gets killed or times out while the daemon runs the fix round on in the
background. Crews read that as daemon death and block on it, and firstmate
had nothing that contradicted them.

Rule 7 of every generated brief now says a drive-call error or a harness
command timeout is not a daemon error, requires `no-mistakes daemon status`
plus `no-mistakes axi status` before a pipeline `blocked:`, and reserves
that report for a refused socket or a run record failed with a daemon
error. The no-mistakes definition of done adds the harness command limit
and the background-and-poll shape that fits inside it.

fm-crew-state gains one classification case: a `blocked:` line blaming the
daemon, a timeout, or unreachability, while the run is running or fixing
AND the pipeline reports fresh activity, now reads as superseded because
the run is alive. Recency comes from the client's own `quiet` marker on
active_steps.last_activity rather than a threshold invented here, and
positive evidence is required, so a run record that outlives a genuinely
dead daemon keeps the plain reading.

stuck-crewmate-recovery gains the inverse-of-a-dead-endpoint playbook:
firstmate reads both statuses itself, steers a reattach, never restarts the
shared daemon on a crew's claim, and escalates only a refused socket.

Nothing here depends on an unshipped no-mistakes capability.

* no-mistakes(review): Prioritize daemon socket failure and narrow unreachable matching

* no-mistakes(review): Honor socket refusal across coarse status and crew guidance

* no-mistakes(test): Replace flaky settle timing assertion with pane-read count

* no-mistakes(document): Document daemon timeout recovery contract

* no-mistakes(ci): Fixed daemon socket failures being suppressed by terminal attributed runs. Positive refused/missing socket evidence now remains blocked regardless of run status. Added a behavioral regression test for terminal failed runs. Verified with fm-crew-state tests, project ShellCheck lint, and git diff checks

* fix(bin): scope the worker role contract for ship and scout launches (#3797)

* fix(brief): scope Firstmate workers to their launch contract

* no-mistakes(document): Clarify supervisor scope and worker contract ownership

* no-mistakes(ci): Removed the heading-based bypass so every ship/scout launch receives the current worker-role contract. Added a regression that failed before the fix and passes afterward. Dispatch, brief, and delivery suites, focused ShellCheck, and git diff --check all passed

* no-mistakes(review): make launch overlay sole owner of worker role contract

* no-mistakes(review): narrow heading test dimension and fix publish error wording

* no-mistakes(review): gate role supersession, fix render guard, drop AGENTS twin

* no-mistakes(document): align architecture AGENTS.md scope and spawn launch-brief header

* fix(bin): stop ringing steering doorbells into dead panes (#3823)

* fix(bin): stop ringing steering doorbells into dead panes

The steering-inbox doorbell was a plain sentence plus Enter typed into a
worker's pane, and the watcher re-rang it on the assumption that a ring is
free. In a pane whose agent has exited that line is a shell command, and the
re-ring ladder kept typing it into a shell that can never acknowledge it.

- Prefix the doorbell with the shell no-op `: ` so a bare shell executes
  nothing while a live worker still reads the same self-describing line.
  `#` is not used because interactive zsh does not treat it as a comment by
  default and the claude harness binds it to memory mode.
- fm_task_inbox_ring skips the pane (return 3) when the backend positively
  classifies the agent as dead; missing, ambiguous, unreadable, and unverified
  endpoints still ring so a blind classifier never starves a live worker.
- The watcher caps the ladder for a dead pane: one stale wake for recovery,
  no ring, no ladder walk, and the durable record stays for
  stuck-crewmate-recovery. fm-send and the remote steer leg report the skip.

Tests cover the no-op in real shells, the dead/live/unclassifiable ring
verdicts, and the single-surfacing watcher path.

* no-mistakes(review): Quote doorbell paths against shell injection

* no-mistakes(review): Reject terminal-control paths before ringing

* no-mistakes(review): Document accepted partial doorbell delivery race

* no-mistakes(review): Skip unavailable endpoints before busy-state handling

* no-mistakes(test): Respect shell startup PATH in environment allowlist test

* no-mistakes(test): Fix doorbell test fixtures for endpoint liveness

* no-mistakes(test): Prioritize confirmed restarts and clean shell test syntax

* no-mistakes(document): Document dead and missing doorbell recovery

* no-mistakes(ci): Fixed the persistence-reply timeout race by rechecking for a correlated reply immediately before falling back to a nudge. Added a deterministic regression covering replies arriving between the preliminary resolution pass and timeout handling. Verified with the targeted restart suite, project lint, coverage guard, bash syntax checks, and diff checks

* fix(bin): wait out transient primary-checkout reads in the spawn worktree poll (#3834)

* fix(spawn): keep the worktree poll from adopting the repository primary

After `treehouse get` is sent, the worktree-discovery poll reads the pane's
foreground-process cwd. While treehouse is still fetching and checking a slot
out, the foreground process is treehouse itself and it reports the repository's
PRIMARY checkout as its cwd for several seconds. The poll accepted any path
that merely differed from the spawning project, so from a linked spawning home
- whose project is itself a worktree of that repository - it adopted the
primary, and the isolation guard then refused a launch whose slot treehouse
went on to create normally.

Screen every candidate with the isolation guard's own conditions, extracted as
spawn_worktree_isolated, so a read the guard would reject stays a transient the
poll keeps waiting through. The two-consecutive-reads rule and the guard as
final backstop are unchanged; a pane that never reaches an isolated worktree
still fails at the existing 60s deadline, now naming the last path it reported.

The already-settled timing assertion counted whole-spawn wall time against a
5s budget and failed on unmodified HEAD on slower machines; it now counts pane
reads, which is what "one confirming read, not an extra cycle" actually means.

* fix(spawn): say which path the worktree wait rejected, and why

Screening every discovery-poll candidate means a host that never reaches an
isolated worktree spends the whole 60s window before refusing. That wait is
deliberate - separating a transient from a terminal misconfiguration needs
machinery this path does not want - so the refusal explains itself instead:
the isolation check records why a candidate failed, and the deadline names the
last path seen together with that reason. Message and diagnostics only; the
poll's control flow is unchanged.

Two suites asserted the guard's wording on paths the poll now rejects rather
than adopts, so their refusal arrives from the deadline instead: realign
fm-tangle-guard's non-git and subdirectory-of-primary cases (each now also
asserting the stated reason, and the second the metadata absence it was
missing) and the herdr projection e2e's forced non-worktree cwd.

* no-mistakes(review): stub poll sleep in tangle-guard spawn isolation test

* no-mistakes(document): document spawn poll isolation screen in fm-spawn header

* test(spawn): make the non-git isolation case non-git anywhere

The refusal-reason assertion for a path outside any repository assumed TMPDIR
is not inside a git repository. Where it is, git walks up from the temporary
directory, finds that repository, and the spawn reports the subdirectory cause
instead - so the case passed or failed on a property of the host rather than on
the behaviour under test.

Build the path under a directory the test then names in GIT_CEILING_DIRECTORIES,
which git documents as not chdir-ing up into a listed directory while looking
for a repository. Git never excludes the directory being searched, so the
ceiling is the parent of the path handed to the spawn.

The assertions pin which cause fired rather than the sentence that explains it,
leaving the operator wording free to improve.

* no-mistakes(document): point spawn poll comment at the isolation screen's comparison

* no-mistakes(ci): Fixed the "Behavior portable serial 2" failure in tests/fm-tangle-guard.test.sh ("non-worktree spawn did not say why the path was rejected (missing: 'not inside a git worktree')"). Root cause, in this PR's code: bin/fm-spawn.sh's spawn_worktree_isolated resolved the git toplevel with `wt_top_real=$(cd "$SPAWN_WT_TOP" ...)`. For a path in no repository, `git rev-parse --show-toplevel` yields empty, and `cd ""` is a SUCCESSFUL no-op on bash before 5.3 (CI's ubuntu-latest ships bash 5.2). The empty toplevel therefore resolved to fm-spawn's own cwd — the CI checkout — so the poll reported "it is a subdirectory of worktree root '/home/runner/work/firstmate/firstmate'" instead of the correct "it is not inside a git worktree". Dev machines with bash 5.3 fail `cd ""`, which is why the suite passed locally and only failed on CI; it is a genuine shell-portability defect in the reason vocabulary this change added, not a test-environment artifact. Fix (smallest root-cause change, 1 line + comment, bin/fm-spawn.sh:2168-2173): guard the empty value so it never reaches `cd` — if [ -n "$SPAWN_WT_TOP" ] && ! wt_top_real=$(cd "$SPAWN_WT_TOP" 2>/dev/null && pwd -P); then No change to the poll's timing or deadline behavior (respecting the recorded refusal-latency and spawn-wt-reason-vocabulary decisions), no new tests, no other files touched. Verification: - Reproduced the exact CI failure locally by putting bash 3.2 (same `cd ""` semantics as CI's 5.2) first on PATH: fails before the fix with the identical message shape, passes after. - tests/fm-tangle-guard.test.sh passes under both bash 3.2 and bash 5.3. - tests/fm-spawn-worktree-settle.test.sh and tests/fm-spawn-pool-base-freshen.test.sh pass; shellcheck -x bin/fm-spawn.sh clean. - bin/fm-test-run.sh --changed: 46 suites completed, every FM_TEST_END exit=0, 1076 passing assertions, 0 "not ok" (including fm-tangle-guard, fm-control-relaunch, fm-lint). The run ended on my own 900s wall-clock cap (rc=124), not on any test failure

* fix: support stock macOS Bash 3.2 paths (#3732)

Co-authored-by: Talon Stark <talonstark@gmail.com>

* fix(bin): read orphaned green ci monitor and daemon-down failed record as not failed (#3846)

* fix(bin): read an orphaned green ci monitor as held-for-merge, not failed

A no-mistakes run held for a captain merge decision keeps its ci step
polling until merged or closed; when the shared daemon restarts under
that poll, the run is recorded failed although every substantive step
completed and GitHub reports the PR green. A monitor whose only
remaining job is to observe a human decision must not convert the
absence of that decision into a failure verdict.

fm-crew-state.sh now reclassifies a terminal failed run as done
(held-for-merge), surfacing the run's PR URL, when the steps table
shows every step completed except exactly ci failed and the ci log's
last recognized marker reads checks green. A genuinely red check, an
unreadable ci log, or a second failed step keeps the failure.

* no-mistakes(review): Read daemon-down coarse failed ledger as unknown, not failed

* no-mistakes(document): docs: align AGENTS.md failed-verdict guidance with crew-state reclassification

* fix(spawn): verify preserved backlog state after interrupted spawn delivery (#3852)

* fix(bin): read preserved spawn state back before the interrupted exit claims it

The deferred-signal exit path asserted the paired task record and
In-flight backlog state were preserved without reading either back,
exactly when a reader is least able to check (fm-yi4j evidence,
2026-09-05). The commit's exit status alone has been observed to agree
with a row that did not actually move.

The exit path now re-reads the record and the row under the same
per-task lock as the commit, repairs a row the commit believed it
moved, and phrases the error as exactly what was verified or attempted
- verified preserved, repaired and verified, or an explicit
preservation-could-not-be-verified with the reason and hand-closeout
instruction. Two behavior tests drive a lying tasks-axi start through a
real interrupted spawn and assert the printed claim and the real
backlog state agree.

* no-mistakes(test): Fix calm suite for Pi 0.85 and pin test umask

* no-mistakes(document): Document interrupted-spawn preservation claim in backlog gate owner

* no-mistakes(ci): Fixed the Greptile P1 in bin/fm-spawn.sh's deferred-signal exit path: during preservation verification, the no-op HUP/INT/TERM re-trap combined with an unresponsive `tasks-axi show`/`start` (bash cannot run traps while a foreground child runs) held the per-task meta lock - and every lifecycle operation waiting on it - indefinitely. Root-cause fix: bound every tasks-axi invocation made under the lock. bin/fm-backlog-transition-lib.sh gains fm_tasks_axi, an exec-based wrapper (GNU timeout, gtimeout fallback) used by fm_backlog_row_show and fm_backlog_mutate that preserves the exact process placement of the plain tasks-axi call; bin/fm-spawn.sh sets FM_TASKS_AXI_TIMEOUT (default 30s) at the commit point so both the commit and the read-back verification are bounded. A timed-out call fails through the existing error plumbing and probe/mutate name the timeout as the reason, so the interrupted exit path prints honest 'preservation could not be verified ... (reason)' wording - never intent phrased as outcome, matching the author's intent. Added a behavior test in tests/fm-backlog-atomicity.test.sh that drives a real interrupted spawn through a lying tasks-axi whose repair start never answers; it asserts the spawn exits promptly (self-bounded by an outer timeout), the attempted wording names the timeout, and the printed claim agrees with the real record/backlog state. Confirmed the test fails on the unfixed tree and passes with the fix. Verified: fm-backlog-atomicity (83 ok), fm-transition-lib, fm-backlog-handoff, fm-captain-hold, fm-teardown, fm-fleet-snapshot-view, fm-secondmate-reconcile, fm-spawn-batch, fm-spawn-dispatch-profile, fm-task-delivery, fm-control-relaunch all pass; bin/fm-lint.sh clean. The fm-bootstrap 'unsplit run lost its local diagnostic' failure reproduces on the pristine base commit and is unrelated to this change

* no-mistakes(ci): Fixed the Greptile P1 in bin/fm-backlog-transition-lib.sh: fm_tasks_axi bounded tasks-axi only through GNU timeout/gtimeout and fell through to an unbounded exec on hosts with neither (stock macOS), so an unresponsive call could hold the per-task meta lock forever during interrupted-spawn verification. Root-cause fix: the bound now has no unbounded path. GNU timeout is preferred, gtimeout next, then a small perl watchdog (fork + waitpid WNOHANG polling at 50ms, TERM on expiry, one bound of grace, then KILL, exit 124 so the callers' existing timeout plumbing reports it; exit statuses and output pass through unchanged). Polling was chosen over alarm+die to avoid perl's platform-dependent syscall-restart semantics. When a bound is requested but no bounding mechanism exists, the call fails closed (exit 127 with a diagnostic) rather than running unbounded, so the interrupted exit path prints honest attempted wording, never intent as outcome. The unbounded exec remains only for the no-bound plain-call case. Added three behavior tests in tests/fm-backlog-atomicity.test.sh driving fm_tasks_axi through a PATH with no timeout binary: a hanging stub must exit 124 within the bound (verified to fail on the pre-fix code), a failing stub's status/output must pass through, and a tool-less PATH must fail closed with the diagnostic. Verified: fm-backlog-atomicity 86/86 ok, fm-transition-lib, fm-backlog-handoff, fm-teardown, fm-spawn-batch, fm-task-delivery, fm-fleet-snapshot-view, fm-secondmate-reconcile, fm-control-relaunch, fm-spawn-dispatch-profile all pass; bin/fm-lint.sh clean

* no-mistakes(ci): Fixed the Greptile P1 on bin/fm-backlog-transition-lib.sh: fm_tasks_axi's GNU timeout and gtimeout paths sent TERM at the bound but had no kill-after, so a tasks-axi that ignores SIGTERM kept the bounded call - and the per-task meta lock - held indefinitely during interrupted-spawn verification. Root-cause fix: both GNU execs now carry -k "$bound" (TERM at the bound, KILL after one further bound of grace), giving every bounded path the same forced-termination contract the perl watchdog already had. Because GNU timeout exits 137 (128+SIGKILL) when the kill-after fires - versus 124 for a TERM expiry - the probe/mutate timeout detection now goes through a new fm_tasks_axi_timeout_expired helper that treats 124 and 137 alike, so the interrupted exit path still names the timeout as the reason; the helper keeps the bound check in one place. Added a behavior test in tests/fm-backlog-atomicity.test.sh that drives fm_tasks_axi through a real GNU timeout with a tasks-axi stub that traps and ignores TERM (the ignored disposition survives exec into sleep) and asserts a bound-expiry status plus completion within bound+grace; on the pre-fix code the suite hangs until killed, confirming the reproduction. Verified: fm-backlog-atomicity 87/87 ok, fm-transition-lib, fm-backlog-handoff, fm-spawn-batch, fm-task-delivery, fm-teardown, fm-secondmate-reconcile, fm-control-relaunch, fm-spawn-dispatch-profile, fm-captain-hold-lifecycle all pass; bin/fm-lint.sh clean

* docs: correct runtime-backend maturity labels for Herdr (#3821)

* docs: correct stale tmux/herdr backend maturity claims

Herdr now has 21 test files, its own required CI job (tests-herdr)
that installs a pinned build and hard-fails on "skip: herdr not
found", while tmux has 3 test files and is only required as a
dependency of the portable-serial e2e lane. zellij, orca, and cmux
still have no CI lane at all. AGENTS.md and docs/herdr-backend.md
still called Herdr merely "experimental" alongside those three,
misleading every session and reader about actual coverage.

Update AGENTS.md's config/backend entry, the opening lines of
docs/herdr-backend.md and docs/tmux-backend.md, the runtime-backend
section of docs/configuration.md, and the matching claims in
docs/architecture.md, CONTRIBUTING.md, and README.md so they agree
and distinguish tmux (default), herdr (own required CI lane, largest
suite, Windows still spike-only), and zellij/orca/cmux (still
experimental, no CI lane). No behavior, selection order, or
dispatch logic changes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDYsWPEfwTuPNjQ2nGBCcj

* no-mistakes(review): docs: fix stale herdr label and CI-lane wording

* no-mistakes(document): docs: align tmux adapter label in scripts.md

* no-mistakes(review): docs: drop duplicated herdr CI claim from tmux page

* no-mistakes(review): docs: drop windows claim, align contributing backend wording

* no-mistakes(review): docs: trim duplicated CI claim from herdr opening line

* no-mistakes(review): docs: drop unguarded largest-test-suite superlative

* no-mistakes(review): docs: restore tmux verified label and README experimental scope

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* fix: restore intent-targeted no-mistakes validation (#3865)

* fix: delete deterministic no-mistakes test baseline, restore intent-targeted Test

PR #3644 pinned commands.test to a fm-test-run.sh --changed walk of the
repository's 75-162 tests/*.test.sh scripts. no-mistakes runs commands.test
verbatim and unconditionally after every fix round, so that walk multiplied
by round count: measured at 32.7 minutes per validation versus 3.6 minutes
intent-targeted. Delete the pin and restore the 3.6-minute posture.

Add tests/fm-nm-test-contract.test.sh as a regression guard, parsing
.no-mistakes.yaml as YAML (ruby's bundled Psych, matching the parser
tests/fm-test-run.test.sh already uses for ci.yml) rather than grepping its
text, restoring in legal form what PR #823 added and PR #1282 removed.

Record the rule in docs/configuration.md's "Gate defaults" section (the
authoritative owner CONTRIBUTING.md already points at) and strengthen
CONTRIBUTING.md's existing local-Test guidance to state it plainly: never
configure commands.test to a deterministic test command, complete or partial.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSt9JvrQMVc4u3jPyUFCFC

* no-mistakes(review): Centralize no-mistakes test policy and narrow guard

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* fix: verify Treehouse slot ownership before teardown (#3837)

* fix(bin): verify pool-slot ownership before returning a worktree slot

Workers were killed when cleanup returned a Treehouse pool slot that a
different, live task had already taken. Teardown now proves the slot is
genuinely this task's before releasing it: it refuses when another task
record claims the same live worktree path, or when the endpoint's working
directory contradicts the recorded slot, and that refusal holds under
--force. Slot allocation, metadata publication, ownership verification,
and slot return are serialized across linked firstmate homes, and forced
secondmate cleanup verifies descendant slot ownership before returning
any child worktree.

Regression coverage drives the scripts with two task records naming one
slot path and asserts the live worker survives and its slot is not reset.

* no-mistakes(review): Protect slots across cloned Firstmate homes

* no-mistakes(test): Gate teardown locking on genuine Treehouse slots

* no-mistakes(test): Clarify pooled descendant slot gating

* no-mistakes(test): Synchronize watcher re-arm test on process exit

* no-mistakes(test): Wait for watcher cleanup before timeout escalation

* no-mistakes(document): Document pool-slot ownership safeguards

* no-mistakes(ci): Fixed all reported CI issues: normalized bare local Git origins to the same Treehouse project-lock identity as absolute clone origins; resolved ShellCheck SC1091 with explicit conditional sourcing; and taught concurrent Herdr teardown coverage to retry expected Treehouse lock contention. Added behavioral regression coverage for bare/absolute origin lock identity. Verified endpoint-safety tests, watcher tests, full CI lint, and the previously failing Herdr teardown assertion

* fix(bin): resolve relative origins from repository root

* no-mistakes(ci): Fixed teardown so an exact recorded endpoint may change cwd without falsely vetoing cleanup. Removed cwd-based ownership refusal while preserving cross-home record exclusivity and project locking. Updated behavioral coverage for both foreign slot ownership refusal and moved-cwd teardown success. Endpoint-safety, backend, watcher, checkpoint, and targeted lint checks pass. Real Herdr presentation E2E progressed successfully but exceeded the 600s local timeout

* feat(bin): add verified omp (Oh My Pi) harness adapter for crew, secondmate, and primary (#3867)

* feat: add verified omp (Oh My Pi) harness adapter for crew, secondmate, and primary

Add omp as a verified harness: anchored process-name detection with a
Firstmate-owned FM_OMP_HARNESS launch marker that needs real omp ancestry,
the fm-spawn launch template with foreign-marker clearing, the tracked
.omp/fm-worker-overlay.yml posture overlay, --auto-approve, --cwd, and
pre-launch model validation scoped to providers 'omp models --json' lists.
Workers get a state-resident busy-state extension keyed on agent_end
without willContinue (omp has no agent_settled). The primary gets two
tracked .omp/extensions: a turn-end guard that answers omp's blocking
session_stop hook by compelling one continuation per turn, with the
pre-tool seatbelts and Run-tier session-start delivery, and a watcher
extension ported from the Pi one with fm_watch_arm_omp. Control tables,
composer busy footers, omp's status row as a bare-composer boundary, the
extension supervision model with an omp-keyed ownership proof, the
session-start diagnostic, and the supervision protocol snippet follow.

Verified live on omp 18.1.11 with openai-codex/gpt-6-astra: a Herdr scout
through spawn, busy state, steer, interrupt, exit, and teardown, and the
isolated rpc primary lab through extension auto-discovery, digest
delivery, lock identity, watcher arm, successor and wake delivery, and
the compelled guard continuation.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test: prove the omp guard continuation through a guard spy

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* fix(spawn): clear the gemini marker at the omp launch boundary

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): force the guard stage by freezing the watcher and clear lint findings

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): reap the live lab by path and record omp's rpc shutdown as a note

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): spawn a real secondmate for the discovery rule and classify the omp surfaces

Replace the template-extraction check with a genuine --secondmate launch
pinned to the fake tmux backend, assert the worker extension's handler set
through the executable rather than its bytes, classify the two new omp
surfaces in the documentation inventory, and record the Herdr worker
evidence in the runtime-backends verification doc.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* no-mistakes(review): omp: unverify remote routes, narrow busy regex, drop overlay approval pin

* no-mistakes(review): omp: validate config-pinned model, correct remote and marker docs

* no-mistakes(review): omp: pin config-model validation with a test, trim overlay

* no-mistakes(review): omp: sync guard evidence, drop dead param, map quota family

* no-mistakes(review): omp quota: refuse unmapped prefixes, match bare model scopes

* no-mistakes(document): docs: cover omp in cd-guard, quota, continuity, tmux

* no-mistakes(document): docs: add omp subagent-guard row, fix live test header

* no-mistakes(ci): Fixed both failing behavior shards and the Greptile P1 in bin/fm-composer-lib.sh. Root cause of "Behavior portable serial 1" and "Behavior portable parallel 2": the omp busy regex (FM_DELIVERY_OMP_BUSY_REGEX_DEFAULT) and omp status-row furniture regex (FM_COMPOSER_OMP_STATUS_RE_DEFAULT) used the bracket range [⠁-⣿]; BSD grep on macOS accepts it but GNU grep on Linux CI aborts with "Invalid collation character", failing every omp busy/furniture read (3 assertions across fm-omp-harness, fm-tmux-submit-busy, fm-composer-lib). Replaced the range with one shared explicit alternation FM_OMP_SPINNER_FRAMES_RE of omp 18.1.11's unicode-preset spinner frames (status set ⣾⣽⣻⢿⡿⣟⣯⣷ + activity set ⠋⠙⠹⠸⠼⠴⠦⠧⠇⠏, read from the installed binary), the same pattern the Kimi busy regex already uses in CI. For Greptile's finding (the harness-agnostic furniture rule's first alternative matched any 1–4-byte token + ' · ', so wrapped typed input like 'fix · tests' with the cursor on it regressed from pending to unknown; reproduced locally vs base), pinned that alternative to omp's identity cell (π|󰵗|pi, the icon.omp of each preset in the 18.1.11 binary). Tests: fm-composer-lib.test.sh asserts 'fix · tests' is not furniture, a status-set spinner row is furniture, and the wrapped composer screen reads pending under both locales (CAPS_TMUX cursor 3); fm-omp-harness.test.sh asserts a status-set frame reads busy. New negative cases fail against the pre-fix lib and pass after. Verified: fm-omp-harness, fm-tmux-submit-busy pass via bin/fm-test-run.sh; fm-composer-lib passes all cases except one pre-existing, unrelated local failure (Herdr half-block test uses printf '▀', unsupported by macOS bash 3.2; fails identically on a pristine HEAD export, passes on CI bash 5); shellcheck and bin/fm-lint.sh clean. Caveat: GNU grep is unavailable locally, so the Linux compile was not run directly; the fix uses only constructs already proven on CI's GNU grep (multibyte literal alternations, incl. under LC_ALL=C). Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh, tests/fm-omp-harness.test.sh. No docs needed changes (they describe the rule generically)

* no-mistakes(ci): Greptile Review: fixed. The omp status-row furniture regex FM_COMPOSER_OMP_STATUS_RE_DEFAULT in bin/fm-composer-lib.sh still accepted a literal `pi ·` opening, so wrapped composer input beginning with `pi ·` was truncated and misclassified. Read the installed omp 18.1.11 binary: the ascii preset's `icon.omp` is `pi` but its `sep.dot` separator is ` - ` (unicode/nerd use ` · `), so a real ascii status row never contains `pi ·` and that alternative could only ever match typed text. Removal-first fix: dropped `pi` from the identity alternation (now `(π|󰵗)`) and updated the comment to record why the ascii preset is excluded. Tests (tests/fm-composer-lib.test.sh): added a negative furniture case for 'pi · e · phi as the three constants' and a wrapped-screen assertion (CAPS_TMUX, cursor 3) that a continuation row opening `pi ·` reads pending in both locales; the new case fails against the unfixed lib and passes after. Verified: composer test with the half-block case skipped passes all 33 cases including the omp matrix; bin/fm-test-run.sh tests/fm-omp-harness.test.sh passes; shellcheck -x clean on both files; bin/fm-lint.sh clean. The full composer test via the runner fails locally only on the pre-existing half-block case (bash 3.2 printf cannot emit ▀; passes on CI bash 5), identical to before this change. Docs unchanged (they describe the rule generically and never mention the ascii identity cell). PR must be raised via no-mistakes: not caused by code. attestation.head_sha is cdddc60 while the PR head is cd51cf4 because the pipeline's ci-phase push moved the head; the outer executor's re-push will re-bind the attestation. No file change for that check. Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* fix(pi): invoke Bash helpers correctly on native Windows (#3843)

* Fix Pi shell invocation on native Windows

* no-mistakes(document): Document Pi Windows Bash transport

* no-mistakes(ci): Captain, staged a narrow fix: register the Pi Windows regression for both extension paths, make Windows mode emulation non-failing, and enforce LF shell checkouts. Mapping and coverage checks pass; CI/Require no-mistakes were approval-gated externally

* no-mistakes(review): Cover async Windows branch-outcome Bash invocation

* no-mistakes(review): Preserve Cygwin checks and refresh Windows timing

* no-mistakes(test): Invoke OpenCode operational-input owner through Bash on Windows

* validation-fixture

* no-mistakes(document): Document Windows Bash helper invocation

* no-mistakes(ci): Fixed PR-caused changed-selection failure by removing the malformed tracked evidence artifact and allowing deleted, unconsumed source paths to retire cleanly while preserving fail-closed behavior for live unmapped paths. Added regression coverage. Verified native-Windows Pi shell-seam test passes and --changed selects the Windows regression

---------

Co-authored-by: test <test@example.invalid>

* test(bin): pin teardown outcomes for squash-merged rebased branches (#3870)

* fix(bin): recognise squash-merged rebased work as landed at teardown

A pipeline rebase can leave the local worktree on pre-rebase commits while
GitHub squash-merges the rebased head. The landed-work test then compared
those stale commits against a squashed main and refused cleanup of work
that had already landed.

When the forge reports the recorded PR merged and its merge commit is on
the default branch, treat a local branch that only repeats paths from the
pipeline push as stale rather than unlanded. If the forge is unreachable,
the same coverage check runs against a PR head whose content is already
on default. Extra local paths still refuse.

* fix(bin): drop unprovable squash-rebase landed-work coverage

Path-set coverage treated a diverged local branch as landed whenever it
touched the same files as the squash merge. That accepts the reviewer's
failing sequence: same path, different content, work discarded.

git cherry and merge-tree containment were already too strict on the real
rebase-fold case. No remaining check is both safe and permissive enough
to recognise a stale pre-rebase copy without also accepting unlanded
edits, so that case still refuses.

Keep the proofs that hold: a merged PR head that contains local work, or
a clean content-in-default tree match. Tests now refuse same-path
different content and extra unlanded commits, and still allow a local
branch that followed the pipeline rebase.

* no-mistakes(review): drop recorded-pr-head fallback and reverted-design leftovers

* no-mistakes(review): silence squash-merge stdout corrupting test PR head

* no-mistakes(review): make unlanded follow-up commit sole cause of refusal

* no-mistakes(document): correct stale squash-rebase fixture comments in teardown tests

* no-mistakes(ci): Split the three reported checks: - CI (run 34061098467) and Require no-mistakes (run 34061098460) both concluded `action_required` — approval-gated workflow runs that never executed a step. Not caused by this PR's code; no change can clear them. - Greptile Review was a genuine defect in the new tests: the three new refusal cases (tests/fm-teardown.test.sh) asserted only exit status 1 and a REFUSED line, so a teardown regression that destroyed the worktree, branch, and task record before reporting refusal would still pass. Fix (tests only): added one `assert_refusal_retained_task_state` helper and called it from `test_squash_merged_same_file_different_content_refuses`, `test_squash_merged_rebased_local_with_unlanded_commit_refuses`, and `test_squash_merged_stale_local_refuses_when_forge_unreachable`, each capturing the worktree HEAD before `run_teardown`. It pins that the refusal left the isolated copy on disk, the task branch still checked out at the same unlanded commit, and state/task-x1.meta intact. Verification: the four squash tests pass; a sensitivity probe ran the ALLOW fixture (teardown completes) and pointed the same helper at the outcome — it fires, because a completed teardown detaches/deletes the branch and removes the task record, proving the assertions discriminate. Full tests/fm-teardown.test.sh: 83 passing. bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 + actionlint 1.7.12 (plus an explicit --external-sources pass on the changed file). bin/fm-test-run.sh --check-coverage ok. Caveat: test_herdr_flat_teardown_preflight_refuses_before_changes (mode missing-adapter) fails on this machine. Verified it fails identically on base commit f91a950 via `git archive`, so it is a pre-existing local environment difference untouched by this diff; skipped to run the rest of the suite, not modified

---------

Co-authored-by: Morten Gad <mogad@itm8.com>

* fix(bin): keep supervision armed for registered custom checks (#3860)

* fix(bin): keep supervision armed for registered custom checks

A custom check bound by bin/fm-check-register.sh only ever runs inside the
watcher's check sweep, but fm_supervision_status counted in-flight tasks, the
relay poll shim, and process-event sources as supervision need, and not
registered checks. Tearing down the last task therefore stopped every
home-level check silently until the next spawn.

Count a state/<id>.check.sh that carries its state/<id>.check-trust binding as
supervision need. The relay shim keeps its own trust path and task PR polls
carry no such binding and are torn down with their task, so neither arms a home
by accident. Presence of the binding is the whole test: the sweep validates the
bytes at execution time and wakes firstmate when it rejects one, which is the
outcome an idle home needs.

Closes #3856

* no-mistakes(review): name registered checks in turn-end block banner and doc invariant

* no-mistakes(review): narrow PR poll predicate test to what it proves

* no-mistakes(document): point Grok re-arm step at supervision-need owner

* fix(bin): resolve Treehouse locks for remote secondmate homes (#3883)

* fix(bin): resolve the shared Treehouse project lock inside remote secondmate homes

Every spawn and teardown inside a remote-seeded secondmate home refused,
because the project lock's anchor could not be resolved there.

fm_firstmate_root_home walks a home's parent bindings upward to find the
anchor the lock lives in, and treated a remote parent binding as an error.
A remote-seeded home's parent is on another machine, so that walk can never
succeed from there - and neither can the home's own local descendants, whose
chain terminates at the same record. Both fail closed on every Treehouse-backed
spawn and every pool-slot teardown.

A remote parent now terminates the walk at the home holding it, which is the
correct anchor: a lock taken on this filesystem is neither held nor observable
across that boundary, and that home is already the top of the local tree
teardown's collect_local_firstmate_states enumerates, since that walk skips
remote registry entries for the same reason. Mutual exclusion is unchanged -
every home reachable through local parent links still derives one identical
lock file per project, and an unreadable binding, an unsupported route, an
unreachable local parent, a cycle, and an over-deep chain all still refuse.
Origin-less local-only projects keep resolving through their worktree top.

Regression coverage pins the anchor for the main-home layout, a local
secondmate, a remote-seeded home, and its local child; drives teardown
end-to-end in a remote-seeded home; keeps the cross-home slot-ownership
refusal across that boundary; and proves two homes still serialize on the
one shared lock file.

* no-mistakes(document): Clarify machine-local Treehouse lock ownership

* fix(bearings): repair board listening and decision reconciliation (#3872)

* fix(bearings): repair the board's listening, card hygiene, and reconcile path

Three defects made the fleet board go quiet and then lie about what still
needs the captain.

Never arm a poll on a session that is not live. `lavish-axi <file>` exits 0
even when it refuses to reopen a session the captain ended from the browser,
reporting `status: user-ended` with the same session id, so the build's
exit-status check accepted a dead session, printed `already-armed`, and left
the board reading "not listening". The build now proves the session is live
from a fresh authoritative listing immediately before arming - not from the
establish call's status alone, which is already stale by then - reopens once
when it finds the session ended, and refuses rather than arming when it stays
ended. A reopen also replaces the pre-reopen source generation before
reporting success, so a runner on its way out cannot be mistaken for a
listener, and a board whose source is registered but unowned gets a
replacement started before the build returns.

Let a dead generation's ownership actually move. Reclaiming a claim ran its
capture-reservation cleanup first, and that cleanup re-verifies the recorded
state-root identity, so a claim naming a pid and a process group that were
both provably gone could not be cleared: reconcile reported a start while
nothing attached, and retire refused with "cannot release source ownership".
Reservation records are keyed by claim token and every replacement claims a
fresh one, so they are hygiene, not an ownership invariant. Reclamation now
additionally requires the owning process group to be absent independently,
which keeps a reused pid whose poll child still runs from ever reading as a
gone generation. A live owner and a crashed leader whose owned group survives
are still never reclaimed.

Stop carding decisions whose subject already landed. The build drops a
decision card whose work item or PR appears in the payload's own landed rows,
and one whose task is no longer an open captain call, naming each drop on
stderr. A task whose state cannot be established is kept, because a call
wrongly hidden is worse than a card wrongly shown.

Add the reconcile choice, and make it structurally incapable of closing a
call. Every decision card carries a standard `reconcile` option, injected by
the build rather than left to the composer. The board now emits the picked
option and any freeform note as separate structured fields instead of fusing
them, so a reconcile selection is not expressible as an answer value at all -
the defect that let `reconcile - <note>` reach the intake as an ordinary
answer. The adapter routes selections from that structured field, creation of
a reconcile request is bound to a verified board source rather than the shared
keyed-answer intake, and the intake still refuses the reserved value on every
channel. Each authorization is bound to the captain-hold generation that
produced the card, so an obsolete card cannot close a later call, and both
terminal outcomes require a pending request: `reconcile close` records the
evidence under its own `reconciled` mode so it never reads as the captain's
words, and `reconcile note` leaves the call open. Anything unprovable -
an unversioned row, a missing generation, an unreadable…
wonder-media-fleet Bot pushed a commit to wonder-media/firstmate that referenced this pull request Sep 27, 2026
* fix(bin): resolve captain holds and legacy teardowns on non-markdown backends (#3782)

* fix(bin): close legacy records on the Beads backend honestly

Two pre-Beads reads blocked honest closure of leftover records:

1. fm-captain-hold.sh complete/verify resolved attested legacy hold ids
   only against the live backend and the pre-collapse derived identity, so
   a home whose holds fm-hold-migration rehomed under fm- ids failed with
   an empty-name absence message (the resolve failure was swallowed by the
   command substitution feeding verify_hold_durable). Resolution now falls
   back, on the Beads backend only, to the legacy id under the configured
   beads prefix and to the row whose notes carry the exact marker line
   'migrated from data/backlog.md id <legacy id>'; every refusal names the
   id it could not resolve, and the markdown path is unchanged.

2. fm-teardown.sh refused any record without spawn_gen forever. A record
   that predates the field can now be torn down with an explicit
   --legacy-record flag once the recovery-grade endpoint classifier
   confirms the recorded endpoint dead or agent-less; the accepted
   incarnation is stamped into the record right before its close marker
   binds to it and named in the teardown line. Refusals leave the record
   byte-identical, the unlanded-work refusal is not relaxed, and a corrupt
   (multi-valued) spawn_gen is never accepted.

The companion repair this branch carries (follow-up commit) is the
backend-gated --file and markdown-file requirement in the mutate path and
lifecycle gates: fm_backlog_mutate passed --file and required the markdown
backlog file regardless of the resolved backend, and the transition gate
plus row probe required that file before any backend work, so a home on a
non-markdown backend could neither gate, probe, nor close its rows.

Behavior tests: self-contained beads fixtures over a scratch bd graph
(self-skipping on markdown-only tasks-axi installs), legacy meta fixtures
for every teardown gate, and the relocated markdown backlog coverage stays
green.

* no-mistakes(review): fix(review): report migrated-hold scan refusals and guard legacy spawn_gen stamp against newline-less records

* fix(backlog): address the configured backend for lifecycle writes

Completes the fm-backlog-transition-lib repair the first commit's message
claims: on this base fm_backlog_mutate passed --file and required the
markdown backlog file regardless of the resolved backend, and
fm_backlog_transition_applies plus fm_backlog_row_probe required that file
before any backend work, so a home on a non-markdown backend could neither
gate, probe, nor close its backlog rows. All three now gate the markdown
file on the resolved tasks-axi backend: markdown keeps exactly its explicit
<data>/backlog.md behavior, non-markdown homes address the backend their
own configuration selects with no markdown file requirement.
fm_backlog_row_show and fm_backlog_row_list already gated correctly and
are unchanged. docs/configuration.md owns the contract line.

Also extends the same backend gate to fm-captain-hold.sh's own mutation
wrapper - hold/add/update/answer/done append the markdown --file only when
the resolved backend is markdown, so a captain call on a Beads home reaches
the Beads store end to end - and applies the review round's two direct
remedies there: the [beads] graph path resolves against the backlog root
when relative (never the process CWD), and a failed bd graph read reports
bd's own trimmed stderr reason in the refusal.

Coverage: tests/fm-backlog-atomicity.test.sh gains a stub-driven Beads
completion case proving the transition gate applies, the row probe reads,
and done runs without any markdown file or --file override; the relocated
markdown backlog test stays green.

* no-mistakes(review): Document root-tasks.toml-only beads settings for migrated-hold resolution

* test(gotmp): stub fm_tasks_axi_backend so the fixture matches the backend-aware transition lib

The legacy-records change made fm-backlog-transition-lib.sh resolve the
configured backend via fm_tasks_axi_backend before the markdown-only skip.
The gotmp fixture's fm-tasks-axi-lib stub lacked that function, so the
markdown check fell through and teardown hit the incompatible-backend
error with unbound FM_TASKS_AXI_MIN under set -u. Stub the backend as
markdown and define the floor, restoring the intended no-backlog skip.

* fix(teardown): roll the legacy stamp back when the close marker fails

A legacy-record teardown stamps its accepted incarnation into the record
right before the close marker binds to it; when that marker write then
fails, the stamp survived, so a retried teardown sailed past the
dead-or-agent-less endpoint gate the stamp now proved unnecessary. The
failed marker write now truncates the record back to its exact pre-stamp
bytes (verified by size), restoring the byte-identical-refusal invariant;
when the rollback itself fails the operator is told to re-run with
--legacy-record after reconciling the endpoint.

Also completes the recorded review decision's coverage wording: the
beads stub test now drives the answer close end to end (update and done
through the gated wrapper), asserting no markdown file override reaches
either verb.

* fix(review): harden the legacy stamp rollback and resolve derived migrated ids

The legacy-record stamp rollback now uses perl (already in the teardown
curated PATH; truncate is not, and is absent on stock macOS), routes every
failure branch inside the stamp block through the same size-verified
rollback so the byte-identical-refusal invariant holds on those paths too,
and gains behavior coverage: an unrecordable close (an invalid pr= link)
fails the teardown, leaves the record byte-identical, keeps the backlog
row in flight, and a flag-less retry still refuses.

Migrated-hold resolution now probes the derived pre-collapse identity
(<origin>-decision-<entry>) alongside the raw entry - fm-hold-migration
recorded the DERIVED id in every migrated row's marker note - in both the
prefix and the migration-note forms, with the ambiguity refusal naming
every identity tried, plus behavior coverage for a bare decision key
resolved through its derived identity's marker.

Also aligns fm-backlog-transition-lib.sh's header ADDRESSING/SCOPE
paragraphs with the backend-gated contract, drops an unreachable FORCE
validity guard the parser rewrite left behind, and switches the new stub
fixture to the portable sed -i.bak idiom.

* no-mistakes(review): Name the configured backend in teardown's backlog reminder

* no-mistakes(review): Scan migration markers before the prefix guess

* no-mistakes(review): Document marker-first resolution and cover the prefix branch

* no-mistakes(document): Record prefix-attestation audit and marker-line forms

* no-mistakes(ci): Fixed the Greptile P1 on bin/fm-teardown.sh: a failed rollback of the synthetic legacy stamp let a retry bypass the dead-or-agent-less endpoint gate. Root cause: teardown minted `spawn_gen=legacy-<ts>-<pid>` into the task record before the close marker bound to it. When the close-marker write failed AND the rollback also failed, the record retained that token. On the next invocation `fm_backlog_meta_spawn_gen` succeeded, so `TEARDOWN_LEGACY_PENDING` stayed 0 and the endpoint gate was skipped entirely — even with `--legacy-record`. The script's own error text told the operator to "re-run teardown with --legacy-record", advice the code could not honor. Fix (bin/fm-teardown.sh): - A `legacy-*` spawn_gen is now recognized as a stamp this teardown path minted, never one a spawn published (fm-spawn.sh publishes `s<epoch>.<pid>.<random>`). Such a record still reads as the legacy record it is: it re-enters the endpoint gate, and a flag-less retry refuses naming `--legacy-record`. - Acceptance reuses the retained token instead of minting a second one; the append block is skipped when the record already carries it, so no duplicate spawn_gen is written. - The rollback attempt and its "could not be rolled back" message are guarded to runs that actually appended a stamp, so a run that appended nothing never claims a rollback it did not perform. - Usage header documents the retained-stamp rule. Test (tests/fm-teardown.test.sh): added `test_retained_legacy_stamp_still_faces_the_endpoint_gate`, an end-to-end reproduction — a `perl` stub that fails only the rollback's `truncate` (delegating every other perl call to the real interpreter) leaves the stamp behind, then the retry must still hit the gate, must not stamp a second incarnation, must not close the backlog row, and the flag-less retry must refuse. Verification: the new test fails against the pre-fix script on exactly the reported defect ("the retry skipped the dead-or-agent-less endpoint gate") and passes after. Full tests/fm-teardown.test.sh 80 ok / 0 failures / rc=0; tests/fm-backlog-atomicity.test.sh 80 ok / 0 failures / rc=0; bin/fm-lint.sh (pinned ShellCheck 0.11.0 + actionlint 1.7.12) clean

* fix: reduce local ShellCheck source-analysis cost (#3778)

* fix(lint): drop source following on the local changed-file gate

The local lint step was inlining library closures through --external-sources
and peaking above 8 GB on a single root. Keep full analysis in CI, on main,
and without a merge-base; exclude the four cross-file codes from the local
pass so those findings still land in CI.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Run local ShellCheck per root and document measurements

* no-mistakes(review): Correct local source-following telemetry

* no-mistakes(document): Clarify context-sensitive lint documentation

---------

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(pi): route decision-owned wake batches to main (#3776)

* fix(pi): route needs-decision wakes and mixed batches wholly to main

Skip the supervision branch for every needs-decision status append, the
same way a check-kind wake already skips it. A coalesced signal/stale
trigger batch containing any needs-decision row is delivered wholly to
main, not split between the branch and a later main wake - the whole
batch, including any co-present routine rows for a different task,
travels together. Heartbeat and unread-status scans stay independent.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE

* test(pi): cover distinct-file mixed batches and heartbeat independence

Add a regression using two distinct files (not the same status file
twice) in one coalesced trigger so a some-vs-every regression on the
file-list cross-reference cannot hide behind a degenerate same-key
case, and a heartbeat/needs-decision co-presence test proving a
needs-decision row neither vetoes nor rides along with an otherwise
eligible heartbeat scan.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE

* no-mistakes(document): Clarify needs-decision and heartbeat routing

* no-mistakes(ci): Fixed captain-held stale reminders so they bypass supervision and wake main, while unrelated unread rows and heartbeats remain independent. Added routing regressions and updated documentation. Pi watcher tests, strict TypeScript checks, lint, and diff checks pass. The no-mistakes attestation failure was pipeline-state related, not a source defect

* no-mistakes(ci): Fixed CI lint by narrowly suppressing false-positive SC2031 diagnostics where background PIDs are captured immediately in the same shell. Verified with `CI=true bin/fm-lint.sh` and `git diff --check`. The no-mistakes attestation failure is pipeline-state-related (`test` was skipped), not a source defect

* no-mistakes(review): Route stale open decisions directly to main

* no-mistakes(review): Honor configured verbs in stale decision routing

* no-mistakes(review): Route second-mate escalations and configured decisions to main

* no-mistakes(review): Ignore trailing whitespace after captain holds

* no-mistakes(review): Cache stale decision classification per status file

* no-mistakes(review): Document unread decision precedence for later task wakes

* no-mistakes(review): Cache unchanged stale decisions across scope scans

* no-mistakes(review): Resolve decision aliases and reject symlinked statuses

* no-mistakes(review): Route surfaced captain-held signals directly to main

* no-mistakes(document): Document decision-owned main routing

* no-mistakes(ci): Fixed captain-held spans to remain actionable while crew working evidence is positive, ensuring the watcher delivers their main-only marker. Updated the executable regression test to cover this case. Verified with the full fm-watch-triage suite, bash syntax checks, and git diff checks. Shellcheck reported only pre-existing test harness warnings (SC1091/SC2034)

* no-mistakes(ci): Fixed the CI regression: captain-held transfers now retain their established non-actionable stale classification while the signal-routing side-band still surfaces them main-only. Verified with tests/fm-daemon.test.sh, tests/fm-watch-triage.test.sh, bash syntax checks, and git diff --check

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* feat(bin): add opt-in worker launch environment allowlist (#3802)

* feat(spawn): add an opt-in worker environment allowlist

Honor a home-local launch-env-allowlist at the shared worker command
boundary and inherit it into secondmate homes. Preserve the existing
launch behavior when the file is absent. Keep the operational environment
and explicit launch assignments, and account for filtered Muse credentials.

Refs https://github.com/kunchenguid/firstmate/issues/3742

Verification:
- Red on origin/main 1820316b66ac2c68e244dd04a02512859ee8c1f4:
  the new enabled-allowlist regression observed synthetic-unrelated in
  the worker; the absent-file control passed.
- Green: fm-test-run.sh on fm-spawn-dispatch-profile, fm-muse-harness,
  and fm-trace-context-spawn; all three passed without skips.
- Synthetic emitted-command probes ran through sh, stock Bash, and zsh.
- Canonical lint, documentation audience checks, and stock Bash syntax
  checks passed.

* no-mistakes(review): Reject inaccessible launch environment configuration

* no-mistakes(review): Preserve inherited allowlists on source inspection errors

* no-mistakes(document): Clarify worker environment grants and inheritance documentation

* no-mistakes(lint): Fix inheritance test ShellCheck source boundary

* feat(bin): add rovo crewmate/scout adapter with home-path file access (#3575)

* feat(bin): add rovo as a verified crewmate/scout worker harness

Wire the Atlassian Rovo CLI (202609.1.2) into the TUI-under-tmux/herdr
adapter contract: detection with marker-precedence ordering, one-shot
positional launch with --startup-receipt readiness polling instead of
composer scraping, model/effort flags, a screen-scrape busy fallback
scoped like grok's, and crew/scout-only lifecycle control that refuses
secondmate launches. Ships with a portable regression suite, a live PTY
guard against the real binary, a per-harness reference doc, and a dated
verification record covering the silent OAuth refresh, the interrupt-ack
divergence from the originating scout report, and the still-open
composer-ghost and tmux/herdr pane-liveness gaps.

* no-mistakes(review): revert rovo launch to positional brief, drop startup-receipt

* no-mistakes(test): rewire rovo adapter to kimi-style launch-then-send shape

* no-mistakes(document): add rovo to stale worker-harness enumerations in docs

* docs(verification): close the rovo herdr-liveness gap with live isolated-lab evidence

Placement, launch-then-send, and busy/idle rendering are now verified live
in an isolated non-default Herdr lab session (bin/fm-herdr-lab.sh), driven
directly through fm-spawn.sh's/fm-backend.sh's own shared primitives since
the cross-session launcher-identity guard refuses this task's own ambient
Herdr identity for a full fm-spawn.sh run.

fm_backend_agent_state reported dead for a live, responding rovo pane at
every point checked, because herdr's own agent-integration registry has no
rovo entry (herdr integration status), so herdr agent get returns
agent_not_found regardless of whether rovo is actually running. This is
recorded as a Herdr-side integration gap rather than a firstmate bug, left
unpatched to avoid a false-positive alive verdict for other idle shells.

Updates docs/verification/rovo.md's backend-liveness section and its two
cross-references (docs/verification/runtime-backends.md, docs/configuration.md)
accordingly.

* fix(bin): close rovo's failed-spawn leak and busy-scrape false idle

Greptile P1s on PR #3575: a failed rovo readiness/submission/delivery gate
exited without tearing down the just-created endpoint, leaving the launched
--yolo rovo process running as an orphaned agent outside task control.
Separately, the busy classifier's rendered-tail fallback returned definitive
idle whenever the "Rovo is thinking" marker scrolled out of the last 12
nonblank lines of a long turn, which could make supervision wrongly conclude
a still-working worker had gone idle.

fm-spawn.sh: rovo_spawn_fail now calls rovo_endpoint_cleanup, which kills the
created endpoint (tmux/herdr/zellij/cmux) via the same generic fm_backend_kill
dispatch fm-spawn.sh's own orca-abort path already uses; orca's worktree and
terminal remain owned by the separate ORCA_ABORT_CLEANUP trap.

fm-busy-lib.sh: the rovo classifier arm now reports "unknown rovo-regex"
instead of "idle rovo-regex" when the marker is absent, matching how muse and
cursor already express "can't tell" for their own fallbacks. The positive
busy match is unchanged.

Extends tests/fm-rovo-harness.test.sh: the readiness and delivery failure
tests now assert the endpoint is torn down (and the success test asserts it
is not), and a new test drives the busy marker out of the tail window to
confirm the verdict is unknown, never idle. bin/fm-lint.sh is clean on both
changed files.

* test(rovo): align spawn fixture with the launch-brief validation contract

Upstream main now requires a brief's ## Captain's intent and
## Firstmate spec subsections (or a nonempty legacy # Task body)
before spawn, and rewrites ship+no-mistakes briefs into
launch-brief.md. Update the rovo harness fixture and pointer
assertions to match, mirroring the kimi harness fixture.

* no-mistakes(review): align rovo.md delivery-gate note with live herdr evidence

* fix: prevent stale supervision wake loops (#3672)

* fix(bin): stop the supervision branch's stale-ack and ghost-report loops

Clean-slate implementation of the four authorized recommendations from the
supervision-ghost-retrigger analysis (items 1, 2, 3, and 7), in their minimal
form, superseding PR #3604:

- fm_branch_report refuses a task the wake being handled never named. The
  extension fixes the reportable task set from the eligible rows before each
  prompt (signal and stale rows resolve to their tasks, a heartbeat allows any
  task with a live record, fleet is always allowed), so a report typed from
  memory about a task whose records teardown already removed is never stored
  or delivered.
- An acknowledgement that consumes nothing says "nothing was acknowledged
  through N" and prints the exact --ack-through / --recovery-generation
  command for the current presented wake, instead of "re-run the drain",
  which re-fed the same stale acknowledgement in a loop.
- bin/fm-guard.sh no longer tells the branch actor to drain queued wakes
  while it is handling them; it names the granted rows instead.
- Teardown removes state/.<task>.branch-outcome-index for ordinary tasks and
  descendants; the index rebuild and the append-side index write both skip a
  task with neither a live record nor a status log, so the branch's report of
  a teardown it just performed is stored without recreating the index.

No new locking, no spawn-generation binding, and no retired-task refusal: the
branch can still report the outcome of a task it just tore down, and the
teardown test now proves that path end to end.

* fix(bin): narrow the branch report scope and guard silence to the minimal form

Apply the four review decisions on the clean-slate branch:

- A signal or stale prompt may report only the tasks its own rows resolve
  to; fleet is refused there too. A heartbeat review is not scoped by task
  at all, so the extension no longer tracks live task records and refuses
  nothing by task id during a fleet review.
- The outcome-index rebuild no longer skips retired tasks; the append-side
  skip alone keeps a torn-down task's index from being recreated.
- bin/fm-guard.sh keeps the queued-wakes warning silent for the branch actor
  instead of printing a replacement note.

* no-mistakes(document): Align supervision docs with scoped wake handling

* fix(bin): grant rovo the per-task home paths its standard crewmate flow needs

rovo confines every file-tool operation to its worktree by default, and its
bash tool independently refuses the same external paths regardless of any
grant (confirmed live), so a rovo worker could not read its own brief or
steering messages or write its status/report - all of which live in the
firstmate home outside the worktree - without hand-feeding it. Grant
toolPermissions.allowedExternalPaths for exactly the task's brief directory,
steering inbox, and status file at launch time via --config-override,
merged with agent.efficiencyLevel into one JSON object since that flag is
single-value and silently discards a second occurrence.

Extends the live PTY guard to prove, against the real binary, that the
grant lets rovo read an external brief and append to an external status
file, and that the same flow is blocked without the grant.

* no-mistakes(document): align rovo reference Effort row with merged single --config-override

* no-mistakes(document): document rovo file-access grant in harness reference

---------

Co-authored-by: PUNEET PATWARI <ppatwari@atlassian.com>
Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>

* fix(bin): prevent false pipeline blocks after drive timeouts (#3813)

* fix(bin): read a crew's pipeline-death claim against the live run

A crew's no-mistakes drive call blocks until the next gate or outcome,
routinely far longer than its harness lets one command live, so the call
gets killed or times out while the daemon runs the fix round on in the
background. Crews read that as daemon death and block on it, and firstmate
had nothing that contradicted them.

Rule 7 of every generated brief now says a drive-call error or a harness
command timeout is not a daemon error, requires `no-mistakes daemon status`
plus `no-mistakes axi status` before a pipeline `blocked:`, and reserves
that report for a refused socket or a run record failed with a daemon
error. The no-mistakes definition of done adds the harness command limit
and the background-and-poll shape that fits inside it.

fm-crew-state gains one classification case: a `blocked:` line blaming the
daemon, a timeout, or unreachability, while the run is running or fixing
AND the pipeline reports fresh activity, now reads as superseded because
the run is alive. Recency comes from the client's own `quiet` marker on
active_steps.last_activity rather than a threshold invented here, and
positive evidence is required, so a run record that outlives a genuinely
dead daemon keeps the plain reading.

stuck-crewmate-recovery gains the inverse-of-a-dead-endpoint playbook:
firstmate reads both statuses itself, steers a reattach, never restarts the
shared daemon on a crew's claim, and escalates only a refused socket.

Nothing here depends on an unshipped no-mistakes capability.

* no-mistakes(review): Prioritize daemon socket failure and narrow unreachable matching

* no-mistakes(review): Honor socket refusal across coarse status and crew guidance

* no-mistakes(test): Replace flaky settle timing assertion with pane-read count

* no-mistakes(document): Document daemon timeout recovery contract

* no-mistakes(ci): Fixed daemon socket failures being suppressed by terminal attributed runs. Positive refused/missing socket evidence now remains blocked regardless of run status. Added a behavioral regression test for terminal failed runs. Verified with fm-crew-state tests, project ShellCheck lint, and git diff checks

* fix(bin): scope the worker role contract for ship and scout launches (#3797)

* fix(brief): scope Firstmate workers to their launch contract

* no-mistakes(document): Clarify supervisor scope and worker contract ownership

* no-mistakes(ci): Removed the heading-based bypass so every ship/scout launch receives the current worker-role contract. Added a regression that failed before the fix and passes afterward. Dispatch, brief, and delivery suites, focused ShellCheck, and git diff --check all passed

* no-mistakes(review): make launch overlay sole owner of worker role contract

* no-mistakes(review): narrow heading test dimension and fix publish error wording

* no-mistakes(review): gate role supersession, fix render guard, drop AGENTS twin

* no-mistakes(document): align architecture AGENTS.md scope and spawn launch-brief header

* fix(bin): stop ringing steering doorbells into dead panes (#3823)

* fix(bin): stop ringing steering doorbells into dead panes

The steering-inbox doorbell was a plain sentence plus Enter typed into a
worker's pane, and the watcher re-rang it on the assumption that a ring is
free. In a pane whose agent has exited that line is a shell command, and the
re-ring ladder kept typing it into a shell that can never acknowledge it.

- Prefix the doorbell with the shell no-op `: ` so a bare shell executes
  nothing while a live worker still reads the same self-describing line.
  `#` is not used because interactive zsh does not treat it as a comment by
  default and the claude harness binds it to memory mode.
- fm_task_inbox_ring skips the pane (return 3) when the backend positively
  classifies the agent as dead; missing, ambiguous, unreadable, and unverified
  endpoints still ring so a blind classifier never starves a live worker.
- The watcher caps the ladder for a dead pane: one stale wake for recovery,
  no ring, no ladder walk, and the durable record stays for
  stuck-crewmate-recovery. fm-send and the remote steer leg report the skip.

Tests cover the no-op in real shells, the dead/live/unclassifiable ring
verdicts, and the single-surfacing watcher path.

* no-mistakes(review): Quote doorbell paths against shell injection

* no-mistakes(review): Reject terminal-control paths before ringing

* no-mistakes(review): Document accepted partial doorbell delivery race

* no-mistakes(review): Skip unavailable endpoints before busy-state handling

* no-mistakes(test): Respect shell startup PATH in environment allowlist test

* no-mistakes(test): Fix doorbell test fixtures for endpoint liveness

* no-mistakes(test): Prioritize confirmed restarts and clean shell test syntax

* no-mistakes(document): Document dead and missing doorbell recovery

* no-mistakes(ci): Fixed the persistence-reply timeout race by rechecking for a correlated reply immediately before falling back to a nudge. Added a deterministic regression covering replies arriving between the preliminary resolution pass and timeout handling. Verified with the targeted restart suite, project lint, coverage guard, bash syntax checks, and diff checks

* fix(bin): wait out transient primary-checkout reads in the spawn worktree poll (#3834)

* fix(spawn): keep the worktree poll from adopting the repository primary

After `treehouse get` is sent, the worktree-discovery poll reads the pane's
foreground-process cwd. While treehouse is still fetching and checking a slot
out, the foreground process is treehouse itself and it reports the repository's
PRIMARY checkout as its cwd for several seconds. The poll accepted any path
that merely differed from the spawning project, so from a linked spawning home
- whose project is itself a worktree of that repository - it adopted the
primary, and the isolation guard then refused a launch whose slot treehouse
went on to create normally.

Screen every candidate with the isolation guard's own conditions, extracted as
spawn_worktree_isolated, so a read the guard would reject stays a transient the
poll keeps waiting through. The two-consecutive-reads rule and the guard as
final backstop are unchanged; a pane that never reaches an isolated worktree
still fails at the existing 60s deadline, now naming the last path it reported.

The already-settled timing assertion counted whole-spawn wall time against a
5s budget and failed on unmodified HEAD on slower machines; it now counts pane
reads, which is what "one confirming read, not an extra cycle" actually means.

* fix(spawn): say which path the worktree wait rejected, and why

Screening every discovery-poll candidate means a host that never reaches an
isolated worktree spends the whole 60s window before refusing. That wait is
deliberate - separating a transient from a terminal misconfiguration needs
machinery this path does not want - so the refusal explains itself instead:
the isolation check records why a candidate failed, and the deadline names the
last path seen together with that reason. Message and diagnostics only; the
poll's control flow is unchanged.

Two suites asserted the guard's wording on paths the poll now rejects rather
than adopts, so their refusal arrives from the deadline instead: realign
fm-tangle-guard's non-git and subdirectory-of-primary cases (each now also
asserting the stated reason, and the second the metadata absence it was
missing) and the herdr projection e2e's forced non-worktree cwd.

* no-mistakes(review): stub poll sleep in tangle-guard spawn isolation test

* no-mistakes(document): document spawn poll isolation screen in fm-spawn header

* test(spawn): make the non-git isolation case non-git anywhere

The refusal-reason assertion for a path outside any repository assumed TMPDIR
is not inside a git repository. Where it is, git walks up from the temporary
directory, finds that repository, and the spawn reports the subdirectory cause
instead - so the case passed or failed on a property of the host rather than on
the behaviour under test.

Build the path under a directory the test then names in GIT_CEILING_DIRECTORIES,
which git documents as not chdir-ing up into a listed directory while looking
for a repository. Git never excludes the directory being searched, so the
ceiling is the parent of the path handed to the spawn.

The assertions pin which cause fired rather than the sentence that explains it,
leaving the operator wording free to improve.

* no-mistakes(document): point spawn poll comment at the isolation screen's comparison

* no-mistakes(ci): Fixed the "Behavior portable serial 2" failure in tests/fm-tangle-guard.test.sh ("non-worktree spawn did not say why the path was rejected (missing: 'not inside a git worktree')"). Root cause, in this PR's code: bin/fm-spawn.sh's spawn_worktree_isolated resolved the git toplevel with `wt_top_real=$(cd "$SPAWN_WT_TOP" ...)`. For a path in no repository, `git rev-parse --show-toplevel` yields empty, and `cd ""` is a SUCCESSFUL no-op on bash before 5.3 (CI's ubuntu-latest ships bash 5.2). The empty toplevel therefore resolved to fm-spawn's own cwd — the CI checkout — so the poll reported "it is a subdirectory of worktree root '/home/runner/work/firstmate/firstmate'" instead of the correct "it is not inside a git worktree". Dev machines with bash 5.3 fail `cd ""`, which is why the suite passed locally and only failed on CI; it is a genuine shell-portability defect in the reason vocabulary this change added, not a test-environment artifact. Fix (smallest root-cause change, 1 line + comment, bin/fm-spawn.sh:2168-2173): guard the empty value so it never reaches `cd` — if [ -n "$SPAWN_WT_TOP" ] && ! wt_top_real=$(cd "$SPAWN_WT_TOP" 2>/dev/null && pwd -P); then No change to the poll's timing or deadline behavior (respecting the recorded refusal-latency and spawn-wt-reason-vocabulary decisions), no new tests, no other files touched. Verification: - Reproduced the exact CI failure locally by putting bash 3.2 (same `cd ""` semantics as CI's 5.2) first on PATH: fails before the fix with the identical message shape, passes after. - tests/fm-tangle-guard.test.sh passes under both bash 3.2 and bash 5.3. - tests/fm-spawn-worktree-settle.test.sh and tests/fm-spawn-pool-base-freshen.test.sh pass; shellcheck -x bin/fm-spawn.sh clean. - bin/fm-test-run.sh --changed: 46 suites completed, every FM_TEST_END exit=0, 1076 passing assertions, 0 "not ok" (including fm-tangle-guard, fm-control-relaunch, fm-lint). The run ended on my own 900s wall-clock cap (rc=124), not on any test failure

* fix: support stock macOS Bash 3.2 paths (#3732)

Co-authored-by: Talon Stark <talonstark@gmail.com>

* fix(bin): read orphaned green ci monitor and daemon-down failed record as not failed (#3846)

* fix(bin): read an orphaned green ci monitor as held-for-merge, not failed

A no-mistakes run held for a captain merge decision keeps its ci step
polling until merged or closed; when the shared daemon restarts under
that poll, the run is recorded failed although every substantive step
completed and GitHub reports the PR green. A monitor whose only
remaining job is to observe a human decision must not convert the
absence of that decision into a failure verdict.

fm-crew-state.sh now reclassifies a terminal failed run as done
(held-for-merge), surfacing the run's PR URL, when the steps table
shows every step completed except exactly ci failed and the ci log's
last recognized marker reads checks green. A genuinely red check, an
unreadable ci log, or a second failed step keeps the failure.

* no-mistakes(review): Read daemon-down coarse failed ledger as unknown, not failed

* no-mistakes(document): docs: align AGENTS.md failed-verdict guidance with crew-state reclassification

* fix(spawn): verify preserved backlog state after interrupted spawn delivery (#3852)

* fix(bin): read preserved spawn state back before the interrupted exit claims it

The deferred-signal exit path asserted the paired task record and
In-flight backlog state were preserved without reading either back,
exactly when a reader is least able to check (fm-yi4j evidence,
2026-09-05). The commit's exit status alone has been observed to agree
with a row that did not actually move.

The exit path now re-reads the record and the row under the same
per-task lock as the commit, repairs a row the commit believed it
moved, and phrases the error as exactly what was verified or attempted
- verified preserved, repaired and verified, or an explicit
preservation-could-not-be-verified with the reason and hand-closeout
instruction. Two behavior tests drive a lying tasks-axi start through a
real interrupted spawn and assert the printed claim and the real
backlog state agree.

* no-mistakes(test): Fix calm suite for Pi 0.85 and pin test umask

* no-mistakes(document): Document interrupted-spawn preservation claim in backlog gate owner

* no-mistakes(ci): Fixed the Greptile P1 in bin/fm-spawn.sh's deferred-signal exit path: during preservation verification, the no-op HUP/INT/TERM re-trap combined with an unresponsive `tasks-axi show`/`start` (bash cannot run traps while a foreground child runs) held the per-task meta lock - and every lifecycle operation waiting on it - indefinitely. Root-cause fix: bound every tasks-axi invocation made under the lock. bin/fm-backlog-transition-lib.sh gains fm_tasks_axi, an exec-based wrapper (GNU timeout, gtimeout fallback) used by fm_backlog_row_show and fm_backlog_mutate that preserves the exact process placement of the plain tasks-axi call; bin/fm-spawn.sh sets FM_TASKS_AXI_TIMEOUT (default 30s) at the commit point so both the commit and the read-back verification are bounded. A timed-out call fails through the existing error plumbing and probe/mutate name the timeout as the reason, so the interrupted exit path prints honest 'preservation could not be verified ... (reason)' wording - never intent phrased as outcome, matching the author's intent. Added a behavior test in tests/fm-backlog-atomicity.test.sh that drives a real interrupted spawn through a lying tasks-axi whose repair start never answers; it asserts the spawn exits promptly (self-bounded by an outer timeout), the attempted wording names the timeout, and the printed claim agrees with the real record/backlog state. Confirmed the test fails on the unfixed tree and passes with the fix. Verified: fm-backlog-atomicity (83 ok), fm-transition-lib, fm-backlog-handoff, fm-captain-hold, fm-teardown, fm-fleet-snapshot-view, fm-secondmate-reconcile, fm-spawn-batch, fm-spawn-dispatch-profile, fm-task-delivery, fm-control-relaunch all pass; bin/fm-lint.sh clean. The fm-bootstrap 'unsplit run lost its local diagnostic' failure reproduces on the pristine base commit and is unrelated to this change

* no-mistakes(ci): Fixed the Greptile P1 in bin/fm-backlog-transition-lib.sh: fm_tasks_axi bounded tasks-axi only through GNU timeout/gtimeout and fell through to an unbounded exec on hosts with neither (stock macOS), so an unresponsive call could hold the per-task meta lock forever during interrupted-spawn verification. Root-cause fix: the bound now has no unbounded path. GNU timeout is preferred, gtimeout next, then a small perl watchdog (fork + waitpid WNOHANG polling at 50ms, TERM on expiry, one bound of grace, then KILL, exit 124 so the callers' existing timeout plumbing reports it; exit statuses and output pass through unchanged). Polling was chosen over alarm+die to avoid perl's platform-dependent syscall-restart semantics. When a bound is requested but no bounding mechanism exists, the call fails closed (exit 127 with a diagnostic) rather than running unbounded, so the interrupted exit path prints honest attempted wording, never intent as outcome. The unbounded exec remains only for the no-bound plain-call case. Added three behavior tests in tests/fm-backlog-atomicity.test.sh driving fm_tasks_axi through a PATH with no timeout binary: a hanging stub must exit 124 within the bound (verified to fail on the pre-fix code), a failing stub's status/output must pass through, and a tool-less PATH must fail closed with the diagnostic. Verified: fm-backlog-atomicity 86/86 ok, fm-transition-lib, fm-backlog-handoff, fm-teardown, fm-spawn-batch, fm-task-delivery, fm-fleet-snapshot-view, fm-secondmate-reconcile, fm-control-relaunch, fm-spawn-dispatch-profile all pass; bin/fm-lint.sh clean

* no-mistakes(ci): Fixed the Greptile P1 on bin/fm-backlog-transition-lib.sh: fm_tasks_axi's GNU timeout and gtimeout paths sent TERM at the bound but had no kill-after, so a tasks-axi that ignores SIGTERM kept the bounded call - and the per-task meta lock - held indefinitely during interrupted-spawn verification. Root-cause fix: both GNU execs now carry -k "$bound" (TERM at the bound, KILL after one further bound of grace), giving every bounded path the same forced-termination contract the perl watchdog already had. Because GNU timeout exits 137 (128+SIGKILL) when the kill-after fires - versus 124 for a TERM expiry - the probe/mutate timeout detection now goes through a new fm_tasks_axi_timeout_expired helper that treats 124 and 137 alike, so the interrupted exit path still names the timeout as the reason; the helper keeps the bound check in one place. Added a behavior test in tests/fm-backlog-atomicity.test.sh that drives fm_tasks_axi through a real GNU timeout with a tasks-axi stub that traps and ignores TERM (the ignored disposition survives exec into sleep) and asserts a bound-expiry status plus completion within bound+grace; on the pre-fix code the suite hangs until killed, confirming the reproduction. Verified: fm-backlog-atomicity 87/87 ok, fm-transition-lib, fm-backlog-handoff, fm-spawn-batch, fm-task-delivery, fm-teardown, fm-secondmate-reconcile, fm-control-relaunch, fm-spawn-dispatch-profile, fm-captain-hold-lifecycle all pass; bin/fm-lint.sh clean

* docs: correct runtime-backend maturity labels for Herdr (#3821)

* docs: correct stale tmux/herdr backend maturity claims

Herdr now has 21 test files, its own required CI job (tests-herdr)
that installs a pinned build and hard-fails on "skip: herdr not
found", while tmux has 3 test files and is only required as a
dependency of the portable-serial e2e lane. zellij, orca, and cmux
still have no CI lane at all. AGENTS.md and docs/herdr-backend.md
still called Herdr merely "experimental" alongside those three,
misleading every session and reader about actual coverage.

Update AGENTS.md's config/backend entry, the opening lines of
docs/herdr-backend.md and docs/tmux-backend.md, the runtime-backend
section of docs/configuration.md, and the matching claims in
docs/architecture.md, CONTRIBUTING.md, and README.md so they agree
and distinguish tmux (default), herdr (own required CI lane, largest
suite, Windows still spike-only), and zellij/orca/cmux (still
experimental, no CI lane). No behavior, selection order, or
dispatch logic changes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDYsWPEfwTuPNjQ2nGBCcj

* no-mistakes(review): docs: fix stale herdr label and CI-lane wording

* no-mistakes(document): docs: align tmux adapter label in scripts.md

* no-mistakes(review): docs: drop duplicated herdr CI claim from tmux page

* no-mistakes(review): docs: drop windows claim, align contributing backend wording

* no-mistakes(review): docs: trim duplicated CI claim from herdr opening line

* no-mistakes(review): docs: drop unguarded largest-test-suite superlative

* no-mistakes(review): docs: restore tmux verified label and README experimental scope

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* fix: restore intent-targeted no-mistakes validation (#3865)

* fix: delete deterministic no-mistakes test baseline, restore intent-targeted Test

PR #3644 pinned commands.test to a fm-test-run.sh --changed walk of the
repository's 75-162 tests/*.test.sh scripts. no-mistakes runs commands.test
verbatim and unconditionally after every fix round, so that walk multiplied
by round count: measured at 32.7 minutes per validation versus 3.6 minutes
intent-targeted. Delete the pin and restore the 3.6-minute posture.

Add tests/fm-nm-test-contract.test.sh as a regression guard, parsing
.no-mistakes.yaml as YAML (ruby's bundled Psych, matching the parser
tests/fm-test-run.test.sh already uses for ci.yml) rather than grepping its
text, restoring in legal form what PR #823 added and PR #1282 removed.

Record the rule in docs/configuration.md's "Gate defaults" section (the
authoritative owner CONTRIBUTING.md already points at) and strengthen
CONTRIBUTING.md's existing local-Test guidance to state it plainly: never
configure commands.test to a deterministic test command, complete or partial.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSt9JvrQMVc4u3jPyUFCFC

* no-mistakes(review): Centralize no-mistakes test policy and narrow guard

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* fix: verify Treehouse slot ownership before teardown (#3837)

* fix(bin): verify pool-slot ownership before returning a worktree slot

Workers were killed when cleanup returned a Treehouse pool slot that a
different, live task had already taken. Teardown now proves the slot is
genuinely this task's before releasing it: it refuses when another task
record claims the same live worktree path, or when the endpoint's working
directory contradicts the recorded slot, and that refusal holds under
--force. Slot allocation, metadata publication, ownership verification,
and slot return are serialized across linked firstmate homes, and forced
secondmate cleanup verifies descendant slot ownership before returning
any child worktree.

Regression coverage drives the scripts with two task records naming one
slot path and asserts the live worker survives and its slot is not reset.

* no-mistakes(review): Protect slots across cloned Firstmate homes

* no-mistakes(test): Gate teardown locking on genuine Treehouse slots

* no-mistakes(test): Clarify pooled descendant slot gating

* no-mistakes(test): Synchronize watcher re-arm test on process exit

* no-mistakes(test): Wait for watcher cleanup before timeout escalation

* no-mistakes(document): Document pool-slot ownership safeguards

* no-mistakes(ci): Fixed all reported CI issues: normalized bare local Git origins to the same Treehouse project-lock identity as absolute clone origins; resolved ShellCheck SC1091 with explicit conditional sourcing; and taught concurrent Herdr teardown coverage to retry expected Treehouse lock contention. Added behavioral regression coverage for bare/absolute origin lock identity. Verified endpoint-safety tests, watcher tests, full CI lint, and the previously failing Herdr teardown assertion

* fix(bin): resolve relative origins from repository root

* no-mistakes(ci): Fixed teardown so an exact recorded endpoint may change cwd without falsely vetoing cleanup. Removed cwd-based ownership refusal while preserving cross-home record exclusivity and project locking. Updated behavioral coverage for both foreign slot ownership refusal and moved-cwd teardown success. Endpoint-safety, backend, watcher, checkpoint, and targeted lint checks pass. Real Herdr presentation E2E progressed successfully but exceeded the 600s local timeout

* feat(bin): add verified omp (Oh My Pi) harness adapter for crew, secondmate, and primary (#3867)

* feat: add verified omp (Oh My Pi) harness adapter for crew, secondmate, and primary

Add omp as a verified harness: anchored process-name detection with a
Firstmate-owned FM_OMP_HARNESS launch marker that needs real omp ancestry,
the fm-spawn launch template with foreign-marker clearing, the tracked
.omp/fm-worker-overlay.yml posture overlay, --auto-approve, --cwd, and
pre-launch model validation scoped to providers 'omp models --json' lists.
Workers get a state-resident busy-state extension keyed on agent_end
without willContinue (omp has no agent_settled). The primary gets two
tracked .omp/extensions: a turn-end guard that answers omp's blocking
session_stop hook by compelling one continuation per turn, with the
pre-tool seatbelts and Run-tier session-start delivery, and a watcher
extension ported from the Pi one with fm_watch_arm_omp. Control tables,
composer busy footers, omp's status row as a bare-composer boundary, the
extension supervision model with an omp-keyed ownership proof, the
session-start diagnostic, and the supervision protocol snippet follow.

Verified live on omp 18.1.11 with openai-codex/gpt-6-astra: a Herdr scout
through spawn, busy state, steer, interrupt, exit, and teardown, and the
isolated rpc primary lab through extension auto-discovery, digest
delivery, lock identity, watcher arm, successor and wake delivery, and
the compelled guard continuation.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test: prove the omp guard continuation through a guard spy

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* fix(spawn): clear the gemini marker at the omp launch boundary

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): force the guard stage by freezing the watcher and clear lint findings

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): reap the live lab by path and record omp's rpc shutdown as a note

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): spawn a real secondmate for the discovery rule and classify the omp surfaces

Replace the template-extraction check with a genuine --secondmate launch
pinned to the fake tmux backend, assert the worker extension's handler set
through the executable rather than its bytes, classify the two new omp
surfaces in the documentation inventory, and record the Herdr worker
evidence in the runtime-backends verification doc.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* no-mistakes(review): omp: unverify remote routes, narrow busy regex, drop overlay approval pin

* no-mistakes(review): omp: validate config-pinned model, correct remote and marker docs

* no-mistakes(review): omp: pin config-model validation with a test, trim overlay

* no-mistakes(review): omp: sync guard evidence, drop dead param, map quota family

* no-mistakes(review): omp quota: refuse unmapped prefixes, match bare model scopes

* no-mistakes(document): docs: cover omp in cd-guard, quota, continuity, tmux

* no-mistakes(document): docs: add omp subagent-guard row, fix live test header

* no-mistakes(ci): Fixed both failing behavior shards and the Greptile P1 in bin/fm-composer-lib.sh. Root cause of "Behavior portable serial 1" and "Behavior portable parallel 2": the omp busy regex (FM_DELIVERY_OMP_BUSY_REGEX_DEFAULT) and omp status-row furniture regex (FM_COMPOSER_OMP_STATUS_RE_DEFAULT) used the bracket range [⠁-⣿]; BSD grep on macOS accepts it but GNU grep on Linux CI aborts with "Invalid collation character", failing every omp busy/furniture read (3 assertions across fm-omp-harness, fm-tmux-submit-busy, fm-composer-lib). Replaced the range with one shared explicit alternation FM_OMP_SPINNER_FRAMES_RE of omp 18.1.11's unicode-preset spinner frames (status set ⣾⣽⣻⢿⡿⣟⣯⣷ + activity set ⠋⠙⠹⠸⠼⠴⠦⠧⠇⠏, read from the installed binary), the same pattern the Kimi busy regex already uses in CI. For Greptile's finding (the harness-agnostic furniture rule's first alternative matched any 1–4-byte token + ' · ', so wrapped typed input like 'fix · tests' with the cursor on it regressed from pending to unknown; reproduced locally vs base), pinned that alternative to omp's identity cell (π|󰵗|pi, the icon.omp of each preset in the 18.1.11 binary). Tests: fm-composer-lib.test.sh asserts 'fix · tests' is not furniture, a status-set spinner row is furniture, and the wrapped composer screen reads pending under both locales (CAPS_TMUX cursor 3); fm-omp-harness.test.sh asserts a status-set frame reads busy. New negative cases fail against the pre-fix lib and pass after. Verified: fm-omp-harness, fm-tmux-submit-busy pass via bin/fm-test-run.sh; fm-composer-lib passes all cases except one pre-existing, unrelated local failure (Herdr half-block test uses printf '▀', unsupported by macOS bash 3.2; fails identically on a pristine HEAD export, passes on CI bash 5); shellcheck and bin/fm-lint.sh clean. Caveat: GNU grep is unavailable locally, so the Linux compile was not run directly; the fix uses only constructs already proven on CI's GNU grep (multibyte literal alternations, incl. under LC_ALL=C). Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh, tests/fm-omp-harness.test.sh. No docs needed changes (they describe the rule generically)

* no-mistakes(ci): Greptile Review: fixed. The omp status-row furniture regex FM_COMPOSER_OMP_STATUS_RE_DEFAULT in bin/fm-composer-lib.sh still accepted a literal `pi ·` opening, so wrapped composer input beginning with `pi ·` was truncated and misclassified. Read the installed omp 18.1.11 binary: the ascii preset's `icon.omp` is `pi` but its `sep.dot` separator is ` - ` (unicode/nerd use ` · `), so a real ascii status row never contains `pi ·` and that alternative could only ever match typed text. Removal-first fix: dropped `pi` from the identity alternation (now `(π|󰵗)`) and updated the comment to record why the ascii preset is excluded. Tests (tests/fm-composer-lib.test.sh): added a negative furniture case for 'pi · e · phi as the three constants' and a wrapped-screen assertion (CAPS_TMUX, cursor 3) that a continuation row opening `pi ·` reads pending in both locales; the new case fails against the unfixed lib and passes after. Verified: composer test with the half-block case skipped passes all 33 cases including the omp matrix; bin/fm-test-run.sh tests/fm-omp-harness.test.sh passes; shellcheck -x clean on both files; bin/fm-lint.sh clean. The full composer test via the runner fails locally only on the pre-existing half-block case (bash 3.2 printf cannot emit ▀; passes on CI bash 5), identical to before this change. Docs unchanged (they describe the rule generically and never mention the ascii identity cell). PR must be raised via no-mistakes: not caused by code. attestation.head_sha is cdddc60 while the PR head is cd51cf4 because the pipeline's ci-phase push moved the head; the outer executor's re-push will re-bind the attestation. No file change for that check. Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* fix(pi): invoke Bash helpers correctly on native Windows (#3843)

* Fix Pi shell invocation on native Windows

* no-mistakes(document): Document Pi Windows Bash transport

* no-mistakes(ci): Captain, staged a narrow fix: register the Pi Windows regression for both extension paths, make Windows mode emulation non-failing, and enforce LF shell checkouts. Mapping and coverage checks pass; CI/Require no-mistakes were approval-gated externally

* no-mistakes(review): Cover async Windows branch-outcome Bash invocation

* no-mistakes(review): Preserve Cygwin checks and refresh Windows timing

* no-mistakes(test): Invoke OpenCode operational-input owner through Bash on Windows

* validation-fixture

* no-mistakes(document): Document Windows Bash helper invocation

* no-mistakes(ci): Fixed PR-caused changed-selection failure by removing the malformed tracked evidence artifact and allowing deleted, unconsumed source paths to retire cleanly while preserving fail-closed behavior for live unmapped paths. Added regression coverage. Verified native-Windows Pi shell-seam test passes and --changed selects the Windows regression

---------

Co-authored-by: test <test@example.invalid>

* test(bin): pin teardown outcomes for squash-merged rebased branches (#3870)

* fix(bin): recognise squash-merged rebased work as landed at teardown

A pipeline rebase can leave the local worktree on pre-rebase commits while
GitHub squash-merges the rebased head. The landed-work test then compared
those stale commits against a squashed main and refused cleanup of work
that had already landed.

When the forge reports the recorded PR merged and its merge commit is on
the default branch, treat a local branch that only repeats paths from the
pipeline push as stale rather than unlanded. If the forge is unreachable,
the same coverage check runs against a PR head whose content is already
on default. Extra local paths still refuse.

* fix(bin): drop unprovable squash-rebase landed-work coverage

Path-set coverage treated a diverged local branch as landed whenever it
touched the same files as the squash merge. That accepts the reviewer's
failing sequence: same path, different content, work discarded.

git cherry and merge-tree containment were already too strict on the real
rebase-fold case. No remaining check is both safe and permissive enough
to recognise a stale pre-rebase copy without also accepting unlanded
edits, so that case still refuses.

Keep the proofs that hold: a merged PR head that contains local work, or
a clean content-in-default tree match. Tests now refuse same-path
different content and extra unlanded commits, and still allow a local
branch that followed the pipeline rebase.

* no-mistakes(review): drop recorded-pr-head fallback and reverted-design leftovers

* no-mistakes(review): silence squash-merge stdout corrupting test PR head

* no-mistakes(review): make unlanded follow-up commit sole cause of refusal

* no-mistakes(document): correct stale squash-rebase fixture comments in teardown tests

* no-mistakes(ci): Split the three reported checks: - CI (run 34061098467) and Require no-mistakes (run 34061098460) both concluded `action_required` — approval-gated workflow runs that never executed a step. Not caused by this PR's code; no change can clear them. - Greptile Review was a genuine defect in the new tests: the three new refusal cases (tests/fm-teardown.test.sh) asserted only exit status 1 and a REFUSED line, so a teardown regression that destroyed the worktree, branch, and task record before reporting refusal would still pass. Fix (tests only): added one `assert_refusal_retained_task_state` helper and called it from `test_squash_merged_same_file_different_content_refuses`, `test_squash_merged_rebased_local_with_unlanded_commit_refuses`, and `test_squash_merged_stale_local_refuses_when_forge_unreachable`, each capturing the worktree HEAD before `run_teardown`. It pins that the refusal left the isolated copy on disk, the task branch still checked out at the same unlanded commit, and state/task-x1.meta intact. Verification: the four squash tests pass; a sensitivity probe ran the ALLOW fixture (teardown completes) and pointed the same helper at the outcome — it fires, because a completed teardown detaches/deletes the branch and removes the task record, proving the assertions discriminate. Full tests/fm-teardown.test.sh: 83 passing. bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 + actionlint 1.7.12 (plus an explicit --external-sources pass on the changed file). bin/fm-test-run.sh --check-coverage ok. Caveat: test_herdr_flat_teardown_preflight_refuses_before_changes (mode missing-adapter) fails on this machine. Verified it fails identically on base commit f91a950 via `git archive`, so it is a pre-existing local environment difference untouched by this diff; skipped to run the rest of the suite, not modified

---------

Co-authored-by: Morten Gad <mogad@itm8.com>

* fix(bin): keep supervision armed for registered custom checks (#3860)

* fix(bin): keep supervision armed for registered custom checks

A custom check bound by bin/fm-check-register.sh only ever runs inside the
watcher's check sweep, but fm_supervision_status counted in-flight tasks, the
relay poll shim, and process-event sources as supervision need, and not
registered checks. Tearing down the last task therefore stopped every
home-level check silently until the next spawn.

Count a state/<id>.check.sh that carries its state/<id>.check-trust binding as
supervision need. The relay shim keeps its own trust path and task PR polls
carry no such binding and are torn down with their task, so neither arms a home
by accident. Presence of the binding is the whole test: the sweep validates the
bytes at execution time and wakes firstmate when it rejects one, which is the
outcome an idle home needs.

Closes #3856

* no-mistakes(review): name registered checks in turn-end block banner and doc invariant

* no-mistakes(review): narrow PR poll predicate test to what it proves

* no-mistakes(document): point Grok re-arm step at supervision-need owner

* fix(bin): resolve Treehouse locks for remote secondmate homes (#3883)

* fix(bin): resolve the shared Treehouse project lock inside remote secondmate homes

Every spawn and teardown inside a remote-seeded secondmate home refused,
because the project lock's anchor could not be resolved there.

fm_firstmate_root_home walks a home's parent bindings upward to find the
anchor the lock lives in, and treated a remote parent binding as an error.
A remote-seeded home's parent is on another machine, so that walk can never
succeed from there - and neither can the home's own local descendants, whose
chain terminates at the same record. Both fail closed on every Treehouse-backed
spawn and every pool-slot teardown.

A remote parent now terminates the walk at the home holding it, which is the
correct anchor: a lock taken on this filesystem is neither held nor observable
across that boundary, and that home is already the top of the local tree
teardown's collect_local_firstmate_states enumerates, since that walk skips
remote registry entries for the same reason. Mutual exclusion is unchanged -
every home reachable through local parent links still derives one identical
lock file per project, and an unreadable binding, an unsupported route, an
unreachable local parent, a cycle, and an over-deep chain all still refuse.
Origin-less local-only projects keep resolving through their worktree top.

Regression coverage pins the anchor for the main-home layout, a local
secondmate, a remote-seeded home, and its local child; drives teardown
end-to-end in a remote-seeded home; keeps the cross-home slot-ownership
refusal across that boundary; and proves two homes still serialize on the
one shared lock file.

* no-mistakes(document): Clarify machine-local Treehouse lock ownership

* fix(bearings): repair board listening and decision reconciliation (#3872)

* fix(bearings): repair the board's listening, card hygiene, and reconcile path

Three defects made the fleet board go quiet and then lie about what still
needs the captain.

Never arm a poll on a session that is not live. `lavish-axi <file>` exits 0
even when it refuses to reopen a session the captain ended from the browser,
reporting `status: user-ended` with the same session id, so the build's
exit-status check accepted a dead session, printed `already-armed`, and left
the board reading "not listening". The build now proves the session is live
from a fresh authoritative listing immediately before arming - not from the
establish call's status alone, which is already stale by then - reopens once
when it finds the session ended, and refuses rather than arming when it stays
ended. A reopen also replaces the pre-reopen source generation before
reporting success, so a runner on its way out cannot be mistaken for a
listener, and a board whose source is registered but unowned gets a
replacement started before the build returns.

Let a dead generation's ownership actually move. Reclaiming a claim ran its
capture-reservation cleanup first, and that cleanup re-verifies the recorded
state-root identity, so a claim naming a pid and a process group that were
both provably gone could not be cleared: reconcile reported a start while
nothing attached, and retire refused with "cannot release source ownership".
Reservation records are keyed by claim token and every replacement claims a
fresh one, so they are hygiene, not an ownership invariant. Reclamation now
additionally requires the owning process group to be absent independently,
which keeps a reused pid whose poll child still runs from ever reading as a
gone generation. A live owner and a crashed leader whose owned group survives
are still never reclaimed.

Stop carding decisions whose subject already landed. The build drops a
decision card whose work item or PR appears in the payload's own landed rows,
and one whose task is no longer an open captain call, naming each drop on
stderr. A task whose state cannot be established is kept, because a call
wrongly hidden is worse than a card wrongly shown.

Add the reconcile choice, and make it structurally incapable of closing a
call. Every decision card carries a standard `reconcile` option, injected by
the build rather than left to the composer. The board now emits the picked
option and any freeform note as separate structured fields instead of fusing
them, so a reconcile selection is not expressible as an answer value at all -
the defect that let `reconcile - <note>` reach the intake as an ordinary
answer. The adapter routes selections from that structured field, creation of
a reconcile request is bound to a verified board source rather than the shared
keyed-answer intake, and the intake still refuses the reserved value on every
channel. Each authorization is bound to the captain-hold generation that
produced the card, so an obsolete card cannot close a later call, and both
terminal outcomes require a pending request: `reconcile close` records the
evidence under its own `reconciled` mode so it never reads as the captain's
words, and `reconcile note` leaves the call open. Anything unprovable -
an unversioned row, a missing generation, an unreadable state - refuses
rather than acting.

Regression coverage fails without each fix, and pins every leak path: a bare
reconcile, a standalone close or note with no pending request, an any-channel
reconcile, an annotated selection from a freeform card, and a
generation-skewed authorization. An opt-in guard re-proves the lavish-axi
shapes and the reopen against the installed tool.

* fix(bin): quote the done comparison in the reconcile intake

shellcheck SC1010 reads the bare word as the loop keyword. The failed run
never reached its lint step, so this shipped in the recovered content.

* no-mistakes(review): Publish reconciled parent resolution before request retirement

* no-mistakes(review): Clarify committed cleanup and reconcile reservation scope

* no-mistakes(review): Preserve remote cards and legacy answer compatibility

* no-mistakes(test): Separate live claim relea…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant