fix(pi): route decision-owned wake batches to main - #3776
Merged
Merged
Conversation
Skip the supervision branch for every needs-decision status append, the same way a check-kind wake already skips it. A coalesced signal/stale trigger batch containing any needs-decision row is delivered wholly to main, not split between the branch and a later main wake - the whole batch, including any co-present routine rows for a different task, travels together. Heartbeat and unread-status scans stay independent. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE
Add a regression using two distinct files (not the same status file twice) in one coalesced trigger so a some-vs-every regression on the file-list cross-reference cannot hide behind a degenerate same-key case, and a heartbeat/needs-decision co-presence test proving a needs-decision row neither vetoes nor rides along with an otherwise eligible heartbeat scan. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE
…pervision and wake main, while unrelated unread rows and heartbeats remain independent. Added routing regressions and updated documentation. Pi watcher tests, strict TypeScript checks, lint, and diff checks pass. The no-mistakes attestation failure was pipeline-state related, not a source defect
… SC2031 diagnostics where background PIDs are captured immediately in the same shell. Verified with `CI=true bin/fm-lint.sh` and `git diff --check`. The no-mistakes attestation failure is pipeline-state-related (`test` was skipped), not a source defect
Confidence Score: 5/5The PR appears safe to merge. No blocking failure remains. Reviews (3): Last reviewed commit: "no-mistakes(ci): Fixed the CI regression..." | Re-trigger Greptile |
…crew working evidence is positive, ensuring the watcher delivers their main-only marker. Updated the executable regression test to cover this case. Verified with the full fm-watch-triage suite, bash syntax checks, and git diff checks. Shellcheck reported only pre-existing test harness warnings (SC1091/SC2034)
…retain their established non-actionable stale classification while the signal-routing side-band still surfaces them main-only. Verified with tests/fm-daemon.test.sh, tests/fm-watch-triage.test.sh, bash syntax checks, and git diff --check
RooseveltAdvisors
added a commit
to RooseveltAdvisors/firstmate
that referenced
this pull request
Sep 5, 2026
Brings in upstream d6660d7 (kunchenguid#3776 pi decision-owned wake batches) and a29cdce (kunchenguid#3778 reduce local ShellCheck source-analysis cost) so this branch stops conflicting with the default branch. Two conflicts, both resolved by keeping the upstream change and layering this branch's addition on top: - CONTRIBUTING.md: kept kunchenguid#3778's "invoke it with no arguments" wording and its "file-set selection, and analysis flags" sentence, and re-applied this branch's backend-purity clause inside the lint-definition list. - bin/fm-lint.sh: kept kunchenguid#3778's LOCAL_NOX_EXCLUDE definition and its header sentence about dropping --external-sources locally, and re-applied this branch's backend-purity note plus the pwd -P resolution of SELF_DIR and ROOT. LOCAL_NOX_EXCLUDE is consumed at the local ShellCheck call site, so dropping it would have broken kunchenguid#3778. Verified: the diff against origin/main for every file those two upstream commits touched contains only this branch's additions, no upstream deletions. CI=true bin/fm-lint.sh exits 0 (ShellCheck 0.11.0 full analysis, backend-purity, actionlint 1.7.12), and bin/fm-test-run.sh tests/fm-lint.test.sh tests/fm-backlog-atomicity.test.sh exits 0 with 0 failures. Additive merge commit: no rebase, no squash, no history rewritten.
zeeshaanahmad
added a commit
to zeeshaanahmad/firstmate
that referenced
this pull request
Sep 6, 2026
…able serial 1" (specifically tests/fm-afk-inject-e2e.test.sh Scenario E: "the digest was cut but carries no truncation marker"). Root cause: this PR merges upstream batch 12, including commit d6660d7 (kunchenguid#3776, "route decision-owned wake batches to main (Pi)"). That commit changed bin/fm-watch.sh so a decision-owned wake-queue row's payload is deliberately marked "needs-decision:$files" instead of "signal:$files", solely so Pi's branch dispatcher (fm-branch-dispatch.ts) excludes that row from what it may claim (per docs/pi-supervision-branch.md). This fork's own bin/fm-supervise-daemon.sh (handle_wake) independently drains the same durable wake queue and classifies each row by pattern-matching the payload's literal prefix (signal:/stale:/check:/heartbeat) — it had no case for "needs-decision:", so any decision-owned row fell through to classify_unknown, producing a short "unknown wake: needs-decision: <raw file paths>" digest instead of running the real per-status note text through classify_signal. Scenario E creates 12 simultaneous decision-owned statuses; the real distilled text is long and must arrive truncated with a marker, but the misclassified path produced a short, unmarked message, failing the assertion. Confirmed via `gh run view --log` (the actual CI failure line) and via code tracing (fm-watch.sh:1875-1890, fm-supervise-daemon.sh handle_wake/handle_durable_wakes) — a genuine merge-consequence regression, not flakiness or a stale/infra check; none of the touched files appear elsewhere in this PR's diff. Fix (bin/fm-supervise-daemon.sh, handle_wake): added a `needs-decision:*` case mirroring `signal:*` — strips the prefix and routes to classify_signal, setting kind=signal — so main still classifies these rows exactly like ordinary signal rows; only Pi's branch-claim exclusion (the actual intent of the upstream change) is unaffected. No other files changed. Verification (completed): `bash -n` passed. Full local run of tests/fm-afk-inject-e2e.test.sh now passes all 5 scenarios (A-E), including Scenario E which previously failed. Full local run of tests/fm-daemon.test.sh (the unit suite covering handle_wake and related daemon logic) passed 126/126 assertions with 0 failures, confirming no regression
lytv
pushed a commit
to lytv/mymate
that referenced
this pull request
Sep 8, 2026
* fix(pi): route needs-decision wakes and mixed batches wholly to main Skip the supervision branch for every needs-decision status append, the same way a check-kind wake already skips it. A coalesced signal/stale trigger batch containing any needs-decision row is delivered wholly to main, not split between the branch and a later main wake - the whole batch, including any co-present routine rows for a different task, travels together. Heartbeat and unread-status scans stay independent. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE * test(pi): cover distinct-file mixed batches and heartbeat independence Add a regression using two distinct files (not the same status file twice) in one coalesced trigger so a some-vs-every regression on the file-list cross-reference cannot hide behind a degenerate same-key case, and a heartbeat/needs-decision co-presence test proving a needs-decision row neither vetoes nor rides along with an otherwise eligible heartbeat scan. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE * no-mistakes(document): Clarify needs-decision and heartbeat routing * no-mistakes(ci): Fixed captain-held stale reminders so they bypass supervision and wake main, while unrelated unread rows and heartbeats remain independent. Added routing regressions and updated documentation. Pi watcher tests, strict TypeScript checks, lint, and diff checks pass. The no-mistakes attestation failure was pipeline-state related, not a source defect * no-mistakes(ci): Fixed CI lint by narrowly suppressing false-positive SC2031 diagnostics where background PIDs are captured immediately in the same shell. Verified with `CI=true bin/fm-lint.sh` and `git diff --check`. The no-mistakes attestation failure is pipeline-state-related (`test` was skipped), not a source defect * no-mistakes(review): Route stale open decisions directly to main * no-mistakes(review): Honor configured verbs in stale decision routing * no-mistakes(review): Route second-mate escalations and configured decisions to main * no-mistakes(review): Ignore trailing whitespace after captain holds * no-mistakes(review): Cache stale decision classification per status file * no-mistakes(review): Document unread decision precedence for later task wakes * no-mistakes(review): Cache unchanged stale decisions across scope scans * no-mistakes(review): Resolve decision aliases and reject symlinked statuses * no-mistakes(review): Route surfaced captain-held signals directly to main * no-mistakes(document): Document decision-owned main routing * no-mistakes(ci): Fixed captain-held spans to remain actionable while crew working evidence is positive, ensuring the watcher delivers their main-only marker. Updated the executable regression test to cover this case. Verified with the full fm-watch-triage suite, bash syntax checks, and git diff checks. Shellcheck reported only pre-existing test harness warnings (SC1091/SC2034) * no-mistakes(ci): Fixed the CI regression: captain-held transfers now retain their established non-actionable stale classification while the signal-routing side-band still surfaces them main-only. Verified with tests/fm-daemon.test.sh, tests/fm-watch-triage.test.sh, bash syntax checks, and git diff --check --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
dnth
added a commit
to dnth/firstmate
that referenced
this pull request
Sep 9, 2026
* fix(bin): derive watcher beacon staleness grace from poll cadence Port of upstream kunchenguid/firstmate kunchenguid#3946, adapted: the fork has no away-mode daemon predicate, so only the shared derivation and its two in-scope readers change. fm-claude-stop-autoarm.sh required the watcher beacon fresh within the flat FM_GUARD_GRACE default (300s), but a healthy watcher touches the beacon once per cycle and can legitimately age up to FM_POLL seconds between touches, so a long-poll home reads stale mid-wait and the hook misjudges a live cycle as down. fm-wake-lib.sh gains fm_poll_derived_grace, the single owner of the max(300, FM_POLL + 60) formula. fm-watch.sh's own pre-acquisition beacon-staleness check and the auto-arm hook both derive their default grace from it, and the hook now exports its resolved FM_GUARD_GRACE to the fm-watch-arm.sh invocations so arm and check agree on one value. * fix(bin): count only actor-presentable wake rows and retire unusable rows Port of upstream kunchenguid/firstmate kunchenguid#3950, adapted to the fork's SCOPED drain and OMP branch-grant machinery. A branch grant already excluded its rows from a main drain, but fm-guard.sh still counted every queued row as pending for the calling actor, so main was told to run a drain that could only print nothing for as long as the branch held the grant. And a queue row that lost the five appended fields or its numeric sequence was counted as queued while it could never be claimed, presented, or named by an --ack-through cutoff, wedging the queue permanently. fm-wake-lib.sh gains the shared per-actor helpers: fm_wake_grant_rows_valid, fm_wake_branch_owner_matches, fm_wake_branch_grant_live, and fm_wake_actor_pending_count - the last reports a pending row when a queue exists but cannot be counted, so an unreadable queue still raises the alarm. fm-wake-drain.sh and fm-wake-grant.sh now delegate their grant row-list and owner-record reads to the library. A main drain retires structurally unusable rows under the queue lock and reports them verbatim (a branch drain never does), and a main drain left with only branch-held rows names the holder instead of exiting silently. fm-guard.sh warns only for rows the calling actor can itself present or retire, and gives main a distinct held-by-branch advisory when that is the whole non-empty queue. tests/fm-wake-queue.test.sh covers the branch-held warning boundary end to end, the uncountable-queue alarm, and main-only retirement. * fix(bin): gate secondmate wake-loop stall alerts on real queue no-progress Port of upstream kunchenguid/firstmate kunchenguid#3943, which lands the detector's final form on this fork (the fork had no foreign-queue stall check at all, so the whole feature arrives already fixed). The watcher now reads the oldest structurally valid actionable row in every endpoint-recorded local secondmate home's durable wake queue, and times the interval since that position last changed instead of the row's age. A queue that is draining is not stalled, so a moved oldest row - drain progress, or a queue reprovisioned under the same task id that restarts its sequence anywhere - ends the no-progress episode and starts a fresh observation interval, recorded by state/.secondmate-wake-progress-<task> as the same epoch-sequence row identity the stall receipts use. Declared external-wait pause rows are not actionable evidence, and a mate provably inside an active turn (an exact busy verdict bounded by FM_BUSY_TURN_MAX_SECS) never escalates. One keyed check notification covers each no-progress episode across watcher and handling crashes; the foreign queue is only ever read. FM_SECONDMATE_WAKE_STALL_SECS defaults to 180 and is the backstop behind the active-turn gate, not a substitute for it. Fork adaptation: the parent watcher keeps its secondmate home-beacon rule, so the ported tests give the fixture mate a fresh .last-watcher-beat under its (stubbed or real) clock to isolate the stall detector; assertions are unchanged from upstream. tests/fm-wake-queue.test.sh covers progress tracking, once-per-episode alerting, declared-pause exclusion, reprovisioned-generation reset, active-turn deferral, and marker symlink rejection. * fix(bin): never type doorbells into dead panes; recover them instead Port of upstream kunchenguid#3823, adapted: the fork's task-inbox ring reserves return code 6 for a positively dead or missing endpoint (codes 3-5 are OMP-native and Hermes-lock outcomes here). The doorbell line itself is now a shell no-op (": ...") with quoted, printable-only paths, so even a liveness race that loses to an exiting agent runs nothing in a bare shell. The watcher escalates a dead/missing endpoint's unhandled record once through the ordinary stale wake - over stale busy-state evidence - instead of walking the re-ring ladder, and fm-send reports the skipped doorbell as recovery, not a re-ring. Remote secondmate send inherits the rc-6 recovery notice through fm-send without a new branch. Secondmate restart also resolves all arrived persist answers before any timeout decision and rechecks at the decision boundary, so a reply landing at the deadline wins over the fallback nudge. Regression coverage: doorbell shell-noop across real shells, control-byte path rejection, dead/missing ring skip, watcher escalate-once-without- ringing, dead-pane override of stale busy state, and the resolution/timeout boundary race. Fake tmux inventories across the suite now list recorded windows and answer pane_current_command so the recovery-grade endpoint check exercises real paths. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(omp): prove the doorbell turn instead of trusting triggerTurn The task-inbox doorbell claimed `.delivered` the moment sendMessage returned, but OMP can downgrade `triggerTurn` to append-only delivery on an idle, non-streaming session: the steer was accepted and reported received while no turn ever ran, leaving an idle worker stranded with an unhandled durable record. The extension now renames an accepted request to `.awaiting-turn` when no turn is open and settles it to `.delivered` only on a `turn_start` or `agent_start` proving the turn actually began. A steer that lands while a turn is open is still delivered immediately, since it joins that turn. When the bound passes with no turn, the extension re-drives the same instruction through `sendUserMessage` - a user prompt an idle session cannot defer - and a failed re-drive reports `.failed` rather than stranding silently. Requests left `.awaiting-turn` by a dead generation re-queue as `.pending` on the next activation. Installs on runtimes without the event or user-prompt surface keep the prior accept-only semantics. On the shell side, `fm_omp_task_doorbell_request_existing` reports an awaiting-turn request as in-flight, so fm-send honestly reports a downgraded doorbell as queued (exit 4 semantics) rather than received while the extension's own recovery still wakes the worker. Regression coverage drives the downgrade end to end: an accepted sendMessage with no turn event stays unproven, the bounded grace re-drives the identical instruction through sendUserMessage exactly once, a firing turn_start settles without any re-drive, an open turn delivers the steer immediately, a failing re-drive reports failed, and a dead generation's unsettled proof re-enters delivery on activate. The suite's tmux fake also now lists the recorded task window so the recovery-grade endpoint check exercises the live path instead of classifying the fixture missing. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(bin): bound detached procevent runners to their owning home's lease A runner is detached into its own process group so it survives the turn that started it, and on its own that lets a runner outlive its whole home: once reparented to init, nothing bounds its lifetime, so a blocking source child - and everything it spawns - can poll forever after the home that owns the work is gone. This is the upstream kunchenguid#3904 port, adapted to the fork's simpler procevent (no extension subsystem or capture reservations). Every claim now records the canonical physical state root with its device, inode, owner, and mode, and every claimed runner spawns a detached owner guard in a separate process group. The guard re-reads the recorded root's .owner-lease marker - refreshed by register, handled, reconcile, sweep, and an attached start's keepalive - and after two consecutive checks cannot prove both the root identity and a fresh lease, it signals the runner's whole process group. Runner descendants inherit FM_PROCEVENT_IN_RUNNER so they cannot keep their own lease alive; a confused agent that deliberately strips the marker is out of scope. Reconcile no longer signals a leaderless surviving group: once the leader is gone, PID/PGID reuse makes the group's provenance unprovable, so the claim stays owned and uncertain instead of risking a foreign group or a second poller on one canonical source. Retirement and reconcile take their signal through an identity-and-leadership-gated group path, and a launch floor paces relaunches per registration generation so a crash-looping source cannot spin. Fork deviations from upstream: the state-root check requires ownership and canonicality but not a private mode (operational homes use group-readable state roots), the serialized launch-floor section releases the source lock at the stamp write because this caller holds no further locked work, central storage keeps its distinct commit point, and the usage window is widened to keep the durability-boundary line visible. Regression coverage: an orphaned runner's group is stopped by its guard after the home's lease expires, a crashed leader's ambiguous group is preserved rather than signalled or replaced, and the registration replacement test now waits for the old generation's source to actually start so the launch-floor supersede check cannot eat the race. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(omp): latch a provider-broken branch with a bounded cooldown probe A settled branch prompt whose last assistant message carries stopReason "error" is a provider failure OMP reports on the message instead of rejecting prompt(), so the old code never latched: every later wake kept paying for a predictably broken branch turn before its settlement could reject to main. Port of upstream kunchenguid#3497, adapted to the fork's extension shape (generation-only guards, offer.accept(promise) settlement). Two consecutive settled provider errors now latch the branch off and merge a captain-facing health note. While latched, main keeps every wake except one recovery probe after each cooldown, which starts at five minutes and doubles to a one-hour ceiling; a probe that ends in a durable fm_branch_report clears the latch with a second note, and a failed probe reschedules by the current cooldown. A settled prompt that produced no durable outcome now rejects its settlement instead of silently claiming its granted wake rows. The branch record keeps its SessionManager so the settled-prompt scan can read only the entries the prompt appended. The strict typecheck also caught two stale errors from the earlier doorbell work: the doorbell options builder returned a Required<> literal missing turnGraceMs, and its `on` property signature could not accept the runtime's overloaded event subscription - the latter is now a method declaration so the parameter stays bivariant. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(omp): start a fresh supervision branch conversation per main session Port of upstream kunchenguid#3600. A branch conversation that outlives its main session keeps reasoning from an older thread's accumulated memory, which competes with today's generated prompt and current fleet rules. The OMP extension kept one persistent conversation keyed on a durable pointer, so /new, /resume, /fork, and reload all reopened the stale thread. The branch record is now scoped to the extension generation: createBranch reopens only a conversation it built under the current generation, and both session_start and session_switch advance the generation, drop the live handle (beginDispose + dispose), and re-anchor the mirror so the fresh branch receives the new main session's dialog from its start. The durable outcome store, not the conversation, carries unacknowledged outcomes across the boundary. The on-disk pointer remains a record for operators and the effort picker's model lookup, not a reopen source. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(omp): give the supervision branch main's live model registry Port of upstream kunchenguid#3871 by equivalent mechanism. A provider an extension registered into main's runtime at run time is invisible to a freshly discovered registry, so pinning or following such a model built a branch that could not resolve its own model. OMP's createAgentSession already takes a modelRegistry option, so instead of upstream's provider-by-provider config copy into an isolated ModelRuntime, the branch session now receives main's live registry directly - read-only sharing: the branch never installs, converts, or overwrites credentials or registrations. The resolve and picker paths already read that same registry, so the build, the pin resolution, and the picker now agree on one model catalog. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(bin/omp): route decision-owned signal and stale wake batches to main Port of upstream kunchenguid#3776. A needs-decision status append, a captain-held status transfer surfaced through the no-verb signal path, or a blocked pending-reply escalation is owned by the captain and must stay on main. The previous logic only forced the check-kind class to main, so a coalesced signal batch that also contained routine rows could end up split: the decision row to the branch, the rest to main, which loses the single-actor ordering the per-actor acknowledgement contract depends on. The bash side now: - `status_span_first_actionable_record` returns a side-band flag when its span carries a needs-decision, a captain-held declaration, or a pending- reply escalation, without changing the event text it already returns. - `signal_files_actionable` surfaces those files (even when the captain-held line is otherwise non-actionable) and populates `FM_SIGNAL_NEEDS_DECISION_FILES`. - The watcher appends the row payload as `needs-decision:$files` for exactly the flagged files; other files in the same batch keep the ordinary reason. The OMP dispatch side now: - `scopeForUnreadWake` parses each mapped task's `.status` log to detect an open needs-decision or a current captain-held declaration, and excludes those rows from `eligibleSeqs` while populating `needsDecisionKeys`. - `offerWakeToBranch` in `fm-primary-omp.ts` cross-references the current signal or stale trigger keys against every unread decision-owned row by task identity: any overlap forces the entire coalesced batch to main. - Heartbeat handling remains independent, and a decision-owned row in one task does not affect eligibility for other tasks. Also fixed a pre-existing full-lint `SC2031` info note in `tests/fm-afk-launch.test.sh` by disabling it file-wide, so the canonical lint run is green again. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(core/omp/pi): settle watcher delivery on runtime acceptance, consume on turn start Port of upstream kunchenguid#3513. A follow-up queued while main is streaming joins that run as a user message without raising before_agent_start, so waiting for the model to consume the wake stalled the successor chain: no successor started, later wakes were not delivered, and the turn-end guard re-armed by hand. The shared core now: - Treats sendFollowUp resolution as delivery; consumption is tracked only to let a session replacement replay a wake the runtime accepted but had not yet read. - Keeps accepted-but-unconsumed wakes in a per-generation unconsumedWakes map keyed by the pending token. - Consumes a wake when the runtime starts a turn carrying the exact wake text: before_agent_start for an idle main, message_start for a streaming main. - Does not mark a pending row delivered or finish cleanup until the wake is consumed or the branch owns it. - Skips over unconsumed rows when picking the next pending record, so the successor chain keeps moving while waiting for a streaming main to read the prior wake. - Defers a verified successor's failure close when the pipeline is still delivering the wake it was started for; the retry runs once that delivery settles. - Adds armRetired so a child the core itself killed does not earn a deferred retry. The Pi and OMP adapters both gain a message_start handler that extracts the user message text and passes it to watch.acknowledgeWake. Documentation in docs/omp-supervision-branch.md now states the delivery versus consumption boundary for main fall-back follow-ups. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * no-mistakes(review): Make supervision provider latch recoverable after transient failures * no-mistakes(review): Correlate doorbells and latch rejected supervision prompts * no-mistakes(review): Correlate doorbell turn proofs and recover prompt failures * no-mistakes(review): Correlate synchronous turn proofs to each doorbell dispatch * no-mistakes(review): Correlate turn proof and redrive uncorrelated doorbells * no-mistakes(test): Fix procevent sweep claims and OMP turn handler coexistence * no-mistakes(test): Preserve claim-only sources for reliable home sweeps * no-mistakes(test): Disabled unreliable turn observation for generated OMP workers * no-mistakes(document): Updated stale branch-session documentation * no-mistakes(review): Enforce state-root identity before ownership and stale-claim decisions * no-mistakes(review): Requeue unreadable awaiting-turn requests for redelivery * no-mistakes(document): Correct stale supervision and process-event verification docs --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
BenWilcox8
pushed a commit
to BenWilcox8/firstmate
that referenced
this pull request
Sep 12, 2026
* fix(pi): route needs-decision wakes and mixed batches wholly to main Skip the supervision branch for every needs-decision status append, the same way a check-kind wake already skips it. A coalesced signal/stale trigger batch containing any needs-decision row is delivered wholly to main, not split between the branch and a later main wake - the whole batch, including any co-present routine rows for a different task, travels together. Heartbeat and unread-status scans stay independent. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE * test(pi): cover distinct-file mixed batches and heartbeat independence Add a regression using two distinct files (not the same status file twice) in one coalesced trigger so a some-vs-every regression on the file-list cross-reference cannot hide behind a degenerate same-key case, and a heartbeat/needs-decision co-presence test proving a needs-decision row neither vetoes nor rides along with an otherwise eligible heartbeat scan. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE * no-mistakes(document): Clarify needs-decision and heartbeat routing * no-mistakes(ci): Fixed captain-held stale reminders so they bypass supervision and wake main, while unrelated unread rows and heartbeats remain independent. Added routing regressions and updated documentation. Pi watcher tests, strict TypeScript checks, lint, and diff checks pass. The no-mistakes attestation failure was pipeline-state related, not a source defect * no-mistakes(ci): Fixed CI lint by narrowly suppressing false-positive SC2031 diagnostics where background PIDs are captured immediately in the same shell. Verified with `CI=true bin/fm-lint.sh` and `git diff --check`. The no-mistakes attestation failure is pipeline-state-related (`test` was skipped), not a source defect * no-mistakes(review): Route stale open decisions directly to main * no-mistakes(review): Honor configured verbs in stale decision routing * no-mistakes(review): Route second-mate escalations and configured decisions to main * no-mistakes(review): Ignore trailing whitespace after captain holds * no-mistakes(review): Cache stale decision classification per status file * no-mistakes(review): Document unread decision precedence for later task wakes * no-mistakes(review): Cache unchanged stale decisions across scope scans * no-mistakes(review): Resolve decision aliases and reject symlinked statuses * no-mistakes(review): Route surfaced captain-held signals directly to main * no-mistakes(document): Document decision-owned main routing * no-mistakes(ci): Fixed captain-held spans to remain actionable while crew working evidence is positive, ensuring the watcher delivers their main-only marker. Updated the executable regression test to cover this case. Verified with the full fm-watch-triage suite, bash syntax checks, and git diff checks. Shellcheck reported only pre-existing test harness warnings (SC1091/SC2034) * no-mistakes(ci): Fixed the CI regression: captain-held transfers now retain their established non-actionable stale classification while the signal-routing side-band still surfaces them main-only. Verified with tests/fm-daemon.test.sh, tests/fm-watch-triage.test.sh, bash syntax checks, and git diff --check --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
friesentius
pushed a commit
to friesentius/firstmate
that referenced
this pull request
Sep 21, 2026
* fix(pi): route needs-decision wakes and mixed batches wholly to main Skip the supervision branch for every needs-decision status append, the same way a check-kind wake already skips it. A coalesced signal/stale trigger batch containing any needs-decision row is delivered wholly to main, not split between the branch and a later main wake - the whole batch, including any co-present routine rows for a different task, travels together. Heartbeat and unread-status scans stay independent. Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE * test(pi): cover distinct-file mixed batches and heartbeat independence Add a regression using two distinct files (not the same status file twice) in one coalesced trigger so a some-vs-every regression on the file-list cross-reference cannot hide behind a degenerate same-key case, and a heartbeat/needs-decision co-presence test proving a needs-decision row neither vetoes nor rides along with an otherwise eligible heartbeat scan. Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE * no-mistakes(document): Clarify needs-decision and heartbeat routing * no-mistakes(ci): Fixed captain-held stale reminders so they bypass supervision and wake main, while unrelated unread rows and heartbeats remain independent. Added routing regressions and updated documentation. Pi watcher tests, strict TypeScript checks, lint, and diff checks pass. The no-mistakes attestation failure was pipeline-state related, not a source defect * no-mistakes(ci): Fixed CI lint by narrowly suppressing false-positive SC2031 diagnostics where background PIDs are captured immediately in the same shell. Verified with `CI=true bin/fm-lint.sh` and `git diff --check`. The no-mistakes attestation failure is pipeline-state-related (`test` was skipped), not a source defect * no-mistakes(review): Route stale open decisions directly to main * no-mistakes(review): Honor configured verbs in stale decision routing * no-mistakes(review): Route second-mate escalations and configured decisions to main * no-mistakes(review): Ignore trailing whitespace after captain holds * no-mistakes(review): Cache stale decision classification per status file * no-mistakes(review): Document unread decision precedence for later task wakes * no-mistakes(review): Cache unchanged stale decisions across scope scans * no-mistakes(review): Resolve decision aliases and reject symlinked statuses * no-mistakes(review): Route surfaced captain-held signals directly to main * no-mistakes(document): Document decision-owned main routing * no-mistakes(ci): Fixed captain-held spans to remain actionable while crew working evidence is positive, ensuring the watcher delivers their main-only marker. Updated the executable regression test to cover this case. Verified with the full fm-watch-triage suite, bash syntax checks, and git diff checks. Shellcheck reported only pre-existing test harness warnings (SC1091/SC2034) * no-mistakes(ci): Fixed the CI regression: captain-held transfers now retain their established non-actionable stale classification while the signal-routing side-band still surfaces them main-only. Verified with tests/fm-daemon.test.sh, tests/fm-watch-triage.test.sh, bash syntax checks, and git diff --check ---------
friesentius
pushed a commit
to friesentius/firstmate
that referenced
this pull request
Sep 21, 2026
* fix(pi): route needs-decision wakes and mixed batches wholly to main Skip the supervision branch for every needs-decision status append, the same way a check-kind wake already skips it. A coalesced signal/stale trigger batch containing any needs-decision row is delivered wholly to main, not split between the branch and a later main wake - the whole batch, including any co-present routine rows for a different task, travels together. Heartbeat and unread-status scans stay independent. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE * test(pi): cover distinct-file mixed batches and heartbeat independence Add a regression using two distinct files (not the same status file twice) in one coalesced trigger so a some-vs-every regression on the file-list cross-reference cannot hide behind a degenerate same-key case, and a heartbeat/needs-decision co-presence test proving a needs-decision row neither vetoes nor rides along with an otherwise eligible heartbeat scan. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE * no-mistakes(document): Clarify needs-decision and heartbeat routing * no-mistakes(ci): Fixed captain-held stale reminders so they bypass supervision and wake main, while unrelated unread rows and heartbeats remain independent. Added routing regressions and updated documentation. Pi watcher tests, strict TypeScript checks, lint, and diff checks pass. The no-mistakes attestation failure was pipeline-state related, not a source defect * no-mistakes(ci): Fixed CI lint by narrowly suppressing false-positive SC2031 diagnostics where background PIDs are captured immediately in the same shell. Verified with `CI=true bin/fm-lint.sh` and `git diff --check`. The no-mistakes attestation failure is pipeline-state-related (`test` was skipped), not a source defect * no-mistakes(review): Route stale open decisions directly to main * no-mistakes(review): Honor configured verbs in stale decision routing * no-mistakes(review): Route second-mate escalations and configured decisions to main * no-mistakes(review): Ignore trailing whitespace after captain holds * no-mistakes(review): Cache stale decision classification per status file * no-mistakes(review): Document unread decision precedence for later task wakes * no-mistakes(review): Cache unchanged stale decisions across scope scans * no-mistakes(review): Resolve decision aliases and reject symlinked statuses * no-mistakes(review): Route surfaced captain-held signals directly to main * no-mistakes(document): Document decision-owned main routing * no-mistakes(ci): Fixed captain-held spans to remain actionable while crew working evidence is positive, ensuring the watcher delivers their main-only marker. Updated the executable regression test to cover this case. Verified with the full fm-watch-triage suite, bash syntax checks, and git diff checks. Shellcheck reported only pre-existing test harness warnings (SC1091/SC2034) * no-mistakes(ci): Fixed the CI regression: captain-held transfers now retain their established non-actionable stale classification while the signal-routing side-band still surfaces them main-only. Verified with tests/fm-daemon.test.sh, tests/fm-watch-triage.test.sh, bash syntax checks, and git diff --check --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
jordanhindo
pushed a commit
to jordanhindo/firstmate
that referenced
this pull request
Sep 22, 2026
* fix(pi): route needs-decision wakes and mixed batches wholly to main Skip the supervision branch for every needs-decision status append, the same way a check-kind wake already skips it. A coalesced signal/stale trigger batch containing any needs-decision row is delivered wholly to main, not split between the branch and a later main wake - the whole batch, including any co-present routine rows for a different task, travels together. Heartbeat and unread-status scans stay independent. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE * test(pi): cover distinct-file mixed batches and heartbeat independence Add a regression using two distinct files (not the same status file twice) in one coalesced trigger so a some-vs-every regression on the file-list cross-reference cannot hide behind a degenerate same-key case, and a heartbeat/needs-decision co-presence test proving a needs-decision row neither vetoes nor rides along with an otherwise eligible heartbeat scan. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE * no-mistakes(document): Clarify needs-decision and heartbeat routing * no-mistakes(ci): Fixed captain-held stale reminders so they bypass supervision and wake main, while unrelated unread rows and heartbeats remain independent. Added routing regressions and updated documentation. Pi watcher tests, strict TypeScript checks, lint, and diff checks pass. The no-mistakes attestation failure was pipeline-state related, not a source defect * no-mistakes(ci): Fixed CI lint by narrowly suppressing false-positive SC2031 diagnostics where background PIDs are captured immediately in the same shell. Verified with `CI=true bin/fm-lint.sh` and `git diff --check`. The no-mistakes attestation failure is pipeline-state-related (`test` was skipped), not a source defect * no-mistakes(review): Route stale open decisions directly to main * no-mistakes(review): Honor configured verbs in stale decision routing * no-mistakes(review): Route second-mate escalations and configured decisions to main * no-mistakes(review): Ignore trailing whitespace after captain holds * no-mistakes(review): Cache stale decision classification per status file * no-mistakes(review): Document unread decision precedence for later task wakes * no-mistakes(review): Cache unchanged stale decisions across scope scans * no-mistakes(review): Resolve decision aliases and reject symlinked statuses * no-mistakes(review): Route surfaced captain-held signals directly to main * no-mistakes(document): Document decision-owned main routing * no-mistakes(ci): Fixed captain-held spans to remain actionable while crew working evidence is positive, ensuring the watcher delivers their main-only marker. Updated the executable regression test to cover this case. Verified with the full fm-watch-triage suite, bash syntax checks, and git diff checks. Shellcheck reported only pre-existing test harness warnings (SC1091/SC2034) * no-mistakes(ci): Fixed the CI regression: captain-held transfers now retain their established non-actionable stale classification while the signal-routing side-band still surfaces them main-only. Verified with tests/fm-daemon.test.sh, tests/fm-watch-triage.test.sh, bash syntax checks, and git diff --check --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
NewAiCoder
added a commit
to NewAiCoder/firstmate
that referenced
this pull request
Sep 26, 2026
…rt (kunchenguid#28) * fix(bin): preserve subshell lock ownership on Bash 3.2 (#3789) * fix: distinguish subshell wake-lock owners on stock Bash Restore distinct process ownership for issue #3743 using the existing PID helper, consistently across lock publication, reclaim, release, role checks, and bounded handoff. The existing wake-queue regression fails on pristine upstream Bash 3.2 with rc=13. The complete suite now passes on Bash 3.2.57 and Bash 5.3.15, with added coverage for ownership when BASHPID is unset. Canonical lint and stock-Bash syntax checks pass. * no-mistakes(document): Correct lock grace-period documentation * no-mistakes(ci): Captain, fixed all 14 SC2031 false positives with nine ShellCheck source-boundary annotations across three tests. Full CI-mode lint and the complete wake-queue suite on stock Bash 3.2 passed. Runtime behavior is unchanged * fix(bin): resolve captain holds and legacy teardowns on non-markdown backends (#3782) * fix(bin): close legacy records on the Beads backend honestly Two pre-Beads reads blocked honest closure of leftover records: 1. fm-captain-hold.sh complete/verify resolved attested legacy hold ids only against the live backend and the pre-collapse derived identity, so a home whose holds fm-hold-migration rehomed under fm- ids failed with an empty-name absence message (the resolve failure was swallowed by the command substitution feeding verify_hold_durable). Resolution now falls back, on the Beads backend only, to the legacy id under the configured beads prefix and to the row whose notes carry the exact marker line 'migrated from data/backlog.md id <legacy id>'; every refusal names the id it could not resolve, and the markdown path is unchanged. 2. fm-teardown.sh refused any record without spawn_gen forever. A record that predates the field can now be torn down with an explicit --legacy-record flag once the recovery-grade endpoint classifier confirms the recorded endpoint dead or agent-less; the accepted incarnation is stamped into the record right before its close marker binds to it and named in the teardown line. Refusals leave the record byte-identical, the unlanded-work refusal is not relaxed, and a corrupt (multi-valued) spawn_gen is never accepted. The companion repair this branch carries (follow-up commit) is the backend-gated --file and markdown-file requirement in the mutate path and lifecycle gates: fm_backlog_mutate passed --file and required the markdown backlog file regardless of the resolved backend, and the transition gate plus row probe required that file before any backend work, so a home on a non-markdown backend could neither gate, probe, nor close its rows. Behavior tests: self-contained beads fixtures over a scratch bd graph (self-skipping on markdown-only tasks-axi installs), legacy meta fixtures for every teardown gate, and the relocated markdown backlog coverage stays green. * no-mistakes(review): fix(review): report migrated-hold scan refusals and guard legacy spawn_gen stamp against newline-less records * fix(backlog): address the configured backend for lifecycle writes Completes the fm-backlog-transition-lib repair the first commit's message claims: on this base fm_backlog_mutate passed --file and required the markdown backlog file regardless of the resolved backend, and fm_backlog_transition_applies plus fm_backlog_row_probe required that file before any backend work, so a home on a non-markdown backend could neither gate, probe, nor close its backlog rows. All three now gate the markdown file on the resolved tasks-axi backend: markdown keeps exactly its explicit <data>/backlog.md behavior, non-markdown homes address the backend their own configuration selects with no markdown file requirement. fm_backlog_row_show and fm_backlog_row_list already gated correctly and are unchanged. docs/configuration.md owns the contract line. Also extends the same backend gate to fm-captain-hold.sh's own mutation wrapper - hold/add/update/answer/done append the markdown --file only when the resolved backend is markdown, so a captain call on a Beads home reaches the Beads store end to end - and applies the review round's two direct remedies there: the [beads] graph path resolves against the backlog root when relative (never the process CWD), and a failed bd graph read reports bd's own trimmed stderr reason in the refusal. Coverage: tests/fm-backlog-atomicity.test.sh gains a stub-driven Beads completion case proving the transition gate applies, the row probe reads, and done runs without any markdown file or --file override; the relocated markdown backlog test stays green. * no-mistakes(review): Document root-tasks.toml-only beads settings for migrated-hold resolution * test(gotmp): stub fm_tasks_axi_backend so the fixture matches the backend-aware transition lib The legacy-records change made fm-backlog-transition-lib.sh resolve the configured backend via fm_tasks_axi_backend before the markdown-only skip. The gotmp fixture's fm-tasks-axi-lib stub lacked that function, so the markdown check fell through and teardown hit the incompatible-backend error with unbound FM_TASKS_AXI_MIN under set -u. Stub the backend as markdown and define the floor, restoring the intended no-backlog skip. * fix(teardown): roll the legacy stamp back when the close marker fails A legacy-record teardown stamps its accepted incarnation into the record right before the close marker binds to it; when that marker write then fails, the stamp survived, so a retried teardown sailed past the dead-or-agent-less endpoint gate the stamp now proved unnecessary. The failed marker write now truncates the record back to its exact pre-stamp bytes (verified by size), restoring the byte-identical-refusal invariant; when the rollback itself fails the operator is told to re-run with --legacy-record after reconciling the endpoint. Also completes the recorded review decision's coverage wording: the beads stub test now drives the answer close end to end (update and done through the gated wrapper), asserting no markdown file override reaches either verb. * fix(review): harden the legacy stamp rollback and resolve derived migrated ids The legacy-record stamp rollback now uses perl (already in the teardown curated PATH; truncate is not, and is absent on stock macOS), routes every failure branch inside the stamp block through the same size-verified rollback so the byte-identical-refusal invariant holds on those paths too, and gains behavior coverage: an unrecordable close (an invalid pr= link) fails the teardown, leaves the record byte-identical, keeps the backlog row in flight, and a flag-less retry still refuses. Migrated-hold resolution now probes the derived pre-collapse identity (<origin>-decision-<entry>) alongside the raw entry - fm-hold-migration recorded the DERIVED id in every migrated row's marker note - in both the prefix and the migration-note forms, with the ambiguity refusal naming every identity tried, plus behavior coverage for a bare decision key resolved through its derived identity's marker. Also aligns fm-backlog-transition-lib.sh's header ADDRESSING/SCOPE paragraphs with the backend-gated contract, drops an unreachable FORCE validity guard the parser rewrite left behind, and switches the new stub fixture to the portable sed -i.bak idiom. * no-mistakes(review): Name the configured backend in teardown's backlog reminder * no-mistakes(review): Scan migration markers before the prefix guess * no-mistakes(review): Document marker-first resolution and cover the prefix branch * no-mistakes(document): Record prefix-attestation audit and marker-line forms * no-mistakes(ci): Fixed the Greptile P1 on bin/fm-teardown.sh: a failed rollback of the synthetic legacy stamp let a retry bypass the dead-or-agent-less endpoint gate. Root cause: teardown minted `spawn_gen=legacy-<ts>-<pid>` into the task record before the close marker bound to it. When the close-marker write failed AND the rollback also failed, the record retained that token. On the next invocation `fm_backlog_meta_spawn_gen` succeeded, so `TEARDOWN_LEGACY_PENDING` stayed 0 and the endpoint gate was skipped entirely — even with `--legacy-record`. The script's own error text told the operator to "re-run teardown with --legacy-record", advice the code could not honor. Fix (bin/fm-teardown.sh): - A `legacy-*` spawn_gen is now recognized as a stamp this teardown path minted, never one a spawn published (fm-spawn.sh publishes `s<epoch>.<pid>.<random>`). Such a record still reads as the legacy record it is: it re-enters the endpoint gate, and a flag-less retry refuses naming `--legacy-record`. - Acceptance reuses the retained token instead of minting a second one; the append block is skipped when the record already carries it, so no duplicate spawn_gen is written. - The rollback attempt and its "could not be rolled back" message are guarded to runs that actually appended a stamp, so a run that appended nothing never claims a rollback it did not perform. - Usage header documents the retained-stamp rule. Test (tests/fm-teardown.test.sh): added `test_retained_legacy_stamp_still_faces_the_endpoint_gate`, an end-to-end reproduction — a `perl` stub that fails only the rollback's `truncate` (delegating every other perl call to the real interpreter) leaves the stamp behind, then the retry must still hit the gate, must not stamp a second incarnation, must not close the backlog row, and the flag-less retry must refuse. Verification: the new test fails against the pre-fix script on exactly the reported defect ("the retry skipped the dead-or-agent-less endpoint gate") and passes after. Full tests/fm-teardown.test.sh 80 ok / 0 failures / rc=0; tests/fm-backlog-atomicity.test.sh 80 ok / 0 failures / rc=0; bin/fm-lint.sh (pinned ShellCheck 0.11.0 + actionlint 1.7.12) clean * fix: reduce local ShellCheck source-analysis cost (#3778) * fix(lint): drop source following on the local changed-file gate The local lint step was inlining library closures through --external-sources and peaking above 8 GB on a single root. Keep full analysis in CI, on main, and without a merge-base; exclude the four cross-file codes from the local pass so those findings still land in CI. Co-authored-by: Cursor <cursoragent@cursor.com> * no-mistakes(review): Run local ShellCheck per root and document measurements * no-mistakes(review): Correct local source-following telemetry * no-mistakes(document): Clarify context-sensitive lint documentation --------- Co-authored-by: Cursor <cursoragent@cursor.com> * fix(pi): route decision-owned wake batches to main (#3776) * fix(pi): route needs-decision wakes and mixed batches wholly to main Skip the supervision branch for every needs-decision status append, the same way a check-kind wake already skips it. A coalesced signal/stale trigger batch containing any needs-decision row is delivered wholly to main, not split between the branch and a later main wake - the whole batch, including any co-present routine rows for a different task, travels together. Heartbeat and unread-status scans stay independent. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE * test(pi): cover distinct-file mixed batches and heartbeat independence Add a regression using two distinct files (not the same status file twice) in one coalesced trigger so a some-vs-every regression on the file-list cross-reference cannot hide behind a degenerate same-key case, and a heartbeat/needs-decision co-presence test proving a needs-decision row neither vetoes nor rides along with an otherwise eligible heartbeat scan. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE * no-mistakes(document): Clarify needs-decision and heartbeat routing * no-mistakes(ci): Fixed captain-held stale reminders so they bypass supervision and wake main, while unrelated unread rows and heartbeats remain independent. Added routing regressions and updated documentation. Pi watcher tests, strict TypeScript checks, lint, and diff checks pass. The no-mistakes attestation failure was pipeline-state related, not a source defect * no-mistakes(ci): Fixed CI lint by narrowly suppressing false-positive SC2031 diagnostics where background PIDs are captured immediately in the same shell. Verified with `CI=true bin/fm-lint.sh` and `git diff --check`. The no-mistakes attestation failure is pipeline-state-related (`test` was skipped), not a source defect * no-mistakes(review): Route stale open decisions directly to main * no-mistakes(review): Honor configured verbs in stale decision routing * no-mistakes(review): Route second-mate escalations and configured decisions to main * no-mistakes(review): Ignore trailing whitespace after captain holds * no-mistakes(review): Cache stale decision classification per status file * no-mistakes(review): Document unread decision precedence for later task wakes * no-mistakes(review): Cache unchanged stale decisions across scope scans * no-mistakes(review): Resolve decision aliases and reject symlinked statuses * no-mistakes(review): Route surfaced captain-held signals directly to main * no-mistakes(document): Document decision-owned main routing * no-mistakes(ci): Fixed captain-held spans to remain actionable while crew working evidence is positive, ensuring the watcher delivers their main-only marker. Updated the executable regression test to cover this case. Verified with the full fm-watch-triage suite, bash syntax checks, and git diff checks. Shellcheck reported only pre-existing test harness warnings (SC1091/SC2034) * no-mistakes(ci): Fixed the CI regression: captain-held transfers now retain their established non-actionable stale classification while the signal-routing side-band still surfaces them main-only. Verified with tests/fm-daemon.test.sh, tests/fm-watch-triage.test.sh, bash syntax checks, and git diff --check --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> * feat(bin): add opt-in worker launch environment allowlist (#3802) * feat(spawn): add an opt-in worker environment allowlist Honor a home-local launch-env-allowlist at the shared worker command boundary and inherit it into secondmate homes. Preserve the existing launch behavior when the file is absent. Keep the operational environment and explicit launch assignments, and account for filtered Muse credentials. Refs https://github.com/kunchenguid/firstmate/issues/3742 Verification: - Red on origin/main 1820316b66ac2c68e244dd04a02512859ee8c1f4: the new enabled-allowlist regression observed synthetic-unrelated in the worker; the absent-file control passed. - Green: fm-test-run.sh on fm-spawn-dispatch-profile, fm-muse-harness, and fm-trace-context-spawn; all three passed without skips. - Synthetic emitted-command probes ran through sh, stock Bash, and zsh. - Canonical lint, documentation audience checks, and stock Bash syntax checks passed. * no-mistakes(review): Reject inaccessible launch environment configuration * no-mistakes(review): Preserve inherited allowlists on source inspection errors * no-mistakes(document): Clarify worker environment grants and inheritance documentation * no-mistakes(lint): Fix inheritance test ShellCheck source boundary * feat(bin): add rovo crewmate/scout adapter with home-path file access (#3575) * feat(bin): add rovo as a verified crewmate/scout worker harness Wire the Atlassian Rovo CLI (202609.1.2) into the TUI-under-tmux/herdr adapter contract: detection with marker-precedence ordering, one-shot positional launch with --startup-receipt readiness polling instead of composer scraping, model/effort flags, a screen-scrape busy fallback scoped like grok's, and crew/scout-only lifecycle control that refuses secondmate launches. Ships with a portable regression suite, a live PTY guard against the real binary, a per-harness reference doc, and a dated verification record covering the silent OAuth refresh, the interrupt-ack divergence from the originating scout report, and the still-open composer-ghost and tmux/herdr pane-liveness gaps. * no-mistakes(review): revert rovo launch to positional brief, drop startup-receipt * no-mistakes(test): rewire rovo adapter to kimi-style launch-then-send shape * no-mistakes(document): add rovo to stale worker-harness enumerations in docs * docs(verification): close the rovo herdr-liveness gap with live isolated-lab evidence Placement, launch-then-send, and busy/idle rendering are now verified live in an isolated non-default Herdr lab session (bin/fm-herdr-lab.sh), driven directly through fm-spawn.sh's/fm-backend.sh's own shared primitives since the cross-session launcher-identity guard refuses this task's own ambient Herdr identity for a full fm-spawn.sh run. fm_backend_agent_state reported dead for a live, responding rovo pane at every point checked, because herdr's own agent-integration registry has no rovo entry (herdr integration status), so herdr agent get returns agent_not_found regardless of whether rovo is actually running. This is recorded as a Herdr-side integration gap rather than a firstmate bug, left unpatched to avoid a false-positive alive verdict for other idle shells. Updates docs/verification/rovo.md's backend-liveness section and its two cross-references (docs/verification/runtime-backends.md, docs/configuration.md) accordingly. * fix(bin): close rovo's failed-spawn leak and busy-scrape false idle Greptile P1s on PR #3575: a failed rovo readiness/submission/delivery gate exited without tearing down the just-created endpoint, leaving the launched --yolo rovo process running as an orphaned agent outside task control. Separately, the busy classifier's rendered-tail fallback returned definitive idle whenever the "Rovo is thinking" marker scrolled out of the last 12 nonblank lines of a long turn, which could make supervision wrongly conclude a still-working worker had gone idle. fm-spawn.sh: rovo_spawn_fail now calls rovo_endpoint_cleanup, which kills the created endpoint (tmux/herdr/zellij/cmux) via the same generic fm_backend_kill dispatch fm-spawn.sh's own orca-abort path already uses; orca's worktree and terminal remain owned by the separate ORCA_ABORT_CLEANUP trap. fm-busy-lib.sh: the rovo classifier arm now reports "unknown rovo-regex" instead of "idle rovo-regex" when the marker is absent, matching how muse and cursor already express "can't tell" for their own fallbacks. The positive busy match is unchanged. Extends tests/fm-rovo-harness.test.sh: the readiness and delivery failure tests now assert the endpoint is torn down (and the success test asserts it is not), and a new test drives the busy marker out of the tail window to confirm the verdict is unknown, never idle. bin/fm-lint.sh is clean on both changed files. * test(rovo): align spawn fixture with the launch-brief validation contract Upstream main now requires a brief's ## Captain's intent and ## Firstmate spec subsections (or a nonempty legacy # Task body) before spawn, and rewrites ship+no-mistakes briefs into launch-brief.md. Update the rovo harness fixture and pointer assertions to match, mirroring the kimi harness fixture. * no-mistakes(review): align rovo.md delivery-gate note with live herdr evidence * fix: prevent stale supervision wake loops (#3672) * fix(bin): stop the supervision branch's stale-ack and ghost-report loops Clean-slate implementation of the four authorized recommendations from the supervision-ghost-retrigger analysis (items 1, 2, 3, and 7), in their minimal form, superseding PR #3604: - fm_branch_report refuses a task the wake being handled never named. The extension fixes the reportable task set from the eligible rows before each prompt (signal and stale rows resolve to their tasks, a heartbeat allows any task with a live record, fleet is always allowed), so a report typed from memory about a task whose records teardown already removed is never stored or delivered. - An acknowledgement that consumes nothing says "nothing was acknowledged through N" and prints the exact --ack-through / --recovery-generation command for the current presented wake, instead of "re-run the drain", which re-fed the same stale acknowledgement in a loop. - bin/fm-guard.sh no longer tells the branch actor to drain queued wakes while it is handling them; it names the granted rows instead. - Teardown removes state/.<task>.branch-outcome-index for ordinary tasks and descendants; the index rebuild and the append-side index write both skip a task with neither a live record nor a status log, so the branch's report of a teardown it just performed is stored without recreating the index. No new locking, no spawn-generation binding, and no retired-task refusal: the branch can still report the outcome of a task it just tore down, and the teardown test now proves that path end to end. * fix(bin): narrow the branch report scope and guard silence to the minimal form Apply the four review decisions on the clean-slate branch: - A signal or stale prompt may report only the tasks its own rows resolve to; fleet is refused there too. A heartbeat review is not scoped by task at all, so the extension no longer tracks live task records and refuses nothing by task id during a fleet review. - The outcome-index rebuild no longer skips retired tasks; the append-side skip alone keeps a torn-down task's index from being recreated. - bin/fm-guard.sh keeps the queued-wakes warning silent for the branch actor instead of printing a replacement note. * no-mistakes(document): Align supervision docs with scoped wake handling * fix(bin): grant rovo the per-task home paths its standard crewmate flow needs rovo confines every file-tool operation to its worktree by default, and its bash tool independently refuses the same external paths regardless of any grant (confirmed live), so a rovo worker could not read its own brief or steering messages or write its status/report - all of which live in the firstmate home outside the worktree - without hand-feeding it. Grant toolPermissions.allowedExternalPaths for exactly the task's brief directory, steering inbox, and status file at launch time via --config-override, merged with agent.efficiencyLevel into one JSON object since that flag is single-value and silently discards a second occurrence. Extends the live PTY guard to prove, against the real binary, that the grant lets rovo read an external brief and append to an external status file, and that the same flow is blocked without the grant. * no-mistakes(document): align rovo reference Effort row with merged single --config-override * no-mistakes(document): document rovo file-access grant in harness reference --------- Co-authored-by: PUNEET PATWARI <ppatwari@atlassian.com> Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com> * fix(bin): prevent false pipeline blocks after drive timeouts (#3813) * fix(bin): read a crew's pipeline-death claim against the live run A crew's no-mistakes drive call blocks until the next gate or outcome, routinely far longer than its harness lets one command live, so the call gets killed or times out while the daemon runs the fix round on in the background. Crews read that as daemon death and block on it, and firstmate had nothing that contradicted them. Rule 7 of every generated brief now says a drive-call error or a harness command timeout is not a daemon error, requires `no-mistakes daemon status` plus `no-mistakes axi status` before a pipeline `blocked:`, and reserves that report for a refused socket or a run record failed with a daemon error. The no-mistakes definition of done adds the harness command limit and the background-and-poll shape that fits inside it. fm-crew-state gains one classification case: a `blocked:` line blaming the daemon, a timeout, or unreachability, while the run is running or fixing AND the pipeline reports fresh activity, now reads as superseded because the run is alive. Recency comes from the client's own `quiet` marker on active_steps.last_activity rather than a threshold invented here, and positive evidence is required, so a run record that outlives a genuinely dead daemon keeps the plain reading. stuck-crewmate-recovery gains the inverse-of-a-dead-endpoint playbook: firstmate reads both statuses itself, steers a reattach, never restarts the shared daemon on a crew's claim, and escalates only a refused socket. Nothing here depends on an unshipped no-mistakes capability. * no-mistakes(review): Prioritize daemon socket failure and narrow unreachable matching * no-mistakes(review): Honor socket refusal across coarse status and crew guidance * no-mistakes(test): Replace flaky settle timing assertion with pane-read count * no-mistakes(document): Document daemon timeout recovery contract * no-mistakes(ci): Fixed daemon socket failures being suppressed by terminal attributed runs. Positive refused/missing socket evidence now remains blocked regardless of run status. Added a behavioral regression test for terminal failed runs. Verified with fm-crew-state tests, project ShellCheck lint, and git diff checks * fix(bin): scope the worker role contract for ship and scout launches (#3797) * fix(brief): scope Firstmate workers to their launch contract * no-mistakes(document): Clarify supervisor scope and worker contract ownership * no-mistakes(ci): Removed the heading-based bypass so every ship/scout launch receives the current worker-role contract. Added a regression that failed before the fix and passes afterward. Dispatch, brief, and delivery suites, focused ShellCheck, and git diff --check all passed * no-mistakes(review): make launch overlay sole owner of worker role contract * no-mistakes(review): narrow heading test dimension and fix publish error wording * no-mistakes(review): gate role supersession, fix render guard, drop AGENTS twin * no-mistakes(document): align architecture AGENTS.md scope and spawn launch-brief header * fix(bin): stop ringing steering doorbells into dead panes (#3823) * fix(bin): stop ringing steering doorbells into dead panes The steering-inbox doorbell was a plain sentence plus Enter typed into a worker's pane, and the watcher re-rang it on the assumption that a ring is free. In a pane whose agent has exited that line is a shell command, and the re-ring ladder kept typing it into a shell that can never acknowledge it. - Prefix the doorbell with the shell no-op `: ` so a bare shell executes nothing while a live worker still reads the same self-describing line. `#` is not used because interactive zsh does not treat it as a comment by default and the claude harness binds it to memory mode. - fm_task_inbox_ring skips the pane (return 3) when the backend positively classifies the agent as dead; missing, ambiguous, unreadable, and unverified endpoints still ring so a blind classifier never starves a live worker. - The watcher caps the ladder for a dead pane: one stale wake for recovery, no ring, no ladder walk, and the durable record stays for stuck-crewmate-recovery. fm-send and the remote steer leg report the skip. Tests cover the no-op in real shells, the dead/live/unclassifiable ring verdicts, and the single-surfacing watcher path. * no-mistakes(review): Quote doorbell paths against shell injection * no-mistakes(review): Reject terminal-control paths before ringing * no-mistakes(review): Document accepted partial doorbell delivery race * no-mistakes(review): Skip unavailable endpoints before busy-state handling * no-mistakes(test): Respect shell startup PATH in environment allowlist test * no-mistakes(test): Fix doorbell test fixtures for endpoint liveness * no-mistakes(test): Prioritize confirmed restarts and clean shell test syntax * no-mistakes(document): Document dead and missing doorbell recovery * no-mistakes(ci): Fixed the persistence-reply timeout race by rechecking for a correlated reply immediately before falling back to a nudge. Added a deterministic regression covering replies arriving between the preliminary resolution pass and timeout handling. Verified with the targeted restart suite, project lint, coverage guard, bash syntax checks, and diff checks * fix(bin): wait out transient primary-checkout reads in the spawn worktree poll (#3834) * fix(spawn): keep the worktree poll from adopting the repository primary After `treehouse get` is sent, the worktree-discovery poll reads the pane's foreground-process cwd. While treehouse is still fetching and checking a slot out, the foreground process is treehouse itself and it reports the repository's PRIMARY checkout as its cwd for several seconds. The poll accepted any path that merely differed from the spawning project, so from a linked spawning home - whose project is itself a worktree of that repository - it adopted the primary, and the isolation guard then refused a launch whose slot treehouse went on to create normally. Screen every candidate with the isolation guard's own conditions, extracted as spawn_worktree_isolated, so a read the guard would reject stays a transient the poll keeps waiting through. The two-consecutive-reads rule and the guard as final backstop are unchanged; a pane that never reaches an isolated worktree still fails at the existing 60s deadline, now naming the last path it reported. The already-settled timing assertion counted whole-spawn wall time against a 5s budget and failed on unmodified HEAD on slower machines; it now counts pane reads, which is what "one confirming read, not an extra cycle" actually means. * fix(spawn): say which path the worktree wait rejected, and why Screening every discovery-poll candidate means a host that never reaches an isolated worktree spends the whole 60s window before refusing. That wait is deliberate - separating a transient from a terminal misconfiguration needs machinery this path does not want - so the refusal explains itself instead: the isolation check records why a candidate failed, and the deadline names the last path seen together with that reason. Message and diagnostics only; the poll's control flow is unchanged. Two suites asserted the guard's wording on paths the poll now rejects rather than adopts, so their refusal arrives from the deadline instead: realign fm-tangle-guard's non-git and subdirectory-of-primary cases (each now also asserting the stated reason, and the second the metadata absence it was missing) and the herdr projection e2e's forced non-worktree cwd. * no-mistakes(review): stub poll sleep in tangle-guard spawn isolation test * no-mistakes(document): document spawn poll isolation screen in fm-spawn header * test(spawn): make the non-git isolation case non-git anywhere The refusal-reason assertion for a path outside any repository assumed TMPDIR is not inside a git repository. Where it is, git walks up from the temporary directory, finds that repository, and the spawn reports the subdirectory cause instead - so the case passed or failed on a property of the host rather than on the behaviour under test. Build the path under a directory the test then names in GIT_CEILING_DIRECTORIES, which git documents as not chdir-ing up into a listed directory while looking for a repository. Git never excludes the directory being searched, so the ceiling is the parent of the path handed to the spawn. The assertions pin which cause fired rather than the sentence that explains it, leaving the operator wording free to improve. * no-mistakes(document): point spawn poll comment at the isolation screen's comparison * no-mistakes(ci): Fixed the "Behavior portable serial 2" failure in tests/fm-tangle-guard.test.sh ("non-worktree spawn did not say why the path was rejected (missing: 'not inside a git worktree')"). Root cause, in this PR's code: bin/fm-spawn.sh's spawn_worktree_isolated resolved the git toplevel with `wt_top_real=$(cd "$SPAWN_WT_TOP" ...)`. For a path in no repository, `git rev-parse --show-toplevel` yields empty, and `cd ""` is a SUCCESSFUL no-op on bash before 5.3 (CI's ubuntu-latest ships bash 5.2). The empty toplevel therefore resolved to fm-spawn's own cwd — the CI checkout — so the poll reported "it is a subdirectory of worktree root '/home/runner/work/firstmate/firstmate'" instead of the correct "it is not inside a git worktree". Dev machines with bash 5.3 fail `cd ""`, which is why the suite passed locally and only failed on CI; it is a genuine shell-portability defect in the reason vocabulary this change added, not a test-environment artifact. Fix (smallest root-cause change, 1 line + comment, bin/fm-spawn.sh:2168-2173): guard the empty value so it never reaches `cd` — if [ -n "$SPAWN_WT_TOP" ] && ! wt_top_real=$(cd "$SPAWN_WT_TOP" 2>/dev/null && pwd -P); then No change to the poll's timing or deadline behavior (respecting the recorded refusal-latency and spawn-wt-reason-vocabulary decisions), no new tests, no other files touched. Verification: - Reproduced the exact CI failure locally by putting bash 3.2 (same `cd ""` semantics as CI's 5.2) first on PATH: fails before the fix with the identical message shape, passes after. - tests/fm-tangle-guard.test.sh passes under both bash 3.2 and bash 5.3. - tests/fm-spawn-worktree-settle.test.sh and tests/fm-spawn-pool-base-freshen.test.sh pass; shellcheck -x bin/fm-spawn.sh clean. - bin/fm-test-run.sh --changed: 46 suites completed, every FM_TEST_END exit=0, 1076 passing assertions, 0 "not ok" (including fm-tangle-guard, fm-control-relaunch, fm-lint). The run ended on my own 900s wall-clock cap (rc=124), not on any test failure * fix: support stock macOS Bash 3.2 paths (#3732) Co-authored-by: Talon Stark <talonstark@gmail.com> * fix(bin): read orphaned green ci monitor and daemon-down failed record as not failed (#3846) * fix(bin): read an orphaned green ci monitor as held-for-merge, not failed A no-mistakes run held for a captain merge decision keeps its ci step polling until merged or closed; when the shared daemon restarts under that poll, the run is recorded failed although every substantive step completed and GitHub reports the PR green. A monitor whose only remaining job is to observe a human decision must not convert the absence of that decision into a failure verdict. fm-crew-state.sh now reclassifies a terminal failed run as done (held-for-merge), surfacing the run's PR URL, when the steps table shows every step completed except exactly ci failed and the ci log's last recognized marker reads checks green. A genuinely red check, an unreadable ci log, or a second failed step keeps the failure. * no-mistakes(review): Read daemon-down coarse failed ledger as unknown, not failed * no-mistakes(document): docs: align AGENTS.md failed-verdict guidance with crew-state reclassification * fix(spawn): verify preserved backlog state after interrupted spawn delivery (#3852) * fix(bin): read preserved spawn state back before the interrupted exit claims it The deferred-signal exit path asserted the paired task record and In-flight backlog state were preserved without reading either back, exactly when a reader is least able to check (fm-yi4j evidence, 2026-09-05). The commit's exit status alone has been observed to agree with a row that did not actually move. The exit path now re-reads the record and the row under the same per-task lock as the commit, repairs a row the commit believed it moved, and phrases the error as exactly what was verified or attempted - verified preserved, repaired and verified, or an explicit preservation-could-not-be-verified with the reason and hand-closeout instruction. Two behavior tests drive a lying tasks-axi start through a real interrupted spawn and assert the printed claim and the real backlog state agree. * no-mistakes(test): Fix calm suite for Pi 0.85 and pin test umask * no-mistakes(document): Document interrupted-spawn preservation claim in backlog gate owner * no-mistakes(ci): Fixed the Greptile P1 in bin/fm-spawn.sh's deferred-signal exit path: during preservation verification, the no-op HUP/INT/TERM re-trap combined with an unresponsive `tasks-axi show`/`start` (bash cannot run traps while a foreground child runs) held the per-task meta lock - and every lifecycle operation waiting on it - indefinitely. Root-cause fix: bound every tasks-axi invocation made under the lock. bin/fm-backlog-transition-lib.sh gains fm_tasks_axi, an exec-based wrapper (GNU timeout, gtimeout fallback) used by fm_backlog_row_show and fm_backlog_mutate that preserves the exact process placement of the plain tasks-axi call; bin/fm-spawn.sh sets FM_TASKS_AXI_TIMEOUT (default 30s) at the commit point so both the commit and the read-back verification are bounded. A timed-out call fails through the existing error plumbing and probe/mutate name the timeout as the reason, so the interrupted exit path prints honest 'preservation could not be verified ... (reason)' wording - never intent phrased as outcome, matching the author's intent. Added a behavior test in tests/fm-backlog-atomicity.test.sh that drives a real interrupted spawn through a lying tasks-axi whose repair start never answers; it asserts the spawn exits promptly (self-bounded by an outer timeout), the attempted wording names the timeout, and the printed claim agrees with the real record/backlog state. Confirmed the test fails on the unfixed tree and passes with the fix. Verified: fm-backlog-atomicity (83 ok), fm-transition-lib, fm-backlog-handoff, fm-captain-hold, fm-teardown, fm-fleet-snapshot-view, fm-secondmate-reconcile, fm-spawn-batch, fm-spawn-dispatch-profile, fm-task-delivery, fm-control-relaunch all pass; bin/fm-lint.sh clean. The fm-bootstrap 'unsplit run lost its local diagnostic' failure reproduces on the pristine base commit and is unrelated to this change * no-mistakes(ci): Fixed the Greptile P1 in bin/fm-backlog-transition-lib.sh: fm_tasks_axi bounded tasks-axi only through GNU timeout/gtimeout and fell through to an unbounded exec on hosts with neither (stock macOS), so an unresponsive call could hold the per-task meta lock forever during interrupted-spawn verification. Root-cause fix: the bound now has no unbounded path. GNU timeout is preferred, gtimeout next, then a small perl watchdog (fork + waitpid WNOHANG polling at 50ms, TERM on expiry, one bound of grace, then KILL, exit 124 so the callers' existing timeout plumbing reports it; exit statuses and output pass through unchanged). Polling was chosen over alarm+die to avoid perl's platform-dependent syscall-restart semantics. When a bound is requested but no bounding mechanism exists, the call fails closed (exit 127 with a diagnostic) rather than running unbounded, so the interrupted exit path prints honest attempted wording, never intent as outcome. The unbounded exec remains only for the no-bound plain-call case. Added three behavior tests in tests/fm-backlog-atomicity.test.sh driving fm_tasks_axi through a PATH with no timeout binary: a hanging stub must exit 124 within the bound (verified to fail on the pre-fix code), a failing stub's status/output must pass through, and a tool-less PATH must fail closed with the diagnostic. Verified: fm-backlog-atomicity 86/86 ok, fm-transition-lib, fm-backlog-handoff, fm-teardown, fm-spawn-batch, fm-task-delivery, fm-fleet-snapshot-view, fm-secondmate-reconcile, fm-control-relaunch, fm-spawn-dispatch-profile all pass; bin/fm-lint.sh clean * no-mistakes(ci): Fixed the Greptile P1 on bin/fm-backlog-transition-lib.sh: fm_tasks_axi's GNU timeout and gtimeout paths sent TERM at the bound but had no kill-after, so a tasks-axi that ignores SIGTERM kept the bounded call - and the per-task meta lock - held indefinitely during interrupted-spawn verification. Root-cause fix: both GNU execs now carry -k "$bound" (TERM at the bound, KILL after one further bound of grace), giving every bounded path the same forced-termination contract the perl watchdog already had. Because GNU timeout exits 137 (128+SIGKILL) when the kill-after fires - versus 124 for a TERM expiry - the probe/mutate timeout detection now goes through a new fm_tasks_axi_timeout_expired helper that treats 124 and 137 alike, so the interrupted exit path still names the timeout as the reason; the helper keeps the bound check in one place. Added a behavior test in tests/fm-backlog-atomicity.test.sh that drives fm_tasks_axi through a real GNU timeout with a tasks-axi stub that traps and ignores TERM (the ignored disposition survives exec into sleep) and asserts a bound-expiry status plus completion within bound+grace; on the pre-fix code the suite hangs until killed, confirming the reproduction. Verified: fm-backlog-atomicity 87/87 ok, fm-transition-lib, fm-backlog-handoff, fm-spawn-batch, fm-task-delivery, fm-teardown, fm-secondmate-reconcile, fm-control-relaunch, fm-spawn-dispatch-profile, fm-captain-hold-lifecycle all pass; bin/fm-lint.sh clean * docs: correct runtime-backend maturity labels for Herdr (#3821) * docs: correct stale tmux/herdr backend maturity claims Herdr now has 21 test files, its own required CI job (tests-herdr) that installs a pinned build and hard-fails on "skip: herdr not found", while tmux has 3 test files and is only required as a dependency of the portable-serial e2e lane. zellij, orca, and cmux still have no CI lane at all. AGENTS.md and docs/herdr-backend.md still called Herdr merely "experimental" alongside those three, misleading every session and reader about actual coverage. Update AGENTS.md's config/backend entry, the opening lines of docs/herdr-backend.md and docs/tmux-backend.md, the runtime-backend section of docs/configuration.md, and the matching claims in docs/architecture.md, CONTRIBUTING.md, and README.md so they agree and distinguish tmux (default), herdr (own required CI lane, largest suite, Windows still spike-only), and zellij/orca/cmux (still experimental, no CI lane). No behavior, selection order, or dispatch logic changes. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GDYsWPEfwTuPNjQ2nGBCcj * no-mistakes(review): docs: fix stale herdr label and CI-lane wording * no-mistakes(document): docs: align tmux adapter label in scripts.md * no-mistakes(review): docs: drop duplicated herdr CI claim from tmux page * no-mistakes(review): docs: drop windows claim, align contributing backend wording * no-mistakes(review): docs: trim duplicated CI claim from herdr opening line * no-mistakes(review): docs: drop unguarded largest-test-suite superlative * no-mistakes(review): docs: restore tmux verified label and README experimental scope --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> * fix: restore intent-targeted no-mistakes validation (#3865) * fix: delete deterministic no-mistakes test baseline, restore intent-targeted Test PR #3644 pinned commands.test to a fm-test-run.sh --changed walk of the repository's 75-162 tests/*.test.sh scripts. no-mistakes runs commands.test verbatim and unconditionally after every fix round, so that walk multiplied by round count: measured at 32.7 minutes per validation versus 3.6 minutes intent-targeted. Delete the pin and restore the 3.6-minute posture. Add tests/fm-nm-test-contract.test.sh as a regression guard, parsing .no-mistakes.yaml as YAML (ruby's bundled Psych, matching the parser tests/fm-test-run.test.sh already uses for ci.yml) rather than grepping its text, restoring in legal form what PR #823 added and PR #1282 removed. Record the rule in docs/configuration.md's "Gate defaults" section (the authoritative owner CONTRIBUTING.md already points at) and strengthen CONTRIBUTING.md's existing local-Test guidance to state it plainly: never configure commands.test to a deterministic test command, complete or partial. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CSt9JvrQMVc4u3jPyUFCFC * no-mistakes(review): Centralize no-mistakes test policy and narrow guard --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> * fix: verify Treehouse slot ownership before teardown (#3837) * fix(bin): verify pool-slot ownership before returning a worktree slot Workers were killed when cleanup returned a Treehouse pool slot that a different, live task had already taken. Teardown now proves the slot is genuinely this task's before releasing it: it refuses when another task record claims the same live worktree path, or when the endpoint's working directory contradicts the recorded slot, and that refusal holds under --force. Slot allocation, metadata publication, ownership verification, and slot return are serialized across linked firstmate homes, and forced secondmate cleanup verifies descendant slot ownership before returning any child worktree. Regression coverage drives the scripts with two task records naming one slot path and asserts the live worker survives and its slot is not reset. * no-mistakes(review): Protect slots across cloned Firstmate homes * no-mistakes(test): Gate teardown locking on genuine Treehouse slots * no-mistakes(test): Clarify pooled descendant slot gating * no-mistakes(test): Synchronize watcher re-arm test on process exit * no-mistakes(test): Wait for watcher cleanup before timeout escalation * no-mistakes(document): Document pool-slot ownership safeguards * no-mistakes(ci): Fixed all reported CI issues: normalized bare local Git origins to the same Treehouse project-lock identity as absolute clone origins; resolved ShellCheck SC1091 with explicit conditional sourcing; and taught concurrent Herdr teardown coverage to retry expected Treehouse lock contention. Added behavioral regression coverage for bare/absolute origin lock identity. Verified endpoint-safety tests, watcher tests, full CI lint, and the previously failing Herdr teardown assertion * fix(bin): resolve relative origins from repository root * no-mistakes(ci): Fixed teardown so an exact recorded endpoint may change cwd without falsely vetoing cleanup. Removed cwd-based ownership refusal while preserving cross-home record exclusivity and project locking. Updated behavioral coverage for both foreign slot ownership refusal and moved-cwd teardown success. Endpoint-safety, backend, watcher, checkpoint, and targeted lint checks pass. Real Herdr presentation E2E progressed successfully but exceeded the 600s local timeout * feat(bin): add verified omp (Oh My Pi) harness adapter for crew, secondmate, and primary (#3867) * feat: add verified omp (Oh My Pi) harness adapter for crew, secondmate, and primary Add omp as a verified harness: anchored process-name detection with a Firstmate-owned FM_OMP_HARNESS launch marker that needs real omp ancestry, the fm-spawn launch template with foreign-marker clearing, the tracked .omp/fm-worker-overlay.yml posture overlay, --auto-approve, --cwd, and pre-launch model validation scoped to providers 'omp models --json' lists. Workers get a state-resident busy-state extension keyed on agent_end without willContinue (omp has no agent_settled). The primary gets two tracked .omp/extensions: a turn-end guard that answers omp's blocking session_stop hook by compelling one continuation per turn, with the pre-tool seatbelts and Run-tier session-start delivery, and a watcher extension ported from the Pi one with fm_watch_arm_omp. Control tables, composer busy footers, omp's status row as a bare-composer boundary, the extension supervision model with an omp-keyed ownership proof, the session-start diagnostic, and the supervision protocol snippet follow. Verified live on omp 18.1.11 with openai-codex/gpt-6-astra: a Herdr scout through spawn, busy state, steer, interrupt, exit, and teardown, and the isolated rpc primary lab through extension auto-discovery, digest delivery, lock identity, watcher arm, successor and wake delivery, and the compelled guard continuation. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh * test: prove the omp guard continuation through a guard spy Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh * fix(spawn): clear the gemini marker at the omp launch boundary Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh * test(omp): force the guard stage by freezing the watcher and clear lint findings Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh * test(omp): reap the live lab by path and record omp's rpc shutdown as a note Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh * test(omp): spawn a real secondmate for the discovery rule and classify the omp surfaces Replace the template-extraction check with a genuine --secondmate launch pinned to the fake tmux backend, assert the worker extension's handler set through the executable rather than its bytes, classify the two new omp surfaces in the documentation inventory, and record the Herdr worker evidence in the runtime-backends verification doc. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh * no-mistakes(review): omp: unverify remote routes, narrow busy regex, drop overlay approval pin * no-mistakes(review): omp: validate config-pinned model, correct remote and marker docs * no-mistakes(review): omp: pin config-model validation with a test, trim overlay * no-mistakes(review): omp: sync guard evidence, drop dead param, map quota family * no-mistakes(review): omp quota: refuse unmapped prefixes, match bare model scopes * no-mistakes(document): docs: cover omp in cd-guard, quota, continuity, tmux * no-mistakes(document): docs: add omp subagent-guard row, fix live test header * no-mistakes(ci): Fixed both failing behavior shards and the Greptile P1 in bin/fm-composer-lib.sh. Root cause of "Behavior portable serial 1" and "Behavior portable parallel 2": the omp busy regex (FM_DELIVERY_OMP_BUSY_REGEX_DEFAULT) and omp status-row furniture regex (FM_COMPOSER_OMP_STATUS_RE_DEFAULT) used the bracket range [⠁-⣿]; BSD grep on macOS accepts it but GNU grep on Linux CI aborts with "Invalid collation character", failing every omp busy/furniture read (3 assertions across fm-omp-harness, fm-tmux-submit-busy, fm-composer-lib). Replaced the range with one shared explicit alternation FM_OMP_SPINNER_FRAMES_RE of omp 18.1.11's unicode-preset spinner frames (status set ⣾⣽⣻⢿⡿⣟⣯⣷ + activity set ⠋⠙⠹⠸⠼⠴⠦⠧⠇⠏, read from the installed binary), the same pattern the Kimi busy regex already uses in CI. For Greptile's finding (the harness-agnostic furniture rule's first alternative matched any 1–4-byte token + ' · ', so wrapped typed input like 'fix · tests' with the cursor on it regressed from pending to unknown; reproduced locally vs base), pinned that alternative to omp's identity cell (π||pi, the icon.omp of each preset in the 18.1.11 binary). Tests: fm-composer-lib.test.sh asserts 'fix · tests' is not furniture, a status-set spinner row is furniture, and the wrapped composer screen reads pending under both locales (CAPS_TMUX cursor 3); fm-omp-harness.test.sh asserts a status-set frame reads busy. New negative cases fail against the pre-fix lib and pass after. Verified: fm-omp-harness, fm-tmux-submit-busy pass via bin/fm-test-run.sh; fm-composer-lib passes all cases except one pre-existing, unrelated local failure (Herdr half-block test uses printf '▀', unsupported by macOS bash 3.2; fails identically on a pristine HEAD export, passes on CI bash 5); shellcheck and bin/fm-lint.sh clean. Caveat: GNU grep is unavailable locally, so the Linux compile was not run directly; the fix uses only constructs already proven on CI's GNU grep (multibyte literal alternations, incl. under LC_ALL=C). Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh, tests/fm-omp-harness.test.sh. No docs needed changes (they describe the rule generically) * no-mistakes(ci): Greptile Review: fixed. The omp status-row furniture regex FM_COMPOSER_OMP_STATUS_RE_DEFAULT in bin/fm-composer-lib.sh still accepted a literal `pi ·` opening, so wrapped composer input beginning with `pi ·` was truncated and misclassified. Read the installed omp 18.1.11 binary: the ascii preset's `icon.omp` is `pi` but its `sep.dot` separator is ` - ` (unicode/nerd use ` · `), so a real ascii status row never contains `pi ·` and that alternative could only ever match typed text. Removal-first fix: dropped `pi` from the identity alternation (now `(π|)`) and updated the comment to record why the ascii preset is excluded. Tests (tests/fm-composer-lib.test.sh): added a negative furniture case for 'pi · e · phi as the three constants' and a wrapped-screen assertion (CAPS_TMUX, cursor 3) that a continuation row opening `pi ·` reads pending in both locales; the new case fails against the unfixed lib and passes after. Verified: composer test with the half-block case skipped passes all 33 cases including the omp matrix; bin/fm-test-run.sh tests/fm-omp-harness.test.sh passes; shellcheck -x clean on both files; bin/fm-lint.sh clean. The full composer test via the runner fails locally only on the pre-existing half-block case (bash 3.2 printf cannot emit ▀; passes on CI bash 5), identical to before this change. Docs unchanged (they describe the rule generically and never mention the ascii identity cell). PR must be raised via no-mistakes: not caused by code. attestation.head_sha is cdddc60 while the PR head is cd51cf4 because the pipeline's ci-phase push moved the head; the outer executor's re-push will re-bind the attestation. No file change for that check. Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> * fix(pi): invoke Bash helpers correctly on native Windows (#3843) * Fix Pi shell invocation on native Windows * no-mistakes(document): Document Pi Windows Bash transport * no-mistakes(ci): Captain, staged a narrow fix: register the Pi Windows regression for both extension paths, make Windows mode emulation non-failing, and enforce LF shell checkouts. Mapping and coverage checks pass; CI/Require no-mistakes were approval-gated externally * no-mistakes(review): Cover async Windows branch-outcome Bash invocation * no-mistakes(review): Preserve Cygwin checks and refresh Windows timing * no-mistakes(test): Invoke OpenCode operational-input owner through Bash on Windows * validation-fixture * no-mistakes(document): Document Windows Bash helper invocation * no-mistakes(ci): Fixed PR-caused changed-selection failure by removing the malformed tracked evidence artifact and allowing deleted, unconsumed source paths to retire cleanly while preserving fail-closed behavior for live unmapped paths. Added regression coverage. Verified native-Windows Pi shell-seam test passes and --changed selects the Windows regression --------- Co-authored-by: test <test@example.invalid> * test(bin): pin teardown outcomes for squash-merged rebased branches (#3870) * fix(bin): recognise squash-merged rebased work as landed at teardown A pipeline rebase can leave the local worktree on pre-rebase commits while GitHub squash-merges the rebased head. The landed-work test then compared those stale commits against a squashed main and refused cleanup of work that had already landed. When the forge reports the recorded PR merged and its merge commit is on the default branch, treat a local branch that only repeats paths from the pipeline push as stale rather than unlanded. If the forge is unreachable, the same coverage check runs against a PR head whose content is already on default. Extra local paths still refuse. * fix(bin): drop unprovable squash-rebase landed-work coverage Path-set coverage treated a diverged local branch as landed whenever it touched the same files as the squash merge. That accepts the reviewer's failing sequence: same path, different content, work discarded. git cherry and merge-tree containment were already too strict on the real rebase-fold case. No remaining check is both safe and permissive enough to recognise a stale pre-rebase copy without also accepting unlanded edits, so that case still refuses. Keep the proofs that hold: a merged PR head that contains local work, or a clean content-in-default tree match. Tests now refuse same-path different content and extra unlanded commits, and still allow a local branch that followed the pipeline rebase. * no-mistakes(review): drop recorded-pr-head fallback and reverted-design leftovers * no-mistakes(review): silence squash-merge stdout corrupting test PR head * no-mistakes(review): make unlanded follow-up commit sole cause of refusal * no-mistakes(document): correct stale squash-rebase fixture comments in teardown tests * no-mistakes(ci): Split the three reported checks: - CI (run 34061098467) and Require no-mistakes (run 34061098460) both concluded `action_required` — approval-gated workflow runs that never executed a step. Not caused by this PR's code; no change can clear them. - Greptile Review was a genuine defect in the new tests: the three new refusal cases (tests/fm-teardown.test.sh) asserted only exit status 1 and a REFUSED line, so a teardown regression that destroyed the worktree, branch, and task record before reporting refusal would still pass. Fix (tests only): added one `assert_refusal_retained_task_state` helper and called it from `test_squash_merged_same_file_different_content_refuses`, `test_squash_merged_rebased_local_with_unlanded_commit_refuses`, and `test_squash_merged_stale_local_refuses_when_forge_unreachable`, each capturing the worktree HEAD before `run_teardown`. It pins that the refusal left the isolated copy on disk, the task branch still checked out at the same unlanded commit, and state/task-x1.meta intact. Verification: the four squash tests pass; a sensitivity probe ran the ALLOW fixture (teardown completes) and pointed the same helper at the outcome — it fires, because a completed teardown detaches/deletes the branch and removes the task record, proving the assertions discriminate. Full tests/fm-teardown.test.sh: 83 passing. bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 + actionlint 1.7.12 (plus an explicit --external-sources pass on the changed file). bin/fm-test-run.sh --check-coverage ok. Caveat: test_herdr_flat_teardown_preflight_refuses_before_changes (mode missing-adapter) fails on this machine. Verified it fails identically on base commit f91a950 via `git archive`, so it is a pre-existing local environment difference untouched by this diff; skipped to run the rest of the suite, not modified --------- Co-authored-by: Morten Gad <mogad@itm8.com> * fix(bin): keep supervision armed for registered custom checks (#3860) * fix(bin): keep supervision armed for registered custom checks A custom check bound by bin/fm-check-register.sh only ever runs inside the watcher's check sweep, but fm_supervision_status counted in-flight tasks, the relay poll shim, and process-event sources as supervision need, and not registered checks. Tearing down the last task therefore stopped every home-level check silently until the next spawn. Count a state/<id>.check.sh that carries its state/<id>.check-trust binding as supervision need. The relay shim keeps its own trust path and task PR polls carry no such binding and are torn down with their task, so neither arms a home by accident. Presence of the binding is the whole test: the sweep validates the bytes at execution time and wakes firstmate when it rejects one, which is the outcome an idle home needs. Closes #3856 * no-mistakes(review): name registered checks in turn-end block banner and doc invariant * no-mistakes(review): narrow PR poll predicate test to what it proves * no-mistakes(document): point Grok re-arm step at supervision-need owner * fix(bin): resolve Treehouse locks for remote secondmate homes (#3883) * fix(bin): resolve the shared Treehouse project lock inside remote secondmate homes Every spawn and teardown inside a remote-seeded secondmate home refused, because the project lock's anchor could not be resolved there. fm_firstmate_root_home walks a home's parent bindings upward to find the anchor the lock lives in, and treated a remote parent binding as an error. A remote-seeded home's parent is on another machine, so that walk can never succeed from there - and neither can the home's own local descendants, whose chain terminates at the same record. Both fail closed on every Treehouse-backed spawn and every pool-slot teardown. A remote parent now terminates the walk at the home holding it, which is the correct anchor: a lock taken on this filesystem is neither held nor observable across that boundary, and that home is already the top of the local tree teardown's collect_local_firstmate_states enumerates, since that walk skips remote registry entries for the same reason. Mutual exclusion is unchanged - every home reachable through local parent links still derives one identical lock file per project, and an unreadable binding, an unsupported route, an unreachable local parent, a cycle, and an over-deep chain all still refuse. Origin-less local-only projects keep resolving through their worktree top. Regression coverage pins the anchor for the main-home layout, a local secondmate, a remote-seeded home, and its local child; drives teardown end-to-end in a remote-seeded home; keeps the cross-home slot-ownership refusal across that boundary; and proves two homes still serialize on the one shared lock file. * no-mistakes(document): Clarify machine-local Treehouse lock ownership * fix(bearings): repair board listening and decision reconciliation (#3872) * fix(bearings): repair the board's listening, card hygiene, and reconcile path Three defects made the fleet board go quiet and then lie about what still needs the captain. Never arm a poll on a session that is not live. `lavish-axi <file>` exits 0 even when it refuses to reopen a session the captain ended from the browser, reporting `status: user-ended` with the same session id, so the build's exit-status check accepted a dead session, printed `already-armed`, and left the board reading "not listening". The build now proves the session is live from a fresh authoritative listing immediately before arming - not from the establish call's status alone, which is already stale by then - reopens once when it finds the session ended, and refuses rather than arming when it stays ended. A reopen also replaces the pre-reopen source generation before reporting success, so a runner on its way out cannot be mistaken for a listener, and a board whose source is registered but unowned gets a replacement started before the build returns. Let a dead generation's ownership actually move. Reclaiming a claim ran its capture-reservation cleanup first, and that cleanup re-verifies the recorded state-root identity, so a claim naming a pid and a process group that were both provably gone could not be cleared: reconcile reported a start while nothing attached, and retire refused with "cannot release source ownership". Reservation records are keyed by claim token and every replacement claims a fresh one, so they are hygiene, not an ownership invariant. Reclamation now additionally requires the owning process group to be absent independently, which keeps a reused pid whose poll child still runs from ever reading as a gone generation. A live owner and a crashed leader whose owned group survives are still never reclaimed. Stop carding decisions whose subject already landed. The build drops a decision card whose work item or PR appears in the payload's own landed rows, and one whose task is no longer an open captain call, naming each drop on stderr. A task whose state cannot be established is kept, because a call wrongly hidden is worse than a card wrongly shown. Add the reconcile choice, and make it structurally incapable of closing a call. Every decision card carries a standard `reconcile` option, injected by the build rather than left to the composer. The board now emits the picked option and any freeform note as separate structured fields instead of fusing them, so a reconcile selection is not expressible as an answer value at all - the defect that let `reconcile - <note>` reach the intake as an ordinary answer. The adapter routes selections from that structured field, creation of a reconcile request is bound to a verified board source rather than the shared keyed-answer intake, and the intake still refuses the reserved value on every channel. Each authorization is bound to the captain-hold generation that produced the card, so an obsolete card cannot close a later call, and both terminal outcomes require a pending request: `reconcile close` records the evidence under its own `reconciled` mode so it never reads as the captain's words, and `reconcile note` leaves the call open. Anything unprovable - an unversioned row, a missing generation, an unreadable…
wonder-media-fleet Bot
pushed a commit
to wonder-media/firstmate
that referenced
this pull request
Sep 27, 2026
* fix(bin): resolve captain holds and legacy teardowns on non-markdown backends (#3782)
* fix(bin): close legacy records on the Beads backend honestly
Two pre-Beads reads blocked honest closure of leftover records:
1. fm-captain-hold.sh complete/verify resolved attested legacy hold ids
only against the live backend and the pre-collapse derived identity, so
a home whose holds fm-hold-migration rehomed under fm- ids failed with
an empty-name absence message (the resolve failure was swallowed by the
command substitution feeding verify_hold_durable). Resolution now falls
back, on the Beads backend only, to the legacy id under the configured
beads prefix and to the row whose notes carry the exact marker line
'migrated from data/backlog.md id <legacy id>'; every refusal names the
id it could not resolve, and the markdown path is unchanged.
2. fm-teardown.sh refused any record without spawn_gen forever. A record
that predates the field can now be torn down with an explicit
--legacy-record flag once the recovery-grade endpoint classifier
confirms the recorded endpoint dead or agent-less; the accepted
incarnation is stamped into the record right before its close marker
binds to it and named in the teardown line. Refusals leave the record
byte-identical, the unlanded-work refusal is not relaxed, and a corrupt
(multi-valued) spawn_gen is never accepted.
The companion repair this branch carries (follow-up commit) is the
backend-gated --file and markdown-file requirement in the mutate path and
lifecycle gates: fm_backlog_mutate passed --file and required the markdown
backlog file regardless of the resolved backend, and the transition gate
plus row probe required that file before any backend work, so a home on a
non-markdown backend could neither gate, probe, nor close its rows.
Behavior tests: self-contained beads fixtures over a scratch bd graph
(self-skipping on markdown-only tasks-axi installs), legacy meta fixtures
for every teardown gate, and the relocated markdown backlog coverage stays
green.
* no-mistakes(review): fix(review): report migrated-hold scan refusals and guard legacy spawn_gen stamp against newline-less records
* fix(backlog): address the configured backend for lifecycle writes
Completes the fm-backlog-transition-lib repair the first commit's message
claims: on this base fm_backlog_mutate passed --file and required the
markdown backlog file regardless of the resolved backend, and
fm_backlog_transition_applies plus fm_backlog_row_probe required that file
before any backend work, so a home on a non-markdown backend could neither
gate, probe, nor close its backlog rows. All three now gate the markdown
file on the resolved tasks-axi backend: markdown keeps exactly its explicit
<data>/backlog.md behavior, non-markdown homes address the backend their
own configuration selects with no markdown file requirement.
fm_backlog_row_show and fm_backlog_row_list already gated correctly and
are unchanged. docs/configuration.md owns the contract line.
Also extends the same backend gate to fm-captain-hold.sh's own mutation
wrapper - hold/add/update/answer/done append the markdown --file only when
the resolved backend is markdown, so a captain call on a Beads home reaches
the Beads store end to end - and applies the review round's two direct
remedies there: the [beads] graph path resolves against the backlog root
when relative (never the process CWD), and a failed bd graph read reports
bd's own trimmed stderr reason in the refusal.
Coverage: tests/fm-backlog-atomicity.test.sh gains a stub-driven Beads
completion case proving the transition gate applies, the row probe reads,
and done runs without any markdown file or --file override; the relocated
markdown backlog test stays green.
* no-mistakes(review): Document root-tasks.toml-only beads settings for migrated-hold resolution
* test(gotmp): stub fm_tasks_axi_backend so the fixture matches the backend-aware transition lib
The legacy-records change made fm-backlog-transition-lib.sh resolve the
configured backend via fm_tasks_axi_backend before the markdown-only skip.
The gotmp fixture's fm-tasks-axi-lib stub lacked that function, so the
markdown check fell through and teardown hit the incompatible-backend
error with unbound FM_TASKS_AXI_MIN under set -u. Stub the backend as
markdown and define the floor, restoring the intended no-backlog skip.
* fix(teardown): roll the legacy stamp back when the close marker fails
A legacy-record teardown stamps its accepted incarnation into the record
right before the close marker binds to it; when that marker write then
fails, the stamp survived, so a retried teardown sailed past the
dead-or-agent-less endpoint gate the stamp now proved unnecessary. The
failed marker write now truncates the record back to its exact pre-stamp
bytes (verified by size), restoring the byte-identical-refusal invariant;
when the rollback itself fails the operator is told to re-run with
--legacy-record after reconciling the endpoint.
Also completes the recorded review decision's coverage wording: the
beads stub test now drives the answer close end to end (update and done
through the gated wrapper), asserting no markdown file override reaches
either verb.
* fix(review): harden the legacy stamp rollback and resolve derived migrated ids
The legacy-record stamp rollback now uses perl (already in the teardown
curated PATH; truncate is not, and is absent on stock macOS), routes every
failure branch inside the stamp block through the same size-verified
rollback so the byte-identical-refusal invariant holds on those paths too,
and gains behavior coverage: an unrecordable close (an invalid pr= link)
fails the teardown, leaves the record byte-identical, keeps the backlog
row in flight, and a flag-less retry still refuses.
Migrated-hold resolution now probes the derived pre-collapse identity
(<origin>-decision-<entry>) alongside the raw entry - fm-hold-migration
recorded the DERIVED id in every migrated row's marker note - in both the
prefix and the migration-note forms, with the ambiguity refusal naming
every identity tried, plus behavior coverage for a bare decision key
resolved through its derived identity's marker.
Also aligns fm-backlog-transition-lib.sh's header ADDRESSING/SCOPE
paragraphs with the backend-gated contract, drops an unreachable FORCE
validity guard the parser rewrite left behind, and switches the new stub
fixture to the portable sed -i.bak idiom.
* no-mistakes(review): Name the configured backend in teardown's backlog reminder
* no-mistakes(review): Scan migration markers before the prefix guess
* no-mistakes(review): Document marker-first resolution and cover the prefix branch
* no-mistakes(document): Record prefix-attestation audit and marker-line forms
* no-mistakes(ci): Fixed the Greptile P1 on bin/fm-teardown.sh: a failed rollback of the synthetic legacy stamp let a retry bypass the dead-or-agent-less endpoint gate. Root cause: teardown minted `spawn_gen=legacy-<ts>-<pid>` into the task record before the close marker bound to it. When the close-marker write failed AND the rollback also failed, the record retained that token. On the next invocation `fm_backlog_meta_spawn_gen` succeeded, so `TEARDOWN_LEGACY_PENDING` stayed 0 and the endpoint gate was skipped entirely — even with `--legacy-record`. The script's own error text told the operator to "re-run teardown with --legacy-record", advice the code could not honor. Fix (bin/fm-teardown.sh): - A `legacy-*` spawn_gen is now recognized as a stamp this teardown path minted, never one a spawn published (fm-spawn.sh publishes `s<epoch>.<pid>.<random>`). Such a record still reads as the legacy record it is: it re-enters the endpoint gate, and a flag-less retry refuses naming `--legacy-record`. - Acceptance reuses the retained token instead of minting a second one; the append block is skipped when the record already carries it, so no duplicate spawn_gen is written. - The rollback attempt and its "could not be rolled back" message are guarded to runs that actually appended a stamp, so a run that appended nothing never claims a rollback it did not perform. - Usage header documents the retained-stamp rule. Test (tests/fm-teardown.test.sh): added `test_retained_legacy_stamp_still_faces_the_endpoint_gate`, an end-to-end reproduction — a `perl` stub that fails only the rollback's `truncate` (delegating every other perl call to the real interpreter) leaves the stamp behind, then the retry must still hit the gate, must not stamp a second incarnation, must not close the backlog row, and the flag-less retry must refuse. Verification: the new test fails against the pre-fix script on exactly the reported defect ("the retry skipped the dead-or-agent-less endpoint gate") and passes after. Full tests/fm-teardown.test.sh 80 ok / 0 failures / rc=0; tests/fm-backlog-atomicity.test.sh 80 ok / 0 failures / rc=0; bin/fm-lint.sh (pinned ShellCheck 0.11.0 + actionlint 1.7.12) clean
* fix: reduce local ShellCheck source-analysis cost (#3778)
* fix(lint): drop source following on the local changed-file gate
The local lint step was inlining library closures through --external-sources
and peaking above 8 GB on a single root. Keep full analysis in CI, on main,
and without a merge-base; exclude the four cross-file codes from the local
pass so those findings still land in CI.
Co-authored-by: Cursor <cursoragent@cursor.com>
* no-mistakes(review): Run local ShellCheck per root and document measurements
* no-mistakes(review): Correct local source-following telemetry
* no-mistakes(document): Clarify context-sensitive lint documentation
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(pi): route decision-owned wake batches to main (#3776)
* fix(pi): route needs-decision wakes and mixed batches wholly to main
Skip the supervision branch for every needs-decision status append, the
same way a check-kind wake already skips it. A coalesced signal/stale
trigger batch containing any needs-decision row is delivered wholly to
main, not split between the branch and a later main wake - the whole
batch, including any co-present routine rows for a different task,
travels together. Heartbeat and unread-status scans stay independent.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE
* test(pi): cover distinct-file mixed batches and heartbeat independence
Add a regression using two distinct files (not the same status file
twice) in one coalesced trigger so a some-vs-every regression on the
file-list cross-reference cannot hide behind a degenerate same-key
case, and a heartbeat/needs-decision co-presence test proving a
needs-decision row neither vetoes nor rides along with an otherwise
eligible heartbeat scan.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE
* no-mistakes(document): Clarify needs-decision and heartbeat routing
* no-mistakes(ci): Fixed captain-held stale reminders so they bypass supervision and wake main, while unrelated unread rows and heartbeats remain independent. Added routing regressions and updated documentation. Pi watcher tests, strict TypeScript checks, lint, and diff checks pass. The no-mistakes attestation failure was pipeline-state related, not a source defect
* no-mistakes(ci): Fixed CI lint by narrowly suppressing false-positive SC2031 diagnostics where background PIDs are captured immediately in the same shell. Verified with `CI=true bin/fm-lint.sh` and `git diff --check`. The no-mistakes attestation failure is pipeline-state-related (`test` was skipped), not a source defect
* no-mistakes(review): Route stale open decisions directly to main
* no-mistakes(review): Honor configured verbs in stale decision routing
* no-mistakes(review): Route second-mate escalations and configured decisions to main
* no-mistakes(review): Ignore trailing whitespace after captain holds
* no-mistakes(review): Cache stale decision classification per status file
* no-mistakes(review): Document unread decision precedence for later task wakes
* no-mistakes(review): Cache unchanged stale decisions across scope scans
* no-mistakes(review): Resolve decision aliases and reject symlinked statuses
* no-mistakes(review): Route surfaced captain-held signals directly to main
* no-mistakes(document): Document decision-owned main routing
* no-mistakes(ci): Fixed captain-held spans to remain actionable while crew working evidence is positive, ensuring the watcher delivers their main-only marker. Updated the executable regression test to cover this case. Verified with the full fm-watch-triage suite, bash syntax checks, and git diff checks. Shellcheck reported only pre-existing test harness warnings (SC1091/SC2034)
* no-mistakes(ci): Fixed the CI regression: captain-held transfers now retain their established non-actionable stale classification while the signal-routing side-band still surfaces them main-only. Verified with tests/fm-daemon.test.sh, tests/fm-watch-triage.test.sh, bash syntax checks, and git diff --check
---------
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* feat(bin): add opt-in worker launch environment allowlist (#3802)
* feat(spawn): add an opt-in worker environment allowlist
Honor a home-local launch-env-allowlist at the shared worker command
boundary and inherit it into secondmate homes. Preserve the existing
launch behavior when the file is absent. Keep the operational environment
and explicit launch assignments, and account for filtered Muse credentials.
Refs https://github.com/kunchenguid/firstmate/issues/3742
Verification:
- Red on origin/main 1820316b66ac2c68e244dd04a02512859ee8c1f4:
the new enabled-allowlist regression observed synthetic-unrelated in
the worker; the absent-file control passed.
- Green: fm-test-run.sh on fm-spawn-dispatch-profile, fm-muse-harness,
and fm-trace-context-spawn; all three passed without skips.
- Synthetic emitted-command probes ran through sh, stock Bash, and zsh.
- Canonical lint, documentation audience checks, and stock Bash syntax
checks passed.
* no-mistakes(review): Reject inaccessible launch environment configuration
* no-mistakes(review): Preserve inherited allowlists on source inspection errors
* no-mistakes(document): Clarify worker environment grants and inheritance documentation
* no-mistakes(lint): Fix inheritance test ShellCheck source boundary
* feat(bin): add rovo crewmate/scout adapter with home-path file access (#3575)
* feat(bin): add rovo as a verified crewmate/scout worker harness
Wire the Atlassian Rovo CLI (202609.1.2) into the TUI-under-tmux/herdr
adapter contract: detection with marker-precedence ordering, one-shot
positional launch with --startup-receipt readiness polling instead of
composer scraping, model/effort flags, a screen-scrape busy fallback
scoped like grok's, and crew/scout-only lifecycle control that refuses
secondmate launches. Ships with a portable regression suite, a live PTY
guard against the real binary, a per-harness reference doc, and a dated
verification record covering the silent OAuth refresh, the interrupt-ack
divergence from the originating scout report, and the still-open
composer-ghost and tmux/herdr pane-liveness gaps.
* no-mistakes(review): revert rovo launch to positional brief, drop startup-receipt
* no-mistakes(test): rewire rovo adapter to kimi-style launch-then-send shape
* no-mistakes(document): add rovo to stale worker-harness enumerations in docs
* docs(verification): close the rovo herdr-liveness gap with live isolated-lab evidence
Placement, launch-then-send, and busy/idle rendering are now verified live
in an isolated non-default Herdr lab session (bin/fm-herdr-lab.sh), driven
directly through fm-spawn.sh's/fm-backend.sh's own shared primitives since
the cross-session launcher-identity guard refuses this task's own ambient
Herdr identity for a full fm-spawn.sh run.
fm_backend_agent_state reported dead for a live, responding rovo pane at
every point checked, because herdr's own agent-integration registry has no
rovo entry (herdr integration status), so herdr agent get returns
agent_not_found regardless of whether rovo is actually running. This is
recorded as a Herdr-side integration gap rather than a firstmate bug, left
unpatched to avoid a false-positive alive verdict for other idle shells.
Updates docs/verification/rovo.md's backend-liveness section and its two
cross-references (docs/verification/runtime-backends.md, docs/configuration.md)
accordingly.
* fix(bin): close rovo's failed-spawn leak and busy-scrape false idle
Greptile P1s on PR #3575: a failed rovo readiness/submission/delivery gate
exited without tearing down the just-created endpoint, leaving the launched
--yolo rovo process running as an orphaned agent outside task control.
Separately, the busy classifier's rendered-tail fallback returned definitive
idle whenever the "Rovo is thinking" marker scrolled out of the last 12
nonblank lines of a long turn, which could make supervision wrongly conclude
a still-working worker had gone idle.
fm-spawn.sh: rovo_spawn_fail now calls rovo_endpoint_cleanup, which kills the
created endpoint (tmux/herdr/zellij/cmux) via the same generic fm_backend_kill
dispatch fm-spawn.sh's own orca-abort path already uses; orca's worktree and
terminal remain owned by the separate ORCA_ABORT_CLEANUP trap.
fm-busy-lib.sh: the rovo classifier arm now reports "unknown rovo-regex"
instead of "idle rovo-regex" when the marker is absent, matching how muse and
cursor already express "can't tell" for their own fallbacks. The positive
busy match is unchanged.
Extends tests/fm-rovo-harness.test.sh: the readiness and delivery failure
tests now assert the endpoint is torn down (and the success test asserts it
is not), and a new test drives the busy marker out of the tail window to
confirm the verdict is unknown, never idle. bin/fm-lint.sh is clean on both
changed files.
* test(rovo): align spawn fixture with the launch-brief validation contract
Upstream main now requires a brief's ## Captain's intent and
## Firstmate spec subsections (or a nonempty legacy # Task body)
before spawn, and rewrites ship+no-mistakes briefs into
launch-brief.md. Update the rovo harness fixture and pointer
assertions to match, mirroring the kimi harness fixture.
* no-mistakes(review): align rovo.md delivery-gate note with live herdr evidence
* fix: prevent stale supervision wake loops (#3672)
* fix(bin): stop the supervision branch's stale-ack and ghost-report loops
Clean-slate implementation of the four authorized recommendations from the
supervision-ghost-retrigger analysis (items 1, 2, 3, and 7), in their minimal
form, superseding PR #3604:
- fm_branch_report refuses a task the wake being handled never named. The
extension fixes the reportable task set from the eligible rows before each
prompt (signal and stale rows resolve to their tasks, a heartbeat allows any
task with a live record, fleet is always allowed), so a report typed from
memory about a task whose records teardown already removed is never stored
or delivered.
- An acknowledgement that consumes nothing says "nothing was acknowledged
through N" and prints the exact --ack-through / --recovery-generation
command for the current presented wake, instead of "re-run the drain",
which re-fed the same stale acknowledgement in a loop.
- bin/fm-guard.sh no longer tells the branch actor to drain queued wakes
while it is handling them; it names the granted rows instead.
- Teardown removes state/.<task>.branch-outcome-index for ordinary tasks and
descendants; the index rebuild and the append-side index write both skip a
task with neither a live record nor a status log, so the branch's report of
a teardown it just performed is stored without recreating the index.
No new locking, no spawn-generation binding, and no retired-task refusal: the
branch can still report the outcome of a task it just tore down, and the
teardown test now proves that path end to end.
* fix(bin): narrow the branch report scope and guard silence to the minimal form
Apply the four review decisions on the clean-slate branch:
- A signal or stale prompt may report only the tasks its own rows resolve
to; fleet is refused there too. A heartbeat review is not scoped by task
at all, so the extension no longer tracks live task records and refuses
nothing by task id during a fleet review.
- The outcome-index rebuild no longer skips retired tasks; the append-side
skip alone keeps a torn-down task's index from being recreated.
- bin/fm-guard.sh keeps the queued-wakes warning silent for the branch actor
instead of printing a replacement note.
* no-mistakes(document): Align supervision docs with scoped wake handling
* fix(bin): grant rovo the per-task home paths its standard crewmate flow needs
rovo confines every file-tool operation to its worktree by default, and its
bash tool independently refuses the same external paths regardless of any
grant (confirmed live), so a rovo worker could not read its own brief or
steering messages or write its status/report - all of which live in the
firstmate home outside the worktree - without hand-feeding it. Grant
toolPermissions.allowedExternalPaths for exactly the task's brief directory,
steering inbox, and status file at launch time via --config-override,
merged with agent.efficiencyLevel into one JSON object since that flag is
single-value and silently discards a second occurrence.
Extends the live PTY guard to prove, against the real binary, that the
grant lets rovo read an external brief and append to an external status
file, and that the same flow is blocked without the grant.
* no-mistakes(document): align rovo reference Effort row with merged single --config-override
* no-mistakes(document): document rovo file-access grant in harness reference
---------
Co-authored-by: PUNEET PATWARI <ppatwari@atlassian.com>
Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
* fix(bin): prevent false pipeline blocks after drive timeouts (#3813)
* fix(bin): read a crew's pipeline-death claim against the live run
A crew's no-mistakes drive call blocks until the next gate or outcome,
routinely far longer than its harness lets one command live, so the call
gets killed or times out while the daemon runs the fix round on in the
background. Crews read that as daemon death and block on it, and firstmate
had nothing that contradicted them.
Rule 7 of every generated brief now says a drive-call error or a harness
command timeout is not a daemon error, requires `no-mistakes daemon status`
plus `no-mistakes axi status` before a pipeline `blocked:`, and reserves
that report for a refused socket or a run record failed with a daemon
error. The no-mistakes definition of done adds the harness command limit
and the background-and-poll shape that fits inside it.
fm-crew-state gains one classification case: a `blocked:` line blaming the
daemon, a timeout, or unreachability, while the run is running or fixing
AND the pipeline reports fresh activity, now reads as superseded because
the run is alive. Recency comes from the client's own `quiet` marker on
active_steps.last_activity rather than a threshold invented here, and
positive evidence is required, so a run record that outlives a genuinely
dead daemon keeps the plain reading.
stuck-crewmate-recovery gains the inverse-of-a-dead-endpoint playbook:
firstmate reads both statuses itself, steers a reattach, never restarts the
shared daemon on a crew's claim, and escalates only a refused socket.
Nothing here depends on an unshipped no-mistakes capability.
* no-mistakes(review): Prioritize daemon socket failure and narrow unreachable matching
* no-mistakes(review): Honor socket refusal across coarse status and crew guidance
* no-mistakes(test): Replace flaky settle timing assertion with pane-read count
* no-mistakes(document): Document daemon timeout recovery contract
* no-mistakes(ci): Fixed daemon socket failures being suppressed by terminal attributed runs. Positive refused/missing socket evidence now remains blocked regardless of run status. Added a behavioral regression test for terminal failed runs. Verified with fm-crew-state tests, project ShellCheck lint, and git diff checks
* fix(bin): scope the worker role contract for ship and scout launches (#3797)
* fix(brief): scope Firstmate workers to their launch contract
* no-mistakes(document): Clarify supervisor scope and worker contract ownership
* no-mistakes(ci): Removed the heading-based bypass so every ship/scout launch receives the current worker-role contract. Added a regression that failed before the fix and passes afterward. Dispatch, brief, and delivery suites, focused ShellCheck, and git diff --check all passed
* no-mistakes(review): make launch overlay sole owner of worker role contract
* no-mistakes(review): narrow heading test dimension and fix publish error wording
* no-mistakes(review): gate role supersession, fix render guard, drop AGENTS twin
* no-mistakes(document): align architecture AGENTS.md scope and spawn launch-brief header
* fix(bin): stop ringing steering doorbells into dead panes (#3823)
* fix(bin): stop ringing steering doorbells into dead panes
The steering-inbox doorbell was a plain sentence plus Enter typed into a
worker's pane, and the watcher re-rang it on the assumption that a ring is
free. In a pane whose agent has exited that line is a shell command, and the
re-ring ladder kept typing it into a shell that can never acknowledge it.
- Prefix the doorbell with the shell no-op `: ` so a bare shell executes
nothing while a live worker still reads the same self-describing line.
`#` is not used because interactive zsh does not treat it as a comment by
default and the claude harness binds it to memory mode.
- fm_task_inbox_ring skips the pane (return 3) when the backend positively
classifies the agent as dead; missing, ambiguous, unreadable, and unverified
endpoints still ring so a blind classifier never starves a live worker.
- The watcher caps the ladder for a dead pane: one stale wake for recovery,
no ring, no ladder walk, and the durable record stays for
stuck-crewmate-recovery. fm-send and the remote steer leg report the skip.
Tests cover the no-op in real shells, the dead/live/unclassifiable ring
verdicts, and the single-surfacing watcher path.
* no-mistakes(review): Quote doorbell paths against shell injection
* no-mistakes(review): Reject terminal-control paths before ringing
* no-mistakes(review): Document accepted partial doorbell delivery race
* no-mistakes(review): Skip unavailable endpoints before busy-state handling
* no-mistakes(test): Respect shell startup PATH in environment allowlist test
* no-mistakes(test): Fix doorbell test fixtures for endpoint liveness
* no-mistakes(test): Prioritize confirmed restarts and clean shell test syntax
* no-mistakes(document): Document dead and missing doorbell recovery
* no-mistakes(ci): Fixed the persistence-reply timeout race by rechecking for a correlated reply immediately before falling back to a nudge. Added a deterministic regression covering replies arriving between the preliminary resolution pass and timeout handling. Verified with the targeted restart suite, project lint, coverage guard, bash syntax checks, and diff checks
* fix(bin): wait out transient primary-checkout reads in the spawn worktree poll (#3834)
* fix(spawn): keep the worktree poll from adopting the repository primary
After `treehouse get` is sent, the worktree-discovery poll reads the pane's
foreground-process cwd. While treehouse is still fetching and checking a slot
out, the foreground process is treehouse itself and it reports the repository's
PRIMARY checkout as its cwd for several seconds. The poll accepted any path
that merely differed from the spawning project, so from a linked spawning home
- whose project is itself a worktree of that repository - it adopted the
primary, and the isolation guard then refused a launch whose slot treehouse
went on to create normally.
Screen every candidate with the isolation guard's own conditions, extracted as
spawn_worktree_isolated, so a read the guard would reject stays a transient the
poll keeps waiting through. The two-consecutive-reads rule and the guard as
final backstop are unchanged; a pane that never reaches an isolated worktree
still fails at the existing 60s deadline, now naming the last path it reported.
The already-settled timing assertion counted whole-spawn wall time against a
5s budget and failed on unmodified HEAD on slower machines; it now counts pane
reads, which is what "one confirming read, not an extra cycle" actually means.
* fix(spawn): say which path the worktree wait rejected, and why
Screening every discovery-poll candidate means a host that never reaches an
isolated worktree spends the whole 60s window before refusing. That wait is
deliberate - separating a transient from a terminal misconfiguration needs
machinery this path does not want - so the refusal explains itself instead:
the isolation check records why a candidate failed, and the deadline names the
last path seen together with that reason. Message and diagnostics only; the
poll's control flow is unchanged.
Two suites asserted the guard's wording on paths the poll now rejects rather
than adopts, so their refusal arrives from the deadline instead: realign
fm-tangle-guard's non-git and subdirectory-of-primary cases (each now also
asserting the stated reason, and the second the metadata absence it was
missing) and the herdr projection e2e's forced non-worktree cwd.
* no-mistakes(review): stub poll sleep in tangle-guard spawn isolation test
* no-mistakes(document): document spawn poll isolation screen in fm-spawn header
* test(spawn): make the non-git isolation case non-git anywhere
The refusal-reason assertion for a path outside any repository assumed TMPDIR
is not inside a git repository. Where it is, git walks up from the temporary
directory, finds that repository, and the spawn reports the subdirectory cause
instead - so the case passed or failed on a property of the host rather than on
the behaviour under test.
Build the path under a directory the test then names in GIT_CEILING_DIRECTORIES,
which git documents as not chdir-ing up into a listed directory while looking
for a repository. Git never excludes the directory being searched, so the
ceiling is the parent of the path handed to the spawn.
The assertions pin which cause fired rather than the sentence that explains it,
leaving the operator wording free to improve.
* no-mistakes(document): point spawn poll comment at the isolation screen's comparison
* no-mistakes(ci): Fixed the "Behavior portable serial 2" failure in tests/fm-tangle-guard.test.sh ("non-worktree spawn did not say why the path was rejected (missing: 'not inside a git worktree')"). Root cause, in this PR's code: bin/fm-spawn.sh's spawn_worktree_isolated resolved the git toplevel with `wt_top_real=$(cd "$SPAWN_WT_TOP" ...)`. For a path in no repository, `git rev-parse --show-toplevel` yields empty, and `cd ""` is a SUCCESSFUL no-op on bash before 5.3 (CI's ubuntu-latest ships bash 5.2). The empty toplevel therefore resolved to fm-spawn's own cwd — the CI checkout — so the poll reported "it is a subdirectory of worktree root '/home/runner/work/firstmate/firstmate'" instead of the correct "it is not inside a git worktree". Dev machines with bash 5.3 fail `cd ""`, which is why the suite passed locally and only failed on CI; it is a genuine shell-portability defect in the reason vocabulary this change added, not a test-environment artifact. Fix (smallest root-cause change, 1 line + comment, bin/fm-spawn.sh:2168-2173): guard the empty value so it never reaches `cd` — if [ -n "$SPAWN_WT_TOP" ] && ! wt_top_real=$(cd "$SPAWN_WT_TOP" 2>/dev/null && pwd -P); then No change to the poll's timing or deadline behavior (respecting the recorded refusal-latency and spawn-wt-reason-vocabulary decisions), no new tests, no other files touched. Verification: - Reproduced the exact CI failure locally by putting bash 3.2 (same `cd ""` semantics as CI's 5.2) first on PATH: fails before the fix with the identical message shape, passes after. - tests/fm-tangle-guard.test.sh passes under both bash 3.2 and bash 5.3. - tests/fm-spawn-worktree-settle.test.sh and tests/fm-spawn-pool-base-freshen.test.sh pass; shellcheck -x bin/fm-spawn.sh clean. - bin/fm-test-run.sh --changed: 46 suites completed, every FM_TEST_END exit=0, 1076 passing assertions, 0 "not ok" (including fm-tangle-guard, fm-control-relaunch, fm-lint). The run ended on my own 900s wall-clock cap (rc=124), not on any test failure
* fix: support stock macOS Bash 3.2 paths (#3732)
Co-authored-by: Talon Stark <talonstark@gmail.com>
* fix(bin): read orphaned green ci monitor and daemon-down failed record as not failed (#3846)
* fix(bin): read an orphaned green ci monitor as held-for-merge, not failed
A no-mistakes run held for a captain merge decision keeps its ci step
polling until merged or closed; when the shared daemon restarts under
that poll, the run is recorded failed although every substantive step
completed and GitHub reports the PR green. A monitor whose only
remaining job is to observe a human decision must not convert the
absence of that decision into a failure verdict.
fm-crew-state.sh now reclassifies a terminal failed run as done
(held-for-merge), surfacing the run's PR URL, when the steps table
shows every step completed except exactly ci failed and the ci log's
last recognized marker reads checks green. A genuinely red check, an
unreadable ci log, or a second failed step keeps the failure.
* no-mistakes(review): Read daemon-down coarse failed ledger as unknown, not failed
* no-mistakes(document): docs: align AGENTS.md failed-verdict guidance with crew-state reclassification
* fix(spawn): verify preserved backlog state after interrupted spawn delivery (#3852)
* fix(bin): read preserved spawn state back before the interrupted exit claims it
The deferred-signal exit path asserted the paired task record and
In-flight backlog state were preserved without reading either back,
exactly when a reader is least able to check (fm-yi4j evidence,
2026-09-05). The commit's exit status alone has been observed to agree
with a row that did not actually move.
The exit path now re-reads the record and the row under the same
per-task lock as the commit, repairs a row the commit believed it
moved, and phrases the error as exactly what was verified or attempted
- verified preserved, repaired and verified, or an explicit
preservation-could-not-be-verified with the reason and hand-closeout
instruction. Two behavior tests drive a lying tasks-axi start through a
real interrupted spawn and assert the printed claim and the real
backlog state agree.
* no-mistakes(test): Fix calm suite for Pi 0.85 and pin test umask
* no-mistakes(document): Document interrupted-spawn preservation claim in backlog gate owner
* no-mistakes(ci): Fixed the Greptile P1 in bin/fm-spawn.sh's deferred-signal exit path: during preservation verification, the no-op HUP/INT/TERM re-trap combined with an unresponsive `tasks-axi show`/`start` (bash cannot run traps while a foreground child runs) held the per-task meta lock - and every lifecycle operation waiting on it - indefinitely. Root-cause fix: bound every tasks-axi invocation made under the lock. bin/fm-backlog-transition-lib.sh gains fm_tasks_axi, an exec-based wrapper (GNU timeout, gtimeout fallback) used by fm_backlog_row_show and fm_backlog_mutate that preserves the exact process placement of the plain tasks-axi call; bin/fm-spawn.sh sets FM_TASKS_AXI_TIMEOUT (default 30s) at the commit point so both the commit and the read-back verification are bounded. A timed-out call fails through the existing error plumbing and probe/mutate name the timeout as the reason, so the interrupted exit path prints honest 'preservation could not be verified ... (reason)' wording - never intent phrased as outcome, matching the author's intent. Added a behavior test in tests/fm-backlog-atomicity.test.sh that drives a real interrupted spawn through a lying tasks-axi whose repair start never answers; it asserts the spawn exits promptly (self-bounded by an outer timeout), the attempted wording names the timeout, and the printed claim agrees with the real record/backlog state. Confirmed the test fails on the unfixed tree and passes with the fix. Verified: fm-backlog-atomicity (83 ok), fm-transition-lib, fm-backlog-handoff, fm-captain-hold, fm-teardown, fm-fleet-snapshot-view, fm-secondmate-reconcile, fm-spawn-batch, fm-spawn-dispatch-profile, fm-task-delivery, fm-control-relaunch all pass; bin/fm-lint.sh clean. The fm-bootstrap 'unsplit run lost its local diagnostic' failure reproduces on the pristine base commit and is unrelated to this change
* no-mistakes(ci): Fixed the Greptile P1 in bin/fm-backlog-transition-lib.sh: fm_tasks_axi bounded tasks-axi only through GNU timeout/gtimeout and fell through to an unbounded exec on hosts with neither (stock macOS), so an unresponsive call could hold the per-task meta lock forever during interrupted-spawn verification. Root-cause fix: the bound now has no unbounded path. GNU timeout is preferred, gtimeout next, then a small perl watchdog (fork + waitpid WNOHANG polling at 50ms, TERM on expiry, one bound of grace, then KILL, exit 124 so the callers' existing timeout plumbing reports it; exit statuses and output pass through unchanged). Polling was chosen over alarm+die to avoid perl's platform-dependent syscall-restart semantics. When a bound is requested but no bounding mechanism exists, the call fails closed (exit 127 with a diagnostic) rather than running unbounded, so the interrupted exit path prints honest attempted wording, never intent as outcome. The unbounded exec remains only for the no-bound plain-call case. Added three behavior tests in tests/fm-backlog-atomicity.test.sh driving fm_tasks_axi through a PATH with no timeout binary: a hanging stub must exit 124 within the bound (verified to fail on the pre-fix code), a failing stub's status/output must pass through, and a tool-less PATH must fail closed with the diagnostic. Verified: fm-backlog-atomicity 86/86 ok, fm-transition-lib, fm-backlog-handoff, fm-teardown, fm-spawn-batch, fm-task-delivery, fm-fleet-snapshot-view, fm-secondmate-reconcile, fm-control-relaunch, fm-spawn-dispatch-profile all pass; bin/fm-lint.sh clean
* no-mistakes(ci): Fixed the Greptile P1 on bin/fm-backlog-transition-lib.sh: fm_tasks_axi's GNU timeout and gtimeout paths sent TERM at the bound but had no kill-after, so a tasks-axi that ignores SIGTERM kept the bounded call - and the per-task meta lock - held indefinitely during interrupted-spawn verification. Root-cause fix: both GNU execs now carry -k "$bound" (TERM at the bound, KILL after one further bound of grace), giving every bounded path the same forced-termination contract the perl watchdog already had. Because GNU timeout exits 137 (128+SIGKILL) when the kill-after fires - versus 124 for a TERM expiry - the probe/mutate timeout detection now goes through a new fm_tasks_axi_timeout_expired helper that treats 124 and 137 alike, so the interrupted exit path still names the timeout as the reason; the helper keeps the bound check in one place. Added a behavior test in tests/fm-backlog-atomicity.test.sh that drives fm_tasks_axi through a real GNU timeout with a tasks-axi stub that traps and ignores TERM (the ignored disposition survives exec into sleep) and asserts a bound-expiry status plus completion within bound+grace; on the pre-fix code the suite hangs until killed, confirming the reproduction. Verified: fm-backlog-atomicity 87/87 ok, fm-transition-lib, fm-backlog-handoff, fm-spawn-batch, fm-task-delivery, fm-teardown, fm-secondmate-reconcile, fm-control-relaunch, fm-spawn-dispatch-profile, fm-captain-hold-lifecycle all pass; bin/fm-lint.sh clean
* docs: correct runtime-backend maturity labels for Herdr (#3821)
* docs: correct stale tmux/herdr backend maturity claims
Herdr now has 21 test files, its own required CI job (tests-herdr)
that installs a pinned build and hard-fails on "skip: herdr not
found", while tmux has 3 test files and is only required as a
dependency of the portable-serial e2e lane. zellij, orca, and cmux
still have no CI lane at all. AGENTS.md and docs/herdr-backend.md
still called Herdr merely "experimental" alongside those three,
misleading every session and reader about actual coverage.
Update AGENTS.md's config/backend entry, the opening lines of
docs/herdr-backend.md and docs/tmux-backend.md, the runtime-backend
section of docs/configuration.md, and the matching claims in
docs/architecture.md, CONTRIBUTING.md, and README.md so they agree
and distinguish tmux (default), herdr (own required CI lane, largest
suite, Windows still spike-only), and zellij/orca/cmux (still
experimental, no CI lane). No behavior, selection order, or
dispatch logic changes.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDYsWPEfwTuPNjQ2nGBCcj
* no-mistakes(review): docs: fix stale herdr label and CI-lane wording
* no-mistakes(document): docs: align tmux adapter label in scripts.md
* no-mistakes(review): docs: drop duplicated herdr CI claim from tmux page
* no-mistakes(review): docs: drop windows claim, align contributing backend wording
* no-mistakes(review): docs: trim duplicated CI claim from herdr opening line
* no-mistakes(review): docs: drop unguarded largest-test-suite superlative
* no-mistakes(review): docs: restore tmux verified label and README experimental scope
---------
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* fix: restore intent-targeted no-mistakes validation (#3865)
* fix: delete deterministic no-mistakes test baseline, restore intent-targeted Test
PR #3644 pinned commands.test to a fm-test-run.sh --changed walk of the
repository's 75-162 tests/*.test.sh scripts. no-mistakes runs commands.test
verbatim and unconditionally after every fix round, so that walk multiplied
by round count: measured at 32.7 minutes per validation versus 3.6 minutes
intent-targeted. Delete the pin and restore the 3.6-minute posture.
Add tests/fm-nm-test-contract.test.sh as a regression guard, parsing
.no-mistakes.yaml as YAML (ruby's bundled Psych, matching the parser
tests/fm-test-run.test.sh already uses for ci.yml) rather than grepping its
text, restoring in legal form what PR #823 added and PR #1282 removed.
Record the rule in docs/configuration.md's "Gate defaults" section (the
authoritative owner CONTRIBUTING.md already points at) and strengthen
CONTRIBUTING.md's existing local-Test guidance to state it plainly: never
configure commands.test to a deterministic test command, complete or partial.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSt9JvrQMVc4u3jPyUFCFC
* no-mistakes(review): Centralize no-mistakes test policy and narrow guard
---------
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* fix: verify Treehouse slot ownership before teardown (#3837)
* fix(bin): verify pool-slot ownership before returning a worktree slot
Workers were killed when cleanup returned a Treehouse pool slot that a
different, live task had already taken. Teardown now proves the slot is
genuinely this task's before releasing it: it refuses when another task
record claims the same live worktree path, or when the endpoint's working
directory contradicts the recorded slot, and that refusal holds under
--force. Slot allocation, metadata publication, ownership verification,
and slot return are serialized across linked firstmate homes, and forced
secondmate cleanup verifies descendant slot ownership before returning
any child worktree.
Regression coverage drives the scripts with two task records naming one
slot path and asserts the live worker survives and its slot is not reset.
* no-mistakes(review): Protect slots across cloned Firstmate homes
* no-mistakes(test): Gate teardown locking on genuine Treehouse slots
* no-mistakes(test): Clarify pooled descendant slot gating
* no-mistakes(test): Synchronize watcher re-arm test on process exit
* no-mistakes(test): Wait for watcher cleanup before timeout escalation
* no-mistakes(document): Document pool-slot ownership safeguards
* no-mistakes(ci): Fixed all reported CI issues: normalized bare local Git origins to the same Treehouse project-lock identity as absolute clone origins; resolved ShellCheck SC1091 with explicit conditional sourcing; and taught concurrent Herdr teardown coverage to retry expected Treehouse lock contention. Added behavioral regression coverage for bare/absolute origin lock identity. Verified endpoint-safety tests, watcher tests, full CI lint, and the previously failing Herdr teardown assertion
* fix(bin): resolve relative origins from repository root
* no-mistakes(ci): Fixed teardown so an exact recorded endpoint may change cwd without falsely vetoing cleanup. Removed cwd-based ownership refusal while preserving cross-home record exclusivity and project locking. Updated behavioral coverage for both foreign slot ownership refusal and moved-cwd teardown success. Endpoint-safety, backend, watcher, checkpoint, and targeted lint checks pass. Real Herdr presentation E2E progressed successfully but exceeded the 600s local timeout
* feat(bin): add verified omp (Oh My Pi) harness adapter for crew, secondmate, and primary (#3867)
* feat: add verified omp (Oh My Pi) harness adapter for crew, secondmate, and primary
Add omp as a verified harness: anchored process-name detection with a
Firstmate-owned FM_OMP_HARNESS launch marker that needs real omp ancestry,
the fm-spawn launch template with foreign-marker clearing, the tracked
.omp/fm-worker-overlay.yml posture overlay, --auto-approve, --cwd, and
pre-launch model validation scoped to providers 'omp models --json' lists.
Workers get a state-resident busy-state extension keyed on agent_end
without willContinue (omp has no agent_settled). The primary gets two
tracked .omp/extensions: a turn-end guard that answers omp's blocking
session_stop hook by compelling one continuation per turn, with the
pre-tool seatbelts and Run-tier session-start delivery, and a watcher
extension ported from the Pi one with fm_watch_arm_omp. Control tables,
composer busy footers, omp's status row as a bare-composer boundary, the
extension supervision model with an omp-keyed ownership proof, the
session-start diagnostic, and the supervision protocol snippet follow.
Verified live on omp 18.1.11 with openai-codex/gpt-6-astra: a Herdr scout
through spawn, busy state, steer, interrupt, exit, and teardown, and the
isolated rpc primary lab through extension auto-discovery, digest
delivery, lock identity, watcher arm, successor and wake delivery, and
the compelled guard continuation.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh
* test: prove the omp guard continuation through a guard spy
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh
* fix(spawn): clear the gemini marker at the omp launch boundary
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh
* test(omp): force the guard stage by freezing the watcher and clear lint findings
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh
* test(omp): reap the live lab by path and record omp's rpc shutdown as a note
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh
* test(omp): spawn a real secondmate for the discovery rule and classify the omp surfaces
Replace the template-extraction check with a genuine --secondmate launch
pinned to the fake tmux backend, assert the worker extension's handler set
through the executable rather than its bytes, classify the two new omp
surfaces in the documentation inventory, and record the Herdr worker
evidence in the runtime-backends verification doc.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh
* no-mistakes(review): omp: unverify remote routes, narrow busy regex, drop overlay approval pin
* no-mistakes(review): omp: validate config-pinned model, correct remote and marker docs
* no-mistakes(review): omp: pin config-model validation with a test, trim overlay
* no-mistakes(review): omp: sync guard evidence, drop dead param, map quota family
* no-mistakes(review): omp quota: refuse unmapped prefixes, match bare model scopes
* no-mistakes(document): docs: cover omp in cd-guard, quota, continuity, tmux
* no-mistakes(document): docs: add omp subagent-guard row, fix live test header
* no-mistakes(ci): Fixed both failing behavior shards and the Greptile P1 in bin/fm-composer-lib.sh. Root cause of "Behavior portable serial 1" and "Behavior portable parallel 2": the omp busy regex (FM_DELIVERY_OMP_BUSY_REGEX_DEFAULT) and omp status-row furniture regex (FM_COMPOSER_OMP_STATUS_RE_DEFAULT) used the bracket range [⠁-⣿]; BSD grep on macOS accepts it but GNU grep on Linux CI aborts with "Invalid collation character", failing every omp busy/furniture read (3 assertions across fm-omp-harness, fm-tmux-submit-busy, fm-composer-lib). Replaced the range with one shared explicit alternation FM_OMP_SPINNER_FRAMES_RE of omp 18.1.11's unicode-preset spinner frames (status set ⣾⣽⣻⢿⡿⣟⣯⣷ + activity set ⠋⠙⠹⠸⠼⠴⠦⠧⠇⠏, read from the installed binary), the same pattern the Kimi busy regex already uses in CI. For Greptile's finding (the harness-agnostic furniture rule's first alternative matched any 1–4-byte token + ' · ', so wrapped typed input like 'fix · tests' with the cursor on it regressed from pending to unknown; reproduced locally vs base), pinned that alternative to omp's identity cell (π||pi, the icon.omp of each preset in the 18.1.11 binary). Tests: fm-composer-lib.test.sh asserts 'fix · tests' is not furniture, a status-set spinner row is furniture, and the wrapped composer screen reads pending under both locales (CAPS_TMUX cursor 3); fm-omp-harness.test.sh asserts a status-set frame reads busy. New negative cases fail against the pre-fix lib and pass after. Verified: fm-omp-harness, fm-tmux-submit-busy pass via bin/fm-test-run.sh; fm-composer-lib passes all cases except one pre-existing, unrelated local failure (Herdr half-block test uses printf '▀', unsupported by macOS bash 3.2; fails identically on a pristine HEAD export, passes on CI bash 5); shellcheck and bin/fm-lint.sh clean. Caveat: GNU grep is unavailable locally, so the Linux compile was not run directly; the fix uses only constructs already proven on CI's GNU grep (multibyte literal alternations, incl. under LC_ALL=C). Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh, tests/fm-omp-harness.test.sh. No docs needed changes (they describe the rule generically)
* no-mistakes(ci): Greptile Review: fixed. The omp status-row furniture regex FM_COMPOSER_OMP_STATUS_RE_DEFAULT in bin/fm-composer-lib.sh still accepted a literal `pi ·` opening, so wrapped composer input beginning with `pi ·` was truncated and misclassified. Read the installed omp 18.1.11 binary: the ascii preset's `icon.omp` is `pi` but its `sep.dot` separator is ` - ` (unicode/nerd use ` · `), so a real ascii status row never contains `pi ·` and that alternative could only ever match typed text. Removal-first fix: dropped `pi` from the identity alternation (now `(π|)`) and updated the comment to record why the ascii preset is excluded. Tests (tests/fm-composer-lib.test.sh): added a negative furniture case for 'pi · e · phi as the three constants' and a wrapped-screen assertion (CAPS_TMUX, cursor 3) that a continuation row opening `pi ·` reads pending in both locales; the new case fails against the unfixed lib and passes after. Verified: composer test with the half-block case skipped passes all 33 cases including the omp matrix; bin/fm-test-run.sh tests/fm-omp-harness.test.sh passes; shellcheck -x clean on both files; bin/fm-lint.sh clean. The full composer test via the runner fails locally only on the pre-existing half-block case (bash 3.2 printf cannot emit ▀; passes on CI bash 5), identical to before this change. Docs unchanged (they describe the rule generically and never mention the ascii identity cell). PR must be raised via no-mistakes: not caused by code. attestation.head_sha is cdddc60 while the PR head is cd51cf4 because the pipeline's ci-phase push moved the head; the outer executor's re-push will re-bind the attestation. No file change for that check. Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* fix(pi): invoke Bash helpers correctly on native Windows (#3843)
* Fix Pi shell invocation on native Windows
* no-mistakes(document): Document Pi Windows Bash transport
* no-mistakes(ci): Captain, staged a narrow fix: register the Pi Windows regression for both extension paths, make Windows mode emulation non-failing, and enforce LF shell checkouts. Mapping and coverage checks pass; CI/Require no-mistakes were approval-gated externally
* no-mistakes(review): Cover async Windows branch-outcome Bash invocation
* no-mistakes(review): Preserve Cygwin checks and refresh Windows timing
* no-mistakes(test): Invoke OpenCode operational-input owner through Bash on Windows
* validation-fixture
* no-mistakes(document): Document Windows Bash helper invocation
* no-mistakes(ci): Fixed PR-caused changed-selection failure by removing the malformed tracked evidence artifact and allowing deleted, unconsumed source paths to retire cleanly while preserving fail-closed behavior for live unmapped paths. Added regression coverage. Verified native-Windows Pi shell-seam test passes and --changed selects the Windows regression
---------
Co-authored-by: test <test@example.invalid>
* test(bin): pin teardown outcomes for squash-merged rebased branches (#3870)
* fix(bin): recognise squash-merged rebased work as landed at teardown
A pipeline rebase can leave the local worktree on pre-rebase commits while
GitHub squash-merges the rebased head. The landed-work test then compared
those stale commits against a squashed main and refused cleanup of work
that had already landed.
When the forge reports the recorded PR merged and its merge commit is on
the default branch, treat a local branch that only repeats paths from the
pipeline push as stale rather than unlanded. If the forge is unreachable,
the same coverage check runs against a PR head whose content is already
on default. Extra local paths still refuse.
* fix(bin): drop unprovable squash-rebase landed-work coverage
Path-set coverage treated a diverged local branch as landed whenever it
touched the same files as the squash merge. That accepts the reviewer's
failing sequence: same path, different content, work discarded.
git cherry and merge-tree containment were already too strict on the real
rebase-fold case. No remaining check is both safe and permissive enough
to recognise a stale pre-rebase copy without also accepting unlanded
edits, so that case still refuses.
Keep the proofs that hold: a merged PR head that contains local work, or
a clean content-in-default tree match. Tests now refuse same-path
different content and extra unlanded commits, and still allow a local
branch that followed the pipeline rebase.
* no-mistakes(review): drop recorded-pr-head fallback and reverted-design leftovers
* no-mistakes(review): silence squash-merge stdout corrupting test PR head
* no-mistakes(review): make unlanded follow-up commit sole cause of refusal
* no-mistakes(document): correct stale squash-rebase fixture comments in teardown tests
* no-mistakes(ci): Split the three reported checks: - CI (run 34061098467) and Require no-mistakes (run 34061098460) both concluded `action_required` — approval-gated workflow runs that never executed a step. Not caused by this PR's code; no change can clear them. - Greptile Review was a genuine defect in the new tests: the three new refusal cases (tests/fm-teardown.test.sh) asserted only exit status 1 and a REFUSED line, so a teardown regression that destroyed the worktree, branch, and task record before reporting refusal would still pass. Fix (tests only): added one `assert_refusal_retained_task_state` helper and called it from `test_squash_merged_same_file_different_content_refuses`, `test_squash_merged_rebased_local_with_unlanded_commit_refuses`, and `test_squash_merged_stale_local_refuses_when_forge_unreachable`, each capturing the worktree HEAD before `run_teardown`. It pins that the refusal left the isolated copy on disk, the task branch still checked out at the same unlanded commit, and state/task-x1.meta intact. Verification: the four squash tests pass; a sensitivity probe ran the ALLOW fixture (teardown completes) and pointed the same helper at the outcome — it fires, because a completed teardown detaches/deletes the branch and removes the task record, proving the assertions discriminate. Full tests/fm-teardown.test.sh: 83 passing. bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 + actionlint 1.7.12 (plus an explicit --external-sources pass on the changed file). bin/fm-test-run.sh --check-coverage ok. Caveat: test_herdr_flat_teardown_preflight_refuses_before_changes (mode missing-adapter) fails on this machine. Verified it fails identically on base commit f91a950 via `git archive`, so it is a pre-existing local environment difference untouched by this diff; skipped to run the rest of the suite, not modified
---------
Co-authored-by: Morten Gad <mogad@itm8.com>
* fix(bin): keep supervision armed for registered custom checks (#3860)
* fix(bin): keep supervision armed for registered custom checks
A custom check bound by bin/fm-check-register.sh only ever runs inside the
watcher's check sweep, but fm_supervision_status counted in-flight tasks, the
relay poll shim, and process-event sources as supervision need, and not
registered checks. Tearing down the last task therefore stopped every
home-level check silently until the next spawn.
Count a state/<id>.check.sh that carries its state/<id>.check-trust binding as
supervision need. The relay shim keeps its own trust path and task PR polls
carry no such binding and are torn down with their task, so neither arms a home
by accident. Presence of the binding is the whole test: the sweep validates the
bytes at execution time and wakes firstmate when it rejects one, which is the
outcome an idle home needs.
Closes #3856
* no-mistakes(review): name registered checks in turn-end block banner and doc invariant
* no-mistakes(review): narrow PR poll predicate test to what it proves
* no-mistakes(document): point Grok re-arm step at supervision-need owner
* fix(bin): resolve Treehouse locks for remote secondmate homes (#3883)
* fix(bin): resolve the shared Treehouse project lock inside remote secondmate homes
Every spawn and teardown inside a remote-seeded secondmate home refused,
because the project lock's anchor could not be resolved there.
fm_firstmate_root_home walks a home's parent bindings upward to find the
anchor the lock lives in, and treated a remote parent binding as an error.
A remote-seeded home's parent is on another machine, so that walk can never
succeed from there - and neither can the home's own local descendants, whose
chain terminates at the same record. Both fail closed on every Treehouse-backed
spawn and every pool-slot teardown.
A remote parent now terminates the walk at the home holding it, which is the
correct anchor: a lock taken on this filesystem is neither held nor observable
across that boundary, and that home is already the top of the local tree
teardown's collect_local_firstmate_states enumerates, since that walk skips
remote registry entries for the same reason. Mutual exclusion is unchanged -
every home reachable through local parent links still derives one identical
lock file per project, and an unreadable binding, an unsupported route, an
unreachable local parent, a cycle, and an over-deep chain all still refuse.
Origin-less local-only projects keep resolving through their worktree top.
Regression coverage pins the anchor for the main-home layout, a local
secondmate, a remote-seeded home, and its local child; drives teardown
end-to-end in a remote-seeded home; keeps the cross-home slot-ownership
refusal across that boundary; and proves two homes still serialize on the
one shared lock file.
* no-mistakes(document): Clarify machine-local Treehouse lock ownership
* fix(bearings): repair board listening and decision reconciliation (#3872)
* fix(bearings): repair the board's listening, card hygiene, and reconcile path
Three defects made the fleet board go quiet and then lie about what still
needs the captain.
Never arm a poll on a session that is not live. `lavish-axi <file>` exits 0
even when it refuses to reopen a session the captain ended from the browser,
reporting `status: user-ended` with the same session id, so the build's
exit-status check accepted a dead session, printed `already-armed`, and left
the board reading "not listening". The build now proves the session is live
from a fresh authoritative listing immediately before arming - not from the
establish call's status alone, which is already stale by then - reopens once
when it finds the session ended, and refuses rather than arming when it stays
ended. A reopen also replaces the pre-reopen source generation before
reporting success, so a runner on its way out cannot be mistaken for a
listener, and a board whose source is registered but unowned gets a
replacement started before the build returns.
Let a dead generation's ownership actually move. Reclaiming a claim ran its
capture-reservation cleanup first, and that cleanup re-verifies the recorded
state-root identity, so a claim naming a pid and a process group that were
both provably gone could not be cleared: reconcile reported a start while
nothing attached, and retire refused with "cannot release source ownership".
Reservation records are keyed by claim token and every replacement claims a
fresh one, so they are hygiene, not an ownership invariant. Reclamation now
additionally requires the owning process group to be absent independently,
which keeps a reused pid whose poll child still runs from ever reading as a
gone generation. A live owner and a crashed leader whose owned group survives
are still never reclaimed.
Stop carding decisions whose subject already landed. The build drops a
decision card whose work item or PR appears in the payload's own landed rows,
and one whose task is no longer an open captain call, naming each drop on
stderr. A task whose state cannot be established is kept, because a call
wrongly hidden is worse than a card wrongly shown.
Add the reconcile choice, and make it structurally incapable of closing a
call. Every decision card carries a standard `reconcile` option, injected by
the build rather than left to the composer. The board now emits the picked
option and any freeform note as separate structured fields instead of fusing
them, so a reconcile selection is not expressible as an answer value at all -
the defect that let `reconcile - <note>` reach the intake as an ordinary
answer. The adapter routes selections from that structured field, creation of
a reconcile request is bound to a verified board source rather than the shared
keyed-answer intake, and the intake still refuses the reserved value on every
channel. Each authorization is bound to the captain-hold generation that
produced the card, so an obsolete card cannot close a later call, and both
terminal outcomes require a pending request: `reconcile close` records the
evidence under its own `reconciled` mode so it never reads as the captain's
words, and `reconcile note` leaves the call open. Anything unprovable -
an unversioned row, a missing generation, an unreadable state - refuses
rather than acting.
Regression coverage fails without each fix, and pins every leak path: a bare
reconcile, a standalone close or note with no pending request, an any-channel
reconcile, an annotated selection from a freeform card, and a
generation-skewed authorization. An opt-in guard re-proves the lavish-axi
shapes and the reopen against the installed tool.
* fix(bin): quote the done comparison in the reconcile intake
shellcheck SC1010 reads the bare word as the loop keyword. The failed run
never reached its lint step, so this shipped in the recovered content.
* no-mistakes(review): Publish reconciled parent resolution before request retirement
* no-mistakes(review): Clarify committed cleanup and reconcile reservation scope
* no-mistakes(review): Preserve remote cards and legacy answer compatibility
* no-mistakes(test): Separate live claim relea…
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Every needs-decision event must skip the supervision branch and come straight to the main session - not only no-mistakes ask-user findings, but every needs-decision: product calls, destructive choices, captain holds, and second-mate escalations.
The accepted routing rule this change implements: when ONE COALESCED signal or stale trigger batch contains any needs-decision row, deliver the ENTIRE batch to main. Do not split that same batch between the supervision branch and a later main wake. This whole-batch routing is the replacement for an earlier per-row approach that split a mixed batch, reserving routine rows for the supervision branch while sending only the needs-decision rows to main; that earlier approach is superseded and must not be reintroduced.
Ordinary task-local heartbeat, signal, and stale routing stays unchanged, and the implementation stays narrowly limited to the needs-decision signal path.
What Changed
Risk Assessment
✅ Low: The routing changes are narrowly scoped, preserve heartbeat and routine task behavior, and the reviewed signal, stale, alias, cache, and symlink paths satisfy the stated invariants without a substantiated defect.
Testing
The successful baseline was supplemented with focused watcher, branch-dispatch, and Pi-extension tests covering needs-decision, captain-held, second-mate escalation, stale aliases, mixed batches, routine routing, and heartbeat independence. An executable integration artifact confirms a mixed batch is rejected by supervision and delivered intact to main; one unrelated stock-renderer compatibility case skipped because installed Pi 0.80.10 predates 0.84.4.
Evidence: Whole-batch needs-decision routing through the Pi watcher interface
Source: Whole-batch needs-decision routing through the Pi watcher interface
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
🔧 **Review** - 1 issue found → auto-fixed (9) ✅
.pi/extensions/lib/fm-branch-dispatch.ts:202- Intent requires “every needs-decision” and any stale trigger batch containing one to reach main, but stale classification only recognizes a finalcaptain-heldline. With a mapped task whose status endsneeds-decision: choose destructive cleanupand a queuedstalerow, this check does not populateneedsDecisionKeys;scope.eligibleremains true andofferWakeToBranchcan send the wake to supervision. The same bypass occurs when an earlier durable decision remains open behind a later unrelated status line. Classify stale rows using the durable open-decision semantics, while retaining captain-held detection, at this shared boundary.🔧 Fix: Route stale open decisions directly to main
2 issues (1 error, 1 warning) still open:
bin/fm-classify-lib.sh:1657- Intent explicitly requires second-mate escalations to reach main, butfm_pending_reply_maybe_escalatewrites them asblocked [key=pending-reply-…](bin/fm-pending-reply-lib.sh:1239), while the changed classifier sets the main-only marker only when the verb equalsneeds-decision. The resulting signal row retains asignal:payload, remains branch-eligible, andofferWakeToBranchcan deliver it to supervision. Extend the narrowly scoped decision signal classification to recognize this concrete second-mate escalation class and route its whole coalesced batch to main..pi/extensions/lib/fm-branch-dispatch.ts:151- The new stale-path fold hardcodesresolvedandcaptain-held, despite the authoritative status fold supportingFM_CLASSIFY_RESOLVE_VERBandFM_CLASSIFY_CAPTAIN_HELD_VERB. For example, with a configured resolution verb,needs-decision [key=x]followed by that resolution remains falsely open here and routes an ordinary stale wake to main; a configured captain-hold verb can instead be missed and offered to the branch. Use the configured verbs so stale routing agrees with the repository's durable fold semantics.🔧 Fix: Honor configured verbs in stale decision routing
2 errors still open:
bin/fm-classify-lib.sh:1657- Intent requires “every needs-decision” including “second-mate escalations” to reach main. However,fm_pending_reply_maybe_escalatewrites these asblocked [key=pending-reply-…](bin/fm-pending-reply-lib.sh:1239), while the changed classifier marks only theneeds-decisionverb. The resulting signal retains asignal:payload, remains branch-eligible, and can be delivered to supervision. Extend the decision-owned signal classification to recognize this concrete escalation class and route its coalesced batch wholly to main..pi/extensions/lib/fm-branch-dispatch.ts:155- The new stale decision fold hardcodespending-reply-as reserved instead of honoringFM_CLASSIFY_RESERVED_KEY_PREFIXES, unlike the authoritative shell fold. For example, withFM_CLASSIFY_RESERVED_KEY_PREFIXES=secret-,needs-decision [key=pending-reply-x]: choose cleanupis a valid open decision in the authoritative fold, but this code skips it and can offer its stale wake to supervision. Make this fold use the configured reserved-prefix semantics so every supported open decision routes to main and ordinary stale routing remains consistent.🔧 Fix: Route second-mate escalations and configured decisions to main
1 error still open:
.pi/extensions/lib/fm-branch-dispatch.ts:245- Intent requires captain holds to reach main, but status parsing removes only empty strings, not whitespace-only lines. Forcaptain-held [key=route]: awaiting captain\n \n,hasOpenNeedsDecisiontreats the hold as closing the decision and the final-line check examines" ", so the stale row remains branch-eligible. This contradicts the repository’slast_status_linesemantics, which ignore whitespace-only lines. Filter status lines by non-whitespace content before selecting the current declaration.🔧 Fix: Ignore trailing whitespace after captain holds
1 warning still open:
.pi/extensions/lib/fm-branch-dispatch.ts:244- Each queued stale row rereads and refolds its task's entire append-only status log synchronously. Multiple unread reminders for one task therefore multiply work on Pi's UI thread; for example, 50 stale rows for a 20 MB status log parse roughly 1 GB per eligibility check without changing the result. Cache the decision classification per status path withinscopeForUnreadWakeso each task is read and folded at most once per scan.🔧 Fix: Cache stale decision classification per status file
1 error still open:
.pi/extensions/fm-primary-pi-watch.ts:622- Intent requires ordinary signal routing to remain unchanged and whole-batch main routing only when that coalesced batch contains a needs-decision row. This key-only cross-reference also matches older unread rows. If task-a's needs-decision row remains queued and a later routine signal batch for task-a arrives,scope.needsDecisionKeysstill containstask-a.status, so the later routine batch is incorrectly forced to main without error. Distinguishing queue history from the triggering batch may require a batch identity or equivalent protocol change, so the remedy needs user authorization.🔧 Fix: Document unread decision precedence for later task wakes
1 warning still open:
.pi/extensions/lib/fm-branch-dispatch.ts:253- The per-scan cache still rereads each unread stale row's entire append-only status log on everyscopeForUnreadWakecall. Concretely, one unread stale decision for a task with a 20 MB status log makes every later routine signal or heartbeat synchronously read and fold that 20 MB file on Pi's UI thread, even when the file is unchanged. Cache classifications across scans using file identity/size/mtime invalidation, preserving fail-closed behavior when metadata or reads fail.🔧 Fix: Cache unchanged stale decisions across scope scans
2 errors still open:
.pi/extensions/fm-primary-pi-watch.ts:624- The documented precedence says an unread decision forces every later signal or stale trigger for the same task to main, but this compares raw queue keys rather than task identity. For task-a, an unread decision signal is keyedtask-a.status; after the status is resolved, a later stale trigger is keyedfm-window. The stale fold no longer marks the resolved status, the keys do not match, and the stale batch is offered to supervision. Preserve the approved same-task precedence across status-file and window aliases..pi/extensions/lib/fm-branch-dispatch.ts:157- The new stale classifier usesstatSyncandreadFileSync, both of which followtask.statussymlinks. A mapped stale row can therefore make the Pi extension synchronously read an arbitrary symlink target and let that external content control main-versus-branch routing. This violates the repository's authoritative status-fold boundary, which explicitly rejects status-file symlinks instatus_open_decisions. Reject symlinks withlstatSyncbefore versioning or reading, failing the scope closed.🔧 Fix: Resolve decision aliases and reject symlinked statuses
1 error still open:
bin/fm-watch.sh:1771- Intent requires every captain hold to reach main, but signal payloads are marked main-only only whenFM_SIGNAL_NEEDS_DECISION_FILEScontains the file.status_is_captain_relevantexcludescaptain-held, so a captain-held append surfaced through the documented no-verb fallback (for example after the crew ends without positive working evidence) is queued with an ordinarysignal:payload.scopeForUnreadWakethen treats that signal row as branch-eligible and the captain hold can enter supervision. Extend the main-only marker to surfaced captain-held signal files and cover that executable path.🔧 Fix: Route surfaced captain-held signals directly to main
✅ Re-checked - no issues remain.
✅ **Test** - passed
✅ No issues found.
bin/fm-test-run.sh --changed --exclude-family real-herdr-gatedBaseline:bin/fm-test-run.sh --changed --exclude-family real-herdr-gatedbin/fm-test-run.sh tests/fm-watch-triage.test.sh tests/fm-pi-branch-extension.test.sh tests/fm-pi-watch-extension.test.shManual executable Pi integration: queued one needs-decision signal and one routine signal in the same coalesced batch, captured the supervision offer and main-session wake, and verified whole-batch main routingVerified testing left the git worktree clean and removed.test-evidence-tmp✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.