fix(bin): harden PR check artifacts - #556
Merged
Merged
Conversation
kunchenguid
force-pushed
the
fm/fm-prcheck-security-patch-h5
branch
from
July 14, 2026 09:53
eb567bc to
29cb21a
Compare
kirangathani
added a commit
to kirangathani/firstmate
that referenced
this pull request
Aug 2, 2026
* feat: establish instruction ownership foundation (kunchenguid#619) * Add instruction owners foundation * no-mistakes(document): Refresh project-management owner pointers * fix: compress Firstmate contract and enforce delivery rigor ownership (kunchenguid#626) * docs: compress firstmate operating contract * docs: make delivery rigor single-owner PR B already removed personal and stacked review requirements, but it did not explicitly assign rigor to the selected delivery path or forbid risk-based manual clean gates. That gap still permitted the Hi Bit inversion. * no-mistakes(review): Honor configured merge authority across faster delivery paths * no-mistakes(document): Align docs with compressed operating contract * feat: add durable captain decision holds (kunchenguid#593) * Add durable captain decision holds * no-mistakes(review): Validate decision hold retries and origin paths * no-mistakes(review): Enforce durable decision lifecycle boundaries * no-mistakes(review): Harden decision display and partial retry recovery * no-mistakes(test): Update scout teardown fixtures for decision inventory * no-mistakes(document): Align decision lifecycle and scout teardown documentation * no-mistakes: apply CI fixes * no-mistakes(review): Reconcile terminal decision holds * no-mistakes(document): Align captain decision-hold documentation * fix(bin): harden PR check artifacts (kunchenguid#556) * fix: harden PR check artifacts * fix: close PR check migration gaps * fix: close PR check retry gaps * fix: clarify migration outcomes and ESM boundary * fix: keep failed migrations authoritative * test: use inert PR validation fixtures * no-mistakes(review): Reserve noncanonical PR quarantine namespace * no-mistakes(review): Prevalidate final PR-check teardown artifacts * no-mistakes(review): Preserve X metadata and validate teardown IDs * no-mistakes(review): Initialize migration state before watcher exclusion * no-mistakes(review): Isolate failed poll migrations from bootstrap recovery * no-mistakes(review): Allow safe polling during incomplete private repairs * no-mistakes(review): Authenticate watcher checks at execution time * no-mistakes(review): Preserve custom checks with hash-bound registration * no-mistakes(review): Clean custom check snapshots on watcher signals * no-mistakes(review): Stop watcher checks promptly on signals * no-mistakes(review): Terminate watcher check groups before cleanup * no-mistakes(document): Correct stale X-mode watcher documentation * fix: drain returned watcher check groups * no-mistakes(review): Harden quarantine links and recover validated replacement polls * no-mistakes(review): Preserve X mode across shim version transitions * no-mistakes(review): Refresh legacy X shims before marker short-circuits * no-mistakes(document): Correct persisted PR-check artifact documentation * no-mistakes(document): Correct stale PR-check documentation * fix: bind PR poll repair provenance * no-mistakes(review): Enforce single-link ownership for custom check artifacts * no-mistakes(review): Preserve private checks, X polling, and lifecycle IDs * no-mistakes(review): Separate task creation and legacy teardown validation * no-mistakes(review): Restore safe legacy operations and teardown validation * no-mistakes(review): Disambiguate migration obligations and preserve legacy retries * no-mistakes(review): Preserve fail-closed diagnostics and legacy quarantine evidence * no-mistakes(review): Reconcile legacy migration retries and teardown collisions * no-mistakes(review): Force legacy namespace reconciliation before marker short-circuits * no-mistakes(document): Document private poll artifact safety contracts * no-mistakes(lint): Suppress intentional literal-dollar lint finding * fix: migrate historical X poll identity * fix: harden PR check artifacts * no-mistakes(review): Preserve fail-closed diagnostics and legacy quarantine evidence * no-mistakes(review): Reconcile legacy migration retries and teardown collisions * fix: migrate historical X poll identity * no-mistakes(review): Harden X-mode artifact publication against symlink corruption * no-mistakes(review): Guard X artifact publication * no-mistakes(review): Enforce private X artifact reads * no-mistakes(test): Fix backend compatibility fixture dependencies * no-mistakes(document): Refresh PR-check documentation * no-mistakes(lint): Remove unused x-mode test locals * no-mistakes: apply CI fixes * fix(bin): compact session-start backlog digest (kunchenguid#636) * fix: compact session-start backlog digest * no-mistakes(test): Fix legacy backend fixture helper * no-mistakes(test): Fix watcher exit wait helper * no-mistakes(document): Document compact backlog digest * fix: dedupe stale watcher guard banners (kunchenguid#637) * fix: dedupe stale watcher guard banner * no-mistakes(review): Keep read-only guard state nonmutating * no-mistakes(document): Clarify stale watcher docs * fix(bin): balance bearings landed baseline (kunchenguid#640) * fix: balance bearings landed defaults * no-mistakes(document): Document balanced landed baseline * no-mistakes: apply CI fixes * fix: clarify captain-facing translation contract (kunchenguid#644) * docs: clarify captain-facing translation contract * no-mistakes(review): Restore runtime fallback mandate * no-mistakes(document): Align Bearings translation wording * fix(bin): make bootstrap output and nudges deterministic (kunchenguid#646) * fix: make bootstrap nudges deterministic * no-mistakes(review): Honor state override for bootstrap nudges * no-mistakes(review): Update benign bootstrap documentation labels * no-mistakes(review): Validate bootstrap nudge retry markers * no-mistakes(document): Align bootstrap nudge documentation * no-mistakes: apply CI fixes * docs(secondmate-provisioning): clarify concise registry ownership (kunchenguid#649) * Clarify concise secondmate registry contract * no-mistakes(review): Expand secondmate registry boilerplate coverage * no-mistakes(document): Point route docs to owner * fix(bin): strip quoted blocked_by values during decision hold resolve (kunchenguid#654) * fix(bin): strip quotes on blocked_by in decision-hold resolve tasks-axi quotes multi-entry blocked_by as "a,b,c", so the comma-boundary membership test only matched middle elements. Strip surrounding quotes before matching so first and last hold ids resolve correctly. * no-mistakes(document): Refresh decision-hold regression evidence * feat(secondmate): inherit shared captain preferences (kunchenguid#656) * feat(secondmate): inherit shared captain preferences * no-mistakes(review): Honor shared captain data overrides * no-mistakes(review): Honor bootstrap data override registry * no-mistakes(document): Refresh shared inheritance docs * no-mistakes(document): Clarify inherited local-material docs * feat: gate local agent secret injection (kunchenguid#658) * feat(spawn): gate local agent secret injection * fix(spawn): align final Keychain slot * no-mistakes: apply CI fixes * test: isolate Herdr autodetect smoke sessions (kunchenguid#662) * test: isolate herdr autodetect smoke session * no-mistakes(review): Restored autodetect smoke gate bypass * no-mistakes(test): Harden Herdr lab provisioning * no-mistakes(document): Refresh Herdr lab docs * docs: adopt under way for active work (kunchenguid#666) * Revert "feat: gate local agent secret injection (kunchenguid#658)" (kunchenguid#668) This reverts commit c27135c. * fix(pi): distinguish stale locks when arming watcher (kunchenguid#681) * fix(pi): distinguish stale locks when arming watcher * no-mistakes(test): Stabilize watcher extension async waits * no-mistakes(document): Document Pi lock recovery * fix: accept secondmate house vocabulary (kunchenguid#685) * fix: accept secondmate as house vocabulary * no-mistakes(test): Update captain vocabulary contract test * no-mistakes(document): Align secondmate documentation vocabulary * fix(bin): parse handoff homes after registry parentheticals (kunchenguid#686) * fix: parse secondmate home after pre-field parentheses Registry summaries often include parentheticals before the structured (home: ...) field. Match that field with a greedy prefix so handoff no longer reports "has no home" for those entries. * no-mistakes(document): Refresh handoff test comments * feat: add native session-start nudges (kunchenguid#687) * feat: add native session-start nudges * no-mistakes(document): Document nudge script inventory * docs: call built-in defaults the firstmate repo, not template (kunchenguid#688) Relabel absent-captain and related domain defaults wording so it names the firstmate repo rather than treating "template" as this domain's identity label. Keep the design-tenet "shared template" statements and unrelated launch/PR-poll template uses unchanged. * fix(bin): repair fm-brief.sh parse error and harden set -u array expansion (kunchenguid#205) * fix(bin): use set -u-safe empty-array expansion in pr-merge and spawn Expanding "${arr[@]}" on an empty array under set -u fails on bash < 4.4 (notably macOS bash 3.2). Quote the portable "${arr[@]+"${arr[@]}"}" idiom in fm-pr-merge and fm-spawn batch dispatch so empty arrays expand to nothing. Co-authored-by: Cursor <cursoragent@cursor.com> * test(brief): harden fm-brief regression coverage for parse and scaffolds Tighten bash -n checking, pin literal backtick rendering in the no-mistakes DOD wording assertion, and keep a scout/secondmate scaffold smoke test so the Co-authored-by: Cursor <cursoragent@cursor.com> kunchenguid#166 apostrophe regression cannot return unnoticed. --------- Co-authored-by: Cursor <cursoragent@cursor.com> * fix(bin): keep watcher supervision continuous across child cycles (kunchenguid#693) * fix: make watcher supervision continuous * no-mistakes(review): Bound watcher retries and log attached signals * no-mistakes(review): Add bounded successor-recovery wake fallbacks * no-mistakes(review): Prevent overlapping successor-arm retries * no-mistakes(review): Resume supervision after late arm closes * no-mistakes(review): Bind OpenCode recovery to attempted arm * no-mistakes(test): Synchronize peer beacon regression fixture * no-mistakes(test): Synchronize Pi and OpenCode late-close lifecycle fixtures * no-mistakes(document): Captain: document watcher successor protocol behavior * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * fix: fetch current PR head for review diffs (kunchenguid#722) * fix: always fetch PR head for review diffs Prefer a freshly fetched refs/pull/<n>/head over a reachable recorded pr_head= so reviewers never hold a merge over a "missing" fix that already landed on the remote PR. Recorded SHA is offline fallback only; local branch is last resort with a warning. Store the tip under refs/fm-review/ so a later base-branch fetch cannot clobber the compare tip via FETCH_HEAD. * no-mistakes(test): Isolate session-start nudge tests from gate state * no-mistakes(document): Correct review-diff documentation * docs: resolve five contract contradictions across AGENTS.md, README, and skills (kunchenguid#736) * docs: resolve five contract contradictions * no-mistakes(test): align owner-pointer assertions with reworded docs; skip absent shellcheck * docs(harness): correct Grok exit guidance (kunchenguid#742) * docs(harness): reverify grok exit command * no-mistakes(test): Correct Grok exit resume attribution * fix(watcher): bound stale wakes for parked crew (kunchenguid#743) * fix(watcher): bound stale wakes for exited paused crew * no-mistakes(review): Gate pause suppression on confirmed agent death * no-mistakes(test): Fixed stale pause cadence * no-mistakes(document): Document dead-agent hold cadence * fix(supervision): distinguish ordinary wakes from recovery (kunchenguid#744) * fix(supervision): distinguish ordinary wakes from repair * no-mistakes(review): Make passive guard follow-ups recovery-only * no-mistakes(document): Clarify recovery-only turn-end guard documentation * fix(x-mode): dedupe pending mention wakes (kunchenguid#745) * fix(x-mode): dedupe pending mention wakes * no-mistakes(review): fix x-poll claim error deduplication * no-mistakes(review): separate claim diagnostics from relay recovery * no-mistakes(document): Document X-mode once-only mention wakes * feat(wake): enrich drained signals with bounded status context (kunchenguid#747) * feat(wake): enrich drained signal context * no-mistakes(review): Bound wake enrichment reads * no-mistakes(document): Document wake-drain annotations * docs(wake): explain at-least-once drain boundary * no-mistakes(review): Prevent symlink races in wake annotations * no-mistakes(review): Exercise wake symlink race regression * test: document intentional AFK marker subprocesses * fix(wake): isolate annotation marker state * test: give merge-success security fixtures a gate-verifiable project fm-pr-merge.sh gates merges on fm-assert-tests-kept.sh, which resolves a real project repo and task worktree from the task meta. The upstream security fixtures predate the gate and used bare or missing paths, so every expected-successful merge was refused with 'could not verify'. Add write_gate_project and scope the legacy-loop gate worktree to the merge step so both teardown invocations still exercise the missing-worktree path. * test: mask node beyond fakebin so MISSING-node tests survive system node test_output_ordering_diagnostics_lead and test_composition_invokes_real_scripts forced a MISSING: node diagnostic only by deleting the fakebin node stub, which silently assumed node is absent from the fallback BASE_PATH (/usr/bin:/bin:/usr/sbin:/sbin). That holds on GitHub runners but breaks on any host where apt/nodesource installed /usr/bin/node: command -v node succeeds, bootstrap correctly reports nothing missing, and both assertions fail. Reuse the existing tmux masking idiom from test_herdr_backend_diagnostics_follow_real_session_start: a shared write_node_mask helper emits a BASH_ENV file overriding command -v node to fail regardless of PATH, and both tests pass it on their run_session_start call while keeping the fakebin removal. The assertions themselves are unchanged; only how node is made to look absent is. * fix(tests): pin fixture git identity in two tests that made bare commits tests/fm-pr-check-security.test.sh's write_gate_project and tests/fm-continuity-pretool-check.test.sh's primary-checkout fixture both ran bare git commit without fm_git_identity or an inline -c identity, so they passed on any machine with a global git config and failed exit 128 (empty ident) on a clean CI runner. Call fm_git_identity after sourcing lib.sh, matching the convention the other nineteen callers use. Audited all 86 test files for the same latent defect; these two were the only instances. Full suite passes under a neutralized-identity environment (GIT_CONFIG_GLOBAL=/dev/null, identity vars unset) and normally. --------- Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Co-authored-by: Johans Ballestar <47162770+Ballestar@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Israel Wilson <31997700+ICGNU3@users.noreply.github.com>
This was referenced Aug 3, 2026
levelupself
added a commit
to levelupself/firstmate
that referenced
this pull request
Aug 3, 2026
* fix: prevent no-mistakes gate agents from driving the fleet (#518)
* feat: contain no-mistakes gate agents from driving the fleet
Add bin/fm-gate-refuse-lib.sh, sourced at the top of fm-spawn/fm-send/
fm-teardown before any fleet mutation. It fails closed when NO_MISTAKES_GATE
is set, and via an unspoofable git-common-dir backstop when invoked from a
no-mistakes gate worktree (.no-mistakes/repos/*.git) even with the marker
unset. A normal firstmate session has neither signal and is unaffected.
Set disable_project_settings: true in the tracked .no-mistakes.yaml so the
installed pipeline neutralizes gate agents' project instructions for this repo
(trusted-only, honored from the default branch).
firstmate's own suite runs from a gate worktree during validation, so the
shared test helpers set FM_GATE_REFUSE_BYPASS=1 to exempt it; the dedicated
tests/fm-gate-refuse.test.sh strips it to verify real refusal.
* no-mistakes(review): Captain, refuse empty no-mistakes gate markers
* no-mistakes(document): Document no-mistakes gate authority boundary
* fix: guard secondmate primary sessions from blind turn ends (#505)
* fix: guard secondmate own-home turn ends
Remove the .fm-secondmate-home early-exit in fm-turnend-guard.sh so the
'no turn ends blind' backstop fires in a secondmate's own primary session,
matching the cd-guard's scope: the own home is guarded, child crew/scout
worktrees stay exempt via the retained git-dir/git-common-dir test. This
was pure scoping from the guard's primary-only origin and guarded against
no secondmate-specific hazard.
Add secondmate regression tests (blind-turn block, idle-by-default,
stop_hook_active loop guard, deferred-death recovery loop, child-worktree
exemption) and record the autonomous background-notify re-invoke
measurement (Claude Code 2.1.207, 11s) in docs/turnend-guard.md.
* no-mistakes(document): Correct secondmate guard documentation, captain
* fix: force-include marked secondmate homes in turn-end guard
The prior remove-only form (just deleting the .fm-secondmate-home check)
left the DEFAULT secondmate topology unguarded: a treehouse-leased home is
a linked git worktree (git-dir != git-common-dir), which the retained
git-dir exemption still skipped, so its own primary session could still end
a turn blind. Invert the marker: a genuinely-marked home is force-included
as a guarded primary (treehouse-leased linked OR git-cloned plain), and the
git-dir exemption applies only to UNMARKED child worktrees. Marker
validation (regular non-symlink file, non-empty id-token content) blocks a
stray or empty marker from spoofing inclusion.
Add real linked-worktree regression tests: a treehouse-leased LINKED
secondmate home is guarded, a stray/empty marker stays exempt, and the
unmarked child worktree stays exempt - the topology the plain git-init
fixtures masked. Predicates, in-flight gate, and loop guard untouched.
* fix: force ASCII collation in secondmate marker validation
Add a function-scoped local LC_ALL=C in fm_root_is_secondmate_home so the
[A-Za-z0-9._-] id allowlist matches under C collation, not the ambient
locale - a locale-crafted non-ASCII marker id can no longer slip through
the range match and spoof force-inclusion of a linked child worktree.
Add a regression test proving a non-ASCII marker id is rejected and the
linked worktree stays exempt.
* no-mistakes(test): fix backend baseline gate-refusal dependency
* no-mistakes(document): Correct secondmate turn-end guard documentation
* fix: make bootstrap diagnostics backend-aware (#519)
* fix: make bootstrap required-tool detection backend-aware
Bootstrap demanded tmux and treehouse for every backend except orca, so a
herdr/zellij/cmux home with tmux absent was wrongly told MISSING: tmux.
Required tools now follow the resolved backend via the single-owner
fm_backend_required_tools helper (bin/fm-backend.sh): each backend's own
session-provider CLI, jq for the JSON-emitting adapters (herdr/zellij/cmux),
and treehouse for session-provider-only backends (orca owns its worktree).
The treehouse lease-support check is gated to backends that use treehouse.
Adds install hints for herdr/zellij/cmux, regression tests for the full
backend dependency matrix (herdr-without-tmux repro plus each boundary),
and updates the authoritative Toolchain docs.
* no-mistakes(review): Captain, prevent executing Herdr install guidance
* no-mistakes(review): Captain, harden backend-aware bootstrap diagnostics
* no-mistakes(review): Captain, separate manual dependency remediation
* no-mistakes(review): Captain, align bootstrap diagnostic consumers
* no-mistakes(document): Align backend adapter dependency comments
* fix: preserve follow-up platform context after inbox cleanup (#520)
* fix: recover X/Discord follow-up platform after inbox cleanup
A milestone follow-up posted directly by request_id after the inbox was
drained - and with no task link, because one persistent secondmate's single
x_request slot collides across concurrent requests - resolved platform only
from the local inbox, so a >280 Discord reply silently defaulted to the X
280-char budget and threaded as (1/2).
- fm-x-poll records a durable per-request reply context
(state/x-context/<rid>.json) at stash time, keyed by request_id so
concurrent requests never overwrite each other; it survives inbox cleanup
and restart.
- fm-x-reply resolves platform/budget through registry -> inbox -> relay
(the relay lookup confined to a live follow-up), recovering the original
platform independent of task-link availability.
- Fail-safe: a follow-up whose platform/budget cannot be authoritatively
resolved and that would split is refused (exit 8) and held for retry,
never wrongly split; fm-x-followup keeps the link on that exit.
- fm-x-dismiss clears the durable context for a dismissed mention.
Refactors reply-context extraction into a single owner and adds regression
coverage for all four cases.
* no-mistakes(review): Captain, fail closed on incomplete follow-up context
* no-mistakes(review): Captain, bound X context registry retention
* no-mistakes(review): Captain, align context retention with answer binding
* no-mistakes(document): Align X follow-up context documentation
* no-mistakes(document): Align durable X follow-up documentation
* fix: preserve secondmate routing markers in terminal sends (#533)
* fix: preserve secondmate routing markers
* no-mistakes(review): Captain, preserve trailing newlines in marked secondmate sends
* no-mistakes(test): Captain, tolerate bootstrap timeout elapsed drift
* no-mistakes(document): Refresh Herdr marker documentation
* fix: align Grok effort handling with 0.2.99 (#527)
* fix: align grok effort docs and spawn with 0.2.99 ceiling
grok 0.2.99 accepts only low|medium|high for --reasoning-effort and
rejects both xhigh and max. Omit unsupported values on spawn, flag them
in crew-dispatch validation, and update harness-adapters.
* no-mistakes(test): Captain: refresh gotmp teardown fixture dependencies
* no-mistakes(document): Clarify Grok effort documentation ownership
* fix: derive bearings from authoritative secondmate state (#555)
* fix: make bearings use secondmate home state
* test: anonymize bearings fixtures
* no-mistakes(review): Bound parent activity evidence scans, captain
* no-mistakes(review): Preserve structured secondmate authority and bounds, captain
* no-mistakes(review): Preserve registry completeness and child inventory, captain
* no-mistakes(review): Reconcile parent evidence by verb and key, captain
* no-mistakes(review): Treat unkeyed parent evidence as inconclusive, captain
* no-mistakes(document): Document bearings local snapshot and PR opt-in
* no-mistakes(lint): Fix fleet snapshot ShellCheck findings
* no-mistakes: apply CI fixes
* fix: restore fleet snapshots on stock macOS Bash (#578)
* fix: restore stock macOS snapshot parsing
* no-mistakes(document): Clarify Linux gate and macOS CI coverage
* fix(afk): make Pi escalation and return catch-up reliable (#587)
* fix: close away-mode blocker supervision gap
* no-mistakes(review): Gate teardown retries and verify U+2063 dedupe
* no-mistakes(test): Fail closed on incomplete Pi composer separators
* no-mistakes(document): Document Pi composer recognition and return gating
* feat: support Pi max reasoning profiles (#537)
* support Pi max thinking profiles
* no-mistakes(review): Captain, allow Pi max dispatch profiles
* chore: no-mistakes(document): Clarify yolo response ownership (#595)
* Clarify validation response ownership
* no-mistakes(document): Clarify yolo response ownership
* feat: establish instruction ownership foundation (#619)
* Add instruction owners foundation
* no-mistakes(document): Refresh project-management owner pointers
* fix: compress Firstmate contract and enforce delivery rigor ownership (#626)
* docs: compress firstmate operating contract
* docs: make delivery rigor single-owner
PR B already removed personal and stacked review requirements, but it did not explicitly assign rigor to the selected delivery path or forbid risk-based manual clean gates. That gap still permitted the Hi Bit inversion.
* no-mistakes(review): Honor configured merge authority across faster delivery paths
* no-mistakes(document): Align docs with compressed operating contract
* feat: add durable captain decision holds (#593)
* Add durable captain decision holds
* no-mistakes(review): Validate decision hold retries and origin paths
* no-mistakes(review): Enforce durable decision lifecycle boundaries
* no-mistakes(review): Harden decision display and partial retry recovery
* no-mistakes(test): Update scout teardown fixtures for decision inventory
* no-mistakes(document): Align decision lifecycle and scout teardown documentation
* no-mistakes: apply CI fixes
* no-mistakes(review): Reconcile terminal decision holds
* no-mistakes(document): Align captain decision-hold documentation
* fix(bin): harden PR check artifacts (#556)
* fix: harden PR check artifacts
* fix: close PR check migration gaps
* fix: close PR check retry gaps
* fix: clarify migration outcomes and ESM boundary
* fix: keep failed migrations authoritative
* test: use inert PR validation fixtures
* no-mistakes(review): Reserve noncanonical PR quarantine namespace
* no-mistakes(review): Prevalidate final PR-check teardown artifacts
* no-mistakes(review): Preserve X metadata and validate teardown IDs
* no-mistakes(review): Initialize migration state before watcher exclusion
* no-mistakes(review): Isolate failed poll migrations from bootstrap recovery
* no-mistakes(review): Allow safe polling during incomplete private repairs
* no-mistakes(review): Authenticate watcher checks at execution time
* no-mistakes(review): Preserve custom checks with hash-bound registration
* no-mistakes(review): Clean custom check snapshots on watcher signals
* no-mistakes(review): Stop watcher checks promptly on signals
* no-mistakes(review): Terminate watcher check groups before cleanup
* no-mistakes(document): Correct stale X-mode watcher documentation
* fix: drain returned watcher check groups
* no-mistakes(review): Harden quarantine links and recover validated replacement polls
* no-mistakes(review): Preserve X mode across shim version transitions
* no-mistakes(review): Refresh legacy X shims before marker short-circuits
* no-mistakes(document): Correct persisted PR-check artifact documentation
* no-mistakes(document): Correct stale PR-check documentation
* fix: bind PR poll repair provenance
* no-mistakes(review): Enforce single-link ownership for custom check artifacts
* no-mistakes(review): Preserve private checks, X polling, and lifecycle IDs
* no-mistakes(review): Separate task creation and legacy teardown validation
* no-mistakes(review): Restore safe legacy operations and teardown validation
* no-mistakes(review): Disambiguate migration obligations and preserve legacy retries
* no-mistakes(review): Preserve fail-closed diagnostics and legacy quarantine evidence
* no-mistakes(review): Reconcile legacy migration retries and teardown collisions
* no-mistakes(review): Force legacy namespace reconciliation before marker short-circuits
* no-mistakes(document): Document private poll artifact safety contracts
* no-mistakes(lint): Suppress intentional literal-dollar lint finding
* fix: migrate historical X poll identity
* fix: harden PR check artifacts
* no-mistakes(review): Preserve fail-closed diagnostics and legacy quarantine evidence
* no-mistakes(review): Reconcile legacy migration retries and teardown collisions
* fix: migrate historical X poll identity
* no-mistakes(review): Harden X-mode artifact publication against symlink corruption
* no-mistakes(review): Guard X artifact publication
* no-mistakes(review): Enforce private X artifact reads
* no-mistakes(test): Fix backend compatibility fixture dependencies
* no-mistakes(document): Refresh PR-check documentation
* no-mistakes(lint): Remove unused x-mode test locals
* no-mistakes: apply CI fixes
* fix(bin): compact session-start backlog digest (#636)
* fix: compact session-start backlog digest
* no-mistakes(test): Fix legacy backend fixture helper
* no-mistakes(test): Fix watcher exit wait helper
* no-mistakes(document): Document compact backlog digest
* fix: dedupe stale watcher guard banners (#637)
* fix: dedupe stale watcher guard banner
* no-mistakes(review): Keep read-only guard state nonmutating
* no-mistakes(document): Clarify stale watcher docs
* fix(bin): balance bearings landed baseline (#640)
* fix: balance bearings landed defaults
* no-mistakes(document): Document balanced landed baseline
* no-mistakes: apply CI fixes
* fix: clarify captain-facing translation contract (#644)
* docs: clarify captain-facing translation contract
* no-mistakes(review): Restore runtime fallback mandate
* no-mistakes(document): Align Bearings translation wording
* fix(bin): make bootstrap output and nudges deterministic (#646)
* fix: make bootstrap nudges deterministic
* no-mistakes(review): Honor state override for bootstrap nudges
* no-mistakes(review): Update benign bootstrap documentation labels
* no-mistakes(review): Validate bootstrap nudge retry markers
* no-mistakes(document): Align bootstrap nudge documentation
* no-mistakes: apply CI fixes
* docs(secondmate-provisioning): clarify concise registry ownership (#649)
* Clarify concise secondmate registry contract
* no-mistakes(review): Expand secondmate registry boilerplate coverage
* no-mistakes(document): Point route docs to owner
* fix(bin): strip quoted blocked_by values during decision hold resolve (#654)
* fix(bin): strip quotes on blocked_by in decision-hold resolve
tasks-axi quotes multi-entry blocked_by as "a,b,c", so the comma-boundary
membership test only matched middle elements. Strip surrounding quotes
before matching so first and last hold ids resolve correctly.
* no-mistakes(document): Refresh decision-hold regression evidence
* feat(secondmate): inherit shared captain preferences (#656)
* feat(secondmate): inherit shared captain preferences
* no-mistakes(review): Honor shared captain data overrides
* no-mistakes(review): Honor bootstrap data override registry
* no-mistakes(document): Refresh shared inheritance docs
* no-mistakes(document): Clarify inherited local-material docs
* feat: gate local agent secret injection (#658)
* feat(spawn): gate local agent secret injection
* fix(spawn): align final Keychain slot
* no-mistakes: apply CI fixes
* test: isolate Herdr autodetect smoke sessions (#662)
* test: isolate herdr autodetect smoke session
* no-mistakes(review): Restored autodetect smoke gate bypass
* no-mistakes(test): Harden Herdr lab provisioning
* no-mistakes(document): Refresh Herdr lab docs
* docs: adopt under way for active work (#666)
* Revert "feat: gate local agent secret injection (#658)" (#668)
This reverts commit c27135cd9d35bc4c237d49b3b374da31fbd52eef.
* fix(pi): distinguish stale locks when arming watcher (#681)
* fix(pi): distinguish stale locks when arming watcher
* no-mistakes(test): Stabilize watcher extension async waits
* no-mistakes(document): Document Pi lock recovery
* fix: accept secondmate house vocabulary (#685)
* fix: accept secondmate as house vocabulary
* no-mistakes(test): Update captain vocabulary contract test
* no-mistakes(document): Align secondmate documentation vocabulary
* fix(bin): parse handoff homes after registry parentheticals (#686)
* fix: parse secondmate home after pre-field parentheses
Registry summaries often include parentheticals before the structured
(home: ...) field. Match that field with a greedy prefix so handoff
no longer reports "has no home" for those entries.
* no-mistakes(document): Refresh handoff test comments
* feat: add native session-start nudges (#687)
* feat: add native session-start nudges
* no-mistakes(document): Document nudge script inventory
* docs: call built-in defaults the firstmate repo, not template (#688)
Relabel absent-captain and related domain defaults wording so it names
the firstmate repo rather than treating "template" as this domain's
identity label. Keep the design-tenet "shared template" statements and
unrelated launch/PR-poll template uses unchanged.
* fix(bin): repair fm-brief.sh parse error and harden set -u array expansion (#205)
* fix(bin): use set -u-safe empty-array expansion in pr-merge and spawn
Expanding "${arr[@]}" on an empty array under set -u fails on bash < 4.4
(notably macOS bash 3.2). Quote the portable "${arr[@]+"${arr[@]}"}" idiom
in fm-pr-merge and fm-spawn batch dispatch so empty arrays expand to nothing.
Co-authored-by: Cursor <cursoragent@cursor.com>
* test(brief): harden fm-brief regression coverage for parse and scaffolds
Tighten bash -n checking, pin literal backtick rendering in the no-mistakes
DOD wording assertion, and keep a scout/secondmate scaffold smoke test so the
Co-authored-by: Cursor <cursoragent@cursor.com>
#166 apostrophe regression cannot return unnoticed.
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(bin): keep watcher supervision continuous across child cycles (#693)
* fix: make watcher supervision continuous
* no-mistakes(review): Bound watcher retries and log attached signals
* no-mistakes(review): Add bounded successor-recovery wake fallbacks
* no-mistakes(review): Prevent overlapping successor-arm retries
* no-mistakes(review): Resume supervision after late arm closes
* no-mistakes(review): Bind OpenCode recovery to attempted arm
* no-mistakes(test): Synchronize peer beacon regression fixture
* no-mistakes(test): Synchronize Pi and OpenCode late-close lifecycle fixtures
* no-mistakes(document): Captain: document watcher successor protocol behavior
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* fix: fetch current PR head for review diffs (#722)
* fix: always fetch PR head for review diffs
Prefer a freshly fetched refs/pull/<n>/head over a reachable recorded
pr_head= so reviewers never hold a merge over a "missing" fix that already
landed on the remote PR. Recorded SHA is offline fallback only; local branch
is last resort with a warning. Store the tip under refs/fm-review/ so a later
base-branch fetch cannot clobber the compare tip via FETCH_HEAD.
* no-mistakes(test): Isolate session-start nudge tests from gate state
* no-mistakes(document): Correct review-diff documentation
* docs: resolve five contract contradictions across AGENTS.md, README, and skills (#736)
* docs: resolve five contract contradictions
* no-mistakes(test): align owner-pointer assertions with reworded docs; skip absent shellcheck
* docs(harness): correct Grok exit guidance (#742)
* docs(harness): reverify grok exit command
* no-mistakes(test): Correct Grok exit resume attribution
* fix(watcher): bound stale wakes for parked crew (#743)
* fix(watcher): bound stale wakes for exited paused crew
* no-mistakes(review): Gate pause suppression on confirmed agent death
* no-mistakes(test): Fixed stale pause cadence
* no-mistakes(document): Document dead-agent hold cadence
* fix(supervision): distinguish ordinary wakes from recovery (#744)
* fix(supervision): distinguish ordinary wakes from repair
* no-mistakes(review): Make passive guard follow-ups recovery-only
* no-mistakes(document): Clarify recovery-only turn-end guard documentation
* fix(x-mode): dedupe pending mention wakes (#745)
* fix(x-mode): dedupe pending mention wakes
* no-mistakes(review): fix x-poll claim error deduplication
* no-mistakes(review): separate claim diagnostics from relay recovery
* no-mistakes(document): Document X-mode once-only mention wakes
* feat(wake): enrich drained signals with bounded status context (#747)
* feat(wake): enrich drained signal context
* no-mistakes(review): Bound wake enrichment reads
* no-mistakes(document): Document wake-drain annotations
* docs(wake): explain at-least-once drain boundary
* no-mistakes(review): Prevent symlink races in wake annotations
* no-mistakes(review): Exercise wake symlink race regression
* test: document intentional AFK marker subprocesses
* fix(wake): isolate annotation marker state
* feat(herdr): add optional presentation spaces (#784)
* feat(herdr): add optional presentation spaces
* no-mistakes(review): Harden Herdr projection creation and spawn serialization
* no-mistakes(review): Captain, disarm Herdr cleanup before launch submission
* no-mistakes(test): Correct stale Orca metadata failure fixture
* no-mistakes(document): Document Herdr presentation projection accurately
* fix(send): treat opencode busy-queued composer state as submitted (#775)
* fix(send): treat opencode busy-queued composer state as submitted
When fm-send sends a message to a BUSY opencode crewmate on the tmux
backend, opencode accepts the Enter and queues the message for the next
turn, but leaves the typed text visible in the composer row. The
submit-verification loop sees a pending composer, exhausts retries, and
reports a false "Enter swallowed" failure while the message is actually
delivered.
Fix: after Enter retries are exhausted and the composer still shows
pending, check fm_pane_is_busy. If the pane is busy (agent mid-turn,
footer shows "esc interrupt"), the harness queued the message, so
return "empty" (accepted). On an idle pane, keep returning "pending"
(genuine swallow detection preserved).
Regression tests cover four scenarios:
- busy pane + pending composer -> empty (message queued)
- idle pane + pending composer -> pending (genuine swallow)
- busy pane + composer clears on first Enter -> empty
- idle pane + composer clears on first Enter -> empty (existing path)
* docs: document busy-queued Enter exception across backend docs and skills
Add explanatory comments and backend documentation for the
busy-queued Enter fix (opencode 1.18.4 accepts Enter mid-turn
but keeps typed text in composer until the turn ends):
- bin/fm-tmux-lib.sh: document the busy-aware fallback in the
file header and above fm_tmux_submit_enter_core
- .agents/skills/afk/SKILL.md: daemon-facing policy note
- .agents/skills/harness-adapters/SKILL.md: harness-specific fact
- docs/tmux-backend.md: submit-acknowledgement section with the
busy-queue exception
- docs/herdr-backend.md: record the known gap
- docs/architecture.md: cross-reference in the daemon section
* test(tmux): fix SC2181 and make busy-submit test executable
* fix(spawn): require two stable reads before accepting worktree path (#765)
* fix(spawn): require two stable reads before accepting worktree path
The treehouse-get worktree-detection loop in fm-spawn.sh accepted the
first pane_current_path read that differed from the project path, but
on some tmux/WSL setups a brand-new window transiently reports a
stale-but-real path before the pane actually settles into the
worktree. Since that stale path is itself a real, distinct git
checkout, it also passes validate_spawn_worktree's isolation check,
so the loop silently recorded the wrong worktree in state/<id>.meta
(and, for claude harness spawns, installed the turn-end hook there
too).
Require two consecutive polls to agree on the same non-project path
before accepting it, using the existing inter-poll sleep as the
confirmation gap so an already-settled pane isn't slowed down by an
extra cycle.
* fix(tests): drop unused CASE_DIR read in worktree-settle test
ShellCheck SC2034: CASE_DIR is split out of the case record but
never referenced; discard it with _ instead.
---------
Co-authored-by: Freudator86 <tim@allesknut.de>
* fix(bin): make watcher process identity immune to Linux wall-clock changes (#752)
* fix(watcher): stabilize Linux process identity
* no-mistakes(document): document FM_PROC_ROOT_OVERRIDE and Linux starttime identity rationale
* fix: prevent AFK idle stalls and stale run attribution (#758)
* fix(supervision): verb-aware captain relevance, AFK wedge, head-bound state
Stop free-text tokens like "merged" from promoting nonterminal working: lines
to captain-relevant, so AFK no longer permanently suppresses idle recovery.
Defend wedge aging independently for nonterminal progress verbs, bind
no-mistakes current-state attribution to code identity (not branch alone),
and mark setup-complete as nonterminal in the ship brief scaffold.
* no-mistakes(review): Enforce nonterminal suppression and head-bound run attribution
* no-mistakes(document): Document current-code-bound run attribution
* no-mistakes(test): Wait for stable Herdr shell readiness
* no-mistakes(test): Make Herdr and watcher readiness tests deterministic
* no-mistakes(test): Make tmux capture and watcher lifecycle deterministic
* no-mistakes(document): Document corrected supervision contracts
* fix(bin): allow safe teardown during watcher recovery (#750)
* fix: allow safe teardown during watcher recovery
* no-mistakes(review): Distinguish unsafe-teardown deny guidance via policy reason code
* no-mistakes(document): Sync continuity-gate docs to allow teardown recovery
* test: mark dynamic teardown fixture literal
* no-mistakes(document): docs: add teardown to continuity gate allow list
* feat(herdr): order presentation spaces while preserving focus (#790)
* feat(herdr): order presentation worker spaces
* fix(herdr): preserve focus during projected cleanup
* no-mistakes(review): Serialize Herdr cleanup and protect active seeded tabs
* no-mistakes(review): Serialize Herdr aborts with guarded focus regressions
* no-mistakes(review): Fall back flat when Herdr serialization is unavailable
* no-mistakes(test): Stabilize watcher startup and AFK handoff tests
* no-mistakes(document): Correct Herdr ordering and focus documentation
* fix(bin): send literal config reread nudges after pushes (#809)
* Send literal config reread after inherited config push
When declared inherited config changes under an already-running secondmate,
build a per-home instruction from validated destination post-write bytes and
deliver it on the routed secondmate path. Unchanged config sends nothing;
ABSENT represents removal; captain-shared is never inlined. Covers mid-session
config-push and the locked bootstrap convergence path without hardening spawn
against deliberate runtime choice.
* no-mistakes(review): Fix config reread framing, partial propagation, and respawn order
* no-mistakes(review): Send config rereads via durable single-line pointers
* no-mistakes(review): Make failed config rereads retryable
* no-mistakes(review): Make config reread retries generation-safe
* no-mistakes(review): Make config rereads durable and ordered
* no-mistakes(review): Drain retries, bound history, preserve detect-only read-only mode
* no-mistakes(review): Retain write retries and quarantine stale respawn generations
* no-mistakes(review): Preserve exact config reread retries and delivery order
* no-mistakes(review): Preserve exact retry bytes and bounded quarantine pruning
* no-mistakes(document): Consolidated config-reread documentation
* feat(watch): follow GitLab merge requests to merge (#797)
* feat(watch): follow GitLab merge requests to merge
The merge watch only understood GitHub pull requests, so a task whose
deliverable is a GitLab merge request was never followed to merge.
Generalize the stored poll identity from owner/repository to a
provider-tagged provider/url/host/path/number record. GitLab runs mostly on
self-hosted instances and its projects nest under groups at no fixed depth,
so the host and the full project path are data in the record rather than
constants, and every consumer rebuilds the URL from those parts and refuses
any record that does not reconstruct it exactly.
The GitLab state is read with plain glab, matching the GitHub path's use of
plain gh, so an upstream checkout needs no extra tooling. Two things about
glab were established by running it rather than assumed, because a wrong
invocation here fails silently into a permanent "not merged":
- glab has no field selector, and its JSON would need a JSON processor that
firstmate does not require, so the state is read from glab's own field
output. Only an exact "merged" wakes firstmate, so a changed format stays
silent instead of reporting a merge.
- glab cannot take a merge request URL the way gh can, because that form
resolves through the current git repository and the watcher has none. It
is addressed by project URL and merge request number instead.
An absent glab produces no wake rather than a false merge, and arming
refuses with a clear message since that is the one point where a missing
CLI can still be reported. A GitLab task records no pr_head, which both
consumers already treat as optional. The merge path still addresses GitHub
only and refuses a merge request URL rather than sending it to the wrong
forge.
The record version moves to v2, and the existing non-executing migration
rebuilds an already-armed watch from its recorded URL, so no watch is lost
by upgrading.
docs/gitlab-merge-watch.md records the evidence, taken against the public
fixture project https://gitlab.com/KarotKris/gitlab-merge-watch-fixture.
* no-mistakes(review): Reject github.com host in GitLab MR URL/sidecar validation
* no-mistakes(document): Note GitLab MR URLs are explicitly refused, not just malformed ones, in fm-pr-merge.sh docs
* fix(herdr): group projected children beneath owning parents (#821)
* feat(herdr): correct all-home child presentation topology
Inherit the presentation opt-in to secondmate homes, label new projected
spaces with the approved corner format, insert each child under its owning
parent under one session-scoped lock, and keep flat non-destructive fallback.
* no-mistakes(review): Exclude secondmates from Herdr presentation projection
* no-mistakes(review): Harden shared Herdr locks and ambiguous child ordering
* no-mistakes(review): Use adjacency-only Herdr child ownership
* no-mistakes(review): Reject foreign legacy projections safely
* no-mistakes(review): Validate Herdr session sockets before projection
* no-mistakes(test): Fix Herdr teardown fixture session socket metadata
* fix(herdr): canonicalize presentation lock socket paths
Always resolve the session socket parent directory so symlink parents
such as /tmp -> /private/tmp cannot split the shared cross-home lock
identity. Refuse relative socket paths. Clarify lock-unavailable warnings.
* no-mistakes(test): Fix Bash-compatible GitLab merge request URL parsing
* no-mistakes(document): Document all-home Herdr child topology
* no-mistakes(lint): Quote fallback provenance string for ShellCheck
* fix: keep local no-mistakes tests intent-targeted (#823)
* fix(no-mistakes): drop full-suite local Test override
Local no-mistakes Test is intent-targeted; CI Behavior keeps the broad
tests/*.test.sh suite. Keep commands.lint on bin/fm-lint.sh and add a
focused contract test so the override cannot silently return.
* no-mistakes(lint): Make CI contract assertion ShellCheck-clean
* feat: add canonical timed test runner (#825)
* feat(test): add canonical timed suite runner and honest CI timeout
Introduce bin/fm-test-run.sh as the single serial owner for selecting
one script, a family, a conservative changed-file set, or the explicit
complete suite, with per-script timing markers and a JSON artifact.
Wire CI Behavior through the runner, raise the hang-tripwire timeout to
25 minutes, and document entry points without restoring a full-suite
local no-mistakes Test command.
* no-mistakes(review): Captain: fix changed selection and empty summaries
* no-mistakes(review): Captain: fail closed on unmapped changed sources
* no-mistakes(document): Document canonical timed test entry points
* fix: surface main inventory gaps in Bearings (#830)
* fix: disclose main-home orphan and unstructured inventory gaps
Main Bearings could report an empty fleet while structured in-flight rows
lacked meta or current backlog rows were free-form. Emit main_inventory from
the fleet snapshot, map it into Bearings omitted surfaces and a Charted Next
gate, and keep meta as the only live Underway source.
* no-mistakes(document): Document Bearings inventory-integrity projection
* no-mistakes: apply CI fixes
* feat: add bounded concurrent test isolation proof (#832)
* feat: add concurrent test isolation proof for Phase 2
Prove an audited portable candidate set passes under concurrent
workers with private mode-0700 temp roots, without enabling
production CI sharding or fm-test-run --jobs.
* no-mistakes(review): Pin isolation proof to audited candidate manifest
* feat: guard against missed secondmate reports (#834)
* feat(secondmate): parent-owned guards for missed status reports
Marked parent-to-secondmate requests now create a durable pending-reply
expectation with a privacy-safe correlation id before delivery. Transport
success never resolves it; only a correlated parent status or document
pointer does. After a completed turn with no report, the parent sends one
recovery repost and escalates once if that turn is also missed, without
scraping the secondmate conversation or looping.
* no-mistakes(review): Deduplicate wrong-home pending-reply sightings
* no-mistakes(review): Harden pending-reply recovery and escalation guards
* no-mistakes(review): Bound pending-reply backend polling
* no-mistakes(review): Cache pending-reply status scans
* no-mistakes(review): Protect undelivered pending-reply records from scans
* no-mistakes(review): Close pending-reply delivery durability gaps
* no-mistakes(review): Separate pending-reply transport outcomes
* no-mistakes(review): Escalate stalled pending-reply deliveries once
* no-mistakes(review): Resolve attempted deliveries from correlated reports
* no-mistakes(review): Resolve late reports after delivery escalation
* no-mistakes(document): Document pending-reply grace and ownership
* no-mistakes(lint): Silence intentional pending-reply test fixture lint warnings
* feat: require pinned real-Herdr CI coverage (#838)
* feat: add required pinned Herdr CI lane
Install exact Herdr 0.7.4 and Treehouse 2.0.1 with official assets and
SHA-256 pins, run the real-herdr-gated family serially through
fm-test-run with hard-fail on herdr-not-found, and keep portable
Behavior free of claimed Herdr coverage.
* no-mistakes(document): Consolidate real-Herdr CI documentation ownership
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* feat: shard portable tests and add bounded local parallelism (#841)
* feat: shard portable CI tests after isolation proof
Balance the Phase 2 proven-isolated set into two LPT portable parallel
lanes from Phase 1 timing evidence, keep stateful work in a required
portable serial lane, exclude real Herdr to its dedicated required lane,
and prove complete inventory coverage with a deterministic guard.
Add bounded local --jobs only for the proven set, per-lane timing plus
aggregate artifacts, and reduce the interim portable hang tripwire now
that the serial remainder owns the long wall-clock path.
* no-mistakes(review): Captain, fix CI contracts and completion-order worker scheduling
* no-mistakes(review): Captain, preserve stderr gate-skip detection in parallel tests
* no-mistakes(document): Document portable sharding and timing aggregation
* no-mistakes: apply CI fixes
* fix: block primary-session delegation outside the fleet (#854)
* feat: fence primary-session delegation outside the fleet
A firstmate primary that delegates through Claude Code's built-in
delegation tools creates work with no state/<id>.meta. Because
fm-supervision-lib.sh counts *.meta and fm-turnend-guard.sh exits
silently at zero, such work does not merely go unsupervised: it makes
the whole guard stack structurally inert, and it dies with the primary
session. On 2026-07-22 that cost two workers mid-flight and left
supervision down for 73 minutes unnoticed.
Layer 1, the primary fix: a permissions.deny list in
.claude/settings.json removes the 18 delegation, scheduling, worktree,
and task-tracking tools from the model's schema, so they are never
offered. This is removal rather than interception, so there is no call
to intercept and no fail-open path. The list is flat and in one file so
its width stays reviewable; the captain owns that width.
Layer 2, bin/fm-subagent-pretool-check.sh: a deny list is fail-open
against tools that do not exist yet, and permissions.allow is a
pre-approval list rather than an availability list, so there is no
fail-closed allowlist to use instead. This backstop classifies the tool
NAME by shape rather than against a fixed list, so a delegation tool
that ships before the deny list is updated is still refused. It excludes
mcp__* names and observe-or-stop operations, scopes itself to a genuine
primary home via the shared fm_primary_scope_matches predicate so a
crewmate's task worktree is unaffected, and offers one deliberate
FM_ALLOW_SUBAGENT=1 escape hatch that must be set at launch.
Verified live against Claude Code 2.1.217, including a deny-key A/B with
a nonsense-name control, layer 2 denying an un-denied Workflow call, the
same call allowed in a linked worktree, and the escape hatch. Corrects a
prior finding: both Task and Agent work as deny keys, so both are
pinned. Codex 0.144.1 verified to expose no delegation tool; grok,
opencode, and pi are inspected and documented as not wired because those
binaries are absent from this host and the repo requires live validation
before trusting a harness hook. Evidence in docs/subagent-guard.md.
* no-mistakes(review): Ship scoped Claude delegation guard
* no-mistakes(test): Ship Claude delegation deny list
* no-mistakes(document): Clarify PreToolUse guard ownership
* no-mistakes(lint): Keep Claude deny list local
* fix: install tasks-axi in portable CI shards (#866)
Reproduction: portable-parallel-2 completed successfully without tasks-axi while fm-decision-hold-lifecycle emitted a gate skip in 30 ms. The pre-shard lane installed tasks-axi and exercised the test fully. Installing tasks-axi is the smallest counterfactual and makes the representative shard execute the test with gate_skip=false in about 20 seconds. Both parallel jobs receive symmetric setup, while the exact 91-test inventory and coverage guard remain unchanged.
* feat(bin): make dispatch profiles quota aware (#867)
* feat: make dispatch profiles quota aware
* no-mistakes(review): Fix quota window and Grok product scoping
* no-mistakes(document): Document implicit quota-aware dispatch accurately
* Add built-in ahoy recap skill (#873)
* fix: preserve trustworthy Bearings data in partial snapshots (#875)
* fix: preserve mixed Bearings projections
* no-mistakes(review): Enforce strict invalidity precedence for partial snapshots
* no-mistakes(review): Enforce ownership for unknown child metadata
* no-mistakes(document): Document partial structured Bearings projections
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* feat(pi): add session-local calm mode (#884)
* Add session-local Pi calm mode
* no-mistakes(review): Preserve Pi HTML exports during calm mode
* no-mistakes(review): Preserve calm exports across submit bindings and share
* no-mistakes(document): Document calm-mode feasibility across supported harnesses
* fix(pi): prevent redundant watcher re-arms (#885)
* fix(pi): limit watcher arm tool to recovery
* no-mistakes(review): Strengthen Pi live re-arm regression coverage
* no-mistakes(document): Document Pi first-cycle and recovery-only watcher arming
* fix(pi): clean up Calm transcript rendering (#895)
* fix(pi): clean up Calm transcript rendering
* no-mistakes(review): Captain, preserve Calm exports and classify Pi launch briefs
* no-mistakes(review): Captain, eliminate Calm gaps and verify exported conversations
* no-mistakes(review): Restore Calm rows received while active
* no-mistakes(review): Preserve diagnostics during Calm restoration
* no-mistakes(document): Clarify Calm transcript behavior and injection paths
* fix: execute every PR body compliance event (#898)
* fix: execute every PR body compliance event
* no-mistakes(document): Document independent PR compliance events
* fix: exclude operational injections from ahoy boundaries (#899)
* fix: distinguish operational input in ahoy
* no-mistakes(review): Handle legacy Ahoy operational boundaries
* no-mistakes(review): Narrow legacy Ahoy boundaries with live regressions
* no-mistakes(document): Document Ahoy operational marker ownership
* no-mistakes(lint): Suppress intentional literal fixture lint warnings
* fix: canonically classify operational inputs across harnesses (#909)
* fix: type canonical operational inputs
* no-mistakes(document): Correct canonical operational-input documentation ownership
* fix: avoid generic secondmate start acknowledgements (#926)
* fix: avoid generic secondmate acknowledgements
* no-mistakes(document): Document sparse secondmate acknowledgement behavior
* no-mistakes: apply CI fixes
* fix(pi): make calm mode persistent and gapless (#927)
* fix(pi): preserve calm presentation across sessions
* no-mistakes(review): Fix Calm home fallback persistence
* no-mistakes(document): Clarify Calm gapless and export contracts
* no-mistakes: apply CI fixes
* fix(watch): retire merged PR polls after durable notification (#932)
* fix: retire merged PR polls after notification
* no-mistakes(review): Decouple PR retirement recovery from template updates
* no-mistakes(review): Recover pending PR retirements before poll migration
* no-mistakes(document): Document merged PR poll retirement contracts
* fix: refine scout intake and parallel dispatch (#934)
* Clarify intake evidence and overlap handling
* no-mistakes(review): Align scout guard with intake classification
* no-mistakes(document): Clarify scout documentation and intake ownership
* fix(pi): prevent duplicate assistant replies in Calm (#936)
* fix(pi): preserve operational follow-up semantics in Calm
* no-mistakes(document): Correct Calm operational-row visibility documentation
* perf(bin): shrink the ShellCheck source graph (#939)
* perf(lint): shrink shell source graph
* no-mistakes(review): Ensure lint workers terminate fully on cancellation
* no-mistakes(document): Repair stale lint documentation ownership
* fix(pi): remove Calm hidden-block gaps (#942)
* fix(pi): remove Calm hidden-block gaps
* no-mistakes(review): Validate Calm geometry against current viewport
* no-mistakes(review): Synchronize Calm geometry checks with reload completion
* fix: enforce contract boundaries for ask-user findings (#945)
* fix: escalate ask-user contract expansion
* no-mistakes(document): Point project management to authority owner
* docs: prefer direct operational paths (#946)
* fix(pi): hide operational user rows in Calm mode (#948)
* fix(pi): hide Calm operational user rows
* no-mistakes(review): Narrow Calm operational input suppression
* no-mistakes(review): Avoid Calm replay classifier subprocesses
* no-mistakes(document): docs: point Pi verification to Calm owner
* fix: relaunch missing second mates at session start (#950)
* fix(session-start): relaunch missing second mates
* fix(test): detect completed parallel workers
* no-mistakes(review): Isolate session-start recovery test cleanup
* no-mistakes(review): Complete backend-safe secondmate session recovery
* no-mistakes(review): Resolve Zellij task ownership before recovery
* no-mistakes(review): Recover relocated Zellij ghost tabs safely
* no-mistakes(review): Restore conservative Zellij recovery boundary
* no-mistakes(review): Reject malformed tmux recovery targets
* no-mistakes(document): Align secondmate recovery documentation
* no-mistakes: apply CI fixes
* fix(herdr): reclaim resumed task projections after restart (#967)
* fix(herdr): reclaim resumed task projections safely
* no-mistakes(review): Enforce safe Herdr reclaim close boundaries
* no-mistakes(document): docs: clarify Herdr restart projection contract
* Teach Ahoy to surface open decisions (#968)
* Require shipshape routine acknowledgement (#969)
* docs: separate current guidance from verification evidence (#994)
* docs: separate current guides from verification
* no-mistakes(review): Restore Herdr 0.7.5 restart-reclaim verification evidence
* fix: preserve Claude watcher continuity across Stop hooks (#997)
* feat(claude): Stop-owned tokenless watcher continuity via asyncRewake auto-arm
Claude primaries (main home and marked secondmate homes) no longer depend
on the model remembering to re-arm the watcher after each wake. A tracked
Stop asyncRewake hook (bin/fm-claude-stop-autoarm.sh, timeout 28800s)
fires on every turn end, claims one home-scoped single-flight owner,
foregrounds bin/fm-watch-arm.sh inside the hook-owned process tree, and
translates an actionable close or typed watcher failure into exactly one
exit-2 rewake. The hook scopes to genuine primary checkouts, requires the
session lock to be held by its own harness ancestor, stays inert while
AFK owns triage or the home is idle, and hands AFK transitions mid-cycle
to the daemon without rewaking.
The synchronous turn-end guard gains a --claude cooperative mode: it
ignores stop_hook_active (true on every post-continuation stop, which is
what re-opened the 2026-07-21 blind window), waits briefly for a watcher
health proof, a live auto-arm owner claim, or a fresh rewake epoch, and
re-blocks only when the auto-arm genuinely failed to establish - bounded
to 3 consecutive blocks per session, safely below Claude Code's 8-block
override, then a degraded allow with a visible systemMessage. Codex
keeps the previous one-block loop guard byte-identically, and Pi,
OpenCode, and Grok adapters are untouched.
Continuity PreToolUse gate and durable wake queue are preserved; the
gate's recovery guidance now names the Stop-owned re-arm and reserves
manual background arms for auto-arm failure. Claude supervision protocol,
harness-adapters facts, architecture, configuration, and continuity docs
updated; docs/turnend-guard.md records the 2026-07-24 Claude 2.1.218
contract revalidation (tokenless multi-cycle rewake, no-dedup, timeout
process-group kill, 8-block cap, interactive non-stall) and the 2.1.219
product live E2Es.
Regression matrix: hermetic tests cover scope, identity, AFK, need,
single-flight, translation, guard cooperation, budget, and registration;
the new live E2E proves two full tokenless auto-arm rewake cycles with
zero model arm commands; Pi and OpenCode Option B live E2Es pass
unchanged.
* no-mistakes(review): Fix Claude X-mode auto-arm continuity backstop
* no-mistakes(review): Remove unsupported Claude contract-lab verification claims
* no-mistakes(document): Update Claude auto-arm continuity documentation
* fix(herdr): clean stale projections at session start (#996)
* Clean stale Herdr projections at session start
* no-mistakes(document): Document stale Herdr session-start projection cleanup
* no-mistakes(review): Enforce locked exact Herdr projection cleanup
* no-mistakes(review): Fail closed on unverified session lock ownership
* no-mistakes(review): Serialize session lock acquisition atomically
* no-mistakes(document): Align session-start and Herdr cleanup documentation
* no-mistakes(document): Generalize lock-refusal diagnostics
* no-mistakes(lint): Avoid reserved keyword in concurrency test
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* fix: recover Claude supervision without watcher-status gate (#1001)
* fix: recover Claude supervision at session start
* fix: remove Claude watcher-status command gate
* no-mistakes(document): docs: remove stale continuity gate references
* fix: make quota-aware profile selection agent-owned (#1018)
* Replace quota dispatch selector instructions
* no-mistakes(review): Align bootstrap docs with agent-owned dispatch selection
* fix(bin): remove vestigial dispatch selector (#1026)
* remove vestigial dispatch selector
* no-mistakes(review): Synchronize isolation proof and portable shard evidence
* no-mistakes(review): Correct shard history and proof archive date
* no-mistakes(review): Remove reintroduced selector documentation reference
* no-mistakes(document): Remove stale dispatch strategy documentation
* docs(agents): drop superseded interim quota-window rule (#1039)
quota-axi 0.1.13 emits schemaVersion 2 with a quotaSemantics object per
provider, so the successor named in the interim rule has landed and the
rule's own removal condition is satisfied.
Keep the ownership clause so quota-axi remains the single owner of how
model or product windows relate to bounding account windows, and drop
the interim weakest-headroom instruction. The unknown-semantics case is
already covered by the existing requirement to stop and report a
candidate whose applicable quota data or interpretation cannot be
established.
Drop the matching assertion phrase from
tests/fm-instruction-owners.test.sh; the retained ownership phrase still
asserts.
* fix(tmux): scope busy detection and recognize current Claude turns (#1049)
* fix(tmux): scope Claude busy detection by harness
* no-mistakes(review): Separate verified and fallback busy signatures
* no-mistakes(test): Scope busy signatures to supplied harnesses
* no-mistakes(document): Document harness-scoped busy detection
* feat: add verified Kimi crewmate adapter (#1047)
* Add verified Kimi crewmate harness adapter
* no-mistakes(review): Scope Kimi moon detection to spinner lines
* no-mistakes(review): Match only complete Kimi spinner rows
* no-mistakes(review): Resolve Kimi binary portably before pane creation
* no-mistakes(document): Align Kimi adapter documentation
* no-mistakes(lint): Suppress false-positive ShellCheck warning for sourced watcher override
* Fix Kimi busy spinner detection
* no-mistakes(review): Recognize Kimi session-lock ancestry and holders
* no-mistakes(review): Scope pending-reply Kimi busy detection by harness
* no-mistakes(document): Correct Kimi spinner capture documentation
* no-mistakes(document): Clarify optional Kimi spinner whitespace
* no-mistakes(lint): Silence intentional pending-reply test stub warnings
* test: align rebased Kimi busy fixtures
* no-mistakes: apply CI fixes
* Reconcile Kimi busy detection after per-harness scoping
* no-mistakes(review): Clarify observed Kimi spinner whitespace contract
* no-mistakes(document): Clarify Kimi harness documentation
* fix: harden Kimi submission and spinner matching (#1058)
* fix kimi pointer submission and spinner conformance
* no-mistakes(review): Preserve Kimi submit target ownership guard
* feat(bin): add guarded Kimi turn-end wake (#1059)
* Add guarded Kimi turn-end hook
* no-mistakes(review): Require jq before installing Kimi turn-end hook
* no-mistakes(review): Expose jq inside isolated Kimi test fixtures
* no-mistakes(review): Preserve Kimi config boundaries during hook removal
* no-mistakes(review): Document Kimi removal newline safeguard
* no-mistakes(document): Document Kimi shared-home preservation
* fix(tmux): classify bordered composers across all rows (#1066)
* Fix structural tmux composer reading
* Verify Calm compatibility with Pi 0.82
* no-mistakes(review): Harden structural composer classification boundaries
* no-mistakes(review): Refresh composer and Kimi regression fixtures
* no-mistakes(review): Fail closed on unbounded composer edges
* no-mistakes(review): Enforce aligned composer geometry safely
* no-mistakes(review): Make composer ambiguity locale-safe
* no-mistakes(review): Preserve ambiguity through composer submission
* no-mistakes(review): Carry composer proof through retries
* no-mistakes(document): Document structural tmux composer delivery guarantees
* no-mistakes: apply CI fixes
* feat(bin): add verified pi-signed runtime adapter (#1145)
* feat: add verified pi-signed adapter
* no-mistakes(review): Correct pi-signed maintainer verification date
* no-mistakes(review): Correct remaining pi-signed verification dates
* no-mistakes(review): Preserve authoritative pi-signed runtime identity
* no-mistakes(document): Document pi-signed shared adapter semantics
* no-mistakes: apply CI fixes
* fix(pi): rearm watcher across session transitions (#1166)
* fix(pi): rearm watcher across same-process session transitions
Pi emits session_shutdown for ordinary /new, /resume, and /fork replacement
as well as terminal quit. The primary watcher extension latched a module-level
stopping flag on every shutdown, so a replacement session in the same process
could not arm monitoring until Pi restarted.
Own arm authority per session generation so only the active live generation
may start, stop, or rearm the child. Replacement sessions can arm again without
restarting Pi, stale prior-generation callbacks cannot mutate the active cycle,
and real quit still blocks late rearm.
* no-mistakes(review): Preserve Pi generation isolation and exit cleanup
* no-mistakes(document): Correct Pi watcher transition documentation
* feat: route crew dispatch using quota-window pace (#1172)
* Consume quota-axi pace signals in dispatch profile array selection.
Add quota-array-dispatch as the single owner of the pace-aware candidate
choice, keep AGENTS.md to the intake boundary and load trigger, and cover
the acceptance cases with sanitized schemaVersion 3 fixtures.
* no-mistakes(review): Stop and report genuine quota dispatch ties
* no-mistakes(document): Document quota pace freshness and uncertainty
* fix: adapt Grok Stop continuation and harden endpoint cleanup (#1171)
* fix(grok): adapt Stop continuation to runtime capability
* no-mistakes(review): Reject ambiguous Grok Stop payloads
* no-mistakes(review): Reject duplicate Grok fields and accept spaced tmux sessions
* no-mistakes(review): Enforce exact tmux cleanup selectors
* no-mistakes(test): Fix historical tmux fixture and validate Grok Stop
* no-mistakes: apply CI fixes
* fix: restore stock macOS Bash 3.2 brief scaffolding (#1093)
* fix(brief): make DOD scaffolding parse-safe on stock macOS Bash 3.2
fm-brief.sh built each Definition-of-done block and the not-enabled
Herdr declaration with `VAR=$(cat <<EOF ... EOF)`. On Bash 3.2 (macOS
/bin/bash) the lexer scans for the command substitution's closing `)`
textually and tracks quote state through the heredoc body, so a single
apostrophe, unbalanced quote, or unbalanced paren in that prose breaks
parsing of the whole script. Every ship-brief scaffold (no-mistakes,
direct-PR, local-only) failed with `unexpected EOF while looking for
matching )`. Bash 4+ parses it fine, so the breakage stayed invisible
everywhere except stock macOS.
Replace all four command-substitution heredocs with
`IFS= read -r -d '' VAR <<EOF || true`. That removes the `$(...)`
wrapper and the entire defect class regardless of future prose, and
preserves the variable expansion the direct-PR and local-only bodies
need. `read` keeps the heredoc's trailing newline that `$(...)` used to
strip, so trim one newline to keep every generated brief byte-identical
to prior output.
Guard the structure, not one historical phrase: a new test rejects any
heredoc nested in a command substitution anywhere in fm-brief.sh, where
the old assertion pinned a single apostrophe phrase and so missed the
reintroduction. Extend the stock-macOS Bash CI job from parsing one
script to the whole maintained shell surface (bin/*.sh,
bin/backends/*.sh, tests/*.sh), matching bin/fm-lint.sh's canonical file
set so parse scope and lint scope cannot drift apart.
* no-mistakes(review): Captain: harden Bash structure and inventory guards
* no-mistakes(document): Align stock macOS Bash contributor checks
* no-mistakes(lint): Suppress deliberate SC2016 literal fixture warnings
* test: stabilize tmux teardown conformance baseline (#1209)
* fix(test): pin teardown tmux baseline to historical kill selectors
merge-base HEAD main collapses to HEAD after the exact-selector change
lands on the default branch, so the old teardown fixture was accidentally
exercising current exact targets. Resolve a content-historical permissive
tmux adapter from first-parent history and force that post-squash topology
inside the conformance case so main and feature branches keep the same
old-vs-new contract.
* no-mistakes(lint): Suppress intentional literal-pattern ShellCheck warnings
* docs: slim quota-array-dispatch to the pace selection core (#1197)
Cut the runtime skill to the compact pace-aware selection procedure plus
minimum owner pointers. Keep every distinct decision rule and move expanded
acceptance scenarios to deterministic fixture ownership assertions.
Size: 170/1374/10187 -> 63/544/4068 (about 63%/60%/60% reduction).
* feat(bin): inherit backend config into secondmate homes (#1219)
* Inherit config/backend into secondmate homes with deliberate-override preservation
Add backend to the shared inheritable config allowlist so launch, locked
bootstrap, and config-push converge a primary pin into secondmate homes as each
home local future-spawn default. Track last-inherited bytes in a private state
provenance marker so deliberate per-home overrides survive present and absent
primary convergence, keep --backend and FM_BACKEND stronger, and extend the
existing inheritance tests plus docs and skill claims.
* no-mistakes(review): Preserve equal unprovenanced backend overrides
* no-mistakes(review): Preserve symlink overrides and verify spawn precedence
* no-mistakes(review): Snapshot backend inheritance for consistent provenance
* no-mistakes(review): Simplify backend inheritance to primary-authoritative convergence
* no-mistakes(document): Document inherited backend override preservation
* fix: restore primary-authoritative backend inheritance after document regression
The document step reintroduced provenance and deliberate per-home override
semantics after review had simplified config/backend to plain primary-authoritative
allowlist membership. Restore the primary-always-wins path: present overwrites,
absent removes, no provenance marker, and docs/tests match that contract.
* no-mistakes(review): Add divergent backend precedence regression fixtures
* no-mistakes(document): Document backend inheritance contract
* fix(pi): remove Calm's upper version ceiling (#1226)
* fix(pi): remove Calm's exclusive Pi upper-version ceiling
tests/fm-calm-pi-extension.test.sh gated on a closed PI_COMPAT_VERSIONS
allowlist ("0.81.1 0.82.0") that refused any other installed Pi, and docs
described that range as "supported" rather than verified evidence. The
Calm CHANGELOG shows no API introduced at either version, so there is no
evidence for a real minimum; the presentation adapters already probe the
exact method they patch rather than checking a version.
Replace the allowlist with dated version evidence that never rejects a
newer Pi, and make each presentation adapter degrade independently with
a diagnostic if a future Pi removes its API, instead of the whole Calm
extension failing to load. Rewrite the feasibility doc's "Pi 0.81.1
through 0.82.0" phrasing to state it as verified evidence, not a
ceiling.
* no-mistakes(review): Probe missing Calm adapter exports safely
* no-mistakes(document): Document Calm's unbounded Pi compatibility
* fix(bin): allow session-local todo tools in the subagent guard (#1204)
* fix(guard): allow session-local todo tools in the primary
The delegation-shape guard denied TaskCreate and TaskUpdate because their
normalized names contain the `task` stem. Those tools write only the harness's
session-local todo list, which has no executor: it spawns no agent, allocates
no worktree, registers no schedule, and starts nothing that outlives the
session. That is not the unaccounted work the guard exists to stop, so the stem
match was a false positive, and the deny text told the primary to run
bin/fm-brief.sh and bin/fm-spawn.sh to create a todo entry.
Add a separately-reasoned PLAN_ONLY_TOOLS exact-name exclusion rather than
widening OBSERVE_ONLY_TOOLS, whose documented contract is tools that only
observe or stop existing work. Both lists stay exact-name so neither can widen
by substring.
Tests cover the two allowed names and six near-miss names that a substring or
shortened-stem widening would release; both mutations were watched red.
* no-mistakes(review): drop session-local todo tools from recommended deny list
* no-mistakes: apply CI fixes
* fix(session-lock): resolve Claude bg-spare ancestry to the outermost claude pid (#1206)
* fix(session-lock): resolve Claude bg-spare ancestry to the outermost claude pid
fm_harness_ancestry_pid() previously returned the first ancestor process
whose command matched a verified harness name. Claude Code's Stop hook
fires as a bg-spare worker several levels below the session's actual
lock-owning claude process (hook shell -> claude bg-spare ->
claude bg-pty-host -> claude -> claude(lock)), so the first match was
the bg-spare worker, not the lock owner. fm_session_lock_owned_by_self()
then never matched state/.lock, and the Claude Stop auto-arm silently
treated its own primary session as an unrelated live owner and never
armed the watcher.
The walk now keeps going past a claude-named match, looking for a still
more ancestral claude-named match, and stops the instant a non-match
follows an already-found match (bounding it to a contiguous run rather
than the literal ancestry top, so an unrelated claude-named process
further up the real process tree is never mistaken for part of this
session's own nested chain). Every other harness keeps the original
first-match-wins behavior, since e.g. Pi's shared signed-wrapper
ancestry actually holds the session at the inner engine pid, not an
outer wrapper pid. Hop limit raised from 8 to 16 to cover the deeper
bg-spare chain.
* no-mistakes(review): Add nested-claude-ancestry regression test; fix nudge doc depth claim
* no-mistakes: apply CI fixes
* fix: conferma l'avvio del watcher su Windows/MSYS (#1212)
* fix: confirm watcher startup on MSYS
* no-mistakes(review): gate MSYS arm ready timeout, cache uname, harden locale test
* no-mistakes(review): validate OpenCode ready timeout, make uname cache internal
* fix(spawn): forward CLAUDE_CONFIG_DIR to claude crewmates (#1195)
* fix(spawn): forward firstmate's CLAUDE_CONFIG_DIR to claude crewmates
Crewmate panes are created by a long-lived tmux/herdr daemon that does not
inherit firstmate's current environment. When firstmate runs under a non-default
CLAUDE_CONFIG_DIR (for example a work-vs-personal subscription split), a bare
`claude` in the crewmate pane fell back to the default ~/.claude store and
launched unauthenticated, blocking the crewmate before it could do any work.
fm-spawn now prefixes the claude launch with firstmate's own resolved
CLAUDE_CONFIG_DIR when set, so the crewmate uses the same credential/config
store firstmate is authenticated with. An unset value is the single-store
default and adds no prefix; non-claude harnesses are unaffected.
Adds three tests in fm-spawn-dispatch-profile.test.sh (forwarded-when-set,
omitted-when-unset, non-claude-ignored) and pins CLAUDE_CONFIG_DIR in the test
helper so launch assertions no longer depend on the developer's environment.
* no-mistakes: apply CI fixes
* fix: preserve dispatch identity across authentication checks (#1233)
* fix: preserve dispatch harness identity
* no-mistakes(review): Fix Grok counterfactual tuple validation
* no-mistakes(document): Scope dispatch authentication to selected tuple
* fix: restore dispatch instruction budget
* no-mistakes(review): Scope dispatch authentication after candidate selection
* fix(bin): normalize relative durable paths (#1256)
* fix(bin): handle dash-leading harness process names (#2)
* fix: handle dash-leading harness process names
* no-mistakes(review): Make dash-leading harness regression hermetic
* fix: preserve secondmate reply routes across relative homes
Resolve relative home, data, and state inputs before …
vipentti
pushed a commit
to vipentti/firstmate
that referenced
this pull request
Aug 5, 2026
* fix: harden PR check artifacts * fix: close PR check migration gaps * fix: close PR check retry gaps * fix: clarify migration outcomes and ESM boundary * fix: keep failed migrations authoritative * test: use inert PR validation fixtures * no-mistakes(review): Reserve noncanonical PR quarantine namespace * no-mistakes(review): Prevalidate final PR-check teardown artifacts * no-mistakes(review): Preserve X metadata and validate teardown IDs * no-mistakes(review): Initialize migration state before watcher exclusion * no-mistakes(review): Isolate failed poll migrations from bootstrap recovery * no-mistakes(review): Allow safe polling during incomplete private repairs * no-mistakes(review): Authenticate watcher checks at execution time * no-mistakes(review): Preserve custom checks with hash-bound registration * no-mistakes(review): Clean custom check snapshots on watcher signals * no-mistakes(review): Stop watcher checks promptly on signals * no-mistakes(review): Terminate watcher check groups before cleanup * no-mistakes(document): Correct stale X-mode watcher documentation * fix: drain returned watcher check groups * no-mistakes(review): Harden quarantine links and recover validated replacement polls * no-mistakes(review): Preserve X mode across shim version transitions * no-mistakes(review): Refresh legacy X shims before marker short-circuits * no-mistakes(document): Correct persisted PR-check artifact documentation * no-mistakes(document): Correct stale PR-check documentation * fix: bind PR poll repair provenance * no-mistakes(review): Enforce single-link ownership for custom check artifacts * no-mistakes(review): Preserve private checks, X polling, and lifecycle IDs * no-mistakes(review): Separate task creation and legacy teardown validation * no-mistakes(review): Restore safe legacy operations and teardown validation * no-mistakes(review): Disambiguate migration obligations and preserve legacy retries * no-mistakes(review): Preserve fail-closed diagnostics and legacy quarantine evidence * no-mistakes(review): Reconcile legacy migration retries and teardown collisions * no-mistakes(review): Force legacy namespace reconciliation before marker short-circuits * no-mistakes(document): Document private poll artifact safety contracts * no-mistakes(lint): Suppress intentional literal-dollar lint finding * fix: migrate historical X poll identity * fix: harden PR check artifacts * no-mistakes(review): Preserve fail-closed diagnostics and legacy quarantine evidence * no-mistakes(review): Reconcile legacy migration retries and teardown collisions * fix: migrate historical X poll identity * no-mistakes(review): Harden X-mode artifact publication against symlink corruption * no-mistakes(review): Guard X artifact publication * no-mistakes(review): Enforce private X artifact reads * no-mistakes(test): Fix backend compatibility fixture dependencies * no-mistakes(document): Refresh PR-check documentation * no-mistakes(lint): Remove unused x-mode test locals * no-mistakes: apply CI fixes
elixlabssolutions
added a commit
to xLabs-OS/firstmate
that referenced
this pull request
Aug 6, 2026
…ing both fork commits (#4) * fix: compress Firstmate contract and enforce delivery rigor ownership (#626) * docs: compress firstmate operating contract * docs: make delivery rigor single-owner PR B already removed personal and stacked review requirements, but it did not explicitly assign rigor to the selected delivery path or forbid risk-based manual clean gates. That gap still permitted the Hi Bit inversion. * no-mistakes(review): Honor configured merge authority across faster delivery paths * no-mistakes(document): Align docs with compressed operating contract * feat: add durable captain decision holds (#593) * Add durable captain decision holds * no-mistakes(review): Validate decision hold retries and origin paths * no-mistakes(review): Enforce durable decision lifecycle boundaries * no-mistakes(review): Harden decision display and partial retry recovery * no-mistakes(test): Update scout teardown fixtures for decision inventory * no-mistakes(document): Align decision lifecycle and scout teardown documentation * no-mistakes: apply CI fixes * no-mistakes(review): Reconcile terminal decision holds * no-mistakes(document): Align captain decision-hold documentation * fix(bin): harden PR check artifacts (#556) * fix: harden PR check artifacts * fix: close PR check migration gaps * fix: close PR check retry gaps * fix: clarify migration outcomes and ESM boundary * fix: keep failed migrations authoritative * test: use inert PR validation fixtures * no-mistakes(review): Reserve noncanonical PR quarantine namespace * no-mistakes(review): Prevalidate final PR-check teardown artifacts * no-mistakes(review): Preserve X metadata and validate teardown IDs * no-mistakes(review): Initialize migration state before watcher exclusion * no-mistakes(review): Isolate failed poll migrations from bootstrap recovery * no-mistakes(review): Allow safe polling during incomplete private repairs * no-mistakes(review): Authenticate watcher checks at execution time * no-mistakes(review): Preserve custom checks with hash-bound registration * no-mistakes(review): Clean custom check snapshots on watcher signals * no-mistakes(review): Stop watcher checks promptly on signals * no-mistakes(review): Terminate watcher check groups before cleanup * no-mistakes(document): Correct stale X-mode watcher documentation * fix: drain returned watcher check groups * no-mistakes(review): Harden quarantine links and recover validated replacement polls * no-mistakes(review): Preserve X mode across shim version transitions * no-mistakes(review): Refresh legacy X shims before marker short-circuits * no-mistakes(document): Correct persisted PR-check artifact documentation * no-mistakes(document): Correct stale PR-check documentation * fix: bind PR poll repair provenance * no-mistakes(review): Enforce single-link ownership for custom check artifacts * no-mistakes(review): Preserve private checks, X polling, and lifecycle IDs * no-mistakes(review): Separate task creation and legacy teardown validation * no-mistakes(review): Restore safe legacy operations and teardown validation * no-mistakes(review): Disambiguate migration obligations and preserve legacy retries * no-mistakes(review): Preserve fail-closed diagnostics and legacy quarantine evidence * no-mistakes(review): Reconcile legacy migration retries and teardown collisions * no-mistakes(review): Force legacy namespace reconciliation before marker short-circuits * no-mistakes(document): Document private poll artifact safety contracts * no-mistakes(lint): Suppress intentional literal-dollar lint finding * fix: migrate historical X poll identity * fix: harden PR check artifacts * no-mistakes(review): Preserve fail-closed diagnostics and legacy quarantine evidence * no-mistakes(review): Reconcile legacy migration retries and teardown collisions * fix: migrate historical X poll identity * no-mistakes(review): Harden X-mode artifact publication against symlink corruption * no-mistakes(review): Guard X artifact publication * no-mistakes(review): Enforce private X artifact reads * no-mistakes(test): Fix backend compatibility fixture dependencies * no-mistakes(document): Refresh PR-check documentation * no-mistakes(lint): Remove unused x-mode test locals * no-mistakes: apply CI fixes * fix(bin): compact session-start backlog digest (#636) * fix: compact session-start backlog digest * no-mistakes(test): Fix legacy backend fixture helper * no-mistakes(test): Fix watcher exit wait helper * no-mistakes(document): Document compact backlog digest * fix: dedupe stale watcher guard banners (#637) * fix: dedupe stale watcher guard banner * no-mistakes(review): Keep read-only guard state nonmutating * no-mistakes(document): Clarify stale watcher docs * fix(bin): balance bearings landed baseline (#640) * fix: balance bearings landed defaults * no-mistakes(document): Document balanced landed baseline * no-mistakes: apply CI fixes * fix: clarify captain-facing translation contract (#644) * docs: clarify captain-facing translation contract * no-mistakes(review): Restore runtime fallback mandate * no-mistakes(document): Align Bearings translation wording * fix(bin): make bootstrap output and nudges deterministic (#646) * fix: make bootstrap nudges deterministic * no-mistakes(review): Honor state override for bootstrap nudges * no-mistakes(review): Update benign bootstrap documentation labels * no-mistakes(review): Validate bootstrap nudge retry markers * no-mistakes(document): Align bootstrap nudge documentation * no-mistakes: apply CI fixes * docs(secondmate-provisioning): clarify concise registry ownership (#649) * Clarify concise secondmate registry contract * no-mistakes(review): Expand secondmate registry boilerplate coverage * no-mistakes(document): Point route docs to owner * fix(bin): strip quoted blocked_by values during decision hold resolve (#654) * fix(bin): strip quotes on blocked_by in decision-hold resolve tasks-axi quotes multi-entry blocked_by as "a,b,c", so the comma-boundary membership test only matched middle elements. Strip surrounding quotes before matching so first and last hold ids resolve correctly. * no-mistakes(document): Refresh decision-hold regression evidence * feat(secondmate): inherit shared captain preferences (#656) * feat(secondmate): inherit shared captain preferences * no-mistakes(review): Honor shared captain data overrides * no-mistakes(review): Honor bootstrap data override registry * no-mistakes(document): Refresh shared inheritance docs * no-mistakes(document): Clarify inherited local-material docs * feat: gate local agent secret injection (#658) * feat(spawn): gate local agent secret injection * fix(spawn): align final Keychain slot * no-mistakes: apply CI fixes * test: isolate Herdr autodetect smoke sessions (#662) * test: isolate herdr autodetect smoke session * no-mistakes(review): Restored autodetect smoke gate bypass * no-mistakes(test): Harden Herdr lab provisioning * no-mistakes(document): Refresh Herdr lab docs * docs: adopt under way for active work (#666) * Revert "feat: gate local agent secret injection (#658)" (#668) This reverts commit c27135cd9d35bc4c237d49b3b374da31fbd52eef. * fix(pi): distinguish stale locks when arming watcher (#681) * fix(pi): distinguish stale locks when arming watcher * no-mistakes(test): Stabilize watcher extension async waits * no-mistakes(document): Document Pi lock recovery * fix: accept secondmate house vocabulary (#685) * fix: accept secondmate as house vocabulary * no-mistakes(test): Update captain vocabulary contract test * no-mistakes(document): Align secondmate documentation vocabulary * fix(bin): parse handoff homes after registry parentheticals (#686) * fix: parse secondmate home after pre-field parentheses Registry summaries often include parentheticals before the structured (home: ...) field. Match that field with a greedy prefix so handoff no longer reports "has no home" for those entries. * no-mistakes(document): Refresh handoff test comments * feat: add native session-start nudges (#687) * feat: add native session-start nudges * no-mistakes(document): Document nudge script inventory * docs: call built-in defaults the firstmate repo, not template (#688) Relabel absent-captain and related domain defaults wording so it names the firstmate repo rather than treating "template" as this domain's identity label. Keep the design-tenet "shared template" statements and unrelated launch/PR-poll template uses unchanged. * fix(bin): repair fm-brief.sh parse error and harden set -u array expansion (#205) * fix(bin): use set -u-safe empty-array expansion in pr-merge and spawn Expanding "${arr[@]}" on an empty array under set -u fails on bash < 4.4 (notably macOS bash 3.2). Quote the portable "${arr[@]+"${arr[@]}"}" idiom in fm-pr-merge and fm-spawn batch dispatch so empty arrays expand to nothing. Co-authored-by: Cursor <cursoragent@cursor.com> * test(brief): harden fm-brief regression coverage for parse and scaffolds Tighten bash -n checking, pin literal backtick rendering in the no-mistakes DOD wording assertion, and keep a scout/secondmate scaffold smoke test so the Co-authored-by: Cursor <cursoragent@cursor.com> #166 apostrophe regression cannot return unnoticed. --------- Co-authored-by: Cursor <cursoragent@cursor.com> * fix(bin): keep watcher supervision continuous across child cycles (#693) * fix: make watcher supervision continuous * no-mistakes(review): Bound watcher retries and log attached signals * no-mistakes(review): Add bounded successor-recovery wake fallbacks * no-mistakes(review): Prevent overlapping successor-arm retries * no-mistakes(review): Resume supervision after late arm closes * no-mistakes(review): Bind OpenCode recovery to attempted arm * no-mistakes(test): Synchronize peer beacon regression fixture * no-mistakes(test): Synchronize Pi and OpenCode late-close lifecycle fixtures * no-mistakes(document): Captain: document watcher successor protocol behavior * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * fix: fetch current PR head for review diffs (#722) * fix: always fetch PR head for review diffs Prefer a freshly fetched refs/pull/<n>/head over a reachable recorded pr_head= so reviewers never hold a merge over a "missing" fix that already landed on the remote PR. Recorded SHA is offline fallback only; local branch is last resort with a warning. Store the tip under refs/fm-review/ so a later base-branch fetch cannot clobber the compare tip via FETCH_HEAD. * no-mistakes(test): Isolate session-start nudge tests from gate state * no-mistakes(document): Correct review-diff documentation * docs: resolve five contract contradictions across AGENTS.md, README, and skills (#736) * docs: resolve five contract contradictions * no-mistakes(test): align owner-pointer assertions with reworded docs; skip absent shellcheck * docs(harness): correct Grok exit guidance (#742) * docs(harness): reverify grok exit command * no-mistakes(test): Correct Grok exit resume attribution * fix(watcher): bound stale wakes for parked crew (#743) * fix(watcher): bound stale wakes for exited paused crew * no-mistakes(review): Gate pause suppression on confirmed agent death * no-mistakes(test): Fixed stale pause cadence * no-mistakes(document): Document dead-agent hold cadence * fix(supervision): distinguish ordinary wakes from recovery (#744) * fix(supervision): distinguish ordinary wakes from repair * no-mistakes(review): Make passive guard follow-ups recovery-only * no-mistakes(document): Clarify recovery-only turn-end guard documentation * fix(x-mode): dedupe pending mention wakes (#745) * fix(x-mode): dedupe pending mention wakes * no-mistakes(review): fix x-poll claim error deduplication * no-mistakes(review): separate claim diagnostics from relay recovery * no-mistakes(document): Document X-mode once-only mention wakes * feat(wake): enrich drained signals with bounded status context (#747) * feat(wake): enrich drained signal context * no-mistakes(review): Bound wake enrichment reads * no-mistakes(document): Document wake-drain annotations * docs(wake): explain at-least-once drain boundary * no-mistakes(review): Prevent symlink races in wake annotations * no-mistakes(review): Exercise wake symlink race regression * test: document intentional AFK marker subprocesses * fix(wake): isolate annotation marker state * feat(herdr): add optional presentation spaces (#784) * feat(herdr): add optional presentation spaces * no-mistakes(review): Harden Herdr projection creation and spawn serialization * no-mistakes(review): Captain, disarm Herdr cleanup before launch submission * no-mistakes(test): Correct stale Orca metadata failure fixture * no-mistakes(document): Document Herdr presentation projection accurately * fix(send): treat opencode busy-queued composer state as submitted (#775) * fix(send): treat opencode busy-queued composer state as submitted When fm-send sends a message to a BUSY opencode crewmate on the tmux backend, opencode accepts the Enter and queues the message for the next turn, but leaves the typed text visible in the composer row. The submit-verification loop sees a pending composer, exhausts retries, and reports a false "Enter swallowed" failure while the message is actually delivered. Fix: after Enter retries are exhausted and the composer still shows pending, check fm_pane_is_busy. If the pane is busy (agent mid-turn, footer shows "esc interrupt"), the harness queued the message, so return "empty" (accepted). On an idle pane, keep returning "pending" (genuine swallow detection preserved). Regression tests cover four scenarios: - busy pane + pending composer -> empty (message queued) - idle pane + pending composer -> pending (genuine swallow) - busy pane + composer clears on first Enter -> empty - idle pane + composer clears on first Enter -> empty (existing path) * docs: document busy-queued Enter exception across backend docs and skills Add explanatory comments and backend documentation for the busy-queued Enter fix (opencode 1.18.4 accepts Enter mid-turn but keeps typed text in composer until the turn ends): - bin/fm-tmux-lib.sh: document the busy-aware fallback in the file header and above fm_tmux_submit_enter_core - .agents/skills/afk/SKILL.md: daemon-facing policy note - .agents/skills/harness-adapters/SKILL.md: harness-specific fact - docs/tmux-backend.md: submit-acknowledgement section with the busy-queue exception - docs/herdr-backend.md: record the known gap - docs/architecture.md: cross-reference in the daemon section * test(tmux): fix SC2181 and make busy-submit test executable * fix(spawn): require two stable reads before accepting worktree path (#765) * fix(spawn): require two stable reads before accepting worktree path The treehouse-get worktree-detection loop in fm-spawn.sh accepted the first pane_current_path read that differed from the project path, but on some tmux/WSL setups a brand-new window transiently reports a stale-but-real path before the pane actually settles into the worktree. Since that stale path is itself a real, distinct git checkout, it also passes validate_spawn_worktree's isolation check, so the loop silently recorded the wrong worktree in state/<id>.meta (and, for claude harness spawns, installed the turn-end hook there too). Require two consecutive polls to agree on the same non-project path before accepting it, using the existing inter-poll sleep as the confirmation gap so an already-settled pane isn't slowed down by an extra cycle. * fix(tests): drop unused CASE_DIR read in worktree-settle test ShellCheck SC2034: CASE_DIR is split out of the case record but never referenced; discard it with _ instead. --------- Co-authored-by: Freudator86 <tim@allesknut.de> * fix(bin): make watcher process identity immune to Linux wall-clock changes (#752) * fix(watcher): stabilize Linux process identity * no-mistakes(document): document FM_PROC_ROOT_OVERRIDE and Linux starttime identity rationale * fix: prevent AFK idle stalls and stale run attribution (#758) * fix(supervision): verb-aware captain relevance, AFK wedge, head-bound state Stop free-text tokens like "merged" from promoting nonterminal working: lines to captain-relevant, so AFK no longer permanently suppresses idle recovery. Defend wedge aging independently for nonterminal progress verbs, bind no-mistakes current-state attribution to code identity (not branch alone), and mark setup-complete as nonterminal in the ship brief scaffold. * no-mistakes(review): Enforce nonterminal suppression and head-bound run attribution * no-mistakes(document): Document current-code-bound run attribution * no-mistakes(test): Wait for stable Herdr shell readiness * no-mistakes(test): Make Herdr and watcher readiness tests deterministic * no-mistakes(test): Make tmux capture and watcher lifecycle deterministic * no-mistakes(document): Document corrected supervision contracts * fix(bin): allow safe teardown during watcher recovery (#750) * fix: allow safe teardown during watcher recovery * no-mistakes(review): Distinguish unsafe-teardown deny guidance via policy reason code * no-mistakes(document): Sync continuity-gate docs to allow teardown recovery * test: mark dynamic teardown fixture literal * no-mistakes(document): docs: add teardown to continuity gate allow list * feat(herdr): order presentation spaces while preserving focus (#790) * feat(herdr): order presentation worker spaces * fix(herdr): preserve focus during projected cleanup * no-mistakes(review): Serialize Herdr cleanup and protect active seeded tabs * no-mistakes(review): Serialize Herdr aborts with guarded focus regressions * no-mistakes(review): Fall back flat when Herdr serialization is unavailable * no-mistakes(test): Stabilize watcher startup and AFK handoff tests * no-mistakes(document): Correct Herdr ordering and focus documentation * fix(bin): send literal config reread nudges after pushes (#809) * Send literal config reread after inherited config push When declared inherited config changes under an already-running secondmate, build a per-home instruction from validated destination post-write bytes and deliver it on the routed secondmate path. Unchanged config sends nothing; ABSENT represents removal; captain-shared is never inlined. Covers mid-session config-push and the locked bootstrap convergence path without hardening spawn against deliberate runtime choice. * no-mistakes(review): Fix config reread framing, partial propagation, and respawn order * no-mistakes(review): Send config rereads via durable single-line pointers * no-mistakes(review): Make failed config rereads retryable * no-mistakes(review): Make config reread retries generation-safe * no-mistakes(review): Make config rereads durable and ordered * no-mistakes(review): Drain retries, bound history, preserve detect-only read-only mode * no-mistakes(review): Retain write retries and quarantine stale respawn generations * no-mistakes(review): Preserve exact config reread retries and delivery order * no-mistakes(review): Preserve exact retry bytes and bounded quarantine pruning * no-mistakes(document): Consolidated config-reread documentation * feat(watch): follow GitLab merge requests to merge (#797) * feat(watch): follow GitLab merge requests to merge The merge watch only understood GitHub pull requests, so a task whose deliverable is a GitLab merge request was never followed to merge. Generalize the stored poll identity from owner/repository to a provider-tagged provider/url/host/path/number record. GitLab runs mostly on self-hosted instances and its projects nest under groups at no fixed depth, so the host and the full project path are data in the record rather than constants, and every consumer rebuilds the URL from those parts and refuses any record that does not reconstruct it exactly. The GitLab state is read with plain glab, matching the GitHub path's use of plain gh, so an upstream checkout needs no extra tooling. Two things about glab were established by running it rather than assumed, because a wrong invocation here fails silently into a permanent "not merged": - glab has no field selector, and its JSON would need a JSON processor that firstmate does not require, so the state is read from glab's own field output. Only an exact "merged" wakes firstmate, so a changed format stays silent instead of reporting a merge. - glab cannot take a merge request URL the way gh can, because that form resolves through the current git repository and the watcher has none. It is addressed by project URL and merge request number instead. An absent glab produces no wake rather than a false merge, and arming refuses with a clear message since that is the one point where a missing CLI can still be reported. A GitLab task records no pr_head, which both consumers already treat as optional. The merge path still addresses GitHub only and refuses a merge request URL rather than sending it to the wrong forge. The record version moves to v2, and the existing non-executing migration rebuilds an already-armed watch from its recorded URL, so no watch is lost by upgrading. docs/gitlab-merge-watch.md records the evidence, taken against the public fixture project https://gitlab.com/KarotKris/gitlab-merge-watch-fixture. * no-mistakes(review): Reject github.com host in GitLab MR URL/sidecar validation * no-mistakes(document): Note GitLab MR URLs are explicitly refused, not just malformed ones, in fm-pr-merge.sh docs * fix(herdr): group projected children beneath owning parents (#821) * feat(herdr): correct all-home child presentation topology Inherit the presentation opt-in to secondmate homes, label new projected spaces with the approved corner format, insert each child under its owning parent under one session-scoped lock, and keep flat non-destructive fallback. * no-mistakes(review): Exclude secondmates from Herdr presentation projection * no-mistakes(review): Harden shared Herdr locks and ambiguous child ordering * no-mistakes(review): Use adjacency-only Herdr child ownership * no-mistakes(review): Reject foreign legacy projections safely * no-mistakes(review): Validate Herdr session sockets before projection * no-mistakes(test): Fix Herdr teardown fixture session socket metadata * fix(herdr): canonicalize presentation lock socket paths Always resolve the session socket parent directory so symlink parents such as /tmp -> /private/tmp cannot split the shared cross-home lock identity. Refuse relative socket paths. Clarify lock-unavailable warnings. * no-mistakes(test): Fix Bash-compatible GitLab merge request URL parsing * no-mistakes(document): Document all-home Herdr child topology * no-mistakes(lint): Quote fallback provenance string for ShellCheck * fix: keep local no-mistakes tests intent-targeted (#823) * fix(no-mistakes): drop full-suite local Test override Local no-mistakes Test is intent-targeted; CI Behavior keeps the broad tests/*.test.sh suite. Keep commands.lint on bin/fm-lint.sh and add a focused contract test so the override cannot silently return. * no-mistakes(lint): Make CI contract assertion ShellCheck-clean * feat: add canonical timed test runner (#825) * feat(test): add canonical timed suite runner and honest CI timeout Introduce bin/fm-test-run.sh as the single serial owner for selecting one script, a family, a conservative changed-file set, or the explicit complete suite, with per-script timing markers and a JSON artifact. Wire CI Behavior through the runner, raise the hang-tripwire timeout to 25 minutes, and document entry points without restoring a full-suite local no-mistakes Test command. * no-mistakes(review): Captain: fix changed selection and empty summaries * no-mistakes(review): Captain: fail closed on unmapped changed sources * no-mistakes(document): Document canonical timed test entry points * fix: surface main inventory gaps in Bearings (#830) * fix: disclose main-home orphan and unstructured inventory gaps Main Bearings could report an empty fleet while structured in-flight rows lacked meta or current backlog rows were free-form. Emit main_inventory from the fleet snapshot, map it into Bearings omitted surfaces and a Charted Next gate, and keep meta as the only live Underway source. * no-mistakes(document): Document Bearings inventory-integrity projection * no-mistakes: apply CI fixes * feat: add bounded concurrent test isolation proof (#832) * feat: add concurrent test isolation proof for Phase 2 Prove an audited portable candidate set passes under concurrent workers with private mode-0700 temp roots, without enabling production CI sharding or fm-test-run --jobs. * no-mistakes(review): Pin isolation proof to audited candidate manifest * feat: guard against missed secondmate reports (#834) * feat(secondmate): parent-owned guards for missed status reports Marked parent-to-secondmate requests now create a durable pending-reply expectation with a privacy-safe correlation id before delivery. Transport success never resolves it; only a correlated parent status or document pointer does. After a completed turn with no report, the parent sends one recovery repost and escalates once if that turn is also missed, without scraping the secondmate conversation or looping. * no-mistakes(review): Deduplicate wrong-home pending-reply sightings * no-mistakes(review): Harden pending-reply recovery and escalation guards * no-mistakes(review): Bound pending-reply backend polling * no-mistakes(review): Cache pending-reply status scans * no-mistakes(review): Protect undelivered pending-reply records from scans * no-mistakes(review): Close pending-reply delivery durability gaps * no-mistakes(review): Separate pending-reply transport outcomes * no-mistakes(review): Escalate stalled pending-reply deliveries once * no-mistakes(review): Resolve attempted deliveries from correlated reports * no-mistakes(review): Resolve late reports after delivery escalation * no-mistakes(document): Document pending-reply grace and ownership * no-mistakes(lint): Silence intentional pending-reply test fixture lint warnings * feat: require pinned real-Herdr CI coverage (#838) * feat: add required pinned Herdr CI lane Install exact Herdr 0.7.4 and Treehouse 2.0.1 with official assets and SHA-256 pins, run the real-herdr-gated family serially through fm-test-run with hard-fail on herdr-not-found, and keep portable Behavior free of claimed Herdr coverage. * no-mistakes(document): Consolidate real-Herdr CI documentation ownership * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * feat: shard portable tests and add bounded local parallelism (#841) * feat: shard portable CI tests after isolation proof Balance the Phase 2 proven-isolated set into two LPT portable parallel lanes from Phase 1 timing evidence, keep stateful work in a required portable serial lane, exclude real Herdr to its dedicated required lane, and prove complete inventory coverage with a deterministic guard. Add bounded local --jobs only for the proven set, per-lane timing plus aggregate artifacts, and reduce the interim portable hang tripwire now that the serial remainder owns the long wall-clock path. * no-mistakes(review): Captain, fix CI contracts and completion-order worker scheduling * no-mistakes(review): Captain, preserve stderr gate-skip detection in parallel tests * no-mistakes(document): Document portable sharding and timing aggregation * no-mistakes: apply CI fixes * fix: block primary-session delegation outside the fleet (#854) * feat: fence primary-session delegation outside the fleet A firstmate primary that delegates through Claude Code's built-in delegation tools creates work with no state/<id>.meta. Because fm-supervision-lib.sh counts *.meta and fm-turnend-guard.sh exits silently at zero, such work does not merely go unsupervised: it makes the whole guard stack structurally inert, and it dies with the primary session. On 2026-07-22 that cost two workers mid-flight and left supervision down for 73 minutes unnoticed. Layer 1, the primary fix: a permissions.deny list in .claude/settings.json removes the 18 delegation, scheduling, worktree, and task-tracking tools from the model's schema, so they are never offered. This is removal rather than interception, so there is no call to intercept and no fail-open path. The list is flat and in one file so its width stays reviewable; the captain owns that width. Layer 2, bin/fm-subagent-pretool-check.sh: a deny list is fail-open against tools that do not exist yet, and permissions.allow is a pre-approval list rather than an availability list, so there is no fail-closed allowlist to use instead. This backstop classifies the tool NAME by shape rather than against a fixed list, so a delegation tool that ships before the deny list is updated is still refused. It excludes mcp__* names and observe-or-stop operations, scopes itself to a genuine primary home via the shared fm_primary_scope_matches predicate so a crewmate's task worktree is unaffected, and offers one deliberate FM_ALLOW_SUBAGENT=1 escape hatch that must be set at launch. Verified live against Claude Code 2.1.217, including a deny-key A/B with a nonsense-name control, layer 2 denying an un-denied Workflow call, the same call allowed in a linked worktree, and the escape hatch. Corrects a prior finding: both Task and Agent work as deny keys, so both are pinned. Codex 0.144.1 verified to expose no delegation tool; grok, opencode, and pi are inspected and documented as not wired because those binaries are absent from this host and the repo requires live validation before trusting a harness hook. Evidence in docs/subagent-guard.md. * no-mistakes(review): Ship scoped Claude delegation guard * no-mistakes(test): Ship Claude delegation deny list * no-mistakes(document): Clarify PreToolUse guard ownership * no-mistakes(lint): Keep Claude deny list local * fix: install tasks-axi in portable CI shards (#866) Reproduction: portable-parallel-2 completed successfully without tasks-axi while fm-decision-hold-lifecycle emitted a gate skip in 30 ms. The pre-shard lane installed tasks-axi and exercised the test fully. Installing tasks-axi is the smallest counterfactual and makes the representative shard execute the test with gate_skip=false in about 20 seconds. Both parallel jobs receive symmetric setup, while the exact 91-test inventory and coverage guard remain unchanged. * feat(bin): make dispatch profiles quota aware (#867) * feat: make dispatch profiles quota aware * no-mistakes(review): Fix quota window and Grok product scoping * no-mistakes(document): Document implicit quota-aware dispatch accurately * Add built-in ahoy recap skill (#873) * fix: preserve trustworthy Bearings data in partial snapshots (#875) * fix: preserve mixed Bearings projections * no-mistakes(review): Enforce strict invalidity precedence for partial snapshots * no-mistakes(review): Enforce ownership for unknown child metadata * no-mistakes(document): Document partial structured Bearings projections * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * feat(pi): add session-local calm mode (#884) * Add session-local Pi calm mode * no-mistakes(review): Preserve Pi HTML exports during calm mode * no-mistakes(review): Preserve calm exports across submit bindings and share * no-mistakes(document): Document calm-mode feasibility across supported harnesses * fix(pi): prevent redundant watcher re-arms (#885) * fix(pi): limit watcher arm tool to recovery * no-mistakes(review): Strengthen Pi live re-arm regression coverage * no-mistakes(document): Document Pi first-cycle and recovery-only watcher arming * fix(pi): clean up Calm transcript rendering (#895) * fix(pi): clean up Calm transcript rendering * no-mistakes(review): Captain, preserve Calm exports and classify Pi launch briefs * no-mistakes(review): Captain, eliminate Calm gaps and verify exported conversations * no-mistakes(review): Restore Calm rows received while active * no-mistakes(review): Preserve diagnostics during Calm restoration * no-mistakes(document): Clarify Calm transcript behavior and injection paths * fix: execute every PR body compliance event (#898) * fix: execute every PR body compliance event * no-mistakes(document): Document independent PR compliance events * fix: exclude operational injections from ahoy boundaries (#899) * fix: distinguish operational input in ahoy * no-mistakes(review): Handle legacy Ahoy operational boundaries * no-mistakes(review): Narrow legacy Ahoy boundaries with live regressions * no-mistakes(document): Document Ahoy operational marker ownership * no-mistakes(lint): Suppress intentional literal fixture lint warnings * fix: canonically classify operational inputs across harnesses (#909) * fix: type canonical operational inputs * no-mistakes(document): Correct canonical operational-input documentation ownership * fix: avoid generic secondmate start acknowledgements (#926) * fix: avoid generic secondmate acknowledgements * no-mistakes(document): Document sparse secondmate acknowledgement behavior * no-mistakes: apply CI fixes * fix(pi): make calm mode persistent and gapless (#927) * fix(pi): preserve calm presentation across sessions * no-mistakes(review): Fix Calm home fallback persistence * no-mistakes(document): Clarify Calm gapless and export contracts * no-mistakes: apply CI fixes * fix(watch): retire merged PR polls after durable notification (#932) * fix: retire merged PR polls after notification * no-mistakes(review): Decouple PR retirement recovery from template updates * no-mistakes(review): Recover pending PR retirements before poll migration * no-mistakes(document): Document merged PR poll retirement contracts * fix: refine scout intake and parallel dispatch (#934) * Clarify intake evidence and overlap handling * no-mistakes(review): Align scout guard with intake classification * no-mistakes(document): Clarify scout documentation and intake ownership * fix(pi): prevent duplicate assistant replies in Calm (#936) * fix(pi): preserve operational follow-up semantics in Calm * no-mistakes(document): Correct Calm operational-row visibility documentation * perf(bin): shrink the ShellCheck source graph (#939) * perf(lint): shrink shell source graph * no-mistakes(review): Ensure lint workers terminate fully on cancellation * no-mistakes(document): Repair stale lint documentation ownership * fix(pi): remove Calm hidden-block gaps (#942) * fix(pi): remove Calm hidden-block gaps * no-mistakes(review): Validate Calm geometry against current viewport * no-mistakes(review): Synchronize Calm geometry checks with reload completion * fix: enforce contract boundaries for ask-user findings (#945) * fix: escalate ask-user contract expansion * no-mistakes(document): Point project management to authority owner * docs: prefer direct operational paths (#946) * fix(pi): hide operational user rows in Calm mode (#948) * fix(pi): hide Calm operational user rows * no-mistakes(review): Narrow Calm operational input suppression * no-mistakes(review): Avoid Calm replay classifier subprocesses * no-mistakes(document): docs: point Pi verification to Calm owner * fix: relaunch missing second mates at session start (#950) * fix(session-start): relaunch missing second mates * fix(test): detect completed parallel workers * no-mistakes(review): Isolate session-start recovery test cleanup * no-mistakes(review): Complete backend-safe secondmate session recovery * no-mistakes(review): Resolve Zellij task ownership before recovery * no-mistakes(review): Recover relocated Zellij ghost tabs safely * no-mistakes(review): Restore conservative Zellij recovery boundary * no-mistakes(review): Reject malformed tmux recovery targets * no-mistakes(document): Align secondmate recovery documentation * no-mistakes: apply CI fixes * fix(herdr): reclaim resumed task projections after restart (#967) * fix(herdr): reclaim resumed task projections safely * no-mistakes(review): Enforce safe Herdr reclaim close boundaries * no-mistakes(document): docs: clarify Herdr restart projection contract * Teach Ahoy to surface open decisions (#968) * Require shipshape routine acknowledgement (#969) * docs: separate current guidance from verification evidence (#994) * docs: separate current guides from verification * no-mistakes(review): Restore Herdr 0.7.5 restart-reclaim verification evidence * fix: preserve Claude watcher continuity across Stop hooks (#997) * feat(claude): Stop-owned tokenless watcher continuity via asyncRewake auto-arm Claude primaries (main home and marked secondmate homes) no longer depend on the model remembering to re-arm the watcher after each wake. A tracked Stop asyncRewake hook (bin/fm-claude-stop-autoarm.sh, timeout 28800s) fires on every turn end, claims one home-scoped single-flight owner, foregrounds bin/fm-watch-arm.sh inside the hook-owned process tree, and translates an actionable close or typed watcher failure into exactly one exit-2 rewake. The hook scopes to genuine primary checkouts, requires the session lock to be held by its own harness ancestor, stays inert while AFK owns triage or the home is idle, and hands AFK transitions mid-cycle to the daemon without rewaking. The synchronous turn-end guard gains a --claude cooperative mode: it ignores stop_hook_active (true on every post-continuation stop, which is what re-opened the 2026-07-21 blind window), waits briefly for a watcher health proof, a live auto-arm owner claim, or a fresh rewake epoch, and re-blocks only when the auto-arm genuinely failed to establish - bounded to 3 consecutive blocks per session, safely below Claude Code's 8-block override, then a degraded allow with a visible systemMessage. Codex keeps the previous one-block loop guard byte-identically, and Pi, OpenCode, and Grok adapters are untouched. Continuity PreToolUse gate and durable wake queue are preserved; the gate's recovery guidance now names the Stop-owned re-arm and reserves manual background arms for auto-arm failure. Claude supervision protocol, harness-adapters facts, architecture, configuration, and continuity docs updated; docs/turnend-guard.md records the 2026-07-24 Claude 2.1.218 contract revalidation (tokenless multi-cycle rewake, no-dedup, timeout process-group kill, 8-block cap, interactive non-stall) and the 2.1.219 product live E2Es. Regression matrix: hermetic tests cover scope, identity, AFK, need, single-flight, translation, guard cooperation, budget, and registration; the new live E2E proves two full tokenless auto-arm rewake cycles with zero model arm commands; Pi and OpenCode Option B live E2Es pass unchanged. * no-mistakes(review): Fix Claude X-mode auto-arm continuity backstop * no-mistakes(review): Remove unsupported Claude contract-lab verification claims * no-mistakes(document): Update Claude auto-arm continuity documentation * fix(herdr): clean stale projections at session start (#996) * Clean stale Herdr projections at session start * no-mistakes(document): Document stale Herdr session-start projection cleanup * no-mistakes(review): Enforce locked exact Herdr projection cleanup * no-mistakes(review): Fail closed on unverified session lock ownership * no-mistakes(review): Serialize session lock acquisition atomically * no-mistakes(document): Align session-start and Herdr cleanup documentation * no-mistakes(document): Generalize lock-refusal diagnostics * no-mistakes(lint): Avoid reserved keyword in concurrency test * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * fix: recover Claude supervision without watcher-status gate (#1001) * fix: recover Claude supervision at session start * fix: remove Claude watcher-status command gate * no-mistakes(document): docs: remove stale continuity gate references * fix: make quota-aware profile selection agent-owned (#1018) * Replace quota dispatch selector instructions * no-mistakes(review): Align bootstrap docs with agent-owned dispatch selection * fix(bin): remove vestigial dispatch selector (#1026) * remove vestigial dispatch selector * no-mistakes(review): Synchronize isolation proof and portable shard evidence * no-mistakes(review): Correct shard history and proof archive date * no-mistakes(review): Remove reintroduced selector documentation reference * no-mistakes(document): Remove stale dispatch strategy documentation * docs(agents): drop superseded interim quota-window rule (#1039) quota-axi 0.1.13 emits schemaVersion 2 with a quotaSemantics object per provider, so the successor named in the interim rule has landed and the rule's own removal condition is satisfied. Keep the ownership clause so quota-axi remains the single owner of how model or product windows relate to bounding account windows, and drop the interim weakest-headroom instruction. The unknown-semantics case is already covered by the existing requirement to stop and report a candidate whose applicable quota data or interpretation cannot be established. Drop the matching assertion phrase from tests/fm-instruction-owners.test.sh; the retained ownership phrase still asserts. * fix(tmux): scope busy detection and recognize current Claude turns (#1049) * fix(tmux): scope Claude busy detection by harness * no-mistakes(review): Separate verified and fallback busy signatures * no-mistakes(test): Scope busy signatures to supplied harnesses * no-mistakes(document): Document harness-scoped busy detection * feat: add verified Kimi crewmate adapter (#1047) * Add verified Kimi crewmate harness adapter * no-mistakes(review): Scope Kimi moon detection to spinner lines * no-mistakes(review): Match only complete Kimi spinner rows * no-mistakes(review): Resolve Kimi binary portably before pane creation * no-mistakes(document): Align Kimi adapter documentation * no-mistakes(lint): Suppress false-positive ShellCheck warning for sourced watcher override * Fix Kimi busy spinner detection * no-mistakes(review): Recognize Kimi session-lock ancestry and holders * no-mistakes(review): Scope pending-reply Kimi busy detection by harness * no-mistakes(document): Correct Kimi spinner capture documentation * no-mistakes(document): Clarify optional Kimi spinner whitespace * no-mistakes(lint): Silence intentional pending-reply test stub warnings * test: align rebased Kimi busy fixtures * no-mistakes: apply CI fixes * Reconcile Kimi busy detection after per-harness scoping * no-mistakes(review): Clarify observed Kimi spinner whitespace contract * no-mistakes(document): Clarify Kimi harness documentation * fix: harden Kimi submission and spinner matching (#1058) * fix kimi pointer submission and spinner conformance * no-mistakes(review): Preserve Kimi submit target ownership guard * feat(bin): add guarded Kimi turn-end wake (#1059) * Add guarded Kimi turn-end hook * no-mistakes(review): Require jq before installing Kimi turn-end hook * no-mistakes(review): Expose jq inside isolated Kimi test fixtures * no-mistakes(review): Preserve Kimi config boundaries during hook removal * no-mistakes(review): Document Kimi removal newline safeguard * no-mistakes(document): Document Kimi shared-home preservation * fix(tmux): classify bordered composers across all rows (#1066) * Fix structural tmux composer reading * Verify Calm compatibility with Pi 0.82 * no-mistakes(review): Harden structural composer classification boundaries * no-mistakes(review): Refresh composer and Kimi regression fixtures * no-mistakes(review): Fail closed on unbounded composer edges * no-mistakes(review): Enforce aligned composer geometry safely * no-mistakes(review): Make composer ambiguity locale-safe * no-mistakes(review): Preserve ambiguity through composer submission * no-mistakes(review): Carry composer proof through retries * no-mistakes(document): Document structural tmux composer delivery guarantees * no-mistakes: apply CI fixes * feat(bin): add verified pi-signed runtime adapter (#1145) * feat: add verified pi-signed adapter * no-mistakes(review): Correct pi-signed maintainer verification date * no-mistakes(review): Correct remaining pi-signed verification dates * no-mistakes(review): Preserve authoritative pi-signed runtime identity * no-mistakes(document): Document pi-signed shared adapter semantics * no-mistakes: apply CI fixes * fix(pi): rearm watcher across session transitions (#1166) * fix(pi): rearm watcher across same-process session transitions Pi emits session_shutdown for ordinary /new, /resume, and /fork replacement as well as terminal quit. The primary watcher extension latched a module-level stopping flag on every shutdown, so a replacement session in the same process could not arm monitoring until Pi restarted. Own arm authority per session generation so only the active live generation may start, stop, or rearm the child. Replacement sessions can arm again without restarting Pi, stale prior-generation callbacks cannot mutate the active cycle, and real quit still blocks late rearm. * no-mistakes(review): Preserve Pi generation isolation and exit cleanup * no-mistakes(document): Correct Pi watcher transition documentation * feat: route crew dispatch using quota-window pace (#1172) * Consume quota-axi pace signals in dispatch profile array selection. Add quota-array-dispatch as the single owner of the pace-aware candidate choice, keep AGENTS.md to the intake boundary and load trigger, and cover the acceptance cases with sanitized schemaVersion 3 fixtures. * no-mistakes(review): Stop and report genuine quota dispatch ties * no-mistakes(document): Document quota pace freshness and uncertainty * fix: adapt Grok Stop continuation and harden endpoint cleanup (#1171) * fix(grok): adapt Stop continuation to runtime capability * no-mistakes(review): Reject ambiguous Grok Stop payloads * no-mistakes(review): Reject duplicate Grok fields and accept spaced tmux sessions * no-mistakes(review): Enforce exact tmux cleanup selectors * no-mistakes(test): Fix historical tmux fixture and validate Grok Stop * no-mistakes: apply CI fixes * fix: restore stock macOS Bash 3.2 brief scaffolding (#1093) * fix(brief): make DOD scaffolding parse-safe on stock macOS Bash 3.2 fm-brief.sh built each Definition-of-done block and the not-enabled Herdr declaration with `VAR=$(cat <<EOF ... EOF)`. On Bash 3.2 (macOS /bin/bash) the lexer scans for the command substitution's closing `)` textually and tracks quote state through the heredoc body, so a single apostrophe, unbalanced quote, or unbalanced paren in that prose breaks parsing of the whole script. Every ship-brief scaffold (no-mistakes, direct-PR, local-only) failed with `unexpected EOF while looking for matching )`. Bash 4+ parses it fine, so the breakage stayed invisible everywhere except stock macOS. Replace all four command-substitution heredocs with `IFS= read -r -d '' VAR <<EOF || true`. That removes the `$(...)` wrapper and the entire defect class regardless of future prose, and preserves the variable expansion the direct-PR and local-only bodies need. `read` keeps the heredoc's trailing newline that `$(...)` used to strip, so trim one newline to keep every generated brief byte-identical to prior output. Guard the structure, not one historical phrase: a new test rejects any heredoc nested in a command substitution anywhere in fm-brief.sh, where the old assertion pinned a single apostrophe phrase and so missed the reintroduction. Extend the stock-macOS Bash CI job from parsing one script to the whole maintained shell surface (bin/*.sh, bin/backends/*.sh, tests/*.sh), matching bin/fm-lint.sh's canonical file set so parse scope and lint scope cannot drift apart. * no-mistakes(review): Captain: harden Bash structure and inventory guards * no-mistakes(document): Align stock macOS Bash contributor checks * no-mistakes(lint): Suppress deliberate SC2016 literal fixture warnings * test: stabilize tmux teardown conformance baseline (#1209) * fix(test): pin teardown tmux baseline to historical kill selectors merge-base HEAD main collapses to HEAD after the exact-selector change lands on the default branch, so the old teardown fixture was accidentally exercising current exact targets. Resolve a content-historical permissive tmux adapter from first-parent history and force that post-squash topology inside the conformance case so main and feature branches keep the same old-vs-new contract. * no-mistakes(lint): Suppress intentional literal-pattern ShellCheck warnings * docs: slim quota-array-dispatch to the pace selection core (#1197) Cut the runtime skill to the compact pace-aware selection procedure plus minimum owner pointers. Keep every distinct decision rule and move expanded acceptance scenarios to deterministic fixture ownership assertions. Size: 170/1374/10187 -> 63/544/4068 (about 63%/60%/60% reduction). * feat(bin): inherit backend config into secondmate homes (#1219) * Inherit config/backend into secondmate homes with deliberate-override preservation Add backend to the shared inheritable config allowlist so launch, locked bootstrap, and config-push converge a primary pin into secondmate homes as each home local future-spawn default. Track last-inherited bytes in a private state provenance marker so deliberate per-home overrides survive present and absent primary convergence, keep --backend and FM_BACKEND stronger, and extend the existing inheritance tests plus docs and skill claims. * no-mistakes(review): Preserve equal unprovenanced backend overrides * no-mistakes(review): Preserve symlink overrides and verify spawn precedence * no-mistakes(review): Snapshot backend inheritance for consistent provenance * no-mistakes(review): Simplify backend inheritance to primary-authoritative convergence * no-mistakes(document): Document inherited backend override preservation * fix: restore primary-authoritative backend inheritance after document regression The document step reintroduced provenance and deliberate per-home override semantics after review had simplified config/backend to plain primary-authoritative allowlist membership. Restore the primary-always-wins path: present overwrites, absent removes, no provenance marker, and docs/tests match that contract. * no-mistakes(review): Add divergent backend precedence regression fixtures * no-mistakes(document): Document backend inheritance contract * fix(pi): remove Calm's upper version ceiling (#1226) * fix(pi): remove Calm's exclusive Pi upper-version ceiling tests/fm-calm-pi-extension.test.sh gated on a closed PI_COMPAT_VERSIONS allowlist ("0.81.1 0.82.0") that refused any other installed Pi, and docs described that range as "supported" rather than verified evidence. The Calm CHANGELOG shows no API introduced at either version, so there is no evidence for a real minimum; the presentation adapters already probe the exact method they patch rather than checking a version. Replace the allowlist with dated version evidence that never rejects a newer Pi, and make each presentation adapter degrade independently with a diagnostic if a future Pi removes its API, instead of the whole Calm extension failing to load. Rewrite the feasibility doc's "Pi 0.81.1 through 0.82.0" phrasing to state it as verified evidence, not a ceiling. * no-mistakes(review): Probe missing Calm adapter exports safely * no-mistakes(document): Document Calm's unbounded Pi compatibility * fix(bin): allow session-local todo tools in the subagent guard (#1204) * fix(guard): allow session-local todo tools in the primary The delegation-shape guard denied TaskCreate and TaskUpdate because their normalized names contain the `task` stem. Those tools write only the harness's session-local todo list, which has no executor: it spawns no agent, allocates no worktree, registers no schedule, and starts nothing that outlives the session. That is not the unaccounted work the guard exists to stop, so the stem match was a false positive, and the deny text told the primary to run bin/fm-brief.sh and bin/fm-spawn.sh to create a todo entry. Add a separately-reasoned PLAN_ONLY_TOOLS exact-name exclusion rather than widening OBSERVE_ONLY_TOOLS, whose documented contract is tools that only observe or stop existing work. Both lists stay exact-name so neither can widen by substring. Tests cover the two allowed names and six near-miss names that a substring or shortened-stem widening would release; both mutations were watched red. * no-mistakes(review): drop session-local todo tools from recommended deny list * no-mistakes: apply CI fixes * fix(session-lock): resolve Claude bg-spare ancestry to the outermost claude pid (#1206) * fix(session-lock): resolve Claude bg-spare ancestry to the outermost claude pid fm_harness_ancestry_pid() previously returned the first ancestor process whose command matched a verified harness name. Claude Code's Stop hook fires as a bg-spare worker several levels below the session's actual lock-owning claude process (hook shell -> claude bg-spare -> claude bg-pty-host -> claude -> claude(lock)), so the first match was the bg-spare worker, not the lock owner. fm_session_lock_owned_by_self() then never matched state/.lock, and the Claude Stop auto-arm silently treated its own primary session as an unrelated live owner and never armed the watcher. The walk now keeps going past a claude-named match, looking for a still more ancestral claude-named match, and stops the instant a non-match follows an already-found match (bounding it to a contiguous run rather than the literal ancestry top, so an unrelated claude-named process further up the real process tree is never mistaken for part of this session's own nested chain). Every other harness keeps the original first-match-wins behavior, since e.g. Pi's shared signed-wrapper ancestry actually holds the session at the inner engine pid, not an outer wrapper pid. Hop limit raised from 8 to 16 to cover the deeper bg-spare chain. * no-mistakes(review): Add nested-claude-ancestry regression test; fix nudge doc depth claim * no-mistakes: apply CI fixes * fix: conferma l'avvio del watcher su Windows/MSYS (#1212) * fix: confirm watcher startup on MSYS * no-mistakes(review): gate MSYS arm ready timeout, cache uname, harden locale test * no-mistakes(review): validate OpenCode ready timeout, make uname cache internal * fix(spawn): forward CLAUDE_CONFIG_DIR to claude crewmates (#1195) * fix(spawn): forward firstmate's CLAUDE_CONFIG_DIR to claude crewmates Crewmate panes are created by a long-lived tmux/herdr daemon that does not inherit firstmate's current environment. When firstmate runs under a non-default CLAUDE_CONFIG_DIR (for example a work-vs-personal subscription split), a bare `claude` in the crewmate pane fell back to the default ~/.claude store and launched unauthenticated, blocking the crewmate before it could do any work. fm-spawn now prefixes the claude launch with firstmate's own resolved CLAUDE_CONFIG_DIR when set, so the crewmate uses the same credential/config store firstmate is authenticated with. An unset value is the single-store default and adds no prefix; non-claude harnesses are unaffected. Adds three tests in fm-spawn-dispatch-profile.test.sh (forwarded-when-set, omitted-when-unset, non-claude-ignored) and pins CLAUDE_CONFIG_DIR in the test helper so launch assertions no longer depend on the developer's environment. * no-mistakes: apply CI fixes * fix: preserve dispatch identity across authentication checks (#1233) * fix: preserve dispatch harness identity * no-mistakes(review): Fix Grok counterfactual tuple validation * no-mistakes(document): Scope dispatch authentication to selected tuple * fix: restore dispatch instruction budget * no-mistakes(review): Scope dispatch authentication after candidate selection * fix(bin): normalize relative durable paths (#1256) * fix(bin): handle dash-leading harness process names (#2) * fix: handle dash-leading harness process names * no-mistakes(review): Make dash-leading harness regression hermetic * fix: preserve secondmate reply routes across relative homes Resolve relative home, data, and state inputs before durable charter generation, and fail when caller-relative directories cannot be resolved. Use absolute paths at the related spawn, AFK daemon, and X-mode cross-process handoffs so later processes cannot reinterpret them from another working directory. * no-mistakes(review): Preserve absolute overrides and normalize relative durable paths * no-mistakes(review): Normalize relative home before deriving durable paths * no-mistakes(document): Document relative durable-path normalization * no-mistakes(review): Captain: Ignore inherited CDPATH during relative path normalization * no-mistakes(lint): Fix empty CDPATH assignments for ShellCheck * refactor(skills): make Bearings chat-only by default (#1136) * Add internal status skill * no-mistakes(document): register /status skill in documentation-audiences inventory * no-mistakes(lint): replace grep|wc -l with grep -c in status skill test * test: silence literal status skill patterns * Refactor bearings default to chat-only --------- Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com> * Clarify follow-up routing during validation (#1277) * fix: honor concrete approval for project operations (#1272) * docs: add captain-approved project operation exception to hard rule 1 Firstmate stays read-only over projects by default, but when the captain clearly approves a concrete project operation and scope in the moment, firstmate may perform exactly that approved operation with its own tools. The approval is never inferred, broadened, or standing, and it does not relax the existing force, discard, unlanded-work, or merge-authority boundaries. * no-mistakes(review): Clarify captain-approved project operation boundaries * no-mistakes(document): Clarify captain-approved project operation scope * docs: cover directories and preserve the operation-or-scope alternative Widen the captain-approved project operation exception in AGENTS.md to files or directories, and restore the explicit operation-or-scope alternative that a prior pipeline auto-fix had collapsed into "and". Rework project-management SKILL.md's Remove section, which previously told firstmate to refuse project removal until a guarded helper existed; that helper was never built, so the text directly contradicted the new instruction-only exception. It now points at the exception plus the existing removal preflight it still requires unchanged. Update the one instruction-owners test assertion that hard-coded the sentence removed above, so the suite tracks current, not obsolete, text. * docs: add captain-approved project operation exception to hard rule 1 Firstmate stays read-only over projects by default, but when the captain clearly approves a concrete project operation and scope in the moment, firstmate may perform exactly that approved operation with its own tools. The approval is never inferred, broadened, or standing, and it does not relax the existing force, discard, unlanded-work, or merge-authority boundaries. * no-mistakes(review): Clarify captain-approved project operation boundaries * no-mistakes(document): Clarify captain-approved project operation scope * docs: cover directories and preserve the operation-or-scope alternative Widen the captain-approved project operation exception in AGENTS.md to files or directories, and restore the explicit operation-or-scope alternative that a prior pipeline auto-fix had collapsed into "and". Rework project-management SKILL.md's Remove section, which previously told firstmate to refuse project removal until a guarded helper existed; that helper was never built, so the text directly contradicted the new instruction-only exception. It now points at the exception plus the existing removal preflight it still requires unchanged. Update the one instruction-owners test assertion that hard-coded the sentence removed above, so the suite tracks current, not obsolete, text. * no-mistakes(review): Align project removal preflight with approved exception * no-mistakes(document): Align project removal documentation with approved exception * fix: restore removal test byte-for-byte and preserve the default sentence tests/fm-instruction-owners.test.sh had been changed to assert different text; restore it byte-for-byte to origin/main. project-management SKILL.md's Remove section now keeps the exact default "Never issue a raw removal command from Firstmate." sentence that test still asserts, immediately followed by the already-approved captain-operation-or-scope exception, so the default and the exception both stay explicit and consistent. * no-mistakes(document): Align project-write boundary documentation * fix(skills): route new project intake through secondmate scopes (#1275) * Route project intake through secondmate scopes * no-mistakes(test): Guard all main-home project registry mutations * no-mistakes(document): Consolidate secondmate routing documentation * no-mistakes: apply CI fixes * Restore new-project routing scope * no-mistakes(document): Clarify secondmate routing for new-project intake * no-mistakes: apply CI fixes * fix: scope validation corrections by accepted behavior (#1281) * fix: scope validation corrections by accepted behavior * no-mistakes(review): Classify stale delivery evidence as an autonomous correction * test: replace source assertions with behavioral coverage (#1282) * test: remove source-content assertions * no-mistakes(review): Replace source assertions with runtime behavior coverage * no-mistakes(review): Isolate Kimi task temp runtime coverage * no-mistakes(document): Refresh test cleanup documentation * no-mistakes: apply CI fixes * fix(watch): escalate busy workers with no completed turn (#1286) * fix(watch): bound how long a busy pane may run with no completed turn A busy pane (backend busy state or the harness's rendered footer) was unconditional, unbounded proof of liveness in every escalation path, so a hung foreground tool call behind a busy signature could run for hours undetected (2026-07 hibit-agent-focus-nonsteal-r1 incident: a catastrophic- backtracking regex hung one bash call for 25h behind an unchanging "Working..." footer). FM_BUSY_TURN_MAX_SECS (default 3600s) now bounds how long a busy pane may run with no completed turn (state/<id>.turn-ended, or its spawn record before any turn has completed). Past the bound, busy_turn_over_age routes the pane through the existing wedge_timer_check, reusing the identical stale reason, escalation counter, and demand-deep-inspection marker for human inspection only - never an automatic interrupt, signal, or restart of the worker or its tool process. A completed turn resets the age. Reproduced end-to-end against the real installed Pi TUI: a foreground `sleep 999999` bash call with no timeout renders the actual busy footer, and two captures ~15s apart show the elapsed counter changing the pane hash while the same turn stays unfinished. Running the pre-fix watcher against the real captures showed it never starts a wedge timer no matter how long the pane stays busy; the fixed watcher starts and escalates the timer through the same mechanism, while the real hung process remained untouched and alive throughout. * no-mistakes(review): fix: parse enriched AFK stale reasons * no-mistakes(review): fix: preserve enriched wedges during AFK supervision * no-mistakes(review): fix: route all enriched AFK wedges * no-mistakes(document): Clarify busy-turn age supervision documentation * fix(gitignore): ignore config/ as a directory, not by exact filename (#1261) A name-by-name list of config/ entries silently stops ignoring any new or home-local file placed there, which makes the working tree read as dirty and blocks guarded sync paths that refuse to touch a dirty home. AGENTS.md already documents config/ as captain-private and gitignored as a category; this makes .gitignore match that contract. * fix(tests): replace source-content .gitignore assertion with behavioral coverage (#1304) The second assertion in fm-gitignore-config.test.sh (added by #1261) greps .gitignore for a specific spelling of the config/ ignore pattern. It fails on a semantically equivalent pattern like config/** and does not prove Git actually ignores anything, per the completed source-content-test audit. Replace it with a real git check-ignore control test on a generated unrelated path, and strengthen the existing directory-coverage test with generated unpredictable direct and nested config/ paths. * feat: bound and consolidate startup memory during stow (#1303) * Add bounded startup memory curation * no-mistakes(review): Record reproducible stow verification evidence * no-mistakes(review): Validate inherited secondmate stow evidence * no-mistakes(document): Do…
DereKk8
added a commit
to DereKk8/firstmate
that referenced
this pull request
Aug 15, 2026
* fix: reconcile existing AGENTS.md safely (#405)
* fix(agents-md): inject self-governance section into existing AGENTS.md
fm-ensure-agents-md.sh only appended the canonical "## Maintaining this
file" section on skeleton create or CLAUDE.md promotion, so an existing
AGENTS.md that lacked it exited unchanged and forced hand-copying the
wording during a rollout across existing projects. Call the already-
idempotent ensure_maintenance_section on the existing-AGENTS.md paths and
report whether the file changed; a re-run and an already-complete file stay
byte-identical.
Also fixes #389: refuse a case-variant real memory file (e.g. a lowercase
agents.md) instead of silently emitting a CLAUDE.md symlink whose uppercase
literal target dangles once the tree lands on a case-sensitive filesystem.
Tests extend tests/fm-ensure-agents-md.test.sh; skeleton-create and
CLAUDE.md-promotion regressions still pass. Docs updated to match.
* no-mistakes(review): Captain: preserve CRLF maintenance-section idempotency
* no-mistakes(review): Preserve CRLF during maintenance-section injection
* no-mistakes(review): Captain: harden dangling-symlink regression coverage
* no-mistakes(document): Document agent-memory injection outcomes
* feat: support project-less secondmate homes (#409)
* feat(secondmate): support project-less homes via --no-projects
fm-brief.sh --secondmate and fm-home-seed.sh now accept an explicit
--no-projects signal to scaffold, seed, and register a secondmate home
whose subject is the firstmate repo itself (no clones). The signal is
mutually exclusive with a project list; omitting both still fails loudly
so an accidental omission is never a silent project-less seed. The
registry line renders an empty projects: field, which spawn and the
snapshot already tolerate. Docs updated in the secondmate-provisioning
skill and both script headers.
* no-mistakes(review): Captain: document project-less secondmate flow
* no-mistakes(review): Captain: refuse project-less reseeding of populated homes
* fix(seed): fail closed on unreadable project data
* no-mistakes(review): Captain: reject stale projectful charters
* no-mistakes(review): Captain: fail closed on unsafe project paths
* no-mistakes(review): Captain: validate project-less charter clone sections
* no-mistakes(document): Document project-less secondmate seeding
* fix: delegate backlog handoffs to tasks-axi (#411)
* wip(handoff): record verified delegation design + tasks-axi mv blocker
No production code changed yet. tasks-axi mv (v0.2.1) cannot atomically
move a blocked-by-linked item set across backlogs (deadlocks both orders,
no batch/--force), which fm-secondmate-lifecycle-e2e requires. Parked
pending a tasks-axi connected-set mv enhancement; note captures the
verified design, semantics, test/CI/doc changes, and resume checklist.
* refactor(handoff): delegate the item move to tasks-axi mv
fm-backlog-handoff.sh's two-pass awk was a second parser of the backlog
format and the source of the PR #401 body-orphaning drift. Delete it and
delegate the move to `tasks-axi mv <id>... --to <dest>` (v0.2.2 atomic
multi-id), the single owner of the format: a connected set (blocker plus
dependents) moves together with blocked-by preserved, item blocks stay
byte-exact, and destination section placement holds. The helper keeps only
the fleet-level validation tasks-axi cannot know - secondmate-home
resolution, the seeded-home safety checks, the In-flight refusal, and
idempotent per-key reporting - and is atomic: on any move failure nothing
moves.
Tests: fm-backlog-handoff.test.sh keeps PR #401's regression matrix but now
exercises the delegated path and skips cleanly when tasks-axi is absent; the
two whole-file fixtures move to tasks-axi's canonical whitespace. The
lifecycle-e2e and safety move-cases gain the same skip guard. CI installs
tasks-axi so the delegated path is exercised. Docs state that
config/backlog-backend=manual governs firstmate's own hand-editing, not this
validated helper, which delegates fleet-wide because bootstrap requires
tasks-axi on PATH.
Remove the now-redundant WIP design note.
* no-mistakes(review): Captain: harden atomic backlog handoffs
* no-mistakes(review): Captain: enforce queued-only backlog handoffs
* no-mistakes(review): Captain: harden handoff section parsing
* no-mistakes(document): Document delegated backlog handoffs
* no-mistakes(lint): Silence ShellCheck source diagnostics
* fix: ignore secondmate home marker during sync (#417)
* fix: gitignore the secondmate home marker
bin/fm-home-seed.sh writes an untracked .fm-secondmate-home marker into
every seeded secondmate home. A secondmate home is a worktree of the
firstmate repo, so any plain `git status --porcelain` dirtiness check
counted the untracked marker and the home read as dirty forever:
fleet-sync reported it STUCK and the local fast-forward convergence
sweeps risked leaving it stale on firstmate updates.
Add .fm-secondmate-home to the tracked .gitignore so the marker is
invisible to every dirtiness check uniformly, without weakening
fleet-sync's deliberate untracked-counting for project clones.
Convergence chicken-and-egg: existing homes predate the fix and it only
arrives by fast-forward. The already-present marker-tolerant ff-skip
(ignore_seed_marker=yes, used by the bootstrap sweep, /updatefirstmate,
and spawn pre-launch) advances such a home past the fix commit, after
which .gitignore takes over - no hand intervention.
Tests in tests/fm-secondmate-sync.test.sh cover a freshly seeded home
reading clean, an existing marker-only home converging then reading
clean, and a genuinely dirty home still skipping.
* no-mistakes(review): Captain: document standalone-clone update path
* no-mistakes(document): Document secondmate marker migration
* fix(composer): prevent dead-shell message injection (#416)
* fix(composer): stop reading dead-shell prompts as empty agent composers
Consolidate composer empty/pending/unknown classification into one shared
owner, bin/fm-composer-lib.sh's fm_composer_classify_content, delegated to by
all four backend adapters (tmux via fm-tmux-lib.sh, herdr, orca, cmux). This
replaces four drifting copies of the glyph decision.
Safety fix: a bare shell prompt glyph (> $ % #) on an unstructured row is now
classified unknown (a dead shell, unsafe for injection), not empty. It is only
empty inside a bordered composer box (the harness's own prompt). Agent glyphs
❯ (claude) and › (codex) read empty either way. The away-mode injector
(inject_msg) now requires an affirmatively-empty composer, deferring on pending
or unknown, so an escalation can never be typed into (or executed by) a pane
whose agent exited to its login shell.
Regression coverage: new tests/fm-composer-lib.test.sh pins the shared owner;
per-backend dead-shell tests in fm-daemon (tmux + injector), orca, and the
existing herdr/cmux suites. shellcheck clean; herdr incident regressions stay
green.
* no-mistakes(review): Captain: harden composer safety checks
* no-mistakes(test): Stabilize Herdr prune safety setup
* no-mistakes(document): Document composer injection safety
* no-mistakes(lint): Clean composer safety lint
* no-mistakes: apply CI fixes
* feat(watcher): add paused external-wait supervision (#421)
* feat(watcher): add paused/awaiting-external crew state
A crew (or firstmate steering it) can declare a deliberate wait on a known
external dependency with a paused: <reason> status. Both the always-on watcher
and the away-mode daemon absorb such an idle pane through shared fm-classify-lib.sh
vocabulary instead of tripping the possible-wedge stale escalation, and re-surface
it for a recheck only on a long bounded cadence (FM_PAUSE_RESURFACE_SECS) so a
forgotten pause cannot rot invisibly. fm-crew-state.sh reports state: paused
distinctly. A crew that goes idle without declaring a pause classifies exactly as
before. Docs and brief scaffold state lists updated; tests colocated.
* no-mistakes(review): Captain: fix paused-state transitions
* init
* no-mistakes(review): Captain: fix paused-state supervision transitions
* no-mistakes(review): Captain: fix paused supervision handoffs
* no-mistakes(review): Reconcile paused supervision markers
* no-mistakes(review): Captain: prioritize paused states over captain relevance
* no-mistakes(review): Captain: preserve paused-working wedge timer
* no-mistakes(review): Captain: honor configured pause verb in briefs
* no-mistakes(test): Captain: fix AFK paused watcher handoff
* no-mistakes(document): Document declared external waits
* no-mistakes(lint): Clean paused-state lint
---------
Co-authored-by: fmtest <fmtest@example.invalid>
* fix: preserve X-mode follow-up platform limits (#425)
* fix(x-mode): make follow-up platform splitting immune to link ordering
A ~470-char Discord follow-up posted as a (1/2)(2/2) thread split at ~280
chars because fm-x-link only learned the platform from the inbox payload,
and the fmx-respond ack path can drain that inbox file before the task is
linked. A link recorded after cleanup silently lost the platform and the
splitter defaulted to the X 280-char budget.
Make platform resolution ordering-proof:
- fm-x-link now resolves the platform AUTHORITATIVELY by request_id via a
new fmx_request_relay_context helper (POST /connector/request-context)
when neither the inbox payload nor carry flags carry it. The request_id
survives the inbox drain, so a post-cleanup link still learns the right
split budget. Best-effort: no token/curl or a non-2xx relay degrades to
the loud warning below rather than a silent X default.
- fm-x-link warns loudly when no platform source resolves, so the loss is
never silent.
- The fmx-respond procedure now orders link-before-inbox-cleanup so the
fast local path stays correct without a relay round-trip.
Colocated regression tests: a Discord follow-up >280 <2000 posts as ONE
message even when linked after inbox cleanup, and an unresolvable platform
warns loudly instead of splitting silently. docs/configuration.md documents
the request-context lookup.
The relay endpoint is the companion durable change (see done status); until
it ships, the link-before-cleanup reorder keeps the normal path correct.
* no-mistakes(document): Document X-mode platform recovery
* fix(composer): handle ANSI ghost text safely (#429)
* fix(composer): one ANSI-aware ghost owner covers claude dim + grok truecolor
Away-mode injection wedged all night on the primary claude-on-herdr pane:
the herdr composer classifier never stripped generic dim ghost text (only a
narrow codex bold-wrapped byte-pattern check), so claude's rotating
prompt-suggestion ghost - a bare "❯" then SGR-2 dim text, which herdr's ANSI
pane read preserves - read as real pending input and every escalation deferred
(6524 lifetime "pending input (non-empty composer)" defers; wedge 30623s).
Consolidate ghost extraction into one fleet-wide ANSI-aware owner,
fm_composer_strip_ghost (bin/fm-composer-lib.sh), that drops every
de-emphasised run - dim/faint (SGR 2: claude, codex) AND a dark/muted truecolor
foreground (grok's placeholder, luminance below FM_COMPOSER_GHOST_LUMA_MAX,
default 128, dark-theme assumption). Both ANSI-capable backends route through
it: fm_tmux_composer_state (fm_tmux_strip_ghost is now a thin adapter) and
fm_backend_herdr_composer_state. The herdr-only faint byte-pattern check is
removed and fm_backend_herdr_strip_ansi reduced to a thin adapter over the
shared fm_composer_strip_ansi. Bordered detection now reads the plain row so a
dark box border dropped with the ghost does not lose the composer shape.
This also closes the documented grok TRUECOLOR placeholder gap by the same
mechanism (harness-adapters skill note updated).
Empirical evidence (read-only live capture + isolated tmux, no herdr lifecycle)
and the incident write-up are in docs/herdr-backend.md; deterministic
regressions feed the exact captured bytes through the real classifiers
(tests/fm-backend-herdr.test.sh, tests/fm-composer-ghost.test.sh). Two prior
ghost-test fixtures that used a near-black 38;2;1;2;3 as "real" colored text
(never a realistic real-input color) are corrected to a bright 38;2;224;222;244,
preserving the truecolor payload-skip parser intent.
* no-mistakes(review): Preserve dark shell prompt safety
* no-mistakes(review): Harden erased shell prompt classification
* no-mistakes(document): Document shared composer ghost extraction
* no-mistakes(lint): Normalize tmux comment punctuation
* fix(spawn): make tmux window handling robust under non-default config (#134)
* test: isolate session-start suite from ambient harness markers (#432)
* fix(session-start): isolate harness env markers in suite runner
Neutralize CLAUDECODE, PI_CODING_AGENT, and GROK_AGENT in
run_session_start so ambient interactive shells cannot override the
suite's fake ps harness (local-vs-CI split on the pi supervision case).
* no-mistakes(document): Correct Pi marker documentation
* fix(teardown): retry transient index locks during worktree return (#435)
* fix(teardown): retry treehouse return on transient index.lock
Killed crew git ops can leave a short-lived worktree index.lock that
makes treehouse return fail. Retry on that error signature with a
bounded wait (env-overridable), never force-delete a live lock, and
only then fall back to the existing provably-stale cleanup path.
* no-mistakes(review): Harden teardown retry configuration
* no-mistakes(document): Document teardown index-lock retry behavior
* no-mistakes(lint): Fix empty shell variable assignments
* fix: complete brief help and consolidate documentation (#438)
* docs: de-feature the scripts.md and CONTRIBUTING test inventories
Slice 1 of the documentation redundancy cleanup wave (firstmate scope).
docs/scripts.md: every row is now one purpose clause; script headers
are the declared owner of behavior, flags, and contracts. Coverage
stays 61/61 scripts; bytes drop 19,922 -> 7,958.
CONTRIBUTING.md: the 54-row per-test inventory is gone; contributors
discover tests by listing tests/*.test.sh and reading each script's
own header, and gated tests print their own skip gates. The run
commands, symlink assertions, and watcher smoke line are unchanged.
Lines drop 135 -> 84 (18,797 -> 7,831 bytes).
Two facts that existed only as inventory rows moved into their
owners' headers first: fm-brief.sh's paused-vs-blocked scaffold
distinction and fm-session-start.sh's Pi extension-loaded check.
No instruction-surface or behavior change; AGENTS.md untouched.
* no-mistakes(review): Captain, fix brief help and Grok test discovery
* no-mistakes(review): Captain: document Grok lock-holder test coverage
* fix: detect Git and centralize backend configuration (#445)
* docs: consolidate universal backend contracts into configuration.md
Slice 2 of the documentation redundancy cleanup wave (firstmate scope).
docs/configuration.md is now the declared single owner of three
universal contracts, each with an explicit ownership sentence:
- the universal toolchain list (Toolchain), now also carrying the
per-tool purpose clauses that previously lived only in the tmux guide;
- the task-selector vocabulary (Runtime backend);
- the tasks-axi compatibility definition (Backlog backend).
The five backend guides' prerequisites replace their verbatim
universal-requirements parentheticals (5 full copies) with a pointer
plus only backend-specific items; zellij/cmux selector restatements
and architecture.md's partial copy become pointers or are dropped;
CONTRIBUTING's compatibility sentence becomes a pointer; two
near-verbatim orca-bootstrap restatements (configuration.md Runtime
backend, orca guide) collapse into the Toolchain owner copy.
Backend-specific setup, behavior, target-string shapes, and every
empirical verification record are untouched. AGENTS.md untouched
(slice 3).
* docs: include git and GitHub auth in the toolchain owner list
The review flagged that the new universal-toolchain owner omitted git
and GitHub authentication while every backend guide now defers its
prerequisites here; bootstrap's NEEDS_GH_AUTH check makes them real
universal requirements.
* no-mistakes(review): Detect Git in bootstrap toolchain
* no-mistakes(document): Clarify GitHub CLI and centralize selector documentation
* feat(daemon): add backend-independent wedge alerts (#444)
* feat(daemon): backend-independent active alert for the wedge alarm
When away-mode injection wedges past max-defer, inject_wedge_alarm only
actively signalled via the tmux status-line, which is skipped on non-tmux
backends. A wedged claude-on-herdr primary left only the passive
state/.subsuper-inject-wedged marker (2026-07-10 overnight incident).
Add a config-gated active alert (config/wedge-alarm, local/gitignored;
FM_WEDGE_ALARM_CHANNEL) that reaches the captain even when every pane and
its status-line is unreadable: an OS-level macOS notification (osascript),
a herdr notification, or a captain-supplied command. Default-on (auto) so
the alarm is never silent; each channel best-effort, degrading to the next
and never crashing the daemon loop. The tmux flash and durable marker stay.
The OS notifiers route through a single FM_WEDGE_ALARM_EXEC seam. When the
daemon is sourced (only tests do this; production execs it) the seam
defaults to "discard", and tests/wake-helpers.sh points it at a recorder,
so it is structurally impossible for any test to post a real notification.
Channels verified once manually on macOS 26.5.2 / herdr 0.7.3; see
docs/wedge-alarm.md.
* no-mistakes(review): Bound wedge alarm notifier execution
* no-mistakes(review): Captain: harden wedge alarm notifier safety
* no-mistakes(review): Captain: harden wedge alarm test notifier isolation
* no-mistakes(review): Captain: harden wedge alarm throttling
* no-mistakes(review): Redact wedge alarm directive logs
* no-mistakes(review): Harden wedge alarm notifier safety
* no-mistakes(review): Track notifier process groups through cleanup
* no-mistakes(document): Document wedge-alarm active alert behavior
* docs: centralize firstmate operating contracts (#447)
* docs(agents): extract conditional AGENTS.md material to owned homes
Slice 3 of the documentation redundancy cleanup wave (firstmate scope):
the always-loaded instruction surface drops from 941 lines / 116,733
bytes (~29k tokens per session per fleet member) to 785 / 91,353
(~22.8k tokens), moving only audit-identified conditional and
situational material while preserving every load-bearing invariant at
its trigger point via the inline-stub pattern.
Moves, each to one declared owner plus an inline stub:
- section 3's bootstrap output-line handbook (~44 lines) -> new
agent-only bootstrap-diagnostics skill, added to the section 13
trigger index; the detect-consent-install rule and the
do-not-dispatch gate stay inline as safety-critical.
- section 4's crew-dispatch JSON schema and field semantics ->
docs/configuration.md 'Crew dispatch profiles' (pointer direction
flipped); the intake procedure, precedence, backstop, and
never-select-unverified rules stay inline.
- section 4's quota-balanced algorithm -> bin/fm-dispatch-select.sh
header (now the declared owner; usage() converted to the dynamic
header extraction pattern PR #438 established for fm-brief.sh).
- section 7's spawn resolution narrative and example sprawl ->
bin/fm-spawn.sh header; the isolated-worktree assertion, refusal-is-
a-blocker rule, and post-spawn duties stay inline.
- section 7's teardown landed-work mechanics -> bin/fm-teardown.sh
header (section 1's containment pointer retargeted); the fork benign
case and never-force rule stay inline.
- section 8's watcher classification narrative -> docs/architecture.md
'Event-driven supervision' (already the owner); every operative rule
(one live cycle, no turn ends blind, drain first, wake ladder,
never-pkill, guard responses) stays inline.
- sections 3/4/6/7 secondmate sync, propagation, schema, and handoff
restatements -> secondmate-provisioning skill, now the declared
owner including the literal-file inheritance nuance.
- section 14's X-mode cadence mechanism -> docs/configuration.md
'X mode (.env)', closing issue #363; activation semantics, the
fmx-respond trigger, and the terminal-wake final-follow-up duty
stay inline.
CLAUDE.md stays a symlink; no behavior or test change.
* no-mistakes(document): Centralize contract-owner documentation
* fix(cmux): close last workspace during teardown (#449)
* fix(cmux): close the last/selected workspace in a window at teardown
cmux keeps every window at >=1 workspace, so close-workspace on the only
workspace in a window silently no-ops (returns OK, workspace stays), and a
window holding a live session cannot be closed over the control socket.
That left a selected task workspace open at teardown (the last workspace
in a window is always the selected one).
Add fm_backend_cmux_window_of_workspace and have fm_backend_cmux_kill
create a throwaway default sibling in the target's window before closing
when the target is the last workspace there, so the close lands; the
window keeps a fresh default workspace (cmux's own "closed the last tab"
outcome). Non-last teardown closes directly, as before.
Cover both kill branches plus the helper with fake-CLI unit tests, add a
real-cmux window/count detection smoke assertion, and record the
empirical evidence in docs/cmux-backend.md.
* no-mistakes(review): Derive cmux count from membership snapshot
* no-mistakes(document): Document cmux last-workspace teardown behavior
* fix: recover orphaned packed-refs locks during fleet sync (#453)
* fix(fleet-sync): recover from an orphaned packed-refs.lock
A git ref rewrite (fetch --prune, pack-refs, branch -D) killed after
creating .git/packed-refs.lock but before renaming it - e.g. bootstrap's
timed-out fleet-sync kill or teardown's process kills - leaves a lock that
makes the next sync's fetch fail with "Unable to create
'...packed-refs.lock': File exists", leaving the clone unsynced.
On that signature only, fm-fleet-sync.sh now retries the fetch with a
bounded wait (transient locks self-clear), then removes the lock and
retries once more ONLY when it is provably stale: still present, mtime
age past a threshold, and no lsof holder of the lock file or of the clone
worktree itself (a live git keeps that as its cwd even in the window after
it closes the lock and before it exits). A live lock, a missing lsof, any
failed check, or any other fetch failure keeps today's behavior. Every
wait/retry/removal prints to stderr, and a successful recovery also prints
one "recovered:" summary to stdout so a session-start refresh - which
discards fleet-sync stderr and relays only stdout - still surfaces it.
The shared "is this git lock provably abandoned?" proof is extracted into
bin/fm-lock-lib.sh so it has one owner, used by both fm-teardown.sh and
fm-fleet-sync.sh. Constants are env-overridable knobs. tests/fm-gotmp.test.sh
gains the fm-lock-lib.sh symlink teardown now needs in its fake bin/.
* no-mistakes(review): Captain, remove obsolete teardown wake dependency
* no-mistakes(document): Document packed-refs lock recovery architecture
* feat(herdr): escalate blocked panes immediately (#472)
* feat(herdr): immediate blocked-state escalation via native events.subscribe push
Fold herdr's native pane.agent_status_changed stream into the single watcher so
a crew entering blocked wakes its supervisor sub-second (measured 0.129s)
instead of after the ~240s stale-pane wedge timer.
- bin/fm-transition-lib.sh: backend-neutral normalized-transition record shape
plus the single-owner status->action policy table (blocked=actionable,
working=absorb+clear-dedupe, idle/done=defer, else=fall back to polling).
- bin/backends/herdr.sh + herdr-eventwait.py: a raw AF_UNIX events.subscribe
subscriber over one connection for all this home's herdr panes, subscribing to
ALL statuses, returning the first fresh blocked edge, with a per-pane dedupe
marker and a reconnect level-reconcile. Version/schema capability gate.
- bin/fm-backend.sh: has-push / events-capable / wait-transition dispatchers so
the watcher stays backend-agnostic and the shape+policy are reusable.
- bin/fm-watch.sh: splice the bounded event wait in as the watcher's terminal
wait primitive (replacing the blind sleep POLL for push-capable homes),
behind a source guard so the splice is unit-testable; secondmate/paused
exemptions; map pane->window->task and enqueue a stale wake. No second
watcher process; the single-cycle invariant and every guard/beacon/turn-end
mechanism are unchanged.
- Polling stays the permanent fail-closed backstop: below-capability, subscribe
failure, and repeated runtime failures all degrade to sleep.
- Tests: fake-CLI units (fm-transition-lib, wait/apply/dedupe/reconcile/
fallbacks in fm-backend-herdr, watcher exemptions in fm-supervision-events)
plus an isolated real-herdr idle->blocked smoke. docs/herdr-backend.md carries
the dated evidence and retires the old gap note.
* no-mistakes(review): Captain, fix Herdr disconnect handling and dedupe docs
* no-mistakes(review): Captain, commit markers after wake and reuse capability cache
* no-mistakes(review): Captain, clear stale markers and secure Herdr FIFOs
* no-mistakes(review): Captain, subscribe before Herdr reconciliation
* no-mistakes(review): Captain, make Herdr FIFO handling Bash 3.2-safe
* no-mistakes(test): Captain: include lock library in teardown fixture
* no-mistakes(document): Captain: document Herdr immediate blocked escalation
* fix: clarify shellcheck conditionals
* docs(readme): reposition firstmate as an agent distro (#473)
* feat: add deterministic bounded bearings snapshots (#475)
* feat(bearings): deterministic bearings snapshot + durable decision model
Add bin/fm-bearings-snapshot.sh: a bounded TOON-by-default projection over the
canonical fm-fleet-snapshot. Default is local-only (zero network); live open-PR
discovery and checks happen only under --include-prs, which fails soft. Every
dropped surface is marked in omitted[] with the flag that reveals it, and the
prs: line states when checks were not requested, so absence is never silent.
Fix the unresolved-decision masking bug in the canonical layer. fm-classify-lib
gains status_open_decisions, the one authoritative keyed open/resolved fold over
the whole status stream: needs-decision/blocked opens a keyed entry, only an
explicit keyed resolution (or, for run-backed tasks, run-step advancement)
closes it, so a later unrelated done/paused can no longer mask a still-open
captain decision. fm-fleet-snapshot surfaces hints.open_decisions and derives
pending_decision/blocked_event from it; the canonical schema stays complete.
Point the /bearings skill at the one command; add the resolved: writer line to
ship, scout, and secondmate briefs. Register the script and add regression
tests for the output bound, TOON/JSON parity, local-only default, opt-in PR
fetch, partial-failure degradation, decision durability, and report pointers.
* fix(bearings): completed scout report is a pointer, not a pending decision
A completed scout that raised a needs-decision and then finished (done) without
a keyed resolution falsely surfaced as an open/pending decision (the Lavish-103
case). Root cause: the open-decision reconciliation in bin/fm-fleet-snapshot.sh
cleared a stale decision only for a live run-step/pane activity read, so a
terminal task whose current state is read from the status log (a scout or ship
that reached done/failed) never cleared its stale, never-keyed-resolved
needs-decision, and it lingered as pending.
The open-decision set is still derived purely from the keyed fold - never from a
report body or decision-like prose - and reconciled against the crew lifecycle.
Extend that reconciliation so a terminal done/failed state on a single-owner
task (scout or ship), whose deliverable is its report or PR, also clears the set;
a completed scout now surfaces only as a report pointer. Secondmates are excluded
from the terminal clear (persistent, multiplexed stream), which keeps the
unrelated-event masking fix intact. Add regression tests: a completed scout with
decision-like report prose is a pointer not pending (canonical + end-to-end), and
a scout still parked at a decision stays pending so the terminal clear never
over-fires.
* no-mistakes(review): Captain, preserve keyed decisions across shared status parsing
* no-mistakes(review): Captain, close blockers and harden keyed decision parsing
* no-mistakes(review): Captain, bound GitHub enrichment without coreutils timeout
* no-mistakes(review): Captain, bound bearings sections and fail closed
* no-mistakes(review): Captain, disclose capped per-repository PR results
* no-mistakes(document): Refresh bearings documentation and status contracts
* fix(bearings): avoid ambiguous worktree guard
* fix: enforce deterministic ShellCheck parity (#481)
* fix(lint): one shellcheck owner pinned to 0.11.0 for CI/local parity
Firstmate PRs passed local no-mistakes validation but failed CI's
"Lint shell scripts" job on shellcheck findings (SC2015, SC1007, SC2034).
Two divergences caused it:
1. The no-mistakes gate had no commands.lint, so its lint step never ran
the deterministic shellcheck bin/*.sh bin/backends/*.sh tests/*.sh that
CI runs. Confirmed from state.sqlite: the lint step_result recorded
findings:null with no lint agent invocation.
2. CI's shellcheck floated with the runner image while local ran a newer
build; shellcheck retired SC2015 in 0.11.0, so an older CI shellcheck
rejected an SC2015 that the newer local one no longer emits.
Establish bin/fm-lint.sh as the single owner of the lint definition: the
file set, the config, and the pinned shellcheck version (0.11.0, printed
via --required-version). Both CI (.github/workflows/ci.yml) and the
no-mistakes gate (.no-mistakes.yaml commands.lint) invoke it; CI installs
the exact version it names and logs the resolved version, and fm-lint.sh
refuses to lint under any other version. This is not a CI relaxation: it
adopts shellcheck 0.11.0's rule set consistently, dropping only the
upstream-retired, false-positive-prone SC2015; default severity and every
still-supported finding stay enforced (no severity downgrade, no excludes).
tests/fm-lint.test.sh asserts both gates invoke the owner, that CI installs
and logs the pinned version, that the owner refuses a non-pinned shellcheck,
and that it rejects a real lint defect the old no-op gate passed.
* no-mistakes(review): Captain, harden deterministic ShellCheck parity
* no-mistakes(review): Captain, neutralize ambient ShellCheck overrides
* feat: guard primary shells from persistent cd commands (#483)
* feat: add cd-guard PreToolUse seatbelt for the primary shell
A stray persistent top-level `cd projects/<clone>` in the primary firstmate
shell relocates the shell, so a later firstmate-owned command (a backlog write,
an fm-* lifecycle call, tasks-axi) runs inside a project clone instead of the
home. The cd-guard denies exactly that command shape before it runs, across all
five verified primary harnesses, mirroring the watcher-arm PreToolUse seatbelt.
- bin/fm-cd-command-policy.mjs: sole block/allow decision owner. Reuses the
shell classifier exported from bin/fm-arm-command-policy.mjs (no duplicate
lexer; that file's CLI now runs only when invoked directly).
- bin/fm-cd-pretool-check.sh: transport, strict-superset prefilter, harness
output rendering, and primary-checkout scoping - fires in a secondmate's own
primary session, inert in crew/scout child worktrees and non-firstmate repos.
- Wired into claude, codex, grok, opencode, and pi PreToolUse-equivalents;
per-harness hooks only call the owner.
- Blocks top-level cd/pushd/popd (including cd to an absolute path, X=1 cd,
and command cd). Allows git -C, subshell / bash -c / env -C / make -C /
find -execdir, pipeline and background forms, and cd-as-data. Fails open on
malformed input; agent-mistake threat model.
- tests/fm-cd-pretool-check.test.sh: 43-case x 5-harness-entry-form matrix,
end-to-end cwd-leak regression, scoping, fail-open, prefilter, and wiring.
- docs/cd-guard.md: full contract plus live validation (claude, codex,
opencode, pi blocked end-to-end; grok live run blocked by an API balance
limit, with mechanism parity and deterministic coverage recorded).
* no-mistakes(review): Captain, fix cd-guard classification and prefilter coverage
* no-mistakes(review): Captain, allow path-qualified command wrappers
* no-mistakes(review): Captain, allow non-executing command queries
* no-mistakes(test): Captain, clarify cd-guard safe-path remediation
* docs: clarify cd guard guidance
* no-mistakes(document): Clarify cd-guard safe target guidance
* brief: add no-mistakes shared-daemon rule to ship and scout scaffolds (#267)
Crews must never stop, restart, or update the shared no-mistakes
daemon since one instance serves every firstmate lane/home; a restart
kills other lanes' in-flight pipeline runs and forces expensive
re-runs. Encodes this as a new numbered rule in both the ship-task and
scout-task brief scaffolds.
Co-authored-by: mielyemitchell <249051873+mielyemitchell@users.noreply.github.com>
* feat: make bearings concise and accurate (#485)
* feat(bearings): four-section chat contract, accurate secondmate landed, resolved-event state render
/bearings skill (one owner of the chat-response format): mandate the four
always-present chat sections - Captain's Call, Recently Landed, Underway,
Charted Next - each with an explicit empty-state sentence, no At Anchor,
materially shorter than and linking to the report file. Resolves the ambiguous
Check first / Decisions pending split into one strict captain-action section.
fm-crew-state: the log fallback derives current state only from a real
run-state verb, so a trailing decision-closing resolved: event no longer
renders a healthy idle crew (typically a secondmate) as unknown with the
resolution prose as its detail. The keyed-decision contract in
fm-classify-lib.sh is untouched; map_log_state stays the one verb->state owner.
fm-fleet-snapshot: add a bounded, read-only secondmate_landed roll-up of Done
records from registered secondmate homes, reusing the single backlog parser and
the one secondmate-home enumerator (meta home= with data/secondmates.md
fallback); no network, per-home capped.
fm-bearings-snapshot: landed now merges main-home Done with the secondmate
roll-up, bounded by a per-home cap and an overall cap with omitted[] disclosure
(also fixing the previously-silent landed truncation); --all-landed reveals the
full set.
tests: resolved-event state render, secondmate landed aggregation with caps and
omitted[] disclosure, Captain's Call anti-leak, and the four-section contract.
* no-mistakes(review): Captain, ensure bearings reveals all landed work
* no-mistakes(document): Document bearings accuracy contracts
* fix: harden away-mode daemon lifecycle (#490)
* fix: script-owned non-visible away-daemon launch + stale-artifact lifecycle
Away-mode entry left "make the daemon a tracked background terminal" to the
operator; on a pi/herdr primary that meant splitting the captain's active pane,
which visibly shrank it. Add bin/fm-afk-launch.sh, a single owner that launches
the daemon in a non-visible tracked terminal per backend (herdr dedicated
--no-focus workspace, detached tmux session), never a split, pins the captain
pane as FM_SUPERVISOR_TARGET/FM_SUPERVISOR_BACKEND, records the exact terminal
id, and tears it down or reconciles a leaked one by that id. No shell &.
Extract supervisor-pane discovery into bin/fm-supervisor-target-lib.sh, shared
with the daemon (one owner).
Fix the stale subsuper-artifact leak: clear the prior away session's delivery
cache on a fresh entry (fm_afk_clear_stale_artifacts), and stop the daemon
before clearing state/.afk so its shutdown flush runs instead of being a no-op.
Tests: tests/fm-afk-launch.test.sh (per-backend topology invariant in a lab
session, stale clear-on-entry vs refresh, exit ordering). Docs: /afk SKILL.md,
docs/herdr-backend.md (dated herdr evidence), AGENTS.md exit stub, docs/scripts.md.
* no-mistakes(review): Captain, serialize AFK launcher lifecycle safely
* no-mistakes(review): Captain, harden AFK launcher lifecycle races
* no-mistakes(review): Captain, ensure AFK daemon launch readiness
* no-mistakes(review): Captain, unify AFK lifecycle ownership and teardown
* no-mistakes(review): Captain, preserve AFK reconciliation records uniformly
* no-mistakes(review): Captain, harden AFK recovery state durability
* no-mistakes(review): Captain, harden AFK tmux ownership checks
* no-mistakes(review): Captain, simplify AFK lifecycle failure handling
* no-mistakes(review): Captain, require confirmed AFK daemon shutdown
* no-mistakes(review): Captain, confirm AFK exit by process identity
* no-mistakes(document): Align AFK launcher lifecycle documentation
* no-mistakes: apply CI fixes
* fix: prevent no-mistakes gate agents from driving the fleet (#518)
* feat: contain no-mistakes gate agents from driving the fleet
Add bin/fm-gate-refuse-lib.sh, sourced at the top of fm-spawn/fm-send/
fm-teardown before any fleet mutation. It fails closed when NO_MISTAKES_GATE
is set, and via an unspoofable git-common-dir backstop when invoked from a
no-mistakes gate worktree (.no-mistakes/repos/*.git) even with the marker
unset. A normal firstmate session has neither signal and is unaffected.
Set disable_project_settings: true in the tracked .no-mistakes.yaml so the
installed pipeline neutralizes gate agents' project instructions for this repo
(trusted-only, honored from the default branch).
firstmate's own suite runs from a gate worktree during validation, so the
shared test helpers set FM_GATE_REFUSE_BYPASS=1 to exempt it; the dedicated
tests/fm-gate-refuse.test.sh strips it to verify real refusal.
* no-mistakes(review): Captain, refuse empty no-mistakes gate markers
* no-mistakes(document): Document no-mistakes gate authority boundary
* fix: guard secondmate primary sessions from blind turn ends (#505)
* fix: guard secondmate own-home turn ends
Remove the .fm-secondmate-home early-exit in fm-turnend-guard.sh so the
'no turn ends blind' backstop fires in a secondmate's own primary session,
matching the cd-guard's scope: the own home is guarded, child crew/scout
worktrees stay exempt via the retained git-dir/git-common-dir test. This
was pure scoping from the guard's primary-only origin and guarded against
no secondmate-specific hazard.
Add secondmate regression tests (blind-turn block, idle-by-default,
stop_hook_active loop guard, deferred-death recovery loop, child-worktree
exemption) and record the autonomous background-notify re-invoke
measurement (Claude Code 2.1.207, 11s) in docs/turnend-guard.md.
* no-mistakes(document): Correct secondmate guard documentation, captain
* fix: force-include marked secondmate homes in turn-end guard
The prior remove-only form (just deleting the .fm-secondmate-home check)
left the DEFAULT secondmate topology unguarded: a treehouse-leased home is
a linked git worktree (git-dir != git-common-dir), which the retained
git-dir exemption still skipped, so its own primary session could still end
a turn blind. Invert the marker: a genuinely-marked home is force-included
as a guarded primary (treehouse-leased linked OR git-cloned plain), and the
git-dir exemption applies only to UNMARKED child worktrees. Marker
validation (regular non-symlink file, non-empty id-token content) blocks a
stray or empty marker from spoofing inclusion.
Add real linked-worktree regression tests: a treehouse-leased LINKED
secondmate home is guarded, a stray/empty marker stays exempt, and the
unmarked child worktree stays exempt - the topology the plain git-init
fixtures masked. Predicates, in-flight gate, and loop guard untouched.
* fix: force ASCII collation in secondmate marker validation
Add a function-scoped local LC_ALL=C in fm_root_is_secondmate_home so the
[A-Za-z0-9._-] id allowlist matches under C collation, not the ambient
locale - a locale-crafted non-ASCII marker id can no longer slip through
the range match and spoof force-inclusion of a linked child worktree.
Add a regression test proving a non-ASCII marker id is rejected and the
linked worktree stays exempt.
* no-mistakes(test): fix backend baseline gate-refusal dependency
* no-mistakes(document): Correct secondmate turn-end guard documentation
* fix: make bootstrap diagnostics backend-aware (#519)
* fix: make bootstrap required-tool detection backend-aware
Bootstrap demanded tmux and treehouse for every backend except orca, so a
herdr/zellij/cmux home with tmux absent was wrongly told MISSING: tmux.
Required tools now follow the resolved backend via the single-owner
fm_backend_required_tools helper (bin/fm-backend.sh): each backend's own
session-provider CLI, jq for the JSON-emitting adapters (herdr/zellij/cmux),
and treehouse for session-provider-only backends (orca owns its worktree).
The treehouse lease-support check is gated to backends that use treehouse.
Adds install hints for herdr/zellij/cmux, regression tests for the full
backend dependency matrix (herdr-without-tmux repro plus each boundary),
and updates the authoritative Toolchain docs.
* no-mistakes(review): Captain, prevent executing Herdr install guidance
* no-mistakes(review): Captain, harden backend-aware bootstrap diagnostics
* no-mistakes(review): Captain, separate manual dependency remediation
* no-mistakes(review): Captain, align bootstrap diagnostic consumers
* no-mistakes(document): Align backend adapter dependency comments
* fix: preserve follow-up platform context after inbox cleanup (#520)
* fix: recover X/Discord follow-up platform after inbox cleanup
A milestone follow-up posted directly by request_id after the inbox was
drained - and with no task link, because one persistent secondmate's single
x_request slot collides across concurrent requests - resolved platform only
from the local inbox, so a >280 Discord reply silently defaulted to the X
280-char budget and threaded as (1/2).
- fm-x-poll records a durable per-request reply context
(state/x-context/<rid>.json) at stash time, keyed by request_id so
concurrent requests never overwrite each other; it survives inbox cleanup
and restart.
- fm-x-reply resolves platform/budget through registry -> inbox -> relay
(the relay lookup confined to a live follow-up), recovering the original
platform independent of task-link availability.
- Fail-safe: a follow-up whose platform/budget cannot be authoritatively
resolved and that would split is refused (exit 8) and held for retry,
never wrongly split; fm-x-followup keeps the link on that exit.
- fm-x-dismiss clears the durable context for a dismissed mention.
Refactors reply-context extraction into a single owner and adds regression
coverage for all four cases.
* no-mistakes(review): Captain, fail closed on incomplete follow-up context
* no-mistakes(review): Captain, bound X context registry retention
* no-mistakes(review): Captain, align context retention with answer binding
* no-mistakes(document): Align X follow-up context documentation
* no-mistakes(document): Align durable X follow-up documentation
* fix: preserve secondmate routing markers in terminal sends (#533)
* fix: preserve secondmate routing markers
* no-mistakes(review): Captain, preserve trailing newlines in marked secondmate sends
* no-mistakes(test): Captain, tolerate bootstrap timeout elapsed drift
* no-mistakes(document): Refresh Herdr marker documentation
* fix: align Grok effort handling with 0.2.99 (#527)
* fix: align grok effort docs and spawn with 0.2.99 ceiling
grok 0.2.99 accepts only low|medium|high for --reasoning-effort and
rejects both xhigh and max. Omit unsupported values on spawn, flag them
in crew-dispatch validation, and update harness-adapters.
* no-mistakes(test): Captain: refresh gotmp teardown fixture dependencies
* no-mistakes(document): Clarify Grok effort documentation ownership
* fix: derive bearings from authoritative secondmate state (#555)
* fix: make bearings use secondmate home state
* test: anonymize bearings fixtures
* no-mistakes(review): Bound parent activity evidence scans, captain
* no-mistakes(review): Preserve structured secondmate authority and bounds, captain
* no-mistakes(review): Preserve registry completeness and child inventory, captain
* no-mistakes(review): Reconcile parent evidence by verb and key, captain
* no-mistakes(review): Treat unkeyed parent evidence as inconclusive, captain
* no-mistakes(document): Document bearings local snapshot and PR opt-in
* no-mistakes(lint): Fix fleet snapshot ShellCheck findings
* no-mistakes: apply CI fixes
* fix: restore fleet snapshots on stock macOS Bash (#578)
* fix: restore stock macOS snapshot parsing
* no-mistakes(document): Clarify Linux gate and macOS CI coverage
* fix(afk): make Pi escalation and return catch-up reliable (#587)
* fix: close away-mode blocker supervision gap
* no-mistakes(review): Gate teardown retries and verify U+2063 dedupe
* no-mistakes(test): Fail closed on incomplete Pi composer separators
* no-mistakes(document): Document Pi composer recognition and return gating
* feat: support Pi max reasoning profiles (#537)
* support Pi max thinking profiles
* no-mistakes(review): Captain, allow Pi max dispatch profiles
* chore: no-mistakes(document): Clarify yolo response ownership (#595)
* Clarify validation response ownership
* no-mistakes(document): Clarify yolo response ownership
* feat: establish instruction ownership foundation (#619)
* Add instruction owners foundation
* no-mistakes(document): Refresh project-management owner pointers
* fix: compress Firstmate contract and enforce delivery rigor ownership (#626)
* docs: compress firstmate operating contract
* docs: make delivery rigor single-owner
PR B already removed personal and stacked review requirements, but it did not explicitly assign rigor to the selected delivery path or forbid risk-based manual clean gates. That gap still permitted the Hi Bit inversion.
* no-mistakes(review): Honor configured merge authority across faster delivery paths
* no-mistakes(document): Align docs with compressed operating contract
* feat: add durable captain decision holds (#593)
* Add durable captain decision holds
* no-mistakes(review): Validate decision hold retries and origin paths
* no-mistakes(review): Enforce durable decision lifecycle boundaries
* no-mistakes(review): Harden decision display and partial retry recovery
* no-mistakes(test): Update scout teardown fixtures for decision inventory
* no-mistakes(document): Align decision lifecycle and scout teardown documentation
* no-mistakes: apply CI fixes
* no-mistakes(review): Reconcile terminal decision holds
* no-mistakes(document): Align captain decision-hold documentation
* fix(bin): harden PR check artifacts (#556)
* fix: harden PR check artifacts
* fix: close PR check migration gaps
* fix: close PR check retry gaps
* fix: clarify migration outcomes and ESM boundary
* fix: keep failed migrations authoritative
* test: use inert PR validation fixtures
* no-mistakes(review): Reserve noncanonical PR quarantine namespace
* no-mistakes(review): Prevalidate final PR-check teardown artifacts
* no-mistakes(review): Preserve X metadata and validate teardown IDs
* no-mistakes(review): Initialize migration state before watcher exclusion
* no-mistakes(review): Isolate failed poll migrations from bootstrap recovery
* no-mistakes(review): Allow safe polling during incomplete private repairs
* no-mistakes(review): Authenticate watcher checks at execution time
* no-mistakes(review): Preserve custom checks with hash-bound registration
* no-mistakes(review): Clean custom check snapshots on watcher signals
* no-mistakes(review): Stop watcher checks promptly on signals
* no-mistakes(review): Terminate watcher check groups before cleanup
* no-mistakes(document): Correct stale X-mode watcher documentation
* fix: drain returned watcher check groups
* no-mistakes(review): Harden quarantine links and recover validated replacement polls
* no-mistakes(review): Preserve X mode across shim version transitions
* no-mistakes(review): Refresh legacy X shims before marker short-circuits
* no-mistakes(document): Correct persisted PR-check artifact documentation
* no-mistakes(document): Correct stale PR-check documentation
* fix: bind PR poll repair provenance
* no-mistakes(review): Enforce single-link ownership for custom check artifacts
* no-mistakes(review): Preserve private checks, X polling, and lifecycle IDs
* no-mistakes(review): Separate task creation and legacy teardown validation
* no-mistakes(review): Restore safe legacy operations and teardown validation
* no-mistakes(review): Disambiguate migration obligations and preserve legacy retries
* no-mistakes(review): Preserve fail-closed diagnostics and legacy quarantine evidence
* no-mistakes(review): Reconcile legacy migration retries and teardown collisions
* no-mistakes(review): Force legacy namespace reconciliation before marker short-circuits
* no-mistakes(document): Document private poll artifact safety contracts
* no-mistakes(lint): Suppress intentional literal-dollar lint finding
* fix: migrate historical X poll identity
* fix: harden PR check artifacts
* no-mistakes(review): Preserve fail-closed diagnostics and legacy quarantine evidence
* no-mistakes(review): Reconcile legacy migration retries and teardown collisions
* fix: migrate historical X poll identity
* no-mistakes(review): Harden X-mode artifact publication against symlink corruption
* no-mistakes(review): Guard X artifact publication
* no-mistakes(review): Enforce private X artifact reads
* no-mistakes(test): Fix backend compatibility fixture dependencies
* no-mistakes(document): Refresh PR-check documentation
* no-mistakes(lint): Remove unused x-mode test locals
* no-mistakes: apply CI fixes
* fix(bin): compact session-start backlog digest (#636)
* fix: compact session-start backlog digest
* no-mistakes(test): Fix legacy backend fixture helper
* no-mistakes(test): Fix watcher exit wait helper
* no-mistakes(document): Document compact backlog digest
* fix: dedupe stale watcher guard banners (#637)
* fix: dedupe stale watcher guard banner
* no-mistakes(review): Keep read-only guard state nonmutating
* no-mistakes(document): Clarify stale watcher docs
* fix(bin): balance bearings landed baseline (#640)
* fix: balance bearings landed defaults
* no-mistakes(document): Document balanced landed baseline
* no-mistakes: apply CI fixes
* fix: clarify captain-facing translation contract (#644)
* docs: clarify captain-facing translation contract
* no-mistakes(review): Restore runtime fallback mandate
* no-mistakes(document): Align Bearings translation wording
* fix(bin): make bootstrap output and nudges deterministic (#646)
* fix: make bootstrap nudges deterministic
* no-mistakes(review): Honor state override for bootstrap nudges
* no-mistakes(review): Update benign bootstrap documentation labels
* no-mistakes(review): Validate bootstrap nudge retry markers
* no-mistakes(document): Align bootstrap nudge documentation
* no-mistakes: apply CI fixes
* docs(secondmate-provisioning): clarify concise registry ownership (#649)
* Clarify concise secondmate registry contract
* no-mistakes(review): Expand secondmate registry boilerplate coverage
* no-mistakes(document): Point route docs to owner
* fix(bin): strip quoted blocked_by values during decision hold resolve (#654)
* fix(bin): strip quotes on blocked_by in decision-hold resolve
tasks-axi quotes multi-entry blocked_by as "a,b,c", so the comma-boundary
membership test only matched middle elements. Strip surrounding quotes
before matching so first and last hold ids resolve correctly.
* no-mistakes(document): Refresh decision-hold regression evidence
* feat(secondmate): inherit shared captain preferences (#656)
* feat(secondmate): inherit shared captain preferences
* no-mistakes(review): Honor shared captain data overrides
* no-mistakes(review): Honor bootstrap data override registry
* no-mistakes(document): Refresh shared inheritance docs
* no-mistakes(document): Clarify inherited local-material docs
* feat: gate local agent secret injection (#658)
* feat(spawn): gate local agent secret injection
* fix(spawn): align final Keychain slot
* no-mistakes: apply CI fixes
* test: isolate Herdr autodetect smoke sessions (#662)
* test: isolate herdr autodetect smoke session
* no-mistakes(review): Restored autodetect smoke gate bypass
* no-mistakes(test): Harden Herdr lab provisioning
* no-mistakes(document): Refresh Herdr lab docs
* docs: adopt under way for active work (#666)
* Revert "feat: gate local agent secret injection (#658)" (#668)
This reverts commit c27135cd9d35bc4c237d49b3b374da31fbd52eef.
* fix(pi): distinguish stale locks when arming watcher (#681)
* fix(pi): distinguish stale locks when arming watcher
* no-mistakes(test): Stabilize watcher extension async waits
* no-mistakes(document): Document Pi lock recovery
* fix: accept secondmate house vocabulary (#685)
* fix: accept secondmate as house vocabulary
* no-mistakes(test): Update captain vocabulary contract test
* no-mistakes(document): Align secondmate documentation vocabulary
* fix(bin): parse handoff homes after registry parentheticals (#686)
* fix: parse secondmate home after pre-field parentheses
Registry summaries often include parentheticals before the structured
(home: ...) field. Match that field with a greedy prefix so handoff
no longer reports "has no home" for those entries.
* no-mistakes(document): Refresh handoff test comments
* feat: add native session-start nudges (#687)
* feat: add native session-start nudges
* no-mistakes(document): Document nudge script inventory
* docs: call built-in defaults the firstmate repo, not template (#688)
Relabel absent-captain and related domain defaults wording so it names
the firstmate repo rather than treating "template" as this domain's
identity label. Keep the design-tenet "shared template" statements and
unrelated launch/PR-poll template uses unchanged.
* fix(bin): repair fm-brief.sh parse error and harden set -u array expansion (#205)
* fix(bin): use set -u-safe empty-array expansion in pr-merge and spawn
Expanding "${arr[@]}" on an empty array under set -u fails on bash < 4.4
(notably macOS bash 3.2). Quote the portable "${arr[@]+"${arr[@]}"}" idiom
in fm-pr-merge and fm-spawn batch dispatch so empty arrays expand to nothing.
Co-authored-by: Cursor <cursoragent@cursor.com>
* test(brief): harden fm-brief regression coverage for parse and scaffolds
Tighten bash -n checking, pin literal backtick rendering in the no-mistakes
DOD wording assertion, and keep a scout/secondmate scaffold smoke test so the
Co-authored-by: Cursor <cursoragent@cursor.com>
#166 apostrophe regression cannot return unnoticed.
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(bin): keep watcher supervision continuous across child cycles (#693)
* fix: make watcher supervision continuous
* no-mistakes(review): Bound watcher retries and log attached signals
* no-mistakes(review): Add bounded successor-recovery wake fallbacks
* no-mistakes(review): Prevent overlapping successor-arm retries
* no-mistakes(review): Resume supervision after late arm closes
* no-mistakes(review): Bind OpenCode recovery to attempted arm
* no-mistakes(test): Synchronize peer beacon regression fixture
* no-mistakes(test): Synchronize Pi and OpenCode late-close lifecycle fixtures
* no-mistakes(document): Captain: document watcher successor protocol behavior
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* fix: fetch current PR head for review diffs (#722)
* fix: always fetch PR head for review diffs
Prefer a freshly fetched refs/pull/<n>/head over a reachable recorded
pr_head= so reviewers never hold a merge over a "missing" fix that already
landed on the remote PR. Recorded SHA is offline fallback only; local branch
is last resort with a warning. Store the tip under refs/fm-review/ so a later
base-branch fetch cannot clobber the compare tip via FETCH_HEAD.
* no-mistakes(test): Isolate session-start nudge tests from gate state
* no-mistakes(document): Correct review-diff documentation
* docs: resolve five contract contradictions across AGENTS.md, README, and skills (#736)
* docs: resolve five contract contradictions
* no-mistakes(test): align owner-pointer assertions with reworded docs; skip absent shellcheck
* docs(harness): correct Grok exit guidance (#742)
* docs(harness): reverify grok exit command
* no-mistakes(test): Correct Grok exit resume attribution
* fix(watcher): bound stale wakes for parked crew (#743)
* fix(watcher): bound stale wakes for exited paused crew
* no-mistakes(review): Gate pause suppression on confirmed agent death
* no-mistakes(test): Fixed stale pause cadence
* no-mistakes(document): Document dead-agent hold cadence
* fix(supervision): distinguish ordinary wakes from recovery (#744)
* fix(supervision): distinguish ordinary wakes from repair
* no-mistakes(review): Make passive guard follow-ups recovery-only
* no-mistakes(document): Clarify recovery-only turn-end guard documentation
* fix(x-mode): dedupe pending mention wakes (#745)
* fix(x-mode): dedupe pending mention wakes
* no-mistakes(review): fix x-poll claim error deduplication
* no-mistakes(review): separate claim diagnostics from relay recovery
* no-mistakes(document): Document X-mode once-only mention wakes
* feat(wake): enrich drained signals with bounded status context (#747)
* feat(wake): enrich drained signal context
* no-mistakes(review): Bound wake enrichment reads
* no-mistakes(document): Document wake-drain annotations
* docs(wake): explain at-least-once drain boundary
* no-mistakes(review): Prevent symlink races in wake annotations
* no-mistakes(review): Exercise wake symlink race regression
* test: document intentional AFK marker subprocesses
* fix(wake): isolate annotation marker state
* feat(herdr): add optional presentation spaces (#784)
* feat(herdr): add optional presentation spaces
* no-mistakes(review): Harden Herdr projection creation and spawn serialization
* no-mistakes(review): Captain, disarm Herdr cleanup before launch submission
* no-mistakes(test): Correct stale Orca metadata failure fixture
* no-mistakes(document): Document Herdr presentation projection accurately
* fix(send): treat opencode busy-queued composer state as submitted (#775)
* fix(send): treat opencode busy-queued composer state as submitted
When fm-send sends a message to a BUSY opencode crewmate on the tmux
backend, opencode accepts the Enter and queues the message for the next
turn, but leaves the typed text visible in the composer row. The
submit-verification loop sees a pending composer, exhausts retries, and
reports a false "Enter swallowed" failure while the message is actually
delivered.
Fix: after Enter retries are exhausted and the composer still shows
pending, check fm_pane_is_busy. If the pane is busy (agent mid-turn,
footer shows "esc interrupt"), the harness queued the message, so
return "empty" (accepted). On an idle pane, keep returning "pending"
(genuine swallow detection preserved).
Regression tests cover four scenarios:
- busy pane + pending composer -> empty (message queued)
- idle pane + pending composer -> pending (genuine swallow)
- busy pane + composer clears on first Enter -> empty
- idle pane + composer clears on first Enter -> empty (existing path)
* docs: document busy-queued Enter exception across backend docs and skills
Add explanatory comments and backend documentation for the
busy-queued Enter fix (opencode 1.18.4 accepts Enter mid-turn
but keeps typed text in composer until the turn ends):
- bin/fm-tmux-lib.sh: document the busy-aware fallback in the
file header and above fm_tmux_submit_enter_core
- .agents/skills/afk/SKILL.md: daemon-facing policy note
- .agents/skills/harness-adapters/SKILL.md: harness-specific fact
- docs/tmux-backend.md: submit-acknowledgement section with the
busy-queue exception
- docs/herdr-backend.md: record the known gap
- docs/architecture.md: cross-reference in the daemon section
* test(tmux): fix SC2181 and make busy-submit test executable
* fix(spawn): require two stable reads before accepting worktree path (#765)
* fix(spawn): require two stable reads before accepting worktree path
The treehouse-get worktree-detection loop in fm-spawn.sh accepted the
first pane_current_path read that differed from the project path, but
on some tmux/WSL setups a brand-new window transiently reports a
stale-but-real path before the pane actually settles into the
worktree. Since that stale path is itself a real, distinct git
checkout, it also passes validate_spawn_worktree's isolation check,
so the loop silently recorded the wrong worktree in state/<id>.meta
(and, for claude harness spawns, installed the turn-end hook there
too).
Require two consecutive polls to agree on the same non-project path
before accepting it, using the existing inter-poll sleep as the
confirmation gap so an already-settled pane isn't slowed down by an
extra cycle.
* fix(tests): drop unused CASE_DIR read in worktree-settle test
ShellCheck SC2034: CASE_DIR is split out of the case record but
never referenced; discard it with _ instead.
---------
Co-authored-by: Freudator86 <tim@allesknut.de>
* fix(bin): make watcher process identity immune to Linux wall-clock changes (#752)
* fix(watcher): stabilize Linux process identity
* no-mistakes(document): document FM_PROC_ROOT_OVERRIDE and Linux starttime identity rationale
* fix: prevent AFK idle stalls and stale run attribution (#758)
* fix(supervision): verb-aware captain relevance, AFK wedge, head-bound state
Stop free-text tokens like "merged" from promoting nonterminal working: lines
to captain-relevant, so AFK no longer permanently suppresses idle recovery.
Defend wedge aging independently for nonterminal progress verbs, bind
no-mistakes current-state attribution to code identity (not branch alone),
and mark setup-complete as nonterminal in the ship brief scaffold.
* no-mistakes(review): Enforce nonterminal suppression and head-bound run attribution
* no-mistakes(document): Document current-code-bound run attribution
* no-mistakes(test): Wait for stable Herdr shell readiness
* no-mistakes(test): Make Herdr and watcher readiness tests deterministic
* no-mistakes(test): Make tmux capture and watcher lifecycle deterministic
* no-mistakes(document): Document corrected supervision contracts
* fix(bin): allow safe teardown during watcher recovery (#750)
* fix: allow safe teardown during watcher recovery
* no-mistakes(review): Distinguish unsafe-teardown deny guidance via policy reason code
* no-mistakes(document): Sync continuity-gate docs to allow teardown recovery
* test: mark dynamic teardown fixture literal
* no-mistakes(document): docs: add teardown to continuity gate allow list
* feat(herdr): order presentation spaces while preserving focus (#790)
* feat(herdr): order presentation worker spaces
* fix(herdr): preserve focus during projected cleanup
* no-mistakes(review): Serialize Herdr cleanup and protect active seeded tabs
* no-mistakes(review): Serialize Herdr aborts with guarded focus regressions
* no-mistakes(review): Fall back flat when Herdr serialization is unavailable
* no-mistakes(test): Stabilize watcher startup and AFK handoff tests
* no-mistakes(document): Correct Herdr ordering and focus documentation
* fix(bin): send literal config reread nudges after pushes (#809)
* Send literal config reread after inherited config push
When declared inherited config changes under an already-running secondmate,
build a per-home instruction from validated destination post-write bytes and
deliver it on the routed secondmate path. Unchanged config sends nothing;
ABSENT represents removal; captain-shared is never inlined. Covers mid-session
config-push and the locked bootstrap convergence path without hardening spawn
against deliberate runtime choice.
* no-mistakes(review): Fix config reread framing, partial propagation, and respawn order
* no-mistakes(review): Send config rereads via durable single-line pointers
* no-mistakes(review): Make failed config rereads retryable
* no-mistakes(review): Make config reread retries generation-safe
* no-mistakes(review): Make config rereads durable and ordered
* no-mistakes(review): Drain retries, bound history, preserve detect-only read-only mode
* no-mistakes(review): Retain write retries and quarantine stale respawn generations
* no-mistakes(review): Preserve exact config reread retries and delivery order
* no-mistakes(review): Preserve exact retry bytes and bounded quarantine pruning
* no-mistakes(document):…
7hoenix
added a commit
to 7hoenix/firstmate
that referenced
this pull request
Aug 27, 2026
…signing (#19) * fix: preserve secondmate routing markers in terminal sends (#533) * fix: preserve secondmate routing markers * no-mistakes(review): Captain, preserve trailing newlines in marked secondmate sends * no-mistakes(test): Captain, tolerate bootstrap timeout elapsed drift * no-mistakes(document): Refresh Herdr marker documentation * fix: align Grok effort handling with 0.2.99 (#527) * fix: align grok effort docs and spawn with 0.2.99 ceiling grok 0.2.99 accepts only low|medium|high for --reasoning-effort and rejects both xhigh and max. Omit unsupported values on spawn, flag them in crew-dispatch validation, and update harness-adapters. * no-mistakes(test): Captain: refresh gotmp teardown fixture dependencies * no-mistakes(document): Clarify Grok effort documentation ownership * fix: derive bearings from authoritative secondmate state (#555) * fix: make bearings use secondmate home state * test: anonymize bearings fixtures * no-mistakes(review): Bound parent activity evidence scans, captain * no-mistakes(review): Preserve structured secondmate authority and bounds, captain * no-mistakes(review): Preserve registry completeness and child inventory, captain * no-mistakes(review): Reconcile parent evidence by verb and key, captain * no-mistakes(review): Treat unkeyed parent evidence as inconclusive, captain * no-mistakes(document): Document bearings local snapshot and PR opt-in * no-mistakes(lint): Fix fleet snapshot ShellCheck findings * no-mistakes: apply CI fixes * fix: restore fleet snapshots on stock macOS Bash (#578) * fix: restore stock macOS snapshot parsing * no-mistakes(document): Clarify Linux gate and macOS CI coverage * fix(afk): make Pi escalation and return catch-up reliable (#587) * fix: close away-mode blocker supervision gap * no-mistakes(review): Gate teardown retries and verify U+2063 dedupe * no-mistakes(test): Fail closed on incomplete Pi composer separators * no-mistakes(document): Document Pi composer recognition and return gating * feat: support Pi max reasoning profiles (#537) * support Pi max thinking profiles * no-mistakes(review): Captain, allow Pi max dispatch profiles * chore: no-mistakes(document): Clarify yolo response ownership (#595) * Clarify validation response ownership * no-mistakes(document): Clarify yolo response ownership * feat: establish instruction ownership foundation (#619) * Add instruction owners foundation * no-mistakes(document): Refresh project-management owner pointers * fix: compress Firstmate contract and enforce delivery rigor ownership (#626) * docs: compress firstmate operating contract * docs: make delivery rigor single-owner PR B already removed personal and stacked review requirements, but it did not explicitly assign rigor to the selected delivery path or forbid risk-based manual clean gates. That gap still permitted the Hi Bit inversion. * no-mistakes(review): Honor configured merge authority across faster delivery paths * no-mistakes(document): Align docs with compressed operating contract * feat: add durable captain decision holds (#593) * Add durable captain decision holds * no-mistakes(review): Validate decision hold retries and origin paths * no-mistakes(review): Enforce durable decision lifecycle boundaries * no-mistakes(review): Harden decision display and partial retry recovery * no-mistakes(test): Update scout teardown fixtures for decision inventory * no-mistakes(document): Align decision lifecycle and scout teardown documentation * no-mistakes: apply CI fixes * no-mistakes(review): Reconcile terminal decision holds * no-mistakes(document): Align captain decision-hold documentation * fix(bin): harden PR check artifacts (#556) * fix: harden PR check artifacts * fix: close PR check migration gaps * fix: close PR check retry gaps * fix: clarify migration outcomes and ESM boundary * fix: keep failed migrations authoritative * test: use inert PR validation fixtures * no-mistakes(review): Reserve noncanonical PR quarantine namespace * no-mistakes(review): Prevalidate final PR-check teardown artifacts * no-mistakes(review): Preserve X metadata and validate teardown IDs * no-mistakes(review): Initialize migration state before watcher exclusion * no-mistakes(review): Isolate failed poll migrations from bootstrap recovery * no-mistakes(review): Allow safe polling during incomplete private repairs * no-mistakes(review): Authenticate watcher checks at execution time * no-mistakes(review): Preserve custom checks with hash-bound registration * no-mistakes(review): Clean custom check snapshots on watcher signals * no-mistakes(review): Stop watcher checks promptly on signals * no-mistakes(review): Terminate watcher check groups before cleanup * no-mistakes(document): Correct stale X-mode watcher documentation * fix: drain returned watcher check groups * no-mistakes(review): Harden quarantine links and recover validated replacement polls * no-mistakes(review): Preserve X mode across shim version transitions * no-mistakes(review): Refresh legacy X shims before marker short-circuits * no-mistakes(document): Correct persisted PR-check artifact documentation * no-mistakes(document): Correct stale PR-check documentation * fix: bind PR poll repair provenance * no-mistakes(review): Enforce single-link ownership for custom check artifacts * no-mistakes(review): Preserve private checks, X polling, and lifecycle IDs * no-mistakes(review): Separate task creation and legacy teardown validation * no-mistakes(review): Restore safe legacy operations and teardown validation * no-mistakes(review): Disambiguate migration obligations and preserve legacy retries * no-mistakes(review): Preserve fail-closed diagnostics and legacy quarantine evidence * no-mistakes(review): Reconcile legacy migration retries and teardown collisions * no-mistakes(review): Force legacy namespace reconciliation before marker short-circuits * no-mistakes(document): Document private poll artifact safety contracts * no-mistakes(lint): Suppress intentional literal-dollar lint finding * fix: migrate historical X poll identity * fix: harden PR check artifacts * no-mistakes(review): Preserve fail-closed diagnostics and legacy quarantine evidence * no-mistakes(review): Reconcile legacy migration retries and teardown collisions * fix: migrate historical X poll identity * no-mistakes(review): Harden X-mode artifact publication against symlink corruption * no-mistakes(review): Guard X artifact publication * no-mistakes(review): Enforce private X artifact reads * no-mistakes(test): Fix backend compatibility fixture dependencies * no-mistakes(document): Refresh PR-check documentation * no-mistakes(lint): Remove unused x-mode test locals * no-mistakes: apply CI fixes * fix(bin): compact session-start backlog digest (#636) * fix: compact session-start backlog digest * no-mistakes(test): Fix legacy backend fixture helper * no-mistakes(test): Fix watcher exit wait helper * no-mistakes(document): Document compact backlog digest * fix: dedupe stale watcher guard banners (#637) * fix: dedupe stale watcher guard banner * no-mistakes(review): Keep read-only guard state nonmutating * no-mistakes(document): Clarify stale watcher docs * fix(bin): balance bearings landed baseline (#640) * fix: balance bearings landed defaults * no-mistakes(document): Document balanced landed baseline * no-mistakes: apply CI fixes * fix: clarify captain-facing translation contract (#644) * docs: clarify captain-facing translation contract * no-mistakes(review): Restore runtime fallback mandate * no-mistakes(document): Align Bearings translation wording * fix(bin): make bootstrap output and nudges deterministic (#646) * fix: make bootstrap nudges deterministic * no-mistakes(review): Honor state override for bootstrap nudges * no-mistakes(review): Update benign bootstrap documentation labels * no-mistakes(review): Validate bootstrap nudge retry markers * no-mistakes(document): Align bootstrap nudge documentation * no-mistakes: apply CI fixes * docs(secondmate-provisioning): clarify concise registry ownership (#649) * Clarify concise secondmate registry contract * no-mistakes(review): Expand secondmate registry boilerplate coverage * no-mistakes(document): Point route docs to owner * fix(bin): strip quoted blocked_by values during decision hold resolve (#654) * fix(bin): strip quotes on blocked_by in decision-hold resolve tasks-axi quotes multi-entry blocked_by as "a,b,c", so the comma-boundary membership test only matched middle elements. Strip surrounding quotes before matching so first and last hold ids resolve correctly. * no-mistakes(document): Refresh decision-hold regression evidence * feat(secondmate): inherit shared captain preferences (#656) * feat(secondmate): inherit shared captain preferences * no-mistakes(review): Honor shared captain data overrides * no-mistakes(review): Honor bootstrap data override registry * no-mistakes(document): Refresh shared inheritance docs * no-mistakes(document): Clarify inherited local-material docs * feat: gate local agent secret injection (#658) * feat(spawn): gate local agent secret injection * fix(spawn): align final Keychain slot * no-mistakes: apply CI fixes * test: isolate Herdr autodetect smoke sessions (#662) * test: isolate herdr autodetect smoke session * no-mistakes(review): Restored autodetect smoke gate bypass * no-mistakes(test): Harden Herdr lab provisioning * no-mistakes(document): Refresh Herdr lab docs * docs: adopt under way for active work (#666) * Revert "feat: gate local agent secret injection (#658)" (#668) This reverts commit c27135cd9d35bc4c237d49b3b374da31fbd52eef. * fix(pi): distinguish stale locks when arming watcher (#681) * fix(pi): distinguish stale locks when arming watcher * no-mistakes(test): Stabilize watcher extension async waits * no-mistakes(document): Document Pi lock recovery * fix: accept secondmate house vocabulary (#685) * fix: accept secondmate as house vocabulary * no-mistakes(test): Update captain vocabulary contract test * no-mistakes(document): Align secondmate documentation vocabulary * fix(bin): parse handoff homes after registry parentheticals (#686) * fix: parse secondmate home after pre-field parentheses Registry summaries often include parentheticals before the structured (home: ...) field. Match that field with a greedy prefix so handoff no longer reports "has no home" for those entries. * no-mistakes(document): Refresh handoff test comments * feat: add native session-start nudges (#687) * feat: add native session-start nudges * no-mistakes(document): Document nudge script inventory * docs: call built-in defaults the firstmate repo, not template (#688) Relabel absent-captain and related domain defaults wording so it names the firstmate repo rather than treating "template" as this domain's identity label. Keep the design-tenet "shared template" statements and unrelated launch/PR-poll template uses unchanged. * fix(bin): repair fm-brief.sh parse error and harden set -u array expansion (#205) * fix(bin): use set -u-safe empty-array expansion in pr-merge and spawn Expanding "${arr[@]}" on an empty array under set -u fails on bash < 4.4 (notably macOS bash 3.2). Quote the portable "${arr[@]+"${arr[@]}"}" idiom in fm-pr-merge and fm-spawn batch dispatch so empty arrays expand to nothing. Co-authored-by: Cursor <cursoragent@cursor.com> * test(brief): harden fm-brief regression coverage for parse and scaffolds Tighten bash -n checking, pin literal backtick rendering in the no-mistakes DOD wording assertion, and keep a scout/secondmate scaffold smoke test so the Co-authored-by: Cursor <cursoragent@cursor.com> #166 apostrophe regression cannot return unnoticed. --------- Co-authored-by: Cursor <cursoragent@cursor.com> * fix(bin): keep watcher supervision continuous across child cycles (#693) * fix: make watcher supervision continuous * no-mistakes(review): Bound watcher retries and log attached signals * no-mistakes(review): Add bounded successor-recovery wake fallbacks * no-mistakes(review): Prevent overlapping successor-arm retries * no-mistakes(review): Resume supervision after late arm closes * no-mistakes(review): Bind OpenCode recovery to attempted arm * no-mistakes(test): Synchronize peer beacon regression fixture * no-mistakes(test): Synchronize Pi and OpenCode late-close lifecycle fixtures * no-mistakes(document): Captain: document watcher successor protocol behavior * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * fix: fetch current PR head for review diffs (#722) * fix: always fetch PR head for review diffs Prefer a freshly fetched refs/pull/<n>/head over a reachable recorded pr_head= so reviewers never hold a merge over a "missing" fix that already landed on the remote PR. Recorded SHA is offline fallback only; local branch is last resort with a warning. Store the tip under refs/fm-review/ so a later base-branch fetch cannot clobber the compare tip via FETCH_HEAD. * no-mistakes(test): Isolate session-start nudge tests from gate state * no-mistakes(document): Correct review-diff documentation * docs: resolve five contract contradictions across AGENTS.md, README, and skills (#736) * docs: resolve five contract contradictions * no-mistakes(test): align owner-pointer assertions with reworded docs; skip absent shellcheck * docs(harness): correct Grok exit guidance (#742) * docs(harness): reverify grok exit command * no-mistakes(test): Correct Grok exit resume attribution * fix(watcher): bound stale wakes for parked crew (#743) * fix(watcher): bound stale wakes for exited paused crew * no-mistakes(review): Gate pause suppression on confirmed agent death * no-mistakes(test): Fixed stale pause cadence * no-mistakes(document): Document dead-agent hold cadence * fix(supervision): distinguish ordinary wakes from recovery (#744) * fix(supervision): distinguish ordinary wakes from repair * no-mistakes(review): Make passive guard follow-ups recovery-only * no-mistakes(document): Clarify recovery-only turn-end guard documentation * fix(x-mode): dedupe pending mention wakes (#745) * fix(x-mode): dedupe pending mention wakes * no-mistakes(review): fix x-poll claim error deduplication * no-mistakes(review): separate claim diagnostics from relay recovery * no-mistakes(document): Document X-mode once-only mention wakes * feat(wake): enrich drained signals with bounded status context (#747) * feat(wake): enrich drained signal context * no-mistakes(review): Bound wake enrichment reads * no-mistakes(document): Document wake-drain annotations * docs(wake): explain at-least-once drain boundary * no-mistakes(review): Prevent symlink races in wake annotations * no-mistakes(review): Exercise wake symlink race regression * test: document intentional AFK marker subprocesses * fix(wake): isolate annotation marker state * feat(herdr): add optional presentation spaces (#784) * feat(herdr): add optional presentation spaces * no-mistakes(review): Harden Herdr projection creation and spawn serialization * no-mistakes(review): Captain, disarm Herdr cleanup before launch submission * no-mistakes(test): Correct stale Orca metadata failure fixture * no-mistakes(document): Document Herdr presentation projection accurately * fix(send): treat opencode busy-queued composer state as submitted (#775) * fix(send): treat opencode busy-queued composer state as submitted When fm-send sends a message to a BUSY opencode crewmate on the tmux backend, opencode accepts the Enter and queues the message for the next turn, but leaves the typed text visible in the composer row. The submit-verification loop sees a pending composer, exhausts retries, and reports a false "Enter swallowed" failure while the message is actually delivered. Fix: after Enter retries are exhausted and the composer still shows pending, check fm_pane_is_busy. If the pane is busy (agent mid-turn, footer shows "esc interrupt"), the harness queued the message, so return "empty" (accepted). On an idle pane, keep returning "pending" (genuine swallow detection preserved). Regression tests cover four scenarios: - busy pane + pending composer -> empty (message queued) - idle pane + pending composer -> pending (genuine swallow) - busy pane + composer clears on first Enter -> empty - idle pane + composer clears on first Enter -> empty (existing path) * docs: document busy-queued Enter exception across backend docs and skills Add explanatory comments and backend documentation for the busy-queued Enter fix (opencode 1.18.4 accepts Enter mid-turn but keeps typed text in composer until the turn ends): - bin/fm-tmux-lib.sh: document the busy-aware fallback in the file header and above fm_tmux_submit_enter_core - .agents/skills/afk/SKILL.md: daemon-facing policy note - .agents/skills/harness-adapters/SKILL.md: harness-specific fact - docs/tmux-backend.md: submit-acknowledgement section with the busy-queue exception - docs/herdr-backend.md: record the known gap - docs/architecture.md: cross-reference in the daemon section * test(tmux): fix SC2181 and make busy-submit test executable * fix(spawn): require two stable reads before accepting worktree path (#765) * fix(spawn): require two stable reads before accepting worktree path The treehouse-get worktree-detection loop in fm-spawn.sh accepted the first pane_current_path read that differed from the project path, but on some tmux/WSL setups a brand-new window transiently reports a stale-but-real path before the pane actually settles into the worktree. Since that stale path is itself a real, distinct git checkout, it also passes validate_spawn_worktree's isolation check, so the loop silently recorded the wrong worktree in state/<id>.meta (and, for claude harness spawns, installed the turn-end hook there too). Require two consecutive polls to agree on the same non-project path before accepting it, using the existing inter-poll sleep as the confirmation gap so an already-settled pane isn't slowed down by an extra cycle. * fix(tests): drop unused CASE_DIR read in worktree-settle test ShellCheck SC2034: CASE_DIR is split out of the case record but never referenced; discard it with _ instead. --------- Co-authored-by: Freudator86 <tim@allesknut.de> * fix(bin): make watcher process identity immune to Linux wall-clock changes (#752) * fix(watcher): stabilize Linux process identity * no-mistakes(document): document FM_PROC_ROOT_OVERRIDE and Linux starttime identity rationale * fix: prevent AFK idle stalls and stale run attribution (#758) * fix(supervision): verb-aware captain relevance, AFK wedge, head-bound state Stop free-text tokens like "merged" from promoting nonterminal working: lines to captain-relevant, so AFK no longer permanently suppresses idle recovery. Defend wedge aging independently for nonterminal progress verbs, bind no-mistakes current-state attribution to code identity (not branch alone), and mark setup-complete as nonterminal in the ship brief scaffold. * no-mistakes(review): Enforce nonterminal suppression and head-bound run attribution * no-mistakes(document): Document current-code-bound run attribution * no-mistakes(test): Wait for stable Herdr shell readiness * no-mistakes(test): Make Herdr and watcher readiness tests deterministic * no-mistakes(test): Make tmux capture and watcher lifecycle deterministic * no-mistakes(document): Document corrected supervision contracts * fix(bin): allow safe teardown during watcher recovery (#750) * fix: allow safe teardown during watcher recovery * no-mistakes(review): Distinguish unsafe-teardown deny guidance via policy reason code * no-mistakes(document): Sync continuity-gate docs to allow teardown recovery * test: mark dynamic teardown fixture literal * no-mistakes(document): docs: add teardown to continuity gate allow list * feat(herdr): order presentation spaces while preserving focus (#790) * feat(herdr): order presentation worker spaces * fix(herdr): preserve focus during projected cleanup * no-mistakes(review): Serialize Herdr cleanup and protect active seeded tabs * no-mistakes(review): Serialize Herdr aborts with guarded focus regressions * no-mistakes(review): Fall back flat when Herdr serialization is unavailable * no-mistakes(test): Stabilize watcher startup and AFK handoff tests * no-mistakes(document): Correct Herdr ordering and focus documentation * fix(bin): send literal config reread nudges after pushes (#809) * Send literal config reread after inherited config push When declared inherited config changes under an already-running secondmate, build a per-home instruction from validated destination post-write bytes and deliver it on the routed secondmate path. Unchanged config sends nothing; ABSENT represents removal; captain-shared is never inlined. Covers mid-session config-push and the locked bootstrap convergence path without hardening spawn against deliberate runtime choice. * no-mistakes(review): Fix config reread framing, partial propagation, and respawn order * no-mistakes(review): Send config rereads via durable single-line pointers * no-mistakes(review): Make failed config rereads retryable * no-mistakes(review): Make config reread retries generation-safe * no-mistakes(review): Make config rereads durable and ordered * no-mistakes(review): Drain retries, bound history, preserve detect-only read-only mode * no-mistakes(review): Retain write retries and quarantine stale respawn generations * no-mistakes(review): Preserve exact config reread retries and delivery order * no-mistakes(review): Preserve exact retry bytes and bounded quarantine pruning * no-mistakes(document): Consolidated config-reread documentation * feat(watch): follow GitLab merge requests to merge (#797) * feat(watch): follow GitLab merge requests to merge The merge watch only understood GitHub pull requests, so a task whose deliverable is a GitLab merge request was never followed to merge. Generalize the stored poll identity from owner/repository to a provider-tagged provider/url/host/path/number record. GitLab runs mostly on self-hosted instances and its projects nest under groups at no fixed depth, so the host and the full project path are data in the record rather than constants, and every consumer rebuilds the URL from those parts and refuses any record that does not reconstruct it exactly. The GitLab state is read with plain glab, matching the GitHub path's use of plain gh, so an upstream checkout needs no extra tooling. Two things about glab were established by running it rather than assumed, because a wrong invocation here fails silently into a permanent "not merged": - glab has no field selector, and its JSON would need a JSON processor that firstmate does not require, so the state is read from glab's own field output. Only an exact "merged" wakes firstmate, so a changed format stays silent instead of reporting a merge. - glab cannot take a merge request URL the way gh can, because that form resolves through the current git repository and the watcher has none. It is addressed by project URL and merge request number instead. An absent glab produces no wake rather than a false merge, and arming refuses with a clear message since that is the one point where a missing CLI can still be reported. A GitLab task records no pr_head, which both consumers already treat as optional. The merge path still addresses GitHub only and refuses a merge request URL rather than sending it to the wrong forge. The record version moves to v2, and the existing non-executing migration rebuilds an already-armed watch from its recorded URL, so no watch is lost by upgrading. docs/gitlab-merge-watch.md records the evidence, taken against the public fixture project https://gitlab.com/KarotKris/gitlab-merge-watch-fixture. * no-mistakes(review): Reject github.com host in GitLab MR URL/sidecar validation * no-mistakes(document): Note GitLab MR URLs are explicitly refused, not just malformed ones, in fm-pr-merge.sh docs * fix(herdr): group projected children beneath owning parents (#821) * feat(herdr): correct all-home child presentation topology Inherit the presentation opt-in to secondmate homes, label new projected spaces with the approved corner format, insert each child under its owning parent under one session-scoped lock, and keep flat non-destructive fallback. * no-mistakes(review): Exclude secondmates from Herdr presentation projection * no-mistakes(review): Harden shared Herdr locks and ambiguous child ordering * no-mistakes(review): Use adjacency-only Herdr child ownership * no-mistakes(review): Reject foreign legacy projections safely * no-mistakes(review): Validate Herdr session sockets before projection * no-mistakes(test): Fix Herdr teardown fixture session socket metadata * fix(herdr): canonicalize presentation lock socket paths Always resolve the session socket parent directory so symlink parents such as /tmp -> /private/tmp cannot split the shared cross-home lock identity. Refuse relative socket paths. Clarify lock-unavailable warnings. * no-mistakes(test): Fix Bash-compatible GitLab merge request URL parsing * no-mistakes(document): Document all-home Herdr child topology * no-mistakes(lint): Quote fallback provenance string for ShellCheck * fix: keep local no-mistakes tests intent-targeted (#823) * fix(no-mistakes): drop full-suite local Test override Local no-mistakes Test is intent-targeted; CI Behavior keeps the broad tests/*.test.sh suite. Keep commands.lint on bin/fm-lint.sh and add a focused contract test so the override cannot silently return. * no-mistakes(lint): Make CI contract assertion ShellCheck-clean * feat: add canonical timed test runner (#825) * feat(test): add canonical timed suite runner and honest CI timeout Introduce bin/fm-test-run.sh as the single serial owner for selecting one script, a family, a conservative changed-file set, or the explicit complete suite, with per-script timing markers and a JSON artifact. Wire CI Behavior through the runner, raise the hang-tripwire timeout to 25 minutes, and document entry points without restoring a full-suite local no-mistakes Test command. * no-mistakes(review): Captain: fix changed selection and empty summaries * no-mistakes(review): Captain: fail closed on unmapped changed sources * no-mistakes(document): Document canonical timed test entry points * fix: surface main inventory gaps in Bearings (#830) * fix: disclose main-home orphan and unstructured inventory gaps Main Bearings could report an empty fleet while structured in-flight rows lacked meta or current backlog rows were free-form. Emit main_inventory from the fleet snapshot, map it into Bearings omitted surfaces and a Charted Next gate, and keep meta as the only live Underway source. * no-mistakes(document): Document Bearings inventory-integrity projection * no-mistakes: apply CI fixes * feat: add bounded concurrent test isolation proof (#832) * feat: add concurrent test isolation proof for Phase 2 Prove an audited portable candidate set passes under concurrent workers with private mode-0700 temp roots, without enabling production CI sharding or fm-test-run --jobs. * no-mistakes(review): Pin isolation proof to audited candidate manifest * feat: guard against missed secondmate reports (#834) * feat(secondmate): parent-owned guards for missed status reports Marked parent-to-secondmate requests now create a durable pending-reply expectation with a privacy-safe correlation id before delivery. Transport success never resolves it; only a correlated parent status or document pointer does. After a completed turn with no report, the parent sends one recovery repost and escalates once if that turn is also missed, without scraping the secondmate conversation or looping. * no-mistakes(review): Deduplicate wrong-home pending-reply sightings * no-mistakes(review): Harden pending-reply recovery and escalation guards * no-mistakes(review): Bound pending-reply backend polling * no-mistakes(review): Cache pending-reply status scans * no-mistakes(review): Protect undelivered pending-reply records from scans * no-mistakes(review): Close pending-reply delivery durability gaps * no-mistakes(review): Separate pending-reply transport outcomes * no-mistakes(review): Escalate stalled pending-reply deliveries once * no-mistakes(review): Resolve attempted deliveries from correlated reports * no-mistakes(review): Resolve late reports after delivery escalation * no-mistakes(document): Document pending-reply grace and ownership * no-mistakes(lint): Silence intentional pending-reply test fixture lint warnings * feat: require pinned real-Herdr CI coverage (#838) * feat: add required pinned Herdr CI lane Install exact Herdr 0.7.4 and Treehouse 2.0.1 with official assets and SHA-256 pins, run the real-herdr-gated family serially through fm-test-run with hard-fail on herdr-not-found, and keep portable Behavior free of claimed Herdr coverage. * no-mistakes(document): Consolidate real-Herdr CI documentation ownership * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * feat: shard portable tests and add bounded local parallelism (#841) * feat: shard portable CI tests after isolation proof Balance the Phase 2 proven-isolated set into two LPT portable parallel lanes from Phase 1 timing evidence, keep stateful work in a required portable serial lane, exclude real Herdr to its dedicated required lane, and prove complete inventory coverage with a deterministic guard. Add bounded local --jobs only for the proven set, per-lane timing plus aggregate artifacts, and reduce the interim portable hang tripwire now that the serial remainder owns the long wall-clock path. * no-mistakes(review): Captain, fix CI contracts and completion-order worker scheduling * no-mistakes(review): Captain, preserve stderr gate-skip detection in parallel tests * no-mistakes(document): Document portable sharding and timing aggregation * no-mistakes: apply CI fixes * fix: block primary-session delegation outside the fleet (#854) * feat: fence primary-session delegation outside the fleet A firstmate primary that delegates through Claude Code's built-in delegation tools creates work with no state/<id>.meta. Because fm-supervision-lib.sh counts *.meta and fm-turnend-guard.sh exits silently at zero, such work does not merely go unsupervised: it makes the whole guard stack structurally inert, and it dies with the primary session. On 2026-07-22 that cost two workers mid-flight and left supervision down for 73 minutes unnoticed. Layer 1, the primary fix: a permissions.deny list in .claude/settings.json removes the 18 delegation, scheduling, worktree, and task-tracking tools from the model's schema, so they are never offered. This is removal rather than interception, so there is no call to intercept and no fail-open path. The list is flat and in one file so its width stays reviewable; the captain owns that width. Layer 2, bin/fm-subagent-pretool-check.sh: a deny list is fail-open against tools that do not exist yet, and permissions.allow is a pre-approval list rather than an availability list, so there is no fail-closed allowlist to use instead. This backstop classifies the tool NAME by shape rather than against a fixed list, so a delegation tool that ships before the deny list is updated is still refused. It excludes mcp__* names and observe-or-stop operations, scopes itself to a genuine primary home via the shared fm_primary_scope_matches predicate so a crewmate's task worktree is unaffected, and offers one deliberate FM_ALLOW_SUBAGENT=1 escape hatch that must be set at launch. Verified live against Claude Code 2.1.217, including a deny-key A/B with a nonsense-name control, layer 2 denying an un-denied Workflow call, the same call allowed in a linked worktree, and the escape hatch. Corrects a prior finding: both Task and Agent work as deny keys, so both are pinned. Codex 0.144.1 verified to expose no delegation tool; grok, opencode, and pi are inspected and documented as not wired because those binaries are absent from this host and the repo requires live validation before trusting a harness hook. Evidence in docs/subagent-guard.md. * no-mistakes(review): Ship scoped Claude delegation guard * no-mistakes(test): Ship Claude delegation deny list * no-mistakes(document): Clarify PreToolUse guard ownership * no-mistakes(lint): Keep Claude deny list local * fix: install tasks-axi in portable CI shards (#866) Reproduction: portable-parallel-2 completed successfully without tasks-axi while fm-decision-hold-lifecycle emitted a gate skip in 30 ms. The pre-shard lane installed tasks-axi and exercised the test fully. Installing tasks-axi is the smallest counterfactual and makes the representative shard execute the test with gate_skip=false in about 20 seconds. Both parallel jobs receive symmetric setup, while the exact 91-test inventory and coverage guard remain unchanged. * feat(bin): make dispatch profiles quota aware (#867) * feat: make dispatch profiles quota aware * no-mistakes(review): Fix quota window and Grok product scoping * no-mistakes(document): Document implicit quota-aware dispatch accurately * Add built-in ahoy recap skill (#873) * fix: preserve trustworthy Bearings data in partial snapshots (#875) * fix: preserve mixed Bearings projections * no-mistakes(review): Enforce strict invalidity precedence for partial snapshots * no-mistakes(review): Enforce ownership for unknown child metadata * no-mistakes(document): Document partial structured Bearings projections * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * feat(pi): add session-local calm mode (#884) * Add session-local Pi calm mode * no-mistakes(review): Preserve Pi HTML exports during calm mode * no-mistakes(review): Preserve calm exports across submit bindings and share * no-mistakes(document): Document calm-mode feasibility across supported harnesses * fix(pi): prevent redundant watcher re-arms (#885) * fix(pi): limit watcher arm tool to recovery * no-mistakes(review): Strengthen Pi live re-arm regression coverage * no-mistakes(document): Document Pi first-cycle and recovery-only watcher arming * fix(pi): clean up Calm transcript rendering (#895) * fix(pi): clean up Calm transcript rendering * no-mistakes(review): Captain, preserve Calm exports and classify Pi launch briefs * no-mistakes(review): Captain, eliminate Calm gaps and verify exported conversations * no-mistakes(review): Restore Calm rows received while active * no-mistakes(review): Preserve diagnostics during Calm restoration * no-mistakes(document): Clarify Calm transcript behavior and injection paths * fix: execute every PR body compliance event (#898) * fix: execute every PR body compliance event * no-mistakes(document): Document independent PR compliance events * fix: exclude operational injections from ahoy boundaries (#899) * fix: distinguish operational input in ahoy * no-mistakes(review): Handle legacy Ahoy operational boundaries * no-mistakes(review): Narrow legacy Ahoy boundaries with live regressions * no-mistakes(document): Document Ahoy operational marker ownership * no-mistakes(lint): Suppress intentional literal fixture lint warnings * fix: canonically classify operational inputs across harnesses (#909) * fix: type canonical operational inputs * no-mistakes(document): Correct canonical operational-input documentation ownership * fix: avoid generic secondmate start acknowledgements (#926) * fix: avoid generic secondmate acknowledgements * no-mistakes(document): Document sparse secondmate acknowledgement behavior * no-mistakes: apply CI fixes * fix(pi): make calm mode persistent and gapless (#927) * fix(pi): preserve calm presentation across sessions * no-mistakes(review): Fix Calm home fallback persistence * no-mistakes(document): Clarify Calm gapless and export contracts * no-mistakes: apply CI fixes * fix(watch): retire merged PR polls after durable notification (#932) * fix: retire merged PR polls after notification * no-mistakes(review): Decouple PR retirement recovery from template updates * no-mistakes(review): Recover pending PR retirements before poll migration * no-mistakes(document): Document merged PR poll retirement contracts * fix: refine scout intake and parallel dispatch (#934) * Clarify intake evidence and overlap handling * no-mistakes(review): Align scout guard with intake classification * no-mistakes(document): Clarify scout documentation and intake ownership * fix(pi): prevent duplicate assistant replies in Calm (#936) * fix(pi): preserve operational follow-up semantics in Calm * no-mistakes(document): Correct Calm operational-row visibility documentation * perf(bin): shrink the ShellCheck source graph (#939) * perf(lint): shrink shell source graph * no-mistakes(review): Ensure lint workers terminate fully on cancellation * no-mistakes(document): Repair stale lint documentation ownership * fix(pi): remove Calm hidden-block gaps (#942) * fix(pi): remove Calm hidden-block gaps * no-mistakes(review): Validate Calm geometry against current viewport * no-mistakes(review): Synchronize Calm geometry checks with reload completion * fix: enforce contract boundaries for ask-user findings (#945) * fix: escalate ask-user contract expansion * no-mistakes(document): Point project management to authority owner * docs: prefer direct operational paths (#946) * fix(pi): hide operational user rows in Calm mode (#948) * fix(pi): hide Calm operational user rows * no-mistakes(review): Narrow Calm operational input suppression * no-mistakes(review): Avoid Calm replay classifier subprocesses * no-mistakes(document): docs: point Pi verification to Calm owner * fix: relaunch missing second mates at session start (#950) * fix(session-start): relaunch missing second mates * fix(test): detect completed parallel workers * no-mistakes(review): Isolate session-start recovery test cleanup * no-mistakes(review): Complete backend-safe secondmate session recovery * no-mistakes(review): Resolve Zellij task ownership before recovery * no-mistakes(review): Recover relocated Zellij ghost tabs safely * no-mistakes(review): Restore conservative Zellij recovery boundary * no-mistakes(review): Reject malformed tmux recovery targets * no-mistakes(document): Align secondmate recovery documentation * no-mistakes: apply CI fixes * fix(herdr): reclaim resumed task projections after restart (#967) * fix(herdr): reclaim resumed task projections safely * no-mistakes(review): Enforce safe Herdr reclaim close boundaries * no-mistakes(document): docs: clarify Herdr restart projection contract * Teach Ahoy to surface open decisions (#968) * Require shipshape routine acknowledgement (#969) * docs: separate current guidance from verification evidence (#994) * docs: separate current guides from verification * no-mistakes(review): Restore Herdr 0.7.5 restart-reclaim verification evidence * fix: preserve Claude watcher continuity across Stop hooks (#997) * feat(claude): Stop-owned tokenless watcher continuity via asyncRewake auto-arm Claude primaries (main home and marked secondmate homes) no longer depend on the model remembering to re-arm the watcher after each wake. A tracked Stop asyncRewake hook (bin/fm-claude-stop-autoarm.sh, timeout 28800s) fires on every turn end, claims one home-scoped single-flight owner, foregrounds bin/fm-watch-arm.sh inside the hook-owned process tree, and translates an actionable close or typed watcher failure into exactly one exit-2 rewake. The hook scopes to genuine primary checkouts, requires the session lock to be held by its own harness ancestor, stays inert while AFK owns triage or the home is idle, and hands AFK transitions mid-cycle to the daemon without rewaking. The synchronous turn-end guard gains a --claude cooperative mode: it ignores stop_hook_active (true on every post-continuation stop, which is what re-opened the 2026-07-21 blind window), waits briefly for a watcher health proof, a live auto-arm owner claim, or a fresh rewake epoch, and re-blocks only when the auto-arm genuinely failed to establish - bounded to 3 consecutive blocks per session, safely below Claude Code's 8-block override, then a degraded allow with a visible systemMessage. Codex keeps the previous one-block loop guard byte-identically, and Pi, OpenCode, and Grok adapters are untouched. Continuity PreToolUse gate and durable wake queue are preserved; the gate's recovery guidance now names the Stop-owned re-arm and reserves manual background arms for auto-arm failure. Claude supervision protocol, harness-adapters facts, architecture, configuration, and continuity docs updated; docs/turnend-guard.md records the 2026-07-24 Claude 2.1.218 contract revalidation (tokenless multi-cycle rewake, no-dedup, timeout process-group kill, 8-block cap, interactive non-stall) and the 2.1.219 product live E2Es. Regression matrix: hermetic tests cover scope, identity, AFK, need, single-flight, translation, guard cooperation, budget, and registration; the new live E2E proves two full tokenless auto-arm rewake cycles with zero model arm commands; Pi and OpenCode Option B live E2Es pass unchanged. * no-mistakes(review): Fix Claude X-mode auto-arm continuity backstop * no-mistakes(review): Remove unsupported Claude contract-lab verification claims * no-mistakes(document): Update Claude auto-arm continuity documentation * fix(herdr): clean stale projections at session start (#996) * Clean stale Herdr projections at session start * no-mistakes(document): Document stale Herdr session-start projection cleanup * no-mistakes(review): Enforce locked exact Herdr projection cleanup * no-mistakes(review): Fail closed on unverified session lock ownership * no-mistakes(review): Serialize session lock acquisition atomically * no-mistakes(document): Align session-start and Herdr cleanup documentation * no-mistakes(document): Generalize lock-refusal diagnostics * no-mistakes(lint): Avoid reserved keyword in concurrency test * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * fix: recover Claude supervision without watcher-status gate (#1001) * fix: recover Claude supervision at session start * fix: remove Claude watcher-status command gate * no-mistakes(document): docs: remove stale continuity gate references * fix: make quota-aware profile selection agent-owned (#1018) * Replace quota dispatch selector instructions * no-mistakes(review): Align bootstrap docs with agent-owned dispatch selection * fix(bin): remove vestigial dispatch selector (#1026) * remove vestigial dispatch selector * no-mistakes(review): Synchronize isolation proof and portable shard evidence * no-mistakes(review): Correct shard history and proof archive date * no-mistakes(review): Remove reintroduced selector documentation reference * no-mistakes(document): Remove stale dispatch strategy documentation * docs(agents): drop superseded interim quota-window rule (#1039) quota-axi 0.1.13 emits schemaVersion 2 with a quotaSemantics object per provider, so the successor named in the interim rule has landed and the rule's own removal condition is satisfied. Keep the ownership clause so quota-axi remains the single owner of how model or product windows relate to bounding account windows, and drop the interim weakest-headroom instruction. The unknown-semantics case is already covered by the existing requirement to stop and report a candidate whose applicable quota data or interpretation cannot be established. Drop the matching assertion phrase from tests/fm-instruction-owners.test.sh; the retained ownership phrase still asserts. * fix(tmux): scope busy detection and recognize current Claude turns (#1049) * fix(tmux): scope Claude busy detection by harness * no-mistakes(review): Separate verified and fallback busy signatures * no-mistakes(test): Scope busy signatures to supplied harnesses * no-mistakes(document): Document harness-scoped busy detection * feat: add verified Kimi crewmate adapter (#1047) * Add verified Kimi crewmate harness adapter * no-mistakes(review): Scope Kimi moon detection to spinner lines * no-mistakes(review): Match only complete Kimi spinner rows * no-mistakes(review): Resolve Kimi binary portably before pane creation * no-mistakes(document): Align Kimi adapter documentation * no-mistakes(lint): Suppress false-positive ShellCheck warning for sourced watcher override * Fix Kimi busy spinner detection * no-mistakes(review): Recognize Kimi session-lock ancestry and holders * no-mistakes(review): Scope pending-reply Kimi busy detection by harness * no-mistakes(document): Correct Kimi spinner capture documentation * no-mistakes(document): Clarify optional Kimi spinner whitespace * no-mistakes(lint): Silence intentional pending-reply test stub warnings * test: align rebased Kimi busy fixtures * no-mistakes: apply CI fixes * Reconcile Kimi busy detection after per-harness scoping * no-mistakes(review): Clarify observed Kimi spinner whitespace contract * no-mistakes(document): Clarify Kimi harness documentation * fix: harden Kimi submission and spinner matching (#1058) * fix kimi pointer submission and spinner conformance * no-mistakes(review): Preserve Kimi submit target ownership guard * feat(bin): add guarded Kimi turn-end wake (#1059) * Add guarded Kimi turn-end hook * no-mistakes(review): Require jq before installing Kimi turn-end hook * no-mistakes(review): Expose jq inside isolated Kimi test fixtures * no-mistakes(review): Preserve Kimi config boundaries during hook removal * no-mistakes(review): Document Kimi removal newline safeguard * no-mistakes(document): Document Kimi shared-home preservation * fix(tmux): classify bordered composers across all rows (#1066) * Fix structural tmux composer reading * Verify Calm compatibility with Pi 0.82 * no-mistakes(review): Harden structural composer classification boundaries * no-mistakes(review): Refresh composer and Kimi regression fixtures * no-mistakes(review): Fail closed on unbounded composer edges * no-mistakes(review): Enforce aligned composer geometry safely * no-mistakes(review): Make composer ambiguity locale-safe * no-mistakes(review): Preserve ambiguity through composer submission * no-mistakes(review): Carry composer proof through retries * no-mistakes(document): Document structural tmux composer delivery guarantees * no-mistakes: apply CI fixes * feat(bin): add verified pi-signed runtime adapter (#1145) * feat: add verified pi-signed adapter * no-mistakes(review): Correct pi-signed maintainer verification date * no-mistakes(review): Correct remaining pi-signed verification dates * no-mistakes(review): Preserve authoritative pi-signed runtime identity * no-mistakes(document): Document pi-signed shared adapter semantics * no-mistakes: apply CI fixes * fix(pi): rearm watcher across session transitions (#1166) * fix(pi): rearm watcher across same-process session transitions Pi emits session_shutdown for ordinary /new, /resume, and /fork replacement as well as terminal quit. The primary watcher extension latched a module-level stopping flag on every shutdown, so a replacement session in the same process could not arm monitoring until Pi restarted. Own arm authority per session generation so only the active live generation may start, stop, or rearm the child. Replacement sessions can arm again without restarting Pi, stale prior-generation callbacks cannot mutate the active cycle, and real quit still blocks late rearm. * no-mistakes(review): Preserve Pi generation isolation and exit cleanup * no-mistakes(document): Correct Pi watcher transition documentation * feat: route crew dispatch using quota-window pace (#1172) * Consume quota-axi pace signals in dispatch profile array selection. Add quota-array-dispatch as the single owner of the pace-aware candidate choice, keep AGENTS.md to the intake boundary and load trigger, and cover the acceptance cases with sanitized schemaVersion 3 fixtures. * no-mistakes(review): Stop and report genuine quota dispatch ties * no-mistakes(document): Document quota pace freshness and uncertainty * fix: adapt Grok Stop continuation and harden endpoint cleanup (#1171) * fix(grok): adapt Stop continuation to runtime capability * no-mistakes(review): Reject ambiguous Grok Stop payloads * no-mistakes(review): Reject duplicate Grok fields and accept spaced tmux sessions * no-mistakes(review): Enforce exact tmux cleanup selectors * no-mistakes(test): Fix historical tmux fixture and validate Grok Stop * no-mistakes: apply CI fixes * fix: restore stock macOS Bash 3.2 brief scaffolding (#1093) * fix(brief): make DOD scaffolding parse-safe on stock macOS Bash 3.2 fm-brief.sh built each Definition-of-done block and the not-enabled Herdr declaration with `VAR=$(cat <<EOF ... EOF)`. On Bash 3.2 (macOS /bin/bash) the lexer scans for the command substitution's closing `)` textually and tracks quote state through the heredoc body, so a single apostrophe, unbalanced quote, or unbalanced paren in that prose breaks parsing of the whole script. Every ship-brief scaffold (no-mistakes, direct-PR, local-only) failed with `unexpected EOF while looking for matching )`. Bash 4+ parses it fine, so the breakage stayed invisible everywhere except stock macOS. Replace all four command-substitution heredocs with `IFS= read -r -d '' VAR <<EOF || true`. That removes the `$(...)` wrapper and the entire defect class regardless of future prose, and preserves the variable expansion the direct-PR and local-only bodies need. `read` keeps the heredoc's trailing newline that `$(...)` used to strip, so trim one newline to keep every generated brief byte-identical to prior output. Guard the structure, not one historical phrase: a new test rejects any heredoc nested in a command substitution anywhere in fm-brief.sh, where the old assertion pinned a single apostrophe phrase and so missed the reintroduction. Extend the stock-macOS Bash CI job from parsing one script to the whole maintained shell surface (bin/*.sh, bin/backends/*.sh, tests/*.sh), matching bin/fm-lint.sh's canonical file set so parse scope and lint scope cannot drift apart. * no-mistakes(review): Captain: harden Bash structure and inventory guards * no-mistakes(document): Align stock macOS Bash contributor checks * no-mistakes(lint): Suppress deliberate SC2016 literal fixture warnings * test: stabilize tmux teardown conformance baseline (#1209) * fix(test): pin teardown tmux baseline to historical kill selectors merge-base HEAD main collapses to HEAD after the exact-selector change lands on the default branch, so the old teardown fixture was accidentally exercising current exact targets. Resolve a content-historical permissive tmux adapter from first-parent history and force that post-squash topology inside the conformance case so main and feature branches keep the same old-vs-new contract. * no-mistakes(lint): Suppress intentional literal-pattern ShellCheck warnings * docs: slim quota-array-dispatch to the pace selection core (#1197) Cut the runtime skill to the compact pace-aware selection procedure plus minimum owner pointers. Keep every distinct decision rule and move expanded acceptance scenarios to deterministic fixture ownership assertions. Size: 170/1374/10187 -> 63/544/4068 (about 63%/60%/60% reduction). * feat(bin): inherit backend config into secondmate homes (#1219) * Inherit config/backend into secondmate homes with deliberate-override preservation Add backend to the shared inheritable config allowlist so launch, locked bootstrap, and config-push converge a primary pin into secondmate homes as each home local future-spawn default. Track last-inherited bytes in a private state provenance marker so deliberate per-home overrides survive present and absent primary convergence, keep --backend and FM_BACKEND stronger, and extend the existing inheritance tests plus docs and skill claims. * no-mistakes(review): Preserve equal unprovenanced backend overrides * no-mistakes(review): Preserve symlink overrides and verify spawn precedence * no-mistakes(review): Snapshot backend inheritance for consistent provenance * no-mistakes(review): Simplify backend inheritance to primary-authoritative convergence * no-mistakes(document): Document inherited backend override preservation * fix: restore primary-authoritative backend inheritance after document regression The document step reintroduced provenance and deliberate per-home override semantics after review had simplified config/backend to plain primary-authoritative allowlist membership. Restore the primary-always-wins path: present overwrites, absent removes, no provenance marker, and docs/tests match that contract. * no-mistakes(review): Add divergent backend precedence regression fixtures * no-mistakes(document): Document backend inheritance contract * fix(pi): remove Calm's upper version ceiling (#1226) * fix(pi): remove Calm's exclusive Pi upper-version ceiling tests/fm-calm-pi-extension.test.sh gated on a closed PI_COMPAT_VERSIONS allowlist ("0.81.1 0.82.0") that refused any other installed Pi, and docs described that range as "supported" rather than verified evidence. The Calm CHANGELOG shows no API introduced at either version, so there is no evidence for a real minimum; the presentation adapters already probe the exact method they patch rather than checking a version. Replace the allowlist with dated version evidence that never rejects a newer Pi, and make each presentation adapter degrade independently with a diagnostic if a future Pi removes its API, instead of the whole Calm extension failing to load. Rewrite the feasibility doc's "Pi 0.81.1 through 0.82.0" phrasing to state it as verified evidence, not a ceiling. * no-mistakes(review): Probe missing Calm adapter exports safely * no-mistakes(document): Document Calm's unbounded Pi compatibility * fix(bin): allow session-local todo tools in the subagent guard (#1204) * fix(guard): allow session-local todo tools in the primary The delegation-shape guard denied TaskCreate and TaskUpdate because their normalized names contain the `task` stem. Those tools write only the harness's session-local todo list, which has no executor: it spawns no agent, allocates no worktree, registers no schedule, and starts nothing that outlives the session. That is not the unaccounted work the guard exists to stop, so the stem match was a false positive, and the deny text told the primary to run bin/fm-brief.sh and bin/fm-spawn.sh to create a todo entry. Add a separately-reasoned PLAN_ONLY_TOOLS exact-name exclusion rather than widening OBSERVE_ONLY_TOOLS, whose documented contract is tools that only observe or stop existing work. Both lists stay exact-name so neither can widen by substring. Tests cover the two allowed names and six near-miss names that a substring or shortened-stem widening would release; both mutations were watched red. * no-mistakes(review): drop session-local todo tools from recommended deny list * no-mistakes: apply CI fixes * fix(session-lock): resolve Claude bg-spare ancestry to the outermost claude pid (#1206) * fix(session-lock): resolve Claude bg-spare ancestry to the outermost claude pid fm_harness_ancestry_pid() previously returned the first ancestor process whose command matched a verified harness name. Claude Code's Stop hook fires as a bg-spare worker several levels below the session's actual lock-owning claude process (hook shell -> claude bg-spare -> claude bg-pty-host -> claude -> claude(lock)), so the first match was the bg-spare worker, not the lock owner. fm_session_lock_owned_by_self() then never matched state/.lock, and the Claude Stop auto-arm silently treated its own primary session as an unrelated live owner and never armed the watcher. The walk now keeps going past a claude-named match, looking for a still more ancestral claude-named match, and stops the instant a non-match follows an already-found match (bounding it to a contiguous run rather than the literal ancestry top, so an unrelated claude-named process further up the real process tree is never mistaken for part of this session's own nested chain). Every other harness keeps the original first-match-wins behavior, since e.g. Pi's shared signed-wrapper ancestry actually holds the session at the inner engine pid, not an outer wrapper pid. Hop limit raised from 8 to 16 to cover the deeper bg-spare chain. * no-mistakes(review): Add nested-claude-ancestry regression test; fix nudge doc depth claim * no-mistakes: apply CI fixes * fix: conferma l'avvio del watcher su Windows/MSYS (#1212) * fix: confirm watcher startup on MSYS * no-mistakes(review): gate MSYS arm ready timeout, cache uname, harden locale test * no-mistakes(review): validate OpenCode ready timeout, make uname cache internal * fix(spawn): forward CLAUDE_CONFIG_DIR to claude crewmates (#1195) * fix(spawn): forward firstmate's CLAUDE_CONFIG_DIR to claude crewmates Crewmate panes are created by a long-lived tmux/herdr daemon that does not inherit firstmate's current environment. When firstmate runs under a non-default CLAUDE_CONFIG_DIR (for example a work-vs-personal subscription split), a bare `claude` in the crewmate pane fell back to the default ~/.claude store and launched unauthenticated, blocking the crewmate before it could do any work. fm-spawn now prefixes the claude launch with firstmate's own resolved CLAUDE_CONFIG_DIR when set, so the crewmate uses the same credential/config store firstmate is authenticated with. An unset value is the single-store default and adds no prefix; non-claude harnesses are unaffected. Adds three tests in fm-spawn-dispatch-profile.test.sh (forwarded-when-set, omitted-when-unset, non-claude-ignored) and pins CLAUDE_CONFIG_DIR in the test helper so launch assertions no longer depend on the developer's environment. * no-mistakes: apply CI fixes * fix: preserve dispatch identity across authentication checks (#1233) * fix: preserve dispatch harness identity * no-mistakes(review): Fix Grok counterfactual tuple validation * no-mistakes(document): Scope dispatch authentication to selected tuple * fix: restore dispatch instruction budget * no-mistakes(review): Scope dispatch authentication after candidate selection * fix(bin): normalize relative durable paths (#1256) * fix(bin): handle dash-leading harness process names (#2) * fix: handle dash-leading harness process names * no-mistakes(review): Make dash-leading harness regression hermetic * fix: preserve secondmate reply routes across relative homes Resolve relative home, data, and state inputs before durable charter generation, and fail when caller-relative directories cannot be resolved. Use absolute paths at the related spawn, AFK daemon, and X-mode cross-process handoffs so later processes cannot reinterpret them from another working directory. * no-mistakes(review): Preserve absolute overrides and normalize relative durable paths * no-mistakes(review): Normalize relative home before deriving durable paths * no-mistakes(document): Document relative durable-path normalization * no-mistakes(review): Captain: Ignore inherited CDPATH during relative path normalization * no-mistakes(lint): Fix empty CDPATH assignments for ShellCheck * refactor(skills): make Bearings chat-only by default (#1136) * Add internal status skill * no-mistakes(document): register /status skill in documentation-audiences inventory * no-mistakes(lint): replace grep|wc -l with grep -c in status skill test * test: silence literal status skill patterns * Refactor bearings default to chat-only --------- Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com> * Clarify follow-up routing during validation (#1277) * fix: honor concrete approval for project operations (#1272) * docs: add captain-approved project operation exception to hard rule 1 Firstmate stays read-only over projects by default, but when the captain clearly approves a concrete project operation and scope in the moment, firstmate may perform exactly that approved operation with its own tools. The approval is never inferred, broadened, or standing, and it does not relax the existing force, discard, unlanded-work, or merge-authority boundaries. * no-mistakes(review): Clarify captain-approved project operation boundaries * no-mistakes(document): Clarify captain-approved project operation scope * docs: cover directories and preserve the operation-or-scope alternative Widen the captain-approved project operation exception in AGENTS.md to files or directories, and restore the explicit operation-or-scope alternative that a prior pipeline auto-fix had collapsed into "and". Rework project-management SKILL.md's Remove section, which previously told firstmate to refuse project removal until a guarded helper existed; that helper was never built, so the text directly contradicted the new instruction-only exception. It now points at the exception plus the existing removal preflight it still requires unchanged. Update the one instruction-owners test assertion that hard-coded the sentence removed above, so the suite tracks current, not obsolete, text. * docs: add captain-approved project operation exception to hard rule 1 Firstmate stays read-only over projects by default, but when the captain clearly approves a concrete project operation and scope in the moment, firstmate may perform exactly that approved operation with its own tools. The approval is never inferred, broadened, or standing, and it does not relax the existing force, discard, unlanded-work, or merge-authority boundaries. * no-mistakes(review): Clarify captain-approved project operation boundaries * no-mistakes(document): Clarify captain-approved project operation scope * docs: cover directories and preserve the operation-or-scope alternative Widen the captain-approved project operation exception in AGENTS.md to files or directories, and restore the explicit operation-or-scope alternative that a prior pipeline auto-fix had collapsed into "and". Rework project-management SKILL.md's Remove section, which previously told firstmate to refuse project removal until a guarded helper existed; that helper was never built, so the text directly contradicted the new instruction-only exception. It now points at the exception plus the existing removal preflight it still requires unchanged. Update the one instruction-owners test assertion that hard-coded the sentence removed above, so the suite tracks current, not obsolete, text. * no-mistakes(review): Align project removal preflight with approved exception * no-mistakes(document): Align project removal documentation with approved exception * fix: restore removal test byte-for-byte and preserve the default sentence tests/fm-instruction-owners.test.sh had been changed to assert different text; restore it byte-for-byte to origin/main. project-management SKILL.md's Remove section now keeps the exact default "Never issue a raw removal command from Firstmate." sentence that test still asserts, immediately followed by the already-approved captain-operation-or-scope exception, so the default and the exception both stay explicit and consistent. * no-mistakes(document): Align project-write boundary documentation * fix(skills): route new project intake through secondmate scopes (#1275) * Route project intake through secondmate scopes * no-mistakes(test): Guard all main-home project registry mutations * no-mistakes(document): Consolidate secondmate routing documentation * no-mistakes: apply CI fixes * Restore new-project routing scope * no-mistakes(document): Clarify secondmate routing for new-project intake * no-mistakes: apply CI fixes * fix: scope validation corrections by accepted behavior (#1281) * fix: scope validation corrections by accepted behavior * no-mistakes(review): Classify stale delivery evidence as an autonomous correction * test: replace source assertions with behavioral coverage (#1282) * test: remove source-content assertions * no-mistakes(review): Replace source assertions with runtime behavior coverage * no-mistakes(review): Isolate Kimi task temp runtime coverage * no-mistakes(document): Refresh test cleanup documentation * no-mistakes: apply CI fixes * fix(watch): escalate busy workers with no completed turn (#1286) * fix(watch): bound how long a busy pane may run with no completed turn A busy pane (backend busy state or the harness's rendered footer) was unconditional, unbounded proof of liveness in every escalation path, so a hung foreground tool call behind a busy signature could run for hours undetected (2026-07 hibit-agent-focus-nonsteal-r1 incident: a catastrophic- backtracking regex hung one bash call for 25h behind an unchanging "Working..." footer). FM_BUSY_TURN_MAX_SECS (default 3600s) now bounds how long a busy pane may run with no completed turn (state/<id>…
|
Landed in https://github.com/lquach99/scenely/pull/564, merged to main. The PR body carried no closing keyword, so this stayed open after the merge — closing by hand. NOT YET VERIFIED AGAINST STRIPE: the two Stripe-side checks (e2e Round 5 asserting no live subscription remains after a cancellation, and the repair script's dry run against a real account) could not run in the validation environment and remain outstanding. |
kunchenguid
pushed a commit
that referenced
this pull request
Sep 16, 2026
* fix(bin): re-record PR poll identity after a volume device renumber (Fixes #4260) A volume remount can renumber the state filesystem's st_dev while every inode and byte stays the same; APFS does this across a reboot. A poll registration records its sidecar and check as device:inode, so every poll armed before the remount failed strict validation and the watcher refused all of them as unauthenticated state checks until each was re-armed by hand. There are two device comparisons. fm_pr_private_file_valid compares a live file's device with the state directory's device read in the same invocation: it refuses a file that is not on the state directory's own filesystem and already survives a renumber, so it is unchanged. The registration's recorded identity versus the live identity (from #556, reused by the #932 retirement receipt) binds the registration to the exact files published in its own transaction; its device part is what breaks. When strict capture fails, the watcher now proves the device is the only difference: every other artifact check passes (template bytes, both hashes, private mode, single link, live device, metadata), both recorded identities name one device, and each recorded inode equals its live inode. Only then, under the task's control lock, does it rewrite the two identity lines, repeating the whole proof and comparing the registration's file identity and bytes just before the rename, and then capture strictly again. A swapped, altered, re-moded, relinked, split-device, or foreign-device artifact still fails a proof and is still refused, and a pending retirement receipt blocks the rewrite. Reproduction: on macOS a poll armed on an APFS disk image that was detached and re-attached behind another image moved st_dev 16777239 -> 16777243 with inodes, bytes, mode, and link count unchanged; the real watcher refused it on main and reports its merge with this change. The portable regression test rewrites a real registration's recorded device and drives the watcher. Not changed here: the status presentation cursor keys rows by its own device:inode identity in bin/fm-classify-lib.sh, a different helper that needs its own fix; a retirement receipt left by a reboot between its publication and removal still names the old device and stays refused; custom check trust binds only a content hash and is unaffected. * fix(review): Serialize PR poll publication writers * fix(review): Bound PR poll publication lock scope
rega10
added a commit
to rega10/firstmate
that referenced
this pull request
Sep 17, 2026
* fix(bin): report verified PR state for passed runs (kunchenguid#4624) * fix(bin): derive passed PR state from PR record A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists. For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim. Fixes kunchenguid#4607 * no-mistakes(review): Add bounded GitLab merge-request state reads * no-mistakes(review): Preserve network-free inactive crew-state scans * no-mistakes(document): Document PR record readers in shared library * fix: restore published contribution follow-up (Fixes kunchenguid#4469) (kunchenguid#4627) * fix: restore published contribution follow-up (Fixes kunchenguid#4469) * fix(review): Fix contribution freshness and merge actor routing * fix(review): Restore issue triage and scope contribution follow-up * fix(test): test: assert one wake per contribution signal * fix(document): Document contribution follow-up * fix: restore truthful terminal delivery evidence * fix(review): Disclose unsupported contributions and deduplicate watcher wakes * fix(review): Preserve unmeasured unsupported contributions across Bearings * fix(review): Deduplicate shared contribution wakes and isolate diagnostics * fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower * fix(bin): make remote report transfers explicit and fail-open (kunchenguid#4658) * fix(bin): make a remote-reply document gap self-clearing and re-attemptable A remote mate's undelivered document raised a keyed `blocked` decision that nothing could ever resolve, and any `data/*.md` substring in any mirrored line was an unconditional fetch instruction. A mate announcing a report it had not written yet therefore manufactured a permanent, factually false blocker, and its own explanation of the false alarm manufactured more. The reader has no permanence vocabulary: a report still being written refuses exactly like a path that will never exist. So an undelivered document is now a durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`, re-attempted on the next delta and on the channel's own quiet poll, and retired with a matching `resolved` line naming the local copy once it arrives. The cursor still advances and no delta stalls on one bad pointer. Only a structured `report=data/....md` pointer now offers a document, so a path merely mentioned in prose - including one under another home's mirror tree, which is provably not that mate's to serve - is never fetched. Offers are deduplicated across the whole delta, the escalation names each missing document once and carries the reader's own reason instead of discarding it, and a strictly increasing notice ordinal keeps a later escalation from being swallowed as duplicate bytes. A mirrored line still lands once whichever pointer form it was first written under. * no-mistakes(review): Require structured pointer token boundaries * no-mistakes(review): Unify boundary-safe pointer extraction and rewriting * fix(bin): identify a mirrored line independently of its delivery state Two defects in the boundary-safe pointer work. The at-most-once check compared only the all-remote and all-local renderings of a line, so it could not recognize a mixed one. A line offering two documents where only the first was deliverable mirrored as local-plus-remote; once the second arrived, a cursor-loss whole-log recapture rendered the same line all-local, matched neither alternate, and mirrored a second time. A line's identity is now the canonical form every boundary-valid pointer would take once delivered, derived by the same parser that does extraction and rewriting, so it no longer depends on which documents happened to be deliverable at the time. The pointer map was passed to awk through the process environment. A delta may carry up to the configured 1 MiB bound, and an expanded map of delivered pointers can exceed the platform's exec argument limit, so awk would fail to start; because no caller checked, the empty result would have been appended as blank lines while the cursor advanced past dropped status content. The map now travels in a file, and every call site checks the exit status and stops the ingest rather than committing a delta it could not render. Both passes now run once per stream instead of twice per line. * no-mistakes(review): Abort ingest when document pointer extraction fails * no-mistakes(review): Exclude structured cross-home pointers from document transfer * fix(bin): fail open on an undeliverable remote document instead of tracking it Narrow the remote-reply document fix to the scope the diagnosis actually requires, as decided after measuring a simpler alternative. A document the reader cannot deliver now fails open. The mate's line is mirrored with its own pointer, the cursor advances, and one unkeyed note carries the reader's reason. A note never enters the open-decision fold, so it cannot stand open the way the original keyed block did - which removes the never-clearing false blocker by construction rather than by resolving it. That makes the durable self-clearing obligation unnecessary, so it goes: the per-mate pending-documents record, its notice ordinal and resolved announcements, and the poll-side retry. Canonical line identity goes too, and with it a way to silently drop a genuine status line; mirroring is back to at-most-once on exact bytes. The cross-home exclusion goes as well: under fail-open a cross-home report= either fails harmlessly or is a nested remote report this mate genuinely holds, which is now relayed again. Kept: fetching only on a structured report= pointer, the boundary-correct parser, the file-based rewrite map, and checked extraction and rewrite exit status. The parser now scans behind a sentinel byte so a rejected candidate can no longer give the text right after it a false leading boundary. The reported incident is covered end to end: a report path announced in prose before it exists raises no decision, and the report still arrives through the ledger publisher's structured offer once written. * no-mistakes(review): Preserve source-line identity across remote reply replays * no-mistakes(document): Document remote reply transfer and replay semantics * no-mistakes(lint): Fix staging truncation lint checks * fix(calm): preserve substantive mid-turn responses (kunchenguid#4655) * Preserve substantive Calm mid-turn text * no-mistakes(review): Distinguish newline-preserved replies from short narration * no-mistakes(document): Document Calm mid-turn preservation boundaries * no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check` * fix(bin): preserve PR merge polls across volume remounts (kunchenguid#4656) * fix(bin): re-record PR poll identity after a volume device renumber (Fixes kunchenguid#4260) A volume remount can renumber the state filesystem's st_dev while every inode and byte stays the same; APFS does this across a reboot. A poll registration records its sidecar and check as device:inode, so every poll armed before the remount failed strict validation and the watcher refused all of them as unauthenticated state checks until each was re-armed by hand. There are two device comparisons. fm_pr_private_file_valid compares a live file's device with the state directory's device read in the same invocation: it refuses a file that is not on the state directory's own filesystem and already survives a renumber, so it is unchanged. The registration's recorded identity versus the live identity (from kunchenguid#556, reused by the kunchenguid#932 retirement receipt) binds the registration to the exact files published in its own transaction; its device part is what breaks. When strict capture fails, the watcher now proves the device is the only difference: every other artifact check passes (template bytes, both hashes, private mode, single link, live device, metadata), both recorded identities name one device, and each recorded inode equals its live inode. Only then, under the task's control lock, does it rewrite the two identity lines, repeating the whole proof and comparing the registration's file identity and bytes just before the rename, and then capture strictly again. A swapped, altered, re-moded, relinked, split-device, or foreign-device artifact still fails a proof and is still refused, and a pending retirement receipt blocks the rewrite. Reproduction: on macOS a poll armed on an APFS disk image that was detached and re-attached behind another image moved st_dev 16777239 -> 16777243 with inodes, bytes, mode, and link count unchanged; the real watcher refused it on main and reports its merge with this change. The portable regression test rewrites a real registration's recorded device and drives the watcher. Not changed here: the status presentation cursor keys rows by its own device:inode identity in bin/fm-classify-lib.sh, a different helper that needs its own fix; a retirement receipt left by a reboot between its publication and removal still names the old device and stays refused; custom check trust binds only a content hash and is unaffected. * fix(review): Serialize PR poll publication writers * fix(review): Bound PR poll publication lock scope * fix(bin): keep contribution records when the poll budget runs out (follow-up to kunchenguid#4627) (kunchenguid#4661) A budget that expires partway through an observation no longer records an error or prints the unavailable wake; the URL keeps its prior record and is observed first next poll. forge() flags budget exhaustion at the point it refuses, or when a read is killed at the budget's own deadline, so a genuine forge failure still records the error and wakes. Each distinct URL is now observed once per poll and applied to every owning task. * fix(bin): clear parent pending-replies on local secondmate retirement (kunchenguid#4680) * fix(bin): clear parent pending-replies on local secondmate retirement Local secondmate teardown left resolved parent pending-reply records behind after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced retirement while any reply for that id is still unresolved, and delete every matching record plus its delivery confirmation after a successful local or remote retirement, matching the remote cleanup path. * no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup * no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung * no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern * no-mistakes(review): Pending-replies Basename und corr_id abgleichen * no-mistakes(document): Clarify forced retirement pending-reply cleanup --------- Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc> * fix(bin): accept Orca's composite worktree id when tearing down a task (kunchenguid#4677) * fix(bin): accept Orca's composite worktree id at teardown Teardown refused every Orca-backed task because the endpoint validator checked orca_worktree_id with the simple-atom rule meant for tmux-style window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca returns that id as `<orca id>::<absolute worktree path>`, so the colon and slashes in every real value made validation fail and finished Orca tasks could never be cleaned up. Validate the field as the composite it is: both halves of the first `::` split present, the path half absolute, and no embedded newline, carriage return, or tab. The terminal field keeps the atom check, which is correct for it, and no other backend's validation changes. The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca never returns, which is why the suite passed a check the real value fails. They now carry the composite form, so the tests exercise the real value. * no-mistakes(document): name Orca's repo id in the composite worktree id * no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points * feat(bin): add opt-in typed dispatch resolution (kunchenguid#4692) * feat(bin): add opt-in typed dispatch resolution through typesafe.ai Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or scout profile from a written brief with typesafe.ai's System One model: one Choice question over the rules' `when` texts, then the confidence floor, the rule's `approval` and `floor`, each profile's `provider` and `floor`, one quota-axi snapshot, and the spendPriority argmax all in code. It is off unless TYPESAFE_API_KEY is in the environment or the home's gitignored .env; off means one stderr line, exit 0, and no network call, so firstmate dispatches exactly as before. The key reaches curl on a file descriptor, never argv. Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates the four new optional dispatch fields. Document the schema, the operator contract, the AGENTS.md intake step, and the live and benchmark evidence. * no-mistakes(review): Harden typed dispatch resolution and quota bounds * no-mistakes(review): Validate dispatch floors and ranking evidence * no-mistakes(review): Tighten dispatch response and floor evidence * no-mistakes(review): Neutralize none matching and resolve defaults locally * no-mistakes(review): Preserve providerless profiles outside typed resolution * no-mistakes(review): Validate response usage and reject duplicate profiles * no-mistakes(review): Escalate unverifiable floors and validate probabilities * no-mistakes(review): Validate probability mass and unknown profile floors * no-mistakes(review): Simplify resolver interface and preserve fallback routing * no-mistakes(review): Fix constants and rank partial quota evidence * no-mistakes(review): Add authoritative provider mapping and enforce explicit providers * no-mistakes(review): Declare provider for documented Pi profile * no-mistakes(review): Validate provider identifiers and support Gemini dispatch * no-mistakes(review): Strictly anchor provider identifiers * no-mistakes(review): Validate selectors and preserve fallback candidate evidence * no-mistakes(review): Gate typed validation and harden resolver evidence * no-mistakes(review): Preserve opt-in routing and harden candidate evidence * no-mistakes(review): Prioritize known exhaustion over quota uncertainty * no-mistakes(review): Isolate API secrets and preserve no-key diagnostics * no-mistakes(review): Fallback safely when dispatch rules are absent * no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets * no-mistakes(document): Document typed dispatch safety and fallback behavior * fix(bin): read the latest status event so buried declarations and open decisions aren't lost (kunchenguid#3753) * test: reproduce buried status declarations in shared readers * fix: share status event reads and preserve open blockers * fix: retain terminal scout and ship status declarations * no-mistakes(review): Fix status chronology, legacy completions, and reader performance * no-mistakes(review): Share terminal decision reconciliation across fleet snapshots * no-mistakes(review): Unify terminal supersession across cached folds and consumers * no-mistakes(review): Filter per-key status history while preserving terminal chronology * no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells * no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses * no-mistakes(document): Document latest-event status read and kind-scoped fold cursor * no-mistakes(lint): Quote literal done in test for-lists for SC1010 * ci: expect 19 snapshot/fleet-view tests This branch adds a fleet-snapshot regression, so the stock macOS Bash lane's hardcoded guard of 18 'ok - ' lines fails on the new count. Bump the guard and its message to 19. * no-mistakes(review): Restore multiline child outcome reporting * no-mistakes(review): Select ledger terminal events through bounded shared reader * no-mistakes(review): Report newest open decision instead of preferring blocked * no-mistakes(review): Require colon before ship/scout terminal supersession in fold * no-mistakes(review): Gate socket-down override on latest event; drop lock matrix * no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions * no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold * no-mistakes(test): Update fleet-view expectations to newest-open-decision rule * no-mistakes(document): Align status-read docs with fold-resolved crew state * no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers * no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree * test: fold terminal-cleanup snapshot coverage into the completed-scout case Keep the ship/scout/secondmate supersession assertions without adding a nineteenth top-level fleet-view test, so CI can stay at the upstream suite count. * no-mistakes(document): Clarify socket-down override expiry in architecture doc * ci: retrigger flaky contribution check * fix(bin): launch codex crewmates with codex's hook layer disabled (kunchenguid#4689) * fix(spawn): launch codex crewmates with codex's hook layer disabled A freshly launched Codex worker never reached its instructions. Codex stopped it on an interactive "Hooks need review" modal whose selection sits on "Review hooks", which is neither trusting nor declining. Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow navigation, so the selection cannot be moved, and pre-accepting the prompt by writing Codex's own trust store would record an operator consent that was never given. The hooks are the machine's own ~/.codex/hooks.json plus any project's .codex/hooks.json. A crewmate needs neither: its turn-end signal is the -c notify= program on the same launch, and Firstmate's project hooks are primary-session infrastructure that stands down in a child worktree. Crewmate and scout launches now pass --disable hooks. That is the opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted hooks; disabling the feature runs none of them and leaves the operator's ~/.codex untouched. An unknown feature name is a hard Codex error, so a release that drops the flag fails the launch loudly instead of silently restoring the modal. A secondmate is a primary in its own home and keeps the project hooks its turn-end guard and session-start digest ride on. Verified on codex-cli 0.151.0: the modal is gone and the turn-end notification still lands. This unblocks the second review that every finished pull request is supposed to get. Fixes kunchenguid#4673 * no-mistakes(review): Fix contradictory hook count in Codex verification record * no-mistakes(document): Clarify typed dispatch gate ownership --------- Co-authored-by: Joseph Kim <jokim1@gmail.com> Co-authored-by: Mickaël Rémond <mremond@process-one.net> Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Co-authored-by: Sebastian <80847374+thelad-dev@users.noreply.github.com> Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc> Co-authored-by: Juan José González Giraldo <juanjose.eng@gmail.com> Co-authored-by: Tiago <tiagop@hey.com> Co-authored-by: Cody <72239807+codyjohnsontx@users.noreply.github.com> Co-authored-by: Rene Garza Jr. <rega1011@RG-Mac-mini.local>
vifar
added a commit
to vifar/firstmate
that referenced
this pull request
Sep 17, 2026
… commits, adopt 31 upstream commits) (#13) * fix(bin): translate Stop hook timeout signals into durable auto-arm failure (#4474) * fix(bin): recover Claude auto-arm after timeout * no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list * fix(spawn): establish Claude task channel authority (#4464) * fix(spawn): establish Claude task channel authority * no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference * fix(bin): refuse fm-control.sh exit when the composer holds unproven or pending text (#4458) * fix: guard relaunch exit against pending input * no-mistakes(review): Verifying test run in progress * no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard * no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now * fix(spawn): establish crewmate identity first (#4481) * fix(bin): reconcile redundant secondmate divergence during updates (#4460) * fix: reconcile diverged secondmate updates * no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md * no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md * feat: enable gpt-5.6-luna max reasoning for crew dispatch (#4497) * fix(dispatch): support Codex Luna max effort * no-mistakes(review): use portable CODEX_HOME path in codex effort reference * feat(calm): render smooth Unicode swell with asymmetric two-color sail (#4498) * feat(calm): render smooth Unicode swell * feat(calm): make sails asymmetric * feat(calm): use quarter sail glyph * no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer * no-mistakes(document): docs: sync calm wave phase doc comment * no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。 * fix(bin): supersede stale scout delivery text in brief.md on promotion (#4491) * fix: supersede scout delivery brief on promotion * fix: preserve ship safety contract after promotion * no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch * fix(bin): make captain holds work on hosts with an older JSON::PP, and stop cleanup dropping accents from a held body (#4471) * fix(bin): let captain holds work on hosts with an older JSON::PP Holding a task for the captain, and the cleanup that keeps a captain-held row open, both fail outright on any host whose JSON::PP defaults allow_nonref off - 2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`, but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older library rejects that whole value with "must be object or array". The consequence is fleet-wide on such a host, not one broken command: a worker there cannot formally record a decision for the captain at all. It can only mention the decision in passing in a status line, where it can be missed - which is how a real decision goes unrecorded. The hold reports that the task lost its hold-set stamp; the cleanup cannot return the row to Queued. Both call sites now ask for allow_nonref explicitly rather than inheriting whatever the installed library defaults to. The second one is worth naming: its `/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string case that fails, so the guard selects for the failing input rather than protecting against it. The regression case forces the older default back off for every perl the commands spawn, then drives both paths - holding a task that carries a body, and tearing down a captain-held row whose deliverable must still be appended. It also probes that the simulation genuinely rejects a bare scalar, so the case cannot pass vacuously on a lenient host. Each half was verified failing on its own unfixed call site with that site's real error message. Suites: fm-captain-hold-lifecycle 51 cases, fm-backlog-atomicity 99 cases, 0 failures. Verification limit: the mechanism is reproduced and tested, but neither fix is verified against a real JSON::PP 2.27202 host, because none is in the loop. This laptop runs 4.06, where the bug does not manifest. `bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a brace-delimited object before decoding, so allow_nonref never applies. * fix(bin): stop cleanup silently dropping accented characters from a held body Cleanup rewrites a captain-held row's body to append the finished work's deliverable, and the decoder it reads that body with printed decoded characters to a stream with no `:raw` layer. A character at or below U+00FF then came out as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the accent. `fm_backlog_retain` writes that body straight back through `--body-file`, and nothing reported an error - the character was simply gone from a row still waiting on the captain. The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus `utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already used. Review of the parent commit found this on one of the lines that commit already changed. It predates that change. The test asserts bytes rather than decoded strings, because comparing strings cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any character above U+00FF makes perl print the whole string as UTF-8, so one body carrying both an accent and an em dash passes even unfixed and proves nothing. Verified failing before the fix on the accented row, passing after. Suites: fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures. * no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc * no-mistakes(review): drop whole-file UTF-8 check from retained-body test * no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc * no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc * fix(bin): read codex 0.154's idle braille starfield rows as composer furniture (#4532) * fix(composer): read codex 0.154's idle starfield and status footer as furniture codex-cli 0.154.0 animates a braille "starfield" around its idle composer: on the row above the bold `›` prompt row, on the `›` row behind the SGR-2 dim `Ask Codex to do anything` placeholder, and on the row below it, then draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`). The cells are truecolor greys on both sides of the ghost luminance ceiling, so the brighter ones survive ghost stripping, and the rows below the glyph carry no structural edge. The shared classifier selected the bare `›` shape, extended its wrap region over the two rows beneath the glyph, read the survivors and the footer as wrapped typed input, and answered `pending`; the steering doorbell defers on exactly that verdict, so no doorbell ever reached an idle codex 0.154 pane. bin/fm-composer-lib.sh now recognises that furniture by shape, declared once next to the idle placeholders and reached from the two wrap-region boundary points: - a row whose non-whitespace content is entirely braille cells (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it never counts as wrapped typed content and bounds a bare composer's wrap region; braille behind the glyph row's content is stripped before the emptiness decision when nothing else follows the glyph; a row mixing braille with other text stays typed content; - the codex status footer bounds the wrap region exactly as omp's status row does, anchored on the effort token, a spaced middle dot, and a `~` or `/` path cell, so a typed `fix · tests` stays composer input; - `^Ask Codex to do anything$` joins the verified idle-placeholder set; the ghost strip remains what proves that row empty, and the bare-row rule that bright placeholder text is real input is unchanged. Unchanged: the strict blank-row rule, the styled=0 degradation (a plain cmux/orca capture of this screen still reads `unknown`, never `pending`), FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape. tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte with the divergence (letters in place of the starfield read `pending`) and the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh is the default-on live guard (token-free, skips explicitly without codex or tmux) that launches the installed codex idle and asserts `empty` through both the tmux and the cursorless styled reads, naming codex --version on failure. docs/verification/runtime-backends.md records the dated Herdr evidence: `pending` before, `empty` after, on the captured screen. * no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry --------- Co-authored-by: Todd Billings <todd@usdvcapital.com> * fix(bin): refuse empty text steers in fm-send (#4259) * fix(bin): refuse empty text steers in fm-send A marked secondmate request sent with an empty message delivered only marker and correlation bytes and minted a pending-reply expectation the parent could never see resolved, stalling the fleet with no loud error (#4255). Fail closed on an empty or whitespace-only message on the text path, mirroring the existing --resolve-key refusal. * chore: retain ambient Pi-lens autoformat as its own commit Formatting-only edits produced by ambient Pi-lens autoformat during the msg-loss investigation, kept separate from the behavioural change in c23acba6 so the fix stays reviewable on its own. AGENTS.md is deliberately excluded: its only autoformat edit stripped the trailing space from the documented FM_OPERATIONAL_PREFIX value, which bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11 records as permanent compatibility. Documenting that constant without its trailing space makes the doc wrong about the contract, so that one line was restored rather than retained. * fix(calm): paint the working ship one yellow over all-blue water (#4554) On rose-pine-moon the two-color water (cyan crests over blue troughs) read as a pink stripe over aqua, the yellow left sail and mast clashed with the red right sail, and the hull carried a blue interior run. Every water cell is now blue so the swell reads through glyph height alone, and both sail halves, the mast, and the whole hull are one yellow run. Geometry, cadence, animation, direction flip, resize clamping, and the narrow fallback are unchanged. Update the unit and real-TUI color assertions to the new palette and the Calm docs that described the old one. * fix(bin): stop aging a second mate's active turn from its launch (#4270) * fix(watch): stop aging a second mate's active turn from its launch The parent watcher's second-mate wake-loop stall check exempts a mate that is demonstrably inside an active turn, but secondmate_in_active_turn asked busy_turn_over_age first and returned "not in a turn" whenever that said the bound was crossed. busy_turn_over_age ages from state/<task>.turn-ended, falling back to state/<task>.meta. A second mate's turns end in its own home, so the parent never gets a turn-ended mark for it and the fallback ages the mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago was therefore permanently "over age", the busy pane was never consulted, and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false wake-loop stall. The gate now bounds the busy exemption by <idle> - how long the queue's drain position has not moved - which is evidence this home actually holds. A busy mate stays exempt while the queue has been frozen for less than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so the bound that stops a busy pane from proving liveness forever is kept rather than removed. busy_turn_over_age is untouched; its remaining callers are the ordinary crew busy-pane bound. The regression pins the case that actually broke: a mate whose launch record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn must not escalate, while the same mate with its queue frozen past the bound still publishes exactly one notification. The existing coverage only exercised a freshly launched mate, which passes either way. Reaching that alert now costs a pane capture inside the gate, so the three checkpoints in this suite that assert an alert move from a 1s to a 4s bound - the value the neighbouring active-turn cases already use. The bound is a ceiling, not a wait: the checkpoint returns on the first actionable wake. On a loaded machine a 1s bound missed the alert repeatedly; at 4s it did not miss in 20 runs under the same load. * no-mistakes(review): scope the second-mate active-turn regression test's coverage claim * no-mistakes(document): fix stale second-mate active-turn comments in fm-watch * feat(bin): add read-only PR blocker and reviewer discovery commands (#4278) * feat(bin): add read-only PR blocker and reviewer-discovery commands Two focused, opt-in commands that read GitHub and never write to it. fm-pr-state.sh reports what still blocks one pull request from the author's side: a closed or merged state, draft state, unknown or conflicting mergeability, absent or failing required checks, and a blocking CHANGES_REQUESTED decision explained by each reviewer's latest verdict, marked STALE when it was left at a superseded head. A pull request that only awaits an approval is not reported as blocked, and advisory checks are omitted. Every reading is taken against one exact head; a push that lands mid-read invalidates the whole result rather than mixing two snapshots. fm-pr-reviewers.sh suggests reviewers from the most recent commits to the pull request's exact changed paths, counting each commit once, resolving handles through GitHub's own commit author.login mapping, and excluding the author and Bot accounts. Both stay read-only: no review request, no approval, no merge. Unresolved review-thread state is left unreported because the REST API does not expose it and unattended commands may not use GraphQL. Closes #3731 * no-mistakes(review): accept only PR URLs and stop at terminal state * no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate * no-mistakes(review): stop attributing readings to unverified heads * no-mistakes(review): narrow readiness contract to checks that have reported * no-mistakes(review): read the pull request once, drop the head guard * no-mistakes(document): scope pr-forge isolation proof to its measured members * no-mistakes(document): record uncovered pr-forge members and their pending proof * docs(isolation-proof): re-prove pr-forge at its full membership tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the pr-forge family in this branch, and script_allows_concurrency grants four workers by family membership alone, so both ran concurrently on a proof measured before they existed. Re-proved the family at all eight members: two consecutive runs, 0 failures, each begun with the one-minute load average below 6.0 so the result measures isolation rather than contention. A third run taken between them is disclosed rather than recorded, because it started while the previous run's workers were still decaying. The new durations are not comparable with the six-member measurement above them, so they are not presented as evidence about the two new members, and that record's 1.72x four-worker figure is left as a statement about its own run rather than restated as current. * no-mistakes(review): disclose gh error-text coupling at its matching site and tests * fix(bin): teach validation-round pauses in generated briefs (#2752) * fix(bin): teach validation-round pauses in briefs * no-mistakes(document): Point classifier comments to authoritative pause examples * docs(readme): add star history chart (#4558) * fix(bin): refuse teardown when a task's endpoint close fails (#4510) * fix(teardown): refuse a cleanup whose endpoint close failed bin/fm-teardown.sh discarded both the exit status and the stderr of every fm_backend_kill call, so a close that genuinely failed was indistinguishable from one that succeeded. Teardown continued past it, deleted the task's durable records, returned its worktree, and reported the cleanup as completed. The deleted metadata is the only record of which endpoint belongs to the task, so such a close did not merely leave a stray session behind, it stranded one: nothing was left on disk naming it. The adapters could not carry that signal either. Driven against the real code, every backend arm returned 0 for a genuine failure exactly as it did for an already-exited endpoint, so there was nothing for the four call sites to propagate even once they stopped swallowing it. The tmux arm now resolves a close that did not succeed against the window's exact recorded identity, since kill-window fails the same way for a window that is gone and one that is still there. The Orca arm reports a close its missing CLI never attempted. Both stay silent for an endpoint that is already legitimately gone, and the remaining arms are unchanged: their close-command timing cannot be established without the real Zellij, Orca, and cmux binaries, and a gate that refused ordinary cleanup of an already-exited session would be worse than the defect. docs/verification/runtime-backends.md records what each backend can prove. A reported close failure now reaches teardown's existing retain-and-stop refusal before the records naming the endpoint are removed, matching where the Herdr confirmed-gone gates already sit for the same hazard, and the retained records let a rerun finish once the close works. * no-mistakes(review): refuse unreadable tmux close re-read; honor --force override * no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close * no-mistakes(document): document endpoint-close refusal in its backend and retirement owners * no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree * feat(calm): add flag-gated Claude Code Calm mode (#4565) * feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that brings Calm to Claude Code: the sailboat replaces the stock working row through a Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration, and canonically classified operational user rows draw at zero height. /calm is registered by the hooks module itself and toggles the same per-home config/calm preference the Pi extension uses, so one choice applies on either harness; rows redraw retroactively on toggle and stay hidden across claude --continue. The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS flag is on. Nothing sets that flag in any settings file, and the plugin carries no command file, skill, agent, or classic hook, so it is a complete no-op while the flag is off. The trusted project auto-loads it through an .agents/skills symlink, the only path Claude Code scans for project plugins. Extract the working-ship geometry, bounce track, cadences, and freeze/resume state into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module imports from outside the plugin folder) and have the Pi widget paint that core's frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify operational rows through a port of bin/fm-operational-input.sh's classify command guarded by a corpus parity test against the shell owner. Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering, Raster packing, policy, classifier parity), the mod's own claude plugin test suites behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op, the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272. Docs: record the version-scoped Claude Code evidence and the three bounded gaps in docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md, and make the shared preference, layout, and contributor notes harness-neutral. * no-mistakes(review): Preserve colliding final replies and strengthen parser parity * no-mistakes(review): Preserve final replies and strengthen canonical parity checks * no-mistakes(review): Require exact function-hooks opt-in before Calm activation * no-mistakes(review): Clarify Calm module loading and activation boundaries * no-mistakes(review): Reset Calm presentation state across session starts * no-mistakes(document): Refresh Calm session lifecycle documentation * feat(calm): paint the Claude Code working ship in Claude's own theme colors The captain picked the "Claude native" palette for the Claude Code mod's Raster: every water cell takes the spinner blue of the active theme family (#93a5ff dark, #5769f7 light) and the whole boat takes the Claude orange of the stock spinner (#d77757), one water color and one boat color. The family follows the `theme` setting's prefix, read at load through $.config.list and re-read on a config.set of that row, with `auto` and custom themes falling back to the dark set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte. Rename the shared sprite's color classes from hue names to `water` and `boat`, since each harness now maps them to its own colors; geometry, motion, cadence, and the activation gate are untouched. Tests cover both palettes' packing and the family rule under Node, and the plugin kit drives every theme value, a theme change mid-session, the Calm-off pass-through, and inertness of the menu read while the flag is off. The docs describe the Claude Code colors and record the guard passing on 2.1.273. * no-mistakes(review): Use light palette for unresolved Claude themes * no-mistakes(document): Refresh Claude Calm verification evidence * fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586) * fix(watch): honour a declared wait before wedge-escalating a quiet pane wedge_timer_check escalated on elapsed idle time alone. Nothing asked whether the worker had already said why its pane was quiet, so a lane that declared a bounded external wait climbed the escalation ladder for as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every repeat carried demand-deep-inspection - which by its own wording forbids re-absorbing on the run-step or pane state, so the supervisor could not use the evidence that was there either. The generated brief promises that declaring `paused:` buys the long recheck cadence instead of a wedge, but the timer was still reachable while that declaration stood: a crew that declares a wait and then has an active run or busy pane attributed to it is handed to the timer as provably-working. The declaration is what the worker said about its own silence, so it now outranks a liveness verdict that only says something is running. The consult runs in the at-threshold branch that was about to escalate, beside the worktree walk already there, and costs one status-line read. Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS recheck the declared-wait absorber already uses, so the wait is still rechecked and cannot rot invisibly. Which verb declared it decides the wording, because the two block on different people: a `paused:` wait is owed by an external dependency and asks the reader to confirm it still holds, while a `captain-held:` transfer is owed by the captain reading the recheck and asks them to answer or release the hold. A hold is not rechecked at all while the away-posture record exists, as on every other captain-held path, and that absorb arms no throttle so the recheck is owed in full on return. A declared clearing time that has already passed stops counting, and a lane that never declared one keeps the identical escalation schedule, reason, count and demand-deep-inspection wording, so detection and its worst-case time are unchanged. The deferral restarts the idle timer rather than cancelling it, so a lane that stops waiting escalates again within one threshold. A lane quiet because its own validation run is parked at a gate awaiting a human decision is deliberately out of scope: reading that state needs a signal carrying who the wait is on and what clears it, rather than one inferred from a parked verdict that also covers gates awaiting the crewmate itself. Tests pin both directions for each case and were each confirmed to fail with the consult removed. * no-mistakes(document): docs: honour declared waits in stale-escalation docs * fix(bin): report verified PR state for passed runs (#4624) * fix(bin): derive passed PR state from PR record A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists. For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim. Fixes #4607 * no-mistakes(review): Add bounded GitLab merge-request state reads * no-mistakes(review): Preserve network-free inactive crew-state scans * no-mistakes(document): Document PR record readers in shared library * fix: restore published contribution follow-up (Fixes #4469) (#4627) * fix: restore published contribution follow-up (Fixes #4469) * fix(review): Fix contribution freshness and merge actor routing * fix(review): Restore issue triage and scope contribution follow-up * fix(test): test: assert one wake per contribution signal * fix(document): Document contribution follow-up * fix: restore truthful terminal delivery evidence * fix(review): Disclose unsupported contributions and deduplicate watcher wakes * fix(review): Preserve unmeasured unsupported contributions across Bearings * fix(review): Deduplicate shared contribution wakes and isolate diagnostics * fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower * fix(bin): make remote report transfers explicit and fail-open (#4658) * fix(bin): make a remote-reply document gap self-clearing and re-attemptable A remote mate's undelivered document raised a keyed `blocked` decision that nothing could ever resolve, and any `data/*.md` substring in any mirrored line was an unconditional fetch instruction. A mate announcing a report it had not written yet therefore manufactured a permanent, factually false blocker, and its own explanation of the false alarm manufactured more. The reader has no permanence vocabulary: a report still being written refuses exactly like a path that will never exist. So an undelivered document is now a durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`, re-attempted on the next delta and on the channel's own quiet poll, and retired with a matching `resolved` line naming the local copy once it arrives. The cursor still advances and no delta stalls on one bad pointer. Only a structured `report=data/....md` pointer now offers a document, so a path merely mentioned in prose - including one under another home's mirror tree, which is provably not that mate's to serve - is never fetched. Offers are deduplicated across the whole delta, the escalation names each missing document once and carries the reader's own reason instead of discarding it, and a strictly increasing notice ordinal keeps a later escalation from being swallowed as duplicate bytes. A mirrored line still lands once whichever pointer form it was first written under. * no-mistakes(review): Require structured pointer token boundaries * no-mistakes(review): Unify boundary-safe pointer extraction and rewriting * fix(bin): identify a mirrored line independently of its delivery state Two defects in the boundary-safe pointer work. The at-most-once check compared only the all-remote and all-local renderings of a line, so it could not recognize a mixed one. A line offering two documents where only the first was deliverable mirrored as local-plus-remote; once the second arrived, a cursor-loss whole-log recapture rendered the same line all-local, matched neither alternate, and mirrored a second time. A line's identity is now the canonical form every boundary-valid pointer would take once delivered, derived by the same parser that does extraction and rewriting, so it no longer depends on which documents happened to be deliverable at the time. The pointer map was passed to awk through the process environment. A delta may carry up to the configured 1 MiB bound, and an expanded map of delivered pointers can exceed the platform's exec argument limit, so awk would fail to start; because no caller checked, the empty result would have been appended as blank lines while the cursor advanced past dropped status content. The map now travels in a file, and every call site checks the exit status and stops the ingest rather than committing a delta it could not render. Both passes now run once per stream instead of twice per line. * no-mistakes(review): Abort ingest when document pointer extraction fails * no-mistakes(review): Exclude structured cross-home pointers from document transfer * fix(bin): fail open on an undeliverable remote document instead of tracking it Narrow the remote-reply document fix to the scope the diagnosis actually requires, as decided after measuring a simpler alternative. A document the reader cannot deliver now fails open. The mate's line is mirrored with its own pointer, the cursor advances, and one unkeyed note carries the reader's reason. A note never enters the open-decision fold, so it cannot stand open the way the original keyed block did - which removes the never-clearing false blocker by construction rather than by resolving it. That makes the durable self-clearing obligation unnecessary, so it goes: the per-mate pending-documents record, its notice ordinal and resolved announcements, and the poll-side retry. Canonical line identity goes too, and with it a way to silently drop a genuine status line; mirroring is back to at-most-once on exact bytes. The cross-home exclusion goes as well: under fail-open a cross-home report= either fails harmlessly or is a nested remote report this mate genuinely holds, which is now relayed again. Kept: fetching only on a structured report= pointer, the boundary-correct parser, the file-based rewrite map, and checked extraction and rewrite exit status. The parser now scans behind a sentinel byte so a rejected candidate can no longer give the text right after it a false leading boundary. The reported incident is covered end to end: a report path announced in prose before it exists raises no decision, and the report still arrives through the ledger publisher's structured offer once written. * no-mistakes(review): Preserve source-line identity across remote reply replays * no-mistakes(document): Document remote reply transfer and replay semantics * no-mistakes(lint): Fix staging truncation lint checks * fix(calm): preserve substantive mid-turn responses (#4655) * Preserve substantive Calm mid-turn text * no-mistakes(review): Distinguish newline-preserved replies from short narration * no-mistakes(document): Document Calm mid-turn preservation boundaries * no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check` * fix(bin): preserve PR merge polls across volume remounts (#4656) * fix(bin): re-record PR poll identity after a volume device renumber (Fixes #4260) A volume remount can renumber the state filesystem's st_dev while every inode and byte stays the same; APFS does this across a reboot. A poll registration records its sidecar and check as device:inode, so every poll armed before the remount failed strict validation and the watcher refused all of them as unauthenticated state checks until each was re-armed by hand. There are two device comparisons. fm_pr_private_file_valid compares a live file's device with the state directory's device read in the same invocation: it refuses a file that is not on the state directory's own filesystem and already survives a renumber, so it is unchanged. The registration's recorded identity versus the live identity (from #556, reused by the #932 retirement receipt) binds the registration to the exact files published in its own transaction; its device part is what breaks. When strict capture fails, the watcher now proves the device is the only difference: every other artifact check passes (template bytes, both hashes, private mode, single link, live device, metadata), both recorded identities name one device, and each recorded inode equals its live inode. Only then, under the task's control lock, does it rewrite the two identity lines, repeating the whole proof and comparing the registration's file identity and bytes just before the rename, and then capture strictly again. A swapped, altered, re-moded, relinked, split-device, or foreign-device artifact still fails a proof and is still refused, and a pending retirement receipt blocks the rewrite. Reproduction: on macOS a poll armed on an APFS disk image that was detached and re-attached behind another image moved st_dev 16777239 -> 16777243 with inodes, bytes, mode, and link count unchanged; the real watcher refused it on main and reports its merge with this change. The portable regression test rewrites a real registration's recorded device and drives the watcher. Not changed here: the status presentation cursor keys rows by its own device:inode identity in bin/fm-classify-lib.sh, a different helper that needs its own fix; a retirement receipt left by a reboot between its publication and removal still names the old device and stays refused; custom check trust binds only a content hash and is unaffected. * fix(review): Serialize PR poll publication writers * fix(review): Bound PR poll publication lock scope * fix(bin): keep contribution records when the poll budget runs out (follow-up to #4627) (#4661) A budget that expires partway through an observation no longer records an error or prints the unavailable wake; the URL keeps its prior record and is observed first next poll. forge() flags budget exhaustion at the point it refuses, or when a read is killed at the budget's own deadline, so a genuine forge failure still records the error and wakes. Each distinct URL is now observed once per poll and applied to every owning task. * fix(bin): clear parent pending-replies on local secondmate retirement (#4680) * fix(bin): clear parent pending-replies on local secondmate retirement Local secondmate teardown left resolved parent pending-reply records behind after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced retirement while any reply for that id is still unresolved, and delete every matching record plus its delivery confirmation after a successful local or remote retirement, matching the remote cleanup path. * no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup * no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung * no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern * no-mistakes(review): Pending-replies Basename und corr_id abgleichen * no-mistakes(document): Clarify forced retirement pending-reply cleanup --------- Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc> * fix(bin): accept Orca's composite worktree id when tearing down a task (#4677) * fix(bin): accept Orca's composite worktree id at teardown Teardown refused every Orca-backed task because the endpoint validator checked orca_worktree_id with the simple-atom rule meant for tmux-style window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca returns that id as `<orca id>::<absolute worktree path>`, so the colon and slashes in every real value made validation fail and finished Orca tasks could never be cleaned up. Validate the field as the composite it is: both halves of the first `::` split present, the path half absolute, and no embedded newline, carriage return, or tab. The terminal field keeps the atom check, which is correct for it, and no other backend's validation changes. The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca never returns, which is why the suite passed a check the real value fails. They now carry the composite form, so the tests exercise the real value. * no-mistakes(document): name Orca's repo id in the composite worktree id * no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points * feat(bin): add opt-in typed dispatch resolution (#4692) * feat(bin): add opt-in typed dispatch resolution through typesafe.ai Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or scout profile from a written brief with typesafe.ai's System One model: one Choice question over the rules' `when` texts, then the confidence floor, the rule's `approval` and `floor`, each profile's `provider` and `floor`, one quota-axi snapshot, and the spendPriority argmax all in code. It is off unless TYPESAFE_API_KEY is in the environment or the home's gitignored .env; off means one stderr line, exit 0, and no network call, so firstmate dispatches exactly as before. The key reaches curl on a file descriptor, never argv. Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates the four new optional dispatch fields. Document the schema, the operator contract, the AGENTS.md intake step, and the live and benchmark evidence. * no-mistakes(review): Harden typed dispatch resolution and quota bounds * no-mistakes(review): Validate dispatch floors and ranking evidence * no-mistakes(review): Tighten dispatch response and floor evidence * no-mistakes(review): Neutralize none matching and resolve defaults locally * no-mistakes(review): Preserve providerless profiles outside typed resolution * no-mistakes(review): Validate response usage and reject duplicate profiles * no-mistakes(review): Escalate unverifiable floors and validate probabilities * no-mistakes(review): Validate probability mass and unknown profile floors * no-mistakes(review): Simplify resolver interface and preserve fallback routing * no-mistakes(review): Fix constants and rank partial quota evidence * no-mistakes(review): Add authoritative provider mapping and enforce explicit providers * no-mistakes(review): Declare provider for documented Pi profile * no-mistakes(review): Validate provider identifiers and support Gemini dispatch * no-mistakes(review): Strictly anchor provider identifiers * no-mistakes(review): Validate selectors and preserve fallback candidate evidence * no-mistakes(review): Gate typed validation and harden resolver evidence * no-mistakes(review): Preserve opt-in routing and harden candidate evidence * no-mistakes(review): Prioritize known exhaustion over quota uncertainty * no-mistakes(review): Isolate API secrets and preserve no-key diagnostics * no-mistakes(review): Fallback safely when dispatch rules are absent * no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets * no-mistakes(document): Document typed dispatch safety and fallback behavior * fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753) * test: reproduce buried status declarations in shared readers * fix: share status event reads and preserve open blockers * fix: retain terminal scout and ship status declarations * no-mistakes(review): Fix status chronology, legacy completions, and reader performance * no-mistakes(review): Share terminal decision reconciliation across fleet snapshots * no-mistakes(review): Unify terminal supersession across cached folds and consumers * no-mistakes(review): Filter per-key status history while preserving terminal chronology * no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells * no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses * no-mistakes(document): Document latest-event status read and kind-scoped fold cursor * no-mistakes(lint): Quote literal done in test for-lists for SC1010 * ci: expect 19 snapshot/fleet-view tests This branch adds a fleet-snapshot regression, so the stock macOS Bash lane's hardcoded guard of 18 'ok - ' lines fails on the new count. Bump the guard and its message to 19. * no-mistakes(review): Restore multiline child outcome reporting * no-mistakes(review): Select ledger terminal events through bounded shared reader * no-mistakes(review): Report newest open decision instead of preferring blocked * no-mistakes(review): Require colon before ship/scout terminal supersession in fold * no-mistakes(review): Gate socket-down override on latest event; drop lock matrix * no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions * no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold * no-mistakes(test): Update fleet-view expectations to newest-open-decision rule * no-mistakes(document): Align status-read docs with fold-resolved crew state * no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers * no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree * test: fold terminal-cleanup snapshot coverage into the completed-scout case Keep the ship/scout/secondmate supersession assertions without adding a nineteenth top-level fleet-view test, so CI can stay at the upstream suite count. * no-mistakes(document): Clarify socket-down override expiry in architecture doc * ci: retrigger flaky contribution check * fix(bin): launch codex crewmates with codex's hook layer disabled (#4689) * fix(spawn): launch codex crewmates with codex's hook layer disabled A freshly launched Codex worker never reached its instructions. Codex stopped it on an interactive "Hooks need review" modal whose selection sits on "Review hooks", which is neither trusting nor declining. Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow navigation, so the selection cannot be moved, and pre-accepting the prompt by writing Codex's own trust store would record an operator consent that was never given. The hooks are the machine's own ~/.codex/hooks.json plus any project's .codex/hooks.json. A crewmate needs neither: its turn-end signal is the -c notify= program on the same launch, and Firstmate's project hooks are primary-session infrastructure that stands down in a child worktree. Crewmate and scout launches now pass --disable hooks. That is the opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted hooks; disabling the feature runs none of them and leaves the operator's ~/.codex untouched. An unknown feature name is a hard Codex error, so a release that drops the flag fails the launch loudly instead of silently restoring the modal. A secondmate is a primary in its own home and keeps the project hooks its turn-end guard and session-start digest ride on. Verified on codex-cli 0.151.0: the modal is gone and the turn-end notification still lands. This unblocks the second review that every finished pull request is supposed to get. Fixes kunchenguid/firstmate#4673 * no-mistakes(review): Fix contradictory hook count in Codex verification record * fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710) * fix(bin): settle terminal contributions and wake once per read-failure episode A contribution whose last good observation is merged or closed is final: poll no longer re-reads it, projection keeps it fresh, and a stale error recorded beside it is cleared once. A genuine forge-read failure on an open contribution still records its error on every cycle but prints the unavailable wake only when it starts a failure episode; a successful read ends the episode. Open PRs linked from done tasks keep being observed. The false unavailable beside a complete observation was budget exhaustion mid-observation, already fixed by #4661. * fix(review): Settle terminal contribution owners * fix(review): Deduplicate shared contribution failure episodes * fix(test): Preserve settled terminal contribution records --------- Co-authored-by: Pablo Ontiveros <pablo.ontiveros@gmail.com> Co-authored-by: Umer <umeranjum17@gmail.com> Co-authored-by: Yasuhito Takamiya <yasuhito@hey.com> Co-authored-by: Marsjohn-11 <74795701+Marsjohn-11@users.noreply.github.com> Co-authored-by: tbillings28 <todd@toddbillings.com> Co-authored-by: Todd Billings <todd@usdvcapital.com> Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Co-authored-by: Tiago <tiagop@hey.com> Co-authored-by: Amin Roudaki <roudaky@gmail.com> Co-authored-by: Joseph Kim <jokim1@gmail.com> Co-authored-by: Mickaël Rémond <mremond@process-one.net> Co-authored-by: Sebastian <80847374+thelad-dev@users.noreply.github.com> Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc> Co-authored-by: Juan José González Giraldo <juanjose.eng@gmail.com> Co-authored-by: Cody <72239807+codyjohnsontx@users.noreply.github.com>
vifar
added a commit
to vifar/firstmate
that referenced
this pull request
Sep 18, 2026
* fix(bin): translate Stop hook timeout signals into durable auto-arm failure (#4474)
* fix(bin): recover Claude auto-arm after timeout
* no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list
* fix(spawn): establish Claude task channel authority (#4464)
* fix(spawn): establish Claude task channel authority
* no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference
* fix(bin): refuse fm-control.sh exit when the composer holds unproven or pending text (#4458)
* fix: guard relaunch exit against pending input
* no-mistakes(review): Verifying test run in progress
* no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard
* no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now
* fix(spawn): establish crewmate identity first (#4481)
* fix(bin): reconcile redundant secondmate divergence during updates (#4460)
* fix: reconcile diverged secondmate updates
* no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md
* no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md
* feat: enable gpt-5.6-luna max reasoning for crew dispatch (#4497)
* fix(dispatch): support Codex Luna max effort
* no-mistakes(review): use portable CODEX_HOME path in codex effort reference
* feat(calm): render smooth Unicode swell with asymmetric two-color sail (#4498)
* feat(calm): render smooth Unicode swell
* feat(calm): make sails asymmetric
* feat(calm): use quarter sail glyph
* no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer
* no-mistakes(document): docs: sync calm wave phase doc comment
* no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。
* fix(bin): supersede stale scout delivery text in brief.md on promotion (#4491)
* fix: supersede scout delivery brief on promotion
* fix: preserve ship safety contract after promotion
* no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch
* fix(bin): make captain holds work on hosts with an older JSON::PP, and stop cleanup dropping accents from a held body (#4471)
* fix(bin): let captain holds work on hosts with an older JSON::PP
Holding a task for the captain, and the cleanup that keeps a captain-held row
open, both fail outright on any host whose JSON::PP defaults allow_nonref off -
2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`,
but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older
library rejects that whole value with "must be object or array".
The consequence is fleet-wide on such a host, not one broken command: a worker
there cannot formally record a decision for the captain at all. It can only
mention the decision in passing in a status line, where it can be missed - which
is how a real decision goes unrecorded. The hold reports that the task lost its
hold-set stamp; the cleanup cannot return the row to Queued.
Both call sites now ask for allow_nonref explicitly rather than inheriting
whatever the installed library defaults to. The second one is worth naming: its
`/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string
case that fails, so the guard selects for the failing input rather than
protecting against it.
The regression case forces the older default back off for every perl the commands
spawn, then drives both paths - holding a task that carries a body, and tearing
down a captain-held row whose deliverable must still be appended. It also probes
that the simulation genuinely rejects a bare scalar, so the case cannot pass
vacuously on a lenient host. Each half was verified failing on its own unfixed
call site with that site's real error message. Suites: fm-captain-hold-lifecycle
51 cases, fm-backlog-atomicity 99 cases, 0 failures.
Verification limit: the mechanism is reproduced and tested, but neither fix is
verified against a real JSON::PP 2.27202 host, because none is in the loop. This
laptop runs 4.06, where the bug does not manifest.
`bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a
brace-delimited object before decoding, so allow_nonref never applies.
* fix(bin): stop cleanup silently dropping accented characters from a held body
Cleanup rewrites a captain-held row's body to append the finished work's
deliverable, and the decoder it reads that body with printed decoded characters
to a stream with no `:raw` layer. A character at or below U+00FF then came out
as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the
accent. `fm_backlog_retain` writes that body straight back through
`--body-file`, and nothing reported an error - the character was simply gone
from a row still waiting on the captain.
The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus
`utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already
used.
Review of the parent commit found this on one of the lines that commit already
changed. It predates that change.
The test asserts bytes rather than decoded strings, because comparing strings
cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any
character above U+00FF makes perl print the whole string as UTF-8, so one body
carrying both an accent and an em dash passes even unfixed and proves nothing.
Verified failing before the fix on the accented row, passing after. Suites:
fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures.
* no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc
* no-mistakes(review): drop whole-file UTF-8 check from retained-body test
* no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc
* no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc
* fix(bin): read codex 0.154's idle braille starfield rows as composer furniture (#4532)
* fix(composer): read codex 0.154's idle starfield and status footer as furniture
codex-cli 0.154.0 animates a braille "starfield" around its idle composer:
on the row above the bold `›` prompt row, on the `›` row behind the SGR-2
dim `Ask Codex to do anything` placeholder, and on the row below it, then
draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`).
The cells are truecolor greys on both sides of the ghost luminance ceiling,
so the brighter ones survive ghost stripping, and the rows below the glyph
carry no structural edge. The shared classifier selected the bare `›` shape,
extended its wrap region over the two rows beneath the glyph, read the
survivors and the footer as wrapped typed input, and answered `pending`;
the steering doorbell defers on exactly that verdict, so no doorbell ever
reached an idle codex 0.154 pane.
bin/fm-composer-lib.sh now recognises that furniture by shape, declared
once next to the idle placeholders and reached from the two wrap-region
boundary points:
- a row whose non-whitespace content is entirely braille cells
(U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it
never counts as wrapped typed content and bounds a bare composer's wrap
region; braille behind the glyph row's content is stripped before the
emptiness decision when nothing else follows the glyph; a row mixing
braille with other text stays typed content;
- the codex status footer bounds the wrap region exactly as omp's status
row does, anchored on the effort token, a spaced middle dot, and a `~` or
`/` path cell, so a typed `fix · tests` stays composer input;
- `^Ask Codex to do anything$` joins the verified idle-placeholder set; the
ghost strip remains what proves that row empty, and the bare-row rule that
bright placeholder text is real input is unchanged.
Unchanged: the strict blank-row rule, the styled=0 degradation (a plain
cmux/orca capture of this screen still reads `unknown`, never `pending`),
FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape.
tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte
with the divergence (letters in place of the starfield read `pending`) and
the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh
is the default-on live guard (token-free, skips explicitly without codex or
tmux) that launches the installed codex idle and asserts `empty` through
both the tmux and the cursorless styled reads, naming codex --version on
failure. docs/verification/runtime-backends.md records the dated Herdr
evidence: `pending` before, `empty` after, on the captured screen.
* no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry
---------
Co-authored-by: Todd Billings <todd@usdvcapital.com>
* fix(bin): refuse empty text steers in fm-send (#4259)
* fix(bin): refuse empty text steers in fm-send
A marked secondmate request sent with an empty message delivered only
marker and correlation bytes and minted a pending-reply expectation the
parent could never see resolved, stalling the fleet with no loud error
(#4255). Fail closed on an empty or whitespace-only message on the text
path, mirroring the existing --resolve-key refusal.
* chore: retain ambient Pi-lens autoformat as its own commit
Formatting-only edits produced by ambient Pi-lens autoformat during the
msg-loss investigation, kept separate from the behavioural change in
c23acba6 so the fix stays reviewable on its own.
AGENTS.md is deliberately excluded: its only autoformat edit stripped the
trailing space from the documented FM_OPERATIONAL_PREFIX value, which
bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11
records as permanent compatibility. Documenting that constant without its
trailing space makes the doc wrong about the contract, so that one line was
restored rather than retained.
* fix(calm): paint the working ship one yellow over all-blue water (#4554)
On rose-pine-moon the two-color water (cyan crests over blue troughs) read as
a pink stripe over aqua, the yellow left sail and mast clashed with the red
right sail, and the hull carried a blue interior run. Every water cell is now
blue so the swell reads through glyph height alone, and both sail halves, the
mast, and the whole hull are one yellow run. Geometry, cadence, animation,
direction flip, resize clamping, and the narrow fallback are unchanged.
Update the unit and real-TUI color assertions to the new palette and the Calm
docs that described the old one.
* fix(bin): stop aging a second mate's active turn from its launch (#4270)
* fix(watch): stop aging a second mate's active turn from its launch
The parent watcher's second-mate wake-loop stall check exempts a mate that
is demonstrably inside an active turn, but secondmate_in_active_turn asked
busy_turn_over_age first and returned "not in a turn" whenever that said
the bound was crossed.
busy_turn_over_age ages from state/<task>.turn-ended, falling back to
state/<task>.meta. A second mate's turns end in its own home, so the
parent never gets a turn-ended mark for it and the fallback ages the
mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago
was therefore permanently "over age", the busy pane was never consulted,
and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false
wake-loop stall.
The gate now bounds the busy exemption by <idle> - how long the queue's
drain position has not moved - which is evidence this home actually
holds. A busy mate stays exempt while the queue has been frozen for less
than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so
the bound that stops a busy pane from proving liveness forever is kept
rather than removed. busy_turn_over_age is untouched; its remaining
callers are the ordinary crew busy-pane bound.
The regression pins the case that actually broke: a mate whose launch
record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn
must not escalate, while the same mate with its queue frozen past the
bound still publishes exactly one notification. The existing coverage
only exercised a freshly launched mate, which passes either way.
Reaching that alert now costs a pane capture inside the gate, so the
three checkpoints in this suite that assert an alert move from a 1s to a
4s bound - the value the neighbouring active-turn cases already use. The
bound is a ceiling, not a wait: the checkpoint returns on the first
actionable wake. On a loaded machine a 1s bound missed the alert
repeatedly; at 4s it did not miss in 20 runs under the same load.
* no-mistakes(review): scope the second-mate active-turn regression test's coverage claim
* no-mistakes(document): fix stale second-mate active-turn comments in fm-watch
* feat(bin): add read-only PR blocker and reviewer discovery commands (#4278)
* feat(bin): add read-only PR blocker and reviewer-discovery commands
Two focused, opt-in commands that read GitHub and never write to it.
fm-pr-state.sh reports what still blocks one pull request from the
author's side: a closed or merged state, draft state, unknown or
conflicting mergeability, absent or failing required checks, and a
blocking CHANGES_REQUESTED decision explained by each reviewer's latest
verdict, marked STALE when it was left at a superseded head. A pull
request that only awaits an approval is not reported as blocked, and
advisory checks are omitted. Every reading is taken against one exact
head; a push that lands mid-read invalidates the whole result rather
than mixing two snapshots.
fm-pr-reviewers.sh suggests reviewers from the most recent commits to
the pull request's exact changed paths, counting each commit once,
resolving handles through GitHub's own commit author.login mapping, and
excluding the author and Bot accounts.
Both stay read-only: no review request, no approval, no merge.
Unresolved review-thread state is left unreported because the REST API
does not expose it and unattended commands may not use GraphQL.
Closes #3731
* no-mistakes(review): accept only PR URLs and stop at terminal state
* no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate
* no-mistakes(review): stop attributing readings to unverified heads
* no-mistakes(review): narrow readiness contract to checks that have reported
* no-mistakes(review): read the pull request once, drop the head guard
* no-mistakes(document): scope pr-forge isolation proof to its measured members
* no-mistakes(document): record uncovered pr-forge members and their pending proof
* docs(isolation-proof): re-prove pr-forge at its full membership
tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the
pr-forge family in this branch, and script_allows_concurrency grants
four workers by family membership alone, so both ran concurrently on a
proof measured before they existed.
Re-proved the family at all eight members: two consecutive runs, 0
failures, each begun with the one-minute load average below 6.0 so the
result measures isolation rather than contention. A third run taken
between them is disclosed rather than recorded, because it started
while the previous run's workers were still decaying.
The new durations are not comparable with the six-member measurement
above them, so they are not presented as evidence about the two new
members, and that record's 1.72x four-worker figure is left as a
statement about its own run rather than restated as current.
* no-mistakes(review): disclose gh error-text coupling at its matching site and tests
* fix(bin): teach validation-round pauses in generated briefs (#2752)
* fix(bin): teach validation-round pauses in briefs
* no-mistakes(document): Point classifier comments to authoritative pause examples
* docs(readme): add star history chart (#4558)
* fix(bin): refuse teardown when a task's endpoint close fails (#4510)
* fix(teardown): refuse a cleanup whose endpoint close failed
bin/fm-teardown.sh discarded both the exit status and the stderr of every
fm_backend_kill call, so a close that genuinely failed was indistinguishable
from one that succeeded. Teardown continued past it, deleted the task's durable
records, returned its worktree, and reported the cleanup as completed. The
deleted metadata is the only record of which endpoint belongs to the task, so
such a close did not merely leave a stray session behind, it stranded one:
nothing was left on disk naming it.
The adapters could not carry that signal either. Driven against the real code,
every backend arm returned 0 for a genuine failure exactly as it did for an
already-exited endpoint, so there was nothing for the four call sites to
propagate even once they stopped swallowing it.
The tmux arm now resolves a close that did not succeed against the window's
exact recorded identity, since kill-window fails the same way for a window that
is gone and one that is still there. The Orca arm reports a close its missing
CLI never attempted. Both stay silent for an endpoint that is already
legitimately gone, and the remaining arms are unchanged: their close-command
timing cannot be established without the real Zellij, Orca, and cmux binaries,
and a gate that refused ordinary cleanup of an already-exited session would be
worse than the defect. docs/verification/runtime-backends.md records what each
backend can prove.
A reported close failure now reaches teardown's existing retain-and-stop
refusal before the records naming the endpoint are removed, matching where the
Herdr confirmed-gone gates already sit for the same hazard, and the retained
records let a rerun finish once the close works.
* no-mistakes(review): refuse unreadable tmux close re-read; honor --force override
* no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close
* no-mistakes(document): document endpoint-close refusal in its backend and retirement owners
* no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree
* feat(calm): add flag-gated Claude Code Calm mode (#4565)
* feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag
Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that
brings Calm to Claude Code: the sailboat replaces the stock working row through a
Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration,
and canonically classified operational user rows draw at zero height. /calm is
registered by the hooks module itself and toggles the same per-home config/calm
preference the Pi extension uses, so one choice applies on either harness; rows
redraw retroactively on toggle and stay hidden across claude --continue.
The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS
flag is on. Nothing sets that flag in any settings file, and the plugin carries no
command file, skill, agent, or classic hook, so it is a complete no-op while the
flag is off. The trusted project auto-loads it through an .agents/skills symlink,
the only path Claude Code scans for project plugins.
Extract the working-ship geometry, bounce track, cadences, and freeze/resume state
into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module
imports from outside the plugin folder) and have the Pi widget paint that core's
frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify
operational rows through a port of bin/fm-operational-input.sh's classify command
guarded by a corpus parity test against the shell owner.
Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering,
Raster packing, policy, classifier parity), the mod's own claude plugin test suites
behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op,
the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272.
Docs: record the version-scoped Claude Code evidence and the three bounded gaps in
docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md,
and make the shared preference, layout, and contributor notes harness-neutral.
* no-mistakes(review): Preserve colliding final replies and strengthen parser parity
* no-mistakes(review): Preserve final replies and strengthen canonical parity checks
* no-mistakes(review): Require exact function-hooks opt-in before Calm activation
* no-mistakes(review): Clarify Calm module loading and activation boundaries
* no-mistakes(review): Reset Calm presentation state across session starts
* no-mistakes(document): Refresh Calm session lifecycle documentation
* feat(calm): paint the Claude Code working ship in Claude's own theme colors
The captain picked the "Claude native" palette for the Claude Code mod's Raster:
every water cell takes the spinner blue of the active theme family (#93a5ff dark,
#5769f7 light) and the whole boat takes the Claude orange of the stock spinner
(#d77757), one water color and one boat color. The family follows the `theme`
setting's prefix, read at load through $.config.list and re-read on a
config.set of that row, with `auto` and custom themes falling back to the dark
set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte.
Rename the shared sprite's color classes from hue names to `water` and `boat`,
since each harness now maps them to its own colors; geometry, motion, cadence,
and the activation gate are untouched.
Tests cover both palettes' packing and the family rule under Node, and the
plugin kit drives every theme value, a theme change mid-session, the Calm-off
pass-through, and inertness of the menu read while the flag is off. The docs
describe the Claude Code colors and record the guard passing on 2.1.273.
* no-mistakes(review): Use light palette for unresolved Claude themes
* no-mistakes(document): Refresh Claude Calm verification evidence
* fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586)
* fix(watch): honour a declared wait before wedge-escalating a quiet pane
wedge_timer_check escalated on elapsed idle time alone. Nothing asked
whether the worker had already said why its pane was quiet, so a lane
that declared a bounded external wait climbed the escalation ladder for
as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every
repeat carried demand-deep-inspection - which by its own wording forbids
re-absorbing on the run-step or pane state, so the supervisor could not
use the evidence that was there either.
The generated brief promises that declaring `paused:` buys the long
recheck cadence instead of a wedge, but the timer was still reachable
while that declaration stood: a crew that declares a wait and then has an
active run or busy pane attributed to it is handed to the timer as
provably-working. The declaration is what the worker said about its own
silence, so it now outranks a liveness verdict that only says something
is running.
The consult runs in the at-threshold branch that was about to escalate,
beside the worktree walk already there, and costs one status-line read.
Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS
recheck the declared-wait absorber already uses, so the wait is still
rechecked and cannot rot invisibly. Which verb declared it decides the
wording, because the two block on different people: a `paused:` wait is
owed by an external dependency and asks the reader to confirm it still
holds, while a `captain-held:` transfer is owed by the captain reading
the recheck and asks them to answer or release the hold. A hold is not
rechecked at all while the away-posture record exists, as on every other
captain-held path, and that absorb arms no throttle so the recheck is
owed in full on return.
A declared clearing time that has already passed stops counting, and a
lane that never declared one keeps the identical escalation schedule,
reason, count and demand-deep-inspection wording, so detection and its
worst-case time are unchanged. The deferral restarts the idle timer
rather than cancelling it, so a lane that stops waiting escalates again
within one threshold.
A lane quiet because its own validation run is parked at a gate awaiting
a human decision is deliberately out of scope: reading that state needs a
signal carrying who the wait is on and what clears it, rather than one
inferred from a parked verdict that also covers gates awaiting the
crewmate itself.
Tests pin both directions for each case and were each confirmed to fail
with the consult removed.
* no-mistakes(document): docs: honour declared waits in stale-escalation docs
* fix(bin): report verified PR state for passed runs (#4624)
* fix(bin): derive passed PR state from PR record
A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists.
For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim.
Fixes #4607
* no-mistakes(review): Add bounded GitLab merge-request state reads
* no-mistakes(review): Preserve network-free inactive crew-state scans
* no-mistakes(document): Document PR record readers in shared library
* fix: restore published contribution follow-up (Fixes #4469) (#4627)
* fix: restore published contribution follow-up (Fixes #4469)
* fix(review): Fix contribution freshness and merge actor routing
* fix(review): Restore issue triage and scope contribution follow-up
* fix(test): test: assert one wake per contribution signal
* fix(document): Document contribution follow-up
* fix: restore truthful terminal delivery evidence
* fix(review): Disclose unsupported contributions and deduplicate watcher wakes
* fix(review): Preserve unmeasured unsupported contributions across Bearings
* fix(review): Deduplicate shared contribution wakes and isolate diagnostics
* fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower
* fix(bin): make remote report transfers explicit and fail-open (#4658)
* fix(bin): make a remote-reply document gap self-clearing and re-attemptable
A remote mate's undelivered document raised a keyed `blocked` decision that
nothing could ever resolve, and any `data/*.md` substring in any mirrored line
was an unconditional fetch instruction. A mate announcing a report it had not
written yet therefore manufactured a permanent, factually false blocker, and
its own explanation of the false alarm manufactured more.
The reader has no permanence vocabulary: a report still being written refuses
exactly like a path that will never exist. So an undelivered document is now a
durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`,
re-attempted on the next delta and on the channel's own quiet poll, and retired
with a matching `resolved` line naming the local copy once it arrives. The
cursor still advances and no delta stalls on one bad pointer.
Only a structured `report=data/....md` pointer now offers a document, so a path
merely mentioned in prose - including one under another home's mirror tree,
which is provably not that mate's to serve - is never fetched. Offers are
deduplicated across the whole delta, the escalation names each missing document
once and carries the reader's own reason instead of discarding it, and a
strictly increasing notice ordinal keeps a later escalation from being
swallowed as duplicate bytes. A mirrored line still lands once whichever
pointer form it was first written under.
* no-mistakes(review): Require structured pointer token boundaries
* no-mistakes(review): Unify boundary-safe pointer extraction and rewriting
* fix(bin): identify a mirrored line independently of its delivery state
Two defects in the boundary-safe pointer work.
The at-most-once check compared only the all-remote and all-local renderings
of a line, so it could not recognize a mixed one. A line offering two documents
where only the first was deliverable mirrored as local-plus-remote; once the
second arrived, a cursor-loss whole-log recapture rendered the same line
all-local, matched neither alternate, and mirrored a second time. A line's
identity is now the canonical form every boundary-valid pointer would take once
delivered, derived by the same parser that does extraction and rewriting, so it
no longer depends on which documents happened to be deliverable at the time.
The pointer map was passed to awk through the process environment. A delta may
carry up to the configured 1 MiB bound, and an expanded map of delivered
pointers can exceed the platform's exec argument limit, so awk would fail to
start; because no caller checked, the empty result would have been appended as
blank lines while the cursor advanced past dropped status content. The map now
travels in a file, and every call site checks the exit status and stops the
ingest rather than committing a delta it could not render.
Both passes now run once per stream instead of twice per line.
* no-mistakes(review): Abort ingest when document pointer extraction fails
* no-mistakes(review): Exclude structured cross-home pointers from document transfer
* fix(bin): fail open on an undeliverable remote document instead of tracking it
Narrow the remote-reply document fix to the scope the diagnosis actually
requires, as decided after measuring a simpler alternative.
A document the reader cannot deliver now fails open. The mate's line is
mirrored with its own pointer, the cursor advances, and one unkeyed note
carries the reader's reason. A note never enters the open-decision fold, so it
cannot stand open the way the original keyed block did - which removes the
never-clearing false blocker by construction rather than by resolving it.
That makes the durable self-clearing obligation unnecessary, so it goes: the
per-mate pending-documents record, its notice ordinal and resolved
announcements, and the poll-side retry. Canonical line identity goes too, and
with it a way to silently drop a genuine status line; mirroring is back to
at-most-once on exact bytes. The cross-home exclusion goes as well: under
fail-open a cross-home report= either fails harmlessly or is a nested remote
report this mate genuinely holds, which is now relayed again.
Kept: fetching only on a structured report= pointer, the boundary-correct
parser, the file-based rewrite map, and checked extraction and rewrite exit
status. The parser now scans behind a sentinel byte so a rejected candidate can
no longer give the text right after it a false leading boundary.
The reported incident is covered end to end: a report path announced in prose
before it exists raises no decision, and the report still arrives through the
ledger publisher's structured offer once written.
* no-mistakes(review): Preserve source-line identity across remote reply replays
* no-mistakes(document): Document remote reply transfer and replay semantics
* no-mistakes(lint): Fix staging truncation lint checks
* fix(calm): preserve substantive mid-turn responses (#4655)
* Preserve substantive Calm mid-turn text
* no-mistakes(review): Distinguish newline-preserved replies from short narration
* no-mistakes(document): Document Calm mid-turn preservation boundaries
* no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`
* fix(bin): preserve PR merge polls across volume remounts (#4656)
* fix(bin): re-record PR poll identity after a volume device renumber (Fixes #4260)
A volume remount can renumber the state filesystem's st_dev while every
inode and byte stays the same; APFS does this across a reboot. A poll
registration records its sidecar and check as device:inode, so every poll
armed before the remount failed strict validation and the watcher refused
all of them as unauthenticated state checks until each was re-armed by hand.
There are two device comparisons. fm_pr_private_file_valid compares a live
file's device with the state directory's device read in the same invocation:
it refuses a file that is not on the state directory's own filesystem and
already survives a renumber, so it is unchanged. The registration's recorded
identity versus the live identity (from #556, reused by the #932 retirement
receipt) binds the registration to the exact files published in its own
transaction; its device part is what breaks.
When strict capture fails, the watcher now proves the device is the only
difference: every other artifact check passes (template bytes, both hashes,
private mode, single link, live device, metadata), both recorded identities
name one device, and each recorded inode equals its live inode. Only then,
under the task's control lock, does it rewrite the two identity lines,
repeating the whole proof and comparing the registration's file identity and
bytes just before the rename, and then capture strictly again. A swapped,
altered, re-moded, relinked, split-device, or foreign-device artifact still
fails a proof and is still refused, and a pending retirement receipt blocks
the rewrite.
Reproduction: on macOS a poll armed on an APFS disk image that was detached
and re-attached behind another image moved st_dev 16777239 -> 16777243 with
inodes, bytes, mode, and link count unchanged; the real watcher refused it on
main and reports its merge with this change. The portable regression test
rewrites a real registration's recorded device and drives the watcher.
Not changed here: the status presentation cursor keys rows by its own
device:inode identity in bin/fm-classify-lib.sh, a different helper that
needs its own fix; a retirement receipt left by a reboot between its
publication and removal still names the old device and stays refused; custom
check trust binds only a content hash and is unaffected.
* fix(review): Serialize PR poll publication writers
* fix(review): Bound PR poll publication lock scope
* fix(bin): keep contribution records when the poll budget runs out (follow-up to #4627) (#4661)
A budget that expires partway through an observation no longer records an
error or prints the unavailable wake; the URL keeps its prior record and is
observed first next poll. forge() flags budget exhaustion at the point it
refuses, or when a read is killed at the budget's own deadline, so a genuine
forge failure still records the error and wakes. Each distinct URL is now
observed once per poll and applied to every owning task.
* fix(bin): clear parent pending-replies on local secondmate retirement (#4680)
* fix(bin): clear parent pending-replies on local secondmate retirement
Local secondmate teardown left resolved parent pending-reply records behind
after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced
retirement while any reply for that id is still unresolved, and delete every
matching record plus its delivery confirmation after a successful local or
remote retirement, matching the remote cleanup path.
* no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup
* no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung
* no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern
* no-mistakes(review): Pending-replies Basename und corr_id abgleichen
* no-mistakes(document): Clarify forced retirement pending-reply cleanup
---------
Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>
* fix(bin): accept Orca's composite worktree id when tearing down a task (#4677)
* fix(bin): accept Orca's composite worktree id at teardown
Teardown refused every Orca-backed task because the endpoint validator
checked orca_worktree_id with the simple-atom rule meant for tmux-style
window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca
returns that id as `<orca id>::<absolute worktree path>`, so the colon and
slashes in every real value made validation fail and finished Orca tasks
could never be cleaned up.
Validate the field as the composite it is: both halves of the first `::`
split present, the path half absolute, and no embedded newline, carriage
return, or tab. The terminal field keeps the atom check, which is correct
for it, and no other backend's validation changes.
The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca
never returns, which is why the suite passed a check the real value fails.
They now carry the composite form, so the tests exercise the real value.
* no-mistakes(document): name Orca's repo id in the composite worktree id
* no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points
* feat(bin): add opt-in typed dispatch resolution (#4692)
* feat(bin): add opt-in typed dispatch resolution through typesafe.ai
Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or
scout profile from a written brief with typesafe.ai's System One model:
one Choice question over the rules' `when` texts, then the confidence
floor, the rule's `approval` and `floor`, each profile's `provider` and
`floor`, one quota-axi snapshot, and the spendPriority argmax all in code.
It is off unless TYPESAFE_API_KEY is in the environment or the home's
gitignored .env; off means one stderr line, exit 0, and no network call,
so firstmate dispatches exactly as before. The key reaches curl on a file
descriptor, never argv.
Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and
the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new
tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates
the four new optional dispatch fields. Document the schema, the operator
contract, the AGENTS.md intake step, and the live and benchmark evidence.
* no-mistakes(review): Harden typed dispatch resolution and quota bounds
* no-mistakes(review): Validate dispatch floors and ranking evidence
* no-mistakes(review): Tighten dispatch response and floor evidence
* no-mistakes(review): Neutralize none matching and resolve defaults locally
* no-mistakes(review): Preserve providerless profiles outside typed resolution
* no-mistakes(review): Validate response usage and reject duplicate profiles
* no-mistakes(review): Escalate unverifiable floors and validate probabilities
* no-mistakes(review): Validate probability mass and unknown profile floors
* no-mistakes(review): Simplify resolver interface and preserve fallback routing
* no-mistakes(review): Fix constants and rank partial quota evidence
* no-mistakes(review): Add authoritative provider mapping and enforce explicit providers
* no-mistakes(review): Declare provider for documented Pi profile
* no-mistakes(review): Validate provider identifiers and support Gemini dispatch
* no-mistakes(review): Strictly anchor provider identifiers
* no-mistakes(review): Validate selectors and preserve fallback candidate evidence
* no-mistakes(review): Gate typed validation and harden resolver evidence
* no-mistakes(review): Preserve opt-in routing and harden candidate evidence
* no-mistakes(review): Prioritize known exhaustion over quota uncertainty
* no-mistakes(review): Isolate API secrets and preserve no-key diagnostics
* no-mistakes(review): Fallback safely when dispatch rules are absent
* no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets
* no-mistakes(document): Document typed dispatch safety and fallback behavior
* fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753)
* test: reproduce buried status declarations in shared readers
* fix: share status event reads and preserve open blockers
* fix: retain terminal scout and ship status declarations
* no-mistakes(review): Fix status chronology, legacy completions, and reader performance
* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots
* no-mistakes(review): Unify terminal supersession across cached folds and consumers
* no-mistakes(review): Filter per-key status history while preserving terminal chronology
* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells
* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses
* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor
* no-mistakes(lint): Quote literal done in test for-lists for SC1010
* ci: expect 19 snapshot/fleet-view tests
This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.
* no-mistakes(review): Restore multiline child outcome reporting
* no-mistakes(review): Select ledger terminal events through bounded shared reader
* no-mistakes(review): Report newest open decision instead of preferring blocked
* no-mistakes(review): Require colon before ship/scout terminal supersession in fold
* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix
* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions
* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold
* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule
* no-mistakes(document): Align status-read docs with fold-resolved crew state
* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers
* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree
* test: fold terminal-cleanup snapshot coverage into the completed-scout case
Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.
* no-mistakes(document): Clarify socket-down override expiry in architecture doc
* ci: retrigger flaky contribution check
* fix(bin): launch codex crewmates with codex's hook layer disabled (#4689)
* fix(spawn): launch codex crewmates with codex's hook layer disabled
A freshly launched Codex worker never reached its instructions. Codex
stopped it on an interactive "Hooks need review" modal whose selection
sits on "Review hooks", which is neither trusting nor declining.
Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow
navigation, so the selection cannot be moved, and pre-accepting the
prompt by writing Codex's own trust store would record an operator
consent that was never given.
The hooks are the machine's own ~/.codex/hooks.json plus any project's
.codex/hooks.json. A crewmate needs neither: its turn-end signal is the
-c notify= program on the same launch, and Firstmate's project hooks are
primary-session infrastructure that stands down in a child worktree.
Crewmate and scout launches now pass --disable hooks. That is the
opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted
hooks; disabling the feature runs none of them and leaves the operator's
~/.codex untouched. An unknown feature name is a hard Codex error, so a
release that drops the flag fails the launch loudly instead of silently
restoring the modal. A secondmate is a primary in its own home and keeps
the project hooks its turn-end guard and session-start digest ride on.
Verified on codex-cli 0.151.0: the modal is gone and the turn-end
notification still lands.
This unblocks the second review that every finished pull request is supposed to get.
Fixes kunchenguid/firstmate#4673
* no-mistakes(review): Fix contradictory hook count in Codex verification record
* fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710)
* fix(bin): settle terminal contributions and wake once per read-failure episode
A contribution whose last good observation is merged or closed is final:
poll no longer re-reads it, projection keeps it fresh, and a stale error
recorded beside it is cleared once. A genuine forge-read failure on an open
contribution still records its error on every cycle but prints the
unavailable wake only when it starts a failure episode; a successful read
ends the episode. Open PRs linked from done tasks keep being observed.
The false unavailable beside a complete observation was budget exhaustion
mid-observation, already fixed by #4661.
* fix(review): Settle terminal contribution owners
* fix(review): Deduplicate shared contribution failure episodes
* fix(test): Preserve settled terminal contribution records
* fix: select authoritative no-mistakes runs (#4476)
* fix(crew-state): select authoritative validation runs by identity
Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row.
Refs: https://github.com/kunchenguid/firstmate/issues/3215
* fix(review): Resolve same-branch run identities beyond capped history
* fix(review): Fix run-selection compatibility, races, and worker-state fallbacks
* fix(review): Limit run validation to the requested branch
* fix(test): Anchor AXI fixtures and document remaining live evidence gaps
* fix(document): Clarify run selection documentation and capture ownership
* fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation
* fix: distinguish captain outcomes from no-op updates (#4738)
* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work
MAIN answered a supervision-branch outcome for completed captain-requested
work (implementation done, PR ready for review and merge approval) with
"Captain, shipshape.", reading section 9's no-action reply as covering it
and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no
captain-facing response is owed".
Section 9 now limits the shipshape reply to true no-ops (idle re-read,
empty heartbeat, consequence-free acknowledgement) and requires a short
outcome response naming what finished and what word is needed whenever
requested work finishes or a result needs the captain's word, even when a
transcript entry already shows the substance. The Pi protocol's re-emit
rule now says it bounds repetition only, and carries a worked example of
the ready-for-review outcome whose correct processing turn a shipshape
reply fails.
No executable contract evaluates the content of MAIN's captain-facing
reply, so the regression is the protocol example in the owner doc rather
than a text-match test.
* no-mistakes(document): Clarify captain-facing outcomes versus no-ops
* docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line
The document step condensed the Pi protocol's re-emit rule and dropped the
worked example of a finished, ready-for-review outcome whose correct
processing turn a "Captain, shipshape." reply fails. That example is the
contract's regression: no executable contract evaluates the content of
MAIN's captain-facing reply, so the owner doc's example is the test case.
Restore it directly under the re-emit rule, prefixed as a regression
example that is kept verbatim and never condensed or summarized away.
* no-mistakes(review): Clarify captain outcome and decision-word requirements
* no-mistakes(document): Clarify captain-facing completion outcomes
* docs(pi): require the PR URL in the visible captain-facing outcome reply
Captain review on the regression example: drop the sample reply string
and say only that the ready-for-review outcome requires relaying a
captain-facing outcome response, not just "Captain, shipshape.".
Fold in the visible-PR-handoff failure seen this session: after the
branch outcome reporting this fix green, MAIN's visible reply was only
"Awaiting your merge call." with no PR URL, leaning on the dim anchor.
Section 9's URL rule now also covers a review or merge ask and names the
visible reply as where the URL goes, sourced from the ready status, pr=
metadata, or the supervision branch's summary and never left to a
transcript entry. The Pi protocol adds the same-way failure and places
the captain-facing text in the final visible assistant reply after the
fm_branch_processed call, because Calm hides assistant text emitted in
the same step as a tool call as a working note.
Investigation verdict, evidence in the PR comment: no recent PR caused
the handoff failure; Pi has hidden same-step pre-tool assistant text
since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and
#4658 touched only remote report transfer.
* no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs
* no-mistakes(document): Clarify captain-facing supervision outcomes
* docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule
The consolidated section 9 URL rule narrowed its trigger to a review or
merge ask, dropping the "whenever a PR is mentioned" catch-all from
#3648 that keeps every PR URL copied from a durable record and never
assembled from memory. Restore that trigger as a union with the review
or merge ask so the one consolidated rule covers both.
* fix(bin): let non-owner Claude Stops exit safely (#4777)
* Fix foreign-owner turn-end supervision loop
* no-mistakes(review): Scope foreign-owner safe exit to Claude guard
* no-mistakes(document): Document Claude foreign-owner safe exit
* fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778)
Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty
indexed array as an unbound variable and aborts the shell. In
signal_turnend_panes_churned() the missing_keys loop was reachable with
an empty array whenever every churned key already held a fresh
.churn-since-* marker (a second churning turn-end inside an open
deferral window), so each watcher cycle died about half a minute in and
supervision restarted endlessly. The created_keys rollback loops had the
same latent crash on their error paths.
Audit of bin/ for the same pattern found one more confirmed-reachable
case: remote_handoff's noncanonical-body scan iterates to_move, which is
empty when a retried remote handoff finds every key already staged in
the outbox. All other "${arr[@]}" sites are either count-guarded,
guaranteed non-empty by construction, or unreachable while empty.
Guard the three reachable expansions with the repo's existing
"${arr[@]+...}" idiom. New regression test drives a real watcher
through the all-marked churn path; the macos-stock-bash CI lane runs it
under real /bin/bash 3.2 via FM_TEST_ONLY.
* Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783)
The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: require complete captain-facing final responses (#4779)
* docs: require complete final responses across harnesses
* no-mistakes(document): Document complete final replies for Grok Bot
* docs: point Grok replies to the shared contract owner
* no-mistakes(review): Clarify final recap without batching decision asks
* fix: preserve substantive mid-turn text in Pi Calm (#4788)
* fix(calm): preserve substantive Pi mid-turn text
* no-mistakes(review): Preserve substantive Pi Calm text per block
* no-mistakes(test): Cover shared Calm preservation boundaries behaviorally
* no-mistakes(document): Consolidate Calm preservation documentation
* no-mistakes(ci): Fixed Lint SC2034 in tests/fm-inactive-reconcile.test.sh (unused attempt → _) and tests/fm-spawn-pool-base-freshen.test.sh (shared-record LOCAL_ONLY_TIP disable). Fixed Behavior portable serial 4 by aligning tests/fm-wake-daemon-lifecycle-e2e.test.sh with classify_stale's post-c6e3cc1 contract: recorded meta + FM_FAKE_CREW_STATE working evidence so transient stale self-handles. Verified locally: lifecycle e2e passes; shellcheck -x on the three files exits 0. PR must be raised via no-mistakes is an attestation/pipeline check, not a code defect—no code change for it
* no-mistakes(document): Document Claude foreign-owner session-lock consumers
---------
Co-authored-by: Pablo Ontiveros <pablo.ontiveros@gmail.com>
Co-authored-by: Umer <umeranjum17@gmail.com>
Co-authored-by: Yasuhito Takamiya <yasuhito@hey.com>
Co-authored-by: Marsjohn-11 <74795701+Marsjohn-11@users.noreply.github.com>
Co-authored-by: tbillings28 <todd@toddbillings.com>
Co-authored-by: Todd Billings <todd@usdvcapital.com>
Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: Tiago <tiagop@hey.com>
Co-authored-by: Amin Roudaki <roudaky@gmail.com>
Co-authored-by: Joseph Kim <jokim1@gmail.com>
Co-authored-by: Mickaël Rémond <mremond@process-one.net>
Co-authored-by: Sebastian <80847374+thelad-dev@users.noreply.github.com>
Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>
Co-authored-by: Juan José González Giraldo <juanjose.eng@gmail.com>
Co-authored-by: Cody <72239807+codyjohnsontx@users.noreply.github.com>
Co-authored-by: Pedro Guimarães <21346846+0x7067@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
sanis
added a commit
to sanis/firstmate
that referenced
this pull request
Sep 19, 2026
* fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (#4424)
* fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (#42)
* fix(pr-merge): read a plan-gated 403 on branch rules as no merge queue
github_read_queue_method left status=unreadable for every failed rules
read, including a 403 whose body is GitHub's own "Upgrade to GitHub
Pro or make this repository public" message. A repository whose plan
cannot expose branch rules cannot have a merge_queue rule either, so
that specific 403 now resolves to status=none instead of unreadable -
unblocking the away-merge grant on private repos without GitHub Pro.
Any other failure (auth, rate limit, network, 404, unrelated 403)
still reads as unreadable.
* no-mistakes(document): Update stale away-merge queue-grant comment for plan-gated 403
---------
Co-authored-by: NewAiCoder <claude@theinbtw.com>
* no-mistakes(review): Fix misleading away-queue-grant comment in fm-pr-merge and its test
* no-mistakes(document): Update architecture.md for plan-gated-403 merge queue exception
---------
Co-authored-by: NewAiCoder <claude@theinbtw.com>
* fix(bin): select suites that read a changed top-level test fixture (#4246)
* fix(tests): select readers of a changed top-level test fixture
bin/fm-test-run.sh --changed recognised shared test helpers by an explicit
list, tests/lib.sh|tests/*-helpers.sh|tests/fixtures.sh. A top-level
tests/*-fixture.sh matched none of those, fell through to the tests/*
catch-all, and was marked unmapped, so selection aborted with "no
changed-test mapping for source path" and the run selected nothing at all.
tests/herdr-client-pair-fixture.sh and tests/remote-herdr-fixture.sh are
real shared fixtures with real consumers, so any branch touching one of
them left a validation pipeline driving --changed with a hard abort rather
than a narrowed selection.
Extend the helper arm to tests/*-fixture.sh rather than routing it through
the tests/fixtures/*/* arm. Both arms resolve consumers with the same
reference scan, and that scan is what selects the right suites here: it
finds exactly the tests that read the fixture. The fixtures/ arm adds only
a directory-keying step, which has nothing to key on for a top-level file,
so the helper arm is the same behaviour with no extra machinery. A
tests/ path nothing reads still reaches the catch-all and still refuses
loudly.
Refs https://github.com/kunchenguid/firstmate/issues/4100
* no-mistakes(test): order nested fixtures arm before top-level fixture glob
* no-mistakes(document): document tests/ shared-file mapping contract and arm order
* no-mistakes(review): drop vacuous test phase, correct header claim, restore comment
* fix(bin): treat Claude Code's default external-imports flags as never asked, not declined (#4387)
* fix(bin): read Claude Code's default external-imports flags as never asked, not declined (#4378)
fm-claude-trust.sh refused the whole trust registration whenever the project-root entry
carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code
writes that value only on an explicit "No, disable". Claude Code's default project
entry carries Approved and WarningShown both false before the dialog is ever shown, so
every such project refused every spawn.
Only Approved === false with WarningShown === true — the pair the dialog writes on a
decline — now counts as a decline. false/false behaves like an absent flag: trust is
registered and no import consent is manufactured.
New case test_project_root_entry_default_import_flags_are_not_a_decline fails on
b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31,
bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* no-mistakes(review): Correct harness doc's external-imports decline predicate
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(bin): keep operator-address labels out of no-mistakes intent (#4445)
* fix(brief): keep operator address out of composed intent
Teach raw-word authoring for intent sections and mid-task relays, with a neutral [captain] provenance marker for legacy mixed tasks. Keep headings and contract prose outside the serialized intent body.
The legacy selector already excluded the old speaker labels from its output; preserve that read compatibility. The reproduced leak comes from adding labels inside a modern intent body, not from the legacy selector. Do not scrub actual request content.
Add exact serialized-input and generated-contract regressions, retaining refusal of unmarked legacy tasks and coverage of scout promotion.
Fixes https://github.com/kunchenguid/firstmate/issues/3882
* no-mistakes(review): Refuse operator-address lines in Captain's intent body
* no-mistakes(document): Document operator-address refusal in intent contract comments
* fix: classify OpenCode ellipsis hint as idle (#4451)
* fix(composer): recognize Grok 1.0.5's oversized titled bottom border as a proven empty composer (#4455)
* fix(composer): accept Grok title overhang
* no-mistakes(review): summary: named Grok overhang constant, doc caveat, restored tmux typed-title coverage
* fix(bin): translate Stop hook timeout signals into durable auto-arm failure (#4474)
* fix(bin): recover Claude auto-arm after timeout
* no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list
* fix(spawn): establish Claude task channel authority (#4464)
* fix(spawn): establish Claude task channel authority
* no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference
* fix(bin): refuse fm-control.sh exit when the composer holds unproven or pending text (#4458)
* fix: guard relaunch exit against pending input
* no-mistakes(review): Verifying test run in progress
* no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard
* no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now
* fix(spawn): establish crewmate identity first (#4481)
* fix(bin): reconcile redundant secondmate divergence during updates (#4460)
* fix: reconcile diverged secondmate updates
* no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md
* no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md
* feat: enable gpt-5.6-luna max reasoning for crew dispatch (#4497)
* fix(dispatch): support Codex Luna max effort
* no-mistakes(review): use portable CODEX_HOME path in codex effort reference
* feat(calm): render smooth Unicode swell with asymmetric two-color sail (#4498)
* feat(calm): render smooth Unicode swell
* feat(calm): make sails asymmetric
* feat(calm): use quarter sail glyph
* no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer
* no-mistakes(document): docs: sync calm wave phase doc comment
* no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。
* fix(bin): supersede stale scout delivery text in brief.md on promotion (#4491)
* fix: supersede scout delivery brief on promotion
* fix: preserve ship safety contract after promotion
* no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch
* fix(bin): make captain holds work on hosts with an older JSON::PP, and stop cleanup dropping accents from a held body (#4471)
* fix(bin): let captain holds work on hosts with an older JSON::PP
Holding a task for the captain, and the cleanup that keeps a captain-held row
open, both fail outright on any host whose JSON::PP defaults allow_nonref off -
2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`,
but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older
library rejects that whole value with "must be object or array".
The consequence is fleet-wide on such a host, not one broken command: a worker
there cannot formally record a decision for the captain at all. It can only
mention the decision in passing in a status line, where it can be missed - which
is how a real decision goes unrecorded. The hold reports that the task lost its
hold-set stamp; the cleanup cannot return the row to Queued.
Both call sites now ask for allow_nonref explicitly rather than inheriting
whatever the installed library defaults to. The second one is worth naming: its
`/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string
case that fails, so the guard selects for the failing input rather than
protecting against it.
The regression case forces the older default back off for every perl the commands
spawn, then drives both paths - holding a task that carries a body, and tearing
down a captain-held row whose deliverable must still be appended. It also probes
that the simulation genuinely rejects a bare scalar, so the case cannot pass
vacuously on a lenient host. Each half was verified failing on its own unfixed
call site with that site's real error message. Suites: fm-captain-hold-lifecycle
51 cases, fm-backlog-atomicity 99 cases, 0 failures.
Verification limit: the mechanism is reproduced and tested, but neither fix is
verified against a real JSON::PP 2.27202 host, because none is in the loop. This
laptop runs 4.06, where the bug does not manifest.
`bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a
brace-delimited object before decoding, so allow_nonref never applies.
* fix(bin): stop cleanup silently dropping accented characters from a held body
Cleanup rewrites a captain-held row's body to append the finished work's
deliverable, and the decoder it reads that body with printed decoded characters
to a stream with no `:raw` layer. A character at or below U+00FF then came out
as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the
accent. `fm_backlog_retain` writes that body straight back through
`--body-file`, and nothing reported an error - the character was simply gone
from a row still waiting on the captain.
The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus
`utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already
used.
Review of the parent commit found this on one of the lines that commit already
changed. It predates that change.
The test asserts bytes rather than decoded strings, because comparing strings
cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any
character above U+00FF makes perl print the whole string as UTF-8, so one body
carrying both an accent and an em dash passes even unfixed and proves nothing.
Verified failing before the fix on the accented row, passing after. Suites:
fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures.
* no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc
* no-mistakes(review): drop whole-file UTF-8 check from retained-body test
* no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc
* no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc
* fix(bin): read codex 0.154's idle braille starfield rows as composer furniture (#4532)
* fix(composer): read codex 0.154's idle starfield and status footer as furniture
codex-cli 0.154.0 animates a braille "starfield" around its idle composer:
on the row above the bold `›` prompt row, on the `›` row behind the SGR-2
dim `Ask Codex to do anything` placeholder, and on the row below it, then
draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`).
The cells are truecolor greys on both sides of the ghost luminance ceiling,
so the brighter ones survive ghost stripping, and the rows below the glyph
carry no structural edge. The shared classifier selected the bare `›` shape,
extended its wrap region over the two rows beneath the glyph, read the
survivors and the footer as wrapped typed input, and answered `pending`;
the steering doorbell defers on exactly that verdict, so no doorbell ever
reached an idle codex 0.154 pane.
bin/fm-composer-lib.sh now recognises that furniture by shape, declared
once next to the idle placeholders and reached from the two wrap-region
boundary points:
- a row whose non-whitespace content is entirely braille cells
(U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it
never counts as wrapped typed content and bounds a bare composer's wrap
region; braille behind the glyph row's content is stripped before the
emptiness decision when nothing else follows the glyph; a row mixing
braille with other text stays typed content;
- the codex status footer bounds the wrap region exactly as omp's status
row does, anchored on the effort token, a spaced middle dot, and a `~` or
`/` path cell, so a typed `fix · tests` stays composer input;
- `^Ask Codex to do anything$` joins the verified idle-placeholder set; the
ghost strip remains what proves that row empty, and the bare-row rule that
bright placeholder text is real input is unchanged.
Unchanged: the strict blank-row rule, the styled=0 degradation (a plain
cmux/orca capture of this screen still reads `unknown`, never `pending`),
FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape.
tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte
with the divergence (letters in place of the starfield read `pending`) and
the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh
is the default-on live guard (token-free, skips explicitly without codex or
tmux) that launches the installed codex idle and asserts `empty` through
both the tmux and the cursorless styled reads, naming codex --version on
failure. docs/verification/runtime-backends.md records the dated Herdr
evidence: `pending` before, `empty` after, on the captured screen.
* no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry
---------
Co-authored-by: Todd Billings <todd@usdvcapital.com>
* fix(bin): refuse empty text steers in fm-send (#4259)
* fix(bin): refuse empty text steers in fm-send
A marked secondmate request sent with an empty message delivered only
marker and correlation bytes and minted a pending-reply expectation the
parent could never see resolved, stalling the fleet with no loud error
(#4255). Fail closed on an empty or whitespace-only message on the text
path, mirroring the existing --resolve-key refusal.
* chore: retain ambient Pi-lens autoformat as its own commit
Formatting-only edits produced by ambient Pi-lens autoformat during the
msg-loss investigation, kept separate from the behavioural change in
c23acba6 so the fix stays reviewable on its own.
AGENTS.md is deliberately excluded: its only autoformat edit stripped the
trailing space from the documented FM_OPERATIONAL_PREFIX value, which
bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11
records as permanent compatibility. Documenting that constant without its
trailing space makes the doc wrong about the contract, so that one line was
restored rather than retained.
* fix(calm): paint the working ship one yellow over all-blue water (#4554)
On rose-pine-moon the two-color water (cyan crests over blue troughs) read as
a pink stripe over aqua, the yellow left sail and mast clashed with the red
right sail, and the hull carried a blue interior run. Every water cell is now
blue so the swell reads through glyph height alone, and both sail halves, the
mast, and the whole hull are one yellow run. Geometry, cadence, animation,
direction flip, resize clamping, and the narrow fallback are unchanged.
Update the unit and real-TUI color assertions to the new palette and the Calm
docs that described the old one.
* fix(bin): stop aging a second mate's active turn from its launch (#4270)
* fix(watch): stop aging a second mate's active turn from its launch
The parent watcher's second-mate wake-loop stall check exempts a mate that
is demonstrably inside an active turn, but secondmate_in_active_turn asked
busy_turn_over_age first and returned "not in a turn" whenever that said
the bound was crossed.
busy_turn_over_age ages from state/<task>.turn-ended, falling back to
state/<task>.meta. A second mate's turns end in its own home, so the
parent never gets a turn-ended mark for it and the fallback ages the
mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago
was therefore permanently "over age", the busy pane was never consulted,
and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false
wake-loop stall.
The gate now bounds the busy exemption by <idle> - how long the queue's
drain position has not moved - which is evidence this home actually
holds. A busy mate stays exempt while the queue has been frozen for less
than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so
the bound that stops a busy pane from proving liveness forever is kept
rather than removed. busy_turn_over_age is untouched; its remaining
callers are the ordinary crew busy-pane bound.
The regression pins the case that actually broke: a mate whose launch
record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn
must not escalate, while the same mate with its queue frozen past the
bound still publishes exactly one notification. The existing coverage
only exercised a freshly launched mate, which passes either way.
Reaching that alert now costs a pane capture inside the gate, so the
three checkpoints in this suite that assert an alert move from a 1s to a
4s bound - the value the neighbouring active-turn cases already use. The
bound is a ceiling, not a wait: the checkpoint returns on the first
actionable wake. On a loaded machine a 1s bound missed the alert
repeatedly; at 4s it did not miss in 20 runs under the same load.
* no-mistakes(review): scope the second-mate active-turn regression test's coverage claim
* no-mistakes(document): fix stale second-mate active-turn comments in fm-watch
* feat(bin): add read-only PR blocker and reviewer discovery commands (#4278)
* feat(bin): add read-only PR blocker and reviewer-discovery commands
Two focused, opt-in commands that read GitHub and never write to it.
fm-pr-state.sh reports what still blocks one pull request from the
author's side: a closed or merged state, draft state, unknown or
conflicting mergeability, absent or failing required checks, and a
blocking CHANGES_REQUESTED decision explained by each reviewer's latest
verdict, marked STALE when it was left at a superseded head. A pull
request that only awaits an approval is not reported as blocked, and
advisory checks are omitted. Every reading is taken against one exact
head; a push that lands mid-read invalidates the whole result rather
than mixing two snapshots.
fm-pr-reviewers.sh suggests reviewers from the most recent commits to
the pull request's exact changed paths, counting each commit once,
resolving handles through GitHub's own commit author.login mapping, and
excluding the author and Bot accounts.
Both stay read-only: no review request, no approval, no merge.
Unresolved review-thread state is left unreported because the REST API
does not expose it and unattended commands may not use GraphQL.
Closes #3731
* no-mistakes(review): accept only PR URLs and stop at terminal state
* no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate
* no-mistakes(review): stop attributing readings to unverified heads
* no-mistakes(review): narrow readiness contract to checks that have reported
* no-mistakes(review): read the pull request once, drop the head guard
* no-mistakes(document): scope pr-forge isolation proof to its measured members
* no-mistakes(document): record uncovered pr-forge members and their pending proof
* docs(isolation-proof): re-prove pr-forge at its full membership
tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the
pr-forge family in this branch, and script_allows_concurrency grants
four workers by family membership alone, so both ran concurrently on a
proof measured before they existed.
Re-proved the family at all eight members: two consecutive runs, 0
failures, each begun with the one-minute load average below 6.0 so the
result measures isolation rather than contention. A third run taken
between them is disclosed rather than recorded, because it started
while the previous run's workers were still decaying.
The new durations are not comparable with the six-member measurement
above them, so they are not presented as evidence about the two new
members, and that record's 1.72x four-worker figure is left as a
statement about its own run rather than restated as current.
* no-mistakes(review): disclose gh error-text coupling at its matching site and tests
* fix(bin): teach validation-round pauses in generated briefs (#2752)
* fix(bin): teach validation-round pauses in briefs
* no-mistakes(document): Point classifier comments to authoritative pause examples
* docs(readme): add star history chart (#4558)
* fix(bin): refuse teardown when a task's endpoint close fails (#4510)
* fix(teardown): refuse a cleanup whose endpoint close failed
bin/fm-teardown.sh discarded both the exit status and the stderr of every
fm_backend_kill call, so a close that genuinely failed was indistinguishable
from one that succeeded. Teardown continued past it, deleted the task's durable
records, returned its worktree, and reported the cleanup as completed. The
deleted metadata is the only record of which endpoint belongs to the task, so
such a close did not merely leave a stray session behind, it stranded one:
nothing was left on disk naming it.
The adapters could not carry that signal either. Driven against the real code,
every backend arm returned 0 for a genuine failure exactly as it did for an
already-exited endpoint, so there was nothing for the four call sites to
propagate even once they stopped swallowing it.
The tmux arm now resolves a close that did not succeed against the window's
exact recorded identity, since kill-window fails the same way for a window that
is gone and one that is still there. The Orca arm reports a close its missing
CLI never attempted. Both stay silent for an endpoint that is already
legitimately gone, and the remaining arms are unchanged: their close-command
timing cannot be established without the real Zellij, Orca, and cmux binaries,
and a gate that refused ordinary cleanup of an already-exited session would be
worse than the defect. docs/verification/runtime-backends.md records what each
backend can prove.
A reported close failure now reaches teardown's existing retain-and-stop
refusal before the records naming the endpoint are removed, matching where the
Herdr confirmed-gone gates already sit for the same hazard, and the retained
records let a rerun finish once the close works.
* no-mistakes(review): refuse unreadable tmux close re-read; honor --force override
* no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close
* no-mistakes(document): document endpoint-close refusal in its backend and retirement owners
* no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree
* feat(calm): add flag-gated Claude Code Calm mode (#4565)
* feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag
Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that
brings Calm to Claude Code: the sailboat replaces the stock working row through a
Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration,
and canonically classified operational user rows draw at zero height. /calm is
registered by the hooks module itself and toggles the same per-home config/calm
preference the Pi extension uses, so one choice applies on either harness; rows
redraw retroactively on toggle and stay hidden across claude --continue.
The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS
flag is on. Nothing sets that flag in any settings file, and the plugin carries no
command file, skill, agent, or classic hook, so it is a complete no-op while the
flag is off. The trusted project auto-loads it through an .agents/skills symlink,
the only path Claude Code scans for project plugins.
Extract the working-ship geometry, bounce track, cadences, and freeze/resume state
into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module
imports from outside the plugin folder) and have the Pi widget paint that core's
frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify
operational rows through a port of bin/fm-operational-input.sh's classify command
guarded by a corpus parity test against the shell owner.
Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering,
Raster packing, policy, classifier parity), the mod's own claude plugin test suites
behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op,
the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272.
Docs: record the version-scoped Claude Code evidence and the three bounded gaps in
docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md,
and make the shared preference, layout, and contributor notes harness-neutral.
* no-mistakes(review): Preserve colliding final replies and strengthen parser parity
* no-mistakes(review): Preserve final replies and strengthen canonical parity checks
* no-mistakes(review): Require exact function-hooks opt-in before Calm activation
* no-mistakes(review): Clarify Calm module loading and activation boundaries
* no-mistakes(review): Reset Calm presentation state across session starts
* no-mistakes(document): Refresh Calm session lifecycle documentation
* feat(calm): paint the Claude Code working ship in Claude's own theme colors
The captain picked the "Claude native" palette for the Claude Code mod's Raster:
every water cell takes the spinner blue of the active theme family (#93a5ff dark,
#5769f7 light) and the whole boat takes the Claude orange of the stock spinner
(#d77757), one water color and one boat color. The family follows the `theme`
setting's prefix, read at load through $.config.list and re-read on a
config.set of that row, with `auto` and custom themes falling back to the dark
set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte.
Rename the shared sprite's color classes from hue names to `water` and `boat`,
since each harness now maps them to its own colors; geometry, motion, cadence,
and the activation gate are untouched.
Tests cover both palettes' packing and the family rule under Node, and the
plugin kit drives every theme value, a theme change mid-session, the Calm-off
pass-through, and inertness of the menu read while the flag is off. The docs
describe the Claude Code colors and record the guard passing on 2.1.273.
* no-mistakes(review): Use light palette for unresolved Claude themes
* no-mistakes(document): Refresh Claude Calm verification evidence
* fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586)
* fix(watch): honour a declared wait before wedge-escalating a quiet pane
wedge_timer_check escalated on elapsed idle time alone. Nothing asked
whether the worker had already said why its pane was quiet, so a lane
that declared a bounded external wait climbed the escalation ladder for
as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every
repeat carried demand-deep-inspection - which by its own wording forbids
re-absorbing on the run-step or pane state, so the supervisor could not
use the evidence that was there either.
The generated brief promises that declaring `paused:` buys the long
recheck cadence instead of a wedge, but the timer was still reachable
while that declaration stood: a crew that declares a wait and then has an
active run or busy pane attributed to it is handed to the timer as
provably-working. The declaration is what the worker said about its own
silence, so it now outranks a liveness verdict that only says something
is running.
The consult runs in the at-threshold branch that was about to escalate,
beside the worktree walk already there, and costs one status-line read.
Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS
recheck the declared-wait absorber already uses, so the wait is still
rechecked and cannot rot invisibly. Which verb declared it decides the
wording, because the two block on different people: a `paused:` wait is
owed by an external dependency and asks the reader to confirm it still
holds, while a `captain-held:` transfer is owed by the captain reading
the recheck and asks them to answer or release the hold. A hold is not
rechecked at all while the away-posture record exists, as on every other
captain-held path, and that absorb arms no throttle so the recheck is
owed in full on return.
A declared clearing time that has already passed stops counting, and a
lane that never declared one keeps the identical escalation schedule,
reason, count and demand-deep-inspection wording, so detection and its
worst-case time are unchanged. The deferral restarts the idle timer
rather than cancelling it, so a lane that stops waiting escalates again
within one threshold.
A lane quiet because its own validation run is parked at a gate awaiting
a human decision is deliberately out of scope: reading that state needs a
signal carrying who the wait is on and what clears it, rather than one
inferred from a parked verdict that also covers gates awaiting the
crewmate itself.
Tests pin both directions for each case and were each confirmed to fail
with the consult removed.
* no-mistakes(document): docs: honour declared waits in stale-escalation docs
* fix(bin): report verified PR state for passed runs (#4624)
* fix(bin): derive passed PR state from PR record
A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists.
For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim.
Fixes #4607
* no-mistakes(review): Add bounded GitLab merge-request state reads
* no-mistakes(review): Preserve network-free inactive crew-state scans
* no-mistakes(document): Document PR record readers in shared library
* fix: restore published contribution follow-up (Fixes #4469) (#4627)
* fix: restore published contribution follow-up (Fixes #4469)
* fix(review): Fix contribution freshness and merge actor routing
* fix(review): Restore issue triage and scope contribution follow-up
* fix(test): test: assert one wake per contribution signal
* fix(document): Document contribution follow-up
* fix: restore truthful terminal delivery evidence
* fix(review): Disclose unsupported contributions and deduplicate watcher wakes
* fix(review): Preserve unmeasured unsupported contributions across Bearings
* fix(review): Deduplicate shared contribution wakes and isolate diagnostics
* fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower
* fix(bin): make remote report transfers explicit and fail-open (#4658)
* fix(bin): make a remote-reply document gap self-clearing and re-attemptable
A remote mate's undelivered document raised a keyed `blocked` decision that
nothing could ever resolve, and any `data/*.md` substring in any mirrored line
was an unconditional fetch instruction. A mate announcing a report it had not
written yet therefore manufactured a permanent, factually false blocker, and
its own explanation of the false alarm manufactured more.
The reader has no permanence vocabulary: a report still being written refuses
exactly like a path that will never exist. So an undelivered document is now a
durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`,
re-attempted on the next delta and on the channel's own quiet poll, and retired
with a matching `resolved` line naming the local copy once it arrives. The
cursor still advances and no delta stalls on one bad pointer.
Only a structured `report=data/....md` pointer now offers a document, so a path
merely mentioned in prose - including one under another home's mirror tree,
which is provably not that mate's to serve - is never fetched. Offers are
deduplicated across the whole delta, the escalation names each missing document
once and carries the reader's own reason instead of discarding it, and a
strictly increasing notice ordinal keeps a later escalation from being
swallowed as duplicate bytes. A mirrored line still lands once whichever
pointer form it was first written under.
* no-mistakes(review): Require structured pointer token boundaries
* no-mistakes(review): Unify boundary-safe pointer extraction and rewriting
* fix(bin): identify a mirrored line independently of its delivery state
Two defects in the boundary-safe pointer work.
The at-most-once check compared only the all-remote and all-local renderings
of a line, so it could not recognize a mixed one. A line offering two documents
where only the first was deliverable mirrored as local-plus-remote; once the
second arrived, a cursor-loss whole-log recapture rendered the same line
all-local, matched neither alternate, and mirrored a second time. A line's
identity is now the canonical form every boundary-valid pointer would take once
delivered, derived by the same parser that does extraction and rewriting, so it
no longer depends on which documents happened to be deliverable at the time.
The pointer map was passed to awk through the process environment. A delta may
carry up to the configured 1 MiB bound, and an expanded map of delivered
pointers can exceed the platform's exec argument limit, so awk would fail to
start; because no caller checked, the empty result would have been appended as
blank lines while the cursor advanced past dropped status content. The map now
travels in a file, and every call site checks the exit status and stops the
ingest rather than committing a delta it could not render.
Both passes now run once per stream instead of twice per line.
* no-mistakes(review): Abort ingest when document pointer extraction fails
* no-mistakes(review): Exclude structured cross-home pointers from document transfer
* fix(bin): fail open on an undeliverable remote document instead of tracking it
Narrow the remote-reply document fix to the scope the diagnosis actually
requires, as decided after measuring a simpler alternative.
A document the reader cannot deliver now fails open. The mate's line is
mirrored with its own pointer, the cursor advances, and one unkeyed note
carries the reader's reason. A note never enters the open-decision fold, so it
cannot stand open the way the original keyed block did - which removes the
never-clearing false blocker by construction rather than by resolving it.
That makes the durable self-clearing obligation unnecessary, so it goes: the
per-mate pending-documents record, its notice ordinal and resolved
announcements, and the poll-side retry. Canonical line identity goes too, and
with it a way to silently drop a genuine status line; mirroring is back to
at-most-once on exact bytes. The cross-home exclusion goes as well: under
fail-open a cross-home report= either fails harmlessly or is a nested remote
report this mate genuinely holds, which is now relayed again.
Kept: fetching only on a structured report= pointer, the boundary-correct
parser, the file-based rewrite map, and checked extraction and rewrite exit
status. The parser now scans behind a sentinel byte so a rejected candidate can
no longer give the text right after it a false leading boundary.
The reported incident is covered end to end: a report path announced in prose
before it exists raises no decision, and the report still arrives through the
ledger publisher's structured offer once written.
* no-mistakes(review): Preserve source-line identity across remote reply replays
* no-mistakes(document): Document remote reply transfer and replay semantics
* no-mistakes(lint): Fix staging truncation lint checks
* fix(calm): preserve substantive mid-turn responses (#4655)
* Preserve substantive Calm mid-turn text
* no-mistakes(review): Distinguish newline-preserved replies from short narration
* no-mistakes(document): Document Calm mid-turn preservation boundaries
* no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`
* fix(bin): preserve PR merge polls across volume remounts (#4656)
* fix(bin): re-record PR poll identity after a volume device renumber (Fixes #4260)
A volume remount can renumber the state filesystem's st_dev while every
inode and byte stays the same; APFS does this across a reboot. A poll
registration records its sidecar and check as device:inode, so every poll
armed before the remount failed strict validation and the watcher refused
all of them as unauthenticated state checks until each was re-armed by hand.
There are two device comparisons. fm_pr_private_file_valid compares a live
file's device with the state directory's device read in the same invocation:
it refuses a file that is not on the state directory's own filesystem and
already survives a renumber, so it is unchanged. The registration's recorded
identity versus the live identity (from #556, reused by the #932 retirement
receipt) binds the registration to the exact files published in its own
transaction; its device part is what breaks.
When strict capture fails, the watcher now proves the device is the only
difference: every other artifact check passes (template bytes, both hashes,
private mode, single link, live device, metadata), both recorded identities
name one device, and each recorded inode equals its live inode. Only then,
under the task's control lock, does it rewrite the two identity lines,
repeating the whole proof and comparing the registration's file identity and
bytes just before the rename, and then capture strictly again. A swapped,
altered, re-moded, relinked, split-device, or foreign-device artifact still
fails a proof and is still refused, and a pending retirement receipt blocks
the rewrite.
Reproduction: on macOS a poll armed on an APFS disk image that was detached
and re-attached behind another image moved st_dev 16777239 -> 16777243 with
inodes, bytes, mode, and link count unchanged; the real watcher refused it on
main and reports its merge with this change. The portable regression test
rewrites a real registration's recorded device and drives the watcher.
Not changed here: the status presentation cursor keys rows by its own
device:inode identity in bin/fm-classify-lib.sh, a different helper that
needs its own fix; a retirement receipt left by a reboot between its
publication and removal still names the old device and stays refused; custom
check trust binds only a content hash and is unaffected.
* fix(review): Serialize PR poll publication writers
* fix(review): Bound PR poll publication lock scope
* fix(bin): keep contribution records when the poll budget runs out (follow-up to #4627) (#4661)
A budget that expires partway through an observation no longer records an
error or prints the unavailable wake; the URL keeps its prior record and is
observed first next poll. forge() flags budget exhaustion at the point it
refuses, or when a read is killed at the budget's own deadline, so a genuine
forge failure still records the error and wakes. Each distinct URL is now
observed once per poll and applied to every owning task.
* fix(bin): clear parent pending-replies on local secondmate retirement (#4680)
* fix(bin): clear parent pending-replies on local secondmate retirement
Local secondmate teardown left resolved parent pending-reply records behind
after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced
retirement while any reply for that id is still unresolved, and delete every
matching record plus its delivery confirmation after a successful local or
remote retirement, matching the remote cleanup path.
* no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup
* no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung
* no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern
* no-mistakes(review): Pending-replies Basename und corr_id abgleichen
* no-mistakes(document): Clarify forced retirement pending-reply cleanup
---------
Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>
* fix(bin): accept Orca's composite worktree id when tearing down a task (#4677)
* fix(bin): accept Orca's composite worktree id at teardown
Teardown refused every Orca-backed task because the endpoint validator
checked orca_worktree_id with the simple-atom rule meant for tmux-style
window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca
returns that id as `<orca id>::<absolute worktree path>`, so the colon and
slashes in every real value made validation fail and finished Orca tasks
could never be cleaned up.
Validate the field as the composite it is: both halves of the first `::`
split present, the path half absolute, and no embedded newline, carriage
return, or tab. The terminal field keeps the atom check, which is correct
for it, and no other backend's validation changes.
The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca
never returns, which is why the suite passed a check the real value fails.
They now carry the composite form, so the tests exercise the real value.
* no-mistakes(document): name Orca's repo id in the composite worktree id
* no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points
* feat(bin): add opt-in typed dispatch resolution (#4692)
* feat(bin): add opt-in typed dispatch resolution through typesafe.ai
Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or
scout profile from a written brief with typesafe.ai's System One model:
one Choice question over the rules' `when` texts, then the confidence
floor, the rule's `approval` and `floor`, each profile's `provider` and
`floor`, one quota-axi snapshot, and the spendPriority argmax all in code.
It is off unless TYPESAFE_API_KEY is in the environment or the home's
gitignored .env; off means one stderr line, exit 0, and no network call,
so firstmate dispatches exactly as before. The key reaches curl on a file
descriptor, never argv.
Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and
the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new
tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates
the four new optional dispatch fields. Document the schema, the operator
contract, the AGENTS.md intake step, and the live and benchmark evidence.
* no-mistakes(review): Harden typed dispatch resolution and quota bounds
* no-mistakes(review): Validate dispatch floors and ranking evidence
* no-mistakes(review): Tighten dispatch response and floor evidence
* no-mistakes(review): Neutralize none matching and resolve defaults locally
* no-mistakes(review): Preserve providerless profiles outside typed resolution
* no-mistakes(review): Validate response usage and reject duplicate profiles
* no-mistakes(review): Escalate unverifiable floors and validate probabilities
* no-mistakes(review): Validate probability mass and unknown profile floors
* no-mistakes(review): Simplify resolver interface and preserve fallback routing
* no-mistakes(review): Fix constants and rank partial quota evidence
* no-mistakes(review): Add authoritative provider mapping and enforce explicit providers
* no-mistakes(review): Declare provider for documented Pi profile
* no-mistakes(review): Validate provider identifiers and support Gemini dispatch
* no-mistakes(review): Strictly anchor provider identifiers
* no-mistakes(review): Validate selectors and preserve fallback candidate evidence
* no-mistakes(review): Gate typed validation and harden resolver evidence
* no-mistakes(review): Preserve opt-in routing and harden candidate evidence
* no-mistakes(review): Prioritize known exhaustion over quota uncertainty
* no-mistakes(review): Isolate API secrets and preserve no-key diagnostics
* no-mistakes(review): Fallback safely when dispatch rules are absent
* no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets
* no-mistakes(document): Document typed dispatch safety and fallback behavior
* fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753)
* test: reproduce buried status declarations in shared readers
* fix: share status event reads and preserve open blockers
* fix: retain terminal scout and ship status declarations
* no-mistakes(review): Fix status chronology, legacy completions, and reader performance
* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots
* no-mistakes(review): Unify terminal supersession across cached folds and consumers
* no-mistakes(review): Filter per-key status history while preserving terminal chronology
* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells
* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses
* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor
* no-mistakes(lint): Quote literal done in test for-lists for SC1010
* ci: expect 19 snapshot/fleet-view tests
This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.
* no-mistakes(review): Restore multiline child outcome reporting
* no-mistakes(review): Select ledger terminal events through bounded shared reader
* no-mistakes(review): Report newest open decision instead of preferring blocked
* no-mistakes(review): Require colon before ship/scout terminal supersession in fold
* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix
* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions
* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold
* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule
* no-mistakes(document): Align status-read docs with fold-resolved crew state
* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers
* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree
* test: fold terminal-cleanup snapshot coverage into the completed-scout case
Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.
* no-mistakes(document): Clarify socket-down override expiry in architecture doc
* ci: retrigger flaky contribution check
* fix(bin): launch codex crewmates with codex's hook layer disabled (#4689)
* fix(spawn): launch codex crewmates with codex's hook layer disabled
A freshly launched Codex worker never reached its instructions. Codex
stopped it on an interactive "Hooks need review" modal whose selection
sits on "Review hooks", which is neither trusting nor declining.
Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow
navigation, so the selection cannot be moved, and pre-accepting the
prompt by writing Codex's own trust store would record an operator
consent that was never given.
The hooks are the machine's own ~/.codex/hooks.json plus any project's
.codex/hooks.json. A crewmate needs neither: its turn-end signal is the
-c notify= program on the same launch, and Firstmate's project hooks are
primary-session infrastructure that stands down in a child worktree.
Crewmate and scout launches now pass --disable hooks. That is the
opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted
hooks; disabling the feature runs none of them and leaves the operator's
~/.codex untouched. An unknown feature name is a hard Codex error, so a
release that drops the flag fails the launch loudly instead of silently
restoring the modal. A secondmate is a primary in its own home and keeps
the project hooks its turn-end guard and session-start digest ride on.
Verified on codex-cli 0.151.0: the modal is gone and the turn-end
notification still lands.
This unblocks the second review that every finished pull request is supposed to get.
Fixes kunchenguid/firstmate#4673
* no-mistakes(review): Fix contradictory hook count in Codex verification record
* fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710)
* fix(bin): settle terminal contributions and wake once per read-failure episode
A contribution whose last good observation is merged or closed is final:
poll no longer re-reads it, projection keeps it fresh, and a stale error
recorded beside it is cleared once. A genuine forge-read failure on an open
contribution still records its error on every cycle but prints the
unavailable wake only when it starts a failure episode; a successful read
ends the episode. Open PRs linked from done tasks keep being observed.
The false unavailable beside a complete observation was budget exhaustion
mid-observation, already fixed by #4661.
* fix(review): Settle terminal contribution owners
* fix(review): Deduplicate shared contribution failure episodes
* fix(test): Preserve settled terminal contribution records
* fix: select authoritative no-mistakes runs (#4476)
* fix(crew-state): select authoritative validation runs by identity
Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row.
Refs: https://github.com/kunchenguid/firstmate/issues/3215
* fix(review): Resolve same-branch run identities beyond capped history
* fix(review): Fix run-selection compatibility, races, and worker-state fallbacks
* fix(review): Limit run validation to the requested branch
* fix(test): Anchor AXI fixtures and document remaining live evidence gaps
* fix(document): Clarify run selection documentation and capture ownership
* fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation
* fix: distinguish captain outcomes from no-op updates (#4738)
* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work
MAIN answered a supervision-branch outcome for completed captain-requested
work (implementation done, PR ready for review and merge approval) with
"Captain, shipshape.", reading section 9's no-action reply as covering it
and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no
captain-facing response is owed".
Section 9 now limits the shipshape reply to true no-ops (idle re-read,
empty heartbeat, consequence-free acknowledgement) and requires a short
outcome response naming what finished and what word is needed whenever
requested work finishes or a result needs the captain's word, even when a
transcript entry already shows the substance. The Pi protocol's re-emit
rule now says it bounds repetition only, and carries a worked example of
the ready-for-review outcome whose correct processing turn a shipshape
reply fails.
No executable contract evaluates the content of MAIN's captain-facing
reply, so the regression is the protocol example in the owner doc rather
than a text-match test.
* no-mistakes(document): Clarify captain-facing outcomes versus no-ops
* docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line
The document step condensed the Pi protocol's re-emit rule and dropped the
worked example of a finished, ready-for-review outcome whose correct
processing turn a "Captain, shipshape." reply fails. That example is the
contract's regression: no executable contract evaluates the content of
MAIN's captain-facing reply, so the owner doc's example is the test case.
Restore it directly under the re-emit rule, prefixed as a regression
example that is kept verbatim and never condensed or summarized away.
* no-mistakes(review): Clarify captain outcome and decision-word requirements
* no-mistakes(document): Clarify captain-facing completion outcomes
* docs(pi): require the PR URL in the visible captain-facing outcome reply
Captain review on the regression example: drop the sample reply string
and say only that the ready-for-review outcome requires relaying a
captain-facing outcome response, not just "Captain, shipshape.".
Fold in the visible-PR-handoff failure seen this session: after the
branch outcome reporting this fix green, MAIN's visible reply was only
"Awaiting your merge call." with no PR URL, leaning on the dim anchor.
Section 9's URL rule now also covers a review or merge ask and names the
visible reply as where the URL goes, sourced from the ready status, pr=
metadata, or the supervision branch's summary and never left to a
transcript entry. The Pi protocol adds the same-way failure and places
the captain-facing text in the final visible assistant reply after the
fm_branch_processed call, because Calm hides assistant text emitted in
the same step as a tool call as a working note.
Investigation verdict, evidence in the PR comment: no recent PR caused
the handoff failure; Pi has hidden same-step pre-tool assistant text
since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and
#4658 touched only remote report transfer.
* no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs
* no-mistakes(document): Clarify captain-facing supervision outcomes
* docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule
The consolidated section 9 URL rule narrowed its trigger to a review or
merge ask, dropping the "whenever a PR is mentioned" catch-all from
#3648 that keeps every PR URL copied from a durable record and never
assembled from memory. Restore that trigger as a union with the review
or merge ask so the one consolidated rule covers both.
* fix(bin): let non-owner Claude Stops exit safely (#4777)
* Fix foreign-owner turn-end supervision loop
* no-mistakes(review): Scope foreign-owner safe exit to Claude guard
* no-mistakes(document): Document Claude foreign-owner safe exit
* fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778)
Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty
indexed array as an unbound variable and aborts the shell. In
signal_turnend_panes_churned() the missing_keys loop was reachable with
an empty array whenever every churned key already held a fresh
.churn-since-* marker (a second churning turn-end inside an open
deferral window), so each watcher cycle died about half a minute in and
supervision restarted endlessly. The created_keys rollback loops had the
same latent crash on their error paths.
Audit of bin/ for the same pattern found one more confirmed-reachable
case: remote_handoff's noncanonical-body scan iterates to_move, which is
empty when a retried remote handoff finds every key already staged in
the outbox. All other "${arr[@]}" sites are either count-guarded,
guaranteed non-empty by construction, or unreachable while empty.
Guard the three reachable expansions with the repo's existing
"${arr[@]+...}" idiom. New regression test drives a real watcher
through the all-marked churn path; the macos-stock-bash CI lane runs it
under real /bin/bash 3.2 via FM_TEST_ONLY.
* Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783)
The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: require complete captain-facing final responses (#4779)
* docs: require complete final responses across harnesses
* no-mistakes(document): Document complete final replies for Grok Bot
* docs: point Grok replies to the shared contract owner
* no-mistakes(review): Clarify final recap without batching decision asks
* fix: preserve substantive mid-turn text in Pi Calm (#4788)
* fix(calm): preserve substantive Pi mid-turn text
* no-mistakes(review): Preserve substantive Pi Calm text per block
* no-mistakes(test): Cover shared Calm preservation boundaries behaviorally
* no-mistakes(document): Consolidate Calm preservation documentation
* fix: harden mail checks and rebalance full-coverage CI (#4800)
* Improve CI reliability and rebalance full-coverage validation
* no-mistakes(document): Clarify lint partition documentation
* fix(bin): answer Kimi 2.0.0 folder-trust dialog during spawn (#4799)
* Handle Kimi workspace trust dialog
* no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers
* no-mistakes(review): Gate Kimi ready on any trust marker and clean captures
* no-mistakes(review): Read visible pane for Kimi trust and ready gates
* no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate
* no-mistakes(review): Harden Kimi viewport capture and trust dialog detection
* no-mistakes(document): Document Kimi spawn refusal on cmux and Orca
* fix(bin): report a dead-agent record once instead of escalating forever (#4775)
* fix(bin): report a record whose agent is gone once instead of escalating forever
The wedge escalation path never asked whether there was still an agent to be
wedged. A wedge is something stuck that might recover, so re-alarming it earns
its cost; an agent that is gone never moves again, its pane never churns, the
idle timer never resets, and the escalate path clears its own timer and re-arms
with nothing bounding the count.
Observed on a live fleet: two finished lanes reached 226 and 203 consecutive
escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400
notifications a day from two lanes with no agent running at all. On one,
fm-control.sh exit answered already-stopped and fm-crew-state.sh read
"failed - run failed". Closing the Herdr pane did not stop it either: with the
pane genuinely gone and herdr pane read returning pane_not_found, the count kept
climbing, because the poll is driven by the record's window= line rather than by
the pane. The cost is not the repetition but that it drowns the alarms that
matter.
fm_backend_agent_state already separates a thinking agent from a gone one at
process level. In the branch that wa…
Ye1806431561
pushed a commit
to Ye1806431561/firstmate
that referenced
this pull request
Sep 20, 2026
…#4656) * fix(bin): re-record PR poll identity after a volume device renumber (Fixes kunchenguid#4260) A volume remount can renumber the state filesystem's st_dev while every inode and byte stays the same; APFS does this across a reboot. A poll registration records its sidecar and check as device:inode, so every poll armed before the remount failed strict validation and the watcher refused all of them as unauthenticated state checks until each was re-armed by hand. There are two device comparisons. fm_pr_private_file_valid compares a live file's device with the state directory's device read in the same invocation: it refuses a file that is not on the state directory's own filesystem and already survives a renumber, so it is unchanged. The registration's recorded identity versus the live identity (from kunchenguid#556, reused by the kunchenguid#932 retirement receipt) binds the registration to the exact files published in its own transaction; its device part is what breaks. When strict capture fails, the watcher now proves the device is the only difference: every other artifact check passes (template bytes, both hashes, private mode, single link, live device, metadata), both recorded identities name one device, and each recorded inode equals its live inode. Only then, under the task's control lock, does it rewrite the two identity lines, repeating the whole proof and comparing the registration's file identity and bytes just before the rename, and then capture strictly again. A swapped, altered, re-moded, relinked, split-device, or foreign-device artifact still fails a proof and is still refused, and a pending retirement receipt blocks the rewrite. Reproduction: on macOS a poll armed on an APFS disk image that was detached and re-attached behind another image moved st_dev 16777239 -> 16777243 with inodes, bytes, mode, and link count unchanged; the real watcher refused it on main and reports its merge with this change. The portable regression test rewrites a real registration's recorded device and drives the watcher. Not changed here: the status presentation cursor keys rows by its own device:inode identity in bin/fm-classify-lib.sh, a different helper that needs its own fix; a retirement receipt left by a reboot between its publication and removal still names the old device and stays refused; custom check trust binds only a content hash and is unaffected. * fix(review): Serialize PR poll publication writers * fix(review): Bound PR poll publication lock scope
sanis
added a commit
to sanis/firstmate
that referenced
this pull request
Sep 21, 2026
…dpoint reclaim (#28) * fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (#4424) * fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (#42) * fix(pr-merge): read a plan-gated 403 on branch rules as no merge queue github_read_queue_method left status=unreadable for every failed rules read, including a 403 whose body is GitHub's own "Upgrade to GitHub Pro or make this repository public" message. A repository whose plan cannot expose branch rules cannot have a merge_queue rule either, so that specific 403 now resolves to status=none instead of unreadable - unblocking the away-merge grant on private repos without GitHub Pro. Any other failure (auth, rate limit, network, 404, unrelated 403) still reads as unreadable. * no-mistakes(document): Update stale away-merge queue-grant comment for plan-gated 403 --------- Co-authored-by: NewAiCoder <claude@theinbtw.com> * no-mistakes(review): Fix misleading away-queue-grant comment in fm-pr-merge and its test * no-mistakes(document): Update architecture.md for plan-gated-403 merge queue exception --------- Co-authored-by: NewAiCoder <claude@theinbtw.com> * fix(bin): select suites that read a changed top-level test fixture (#4246) * fix(tests): select readers of a changed top-level test fixture bin/fm-test-run.sh --changed recognised shared test helpers by an explicit list, tests/lib.sh|tests/*-helpers.sh|tests/fixtures.sh. A top-level tests/*-fixture.sh matched none of those, fell through to the tests/* catch-all, and was marked unmapped, so selection aborted with "no changed-test mapping for source path" and the run selected nothing at all. tests/herdr-client-pair-fixture.sh and tests/remote-herdr-fixture.sh are real shared fixtures with real consumers, so any branch touching one of them left a validation pipeline driving --changed with a hard abort rather than a narrowed selection. Extend the helper arm to tests/*-fixture.sh rather than routing it through the tests/fixtures/*/* arm. Both arms resolve consumers with the same reference scan, and that scan is what selects the right suites here: it finds exactly the tests that read the fixture. The fixtures/ arm adds only a directory-keying step, which has nothing to key on for a top-level file, so the helper arm is the same behaviour with no extra machinery. A tests/ path nothing reads still reaches the catch-all and still refuses loudly. Refs https://github.com/kunchenguid/firstmate/issues/4100 * no-mistakes(test): order nested fixtures arm before top-level fixture glob * no-mistakes(document): document tests/ shared-file mapping contract and arm order * no-mistakes(review): drop vacuous test phase, correct header claim, restore comment * fix(bin): treat Claude Code's default external-imports flags as never asked, not declined (#4387) * fix(bin): read Claude Code's default external-imports flags as never asked, not declined (#4378) fm-claude-trust.sh refused the whole trust registration whenever the project-root entry carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code writes that value only on an explicit "No, disable". Claude Code's default project entry carries Approved and WarningShown both false before the dialog is ever shown, so every such project refused every spawn. Only Approved === false with WarningShown === true — the pair the dialog writes on a decline — now counts as a decline. false/false behaves like an absent flag: trust is registered and no import consent is manufactured. New case test_project_root_entry_default_import_flags_are_not_a_decline fails on b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31, bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * no-mistakes(review): Correct harness doc's external-imports decline predicate --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(bin): keep operator-address labels out of no-mistakes intent (#4445) * fix(brief): keep operator address out of composed intent Teach raw-word authoring for intent sections and mid-task relays, with a neutral [captain] provenance marker for legacy mixed tasks. Keep headings and contract prose outside the serialized intent body. The legacy selector already excluded the old speaker labels from its output; preserve that read compatibility. The reproduced leak comes from adding labels inside a modern intent body, not from the legacy selector. Do not scrub actual request content. Add exact serialized-input and generated-contract regressions, retaining refusal of unmarked legacy tasks and coverage of scout promotion. Fixes https://github.com/kunchenguid/firstmate/issues/3882 * no-mistakes(review): Refuse operator-address lines in Captain's intent body * no-mistakes(document): Document operator-address refusal in intent contract comments * fix: classify OpenCode ellipsis hint as idle (#4451) * fix(composer): recognize Grok 1.0.5's oversized titled bottom border as a proven empty composer (#4455) * fix(composer): accept Grok title overhang * no-mistakes(review): summary: named Grok overhang constant, doc caveat, restored tmux typed-title coverage * fix(bin): translate Stop hook timeout signals into durable auto-arm failure (#4474) * fix(bin): recover Claude auto-arm after timeout * no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list * fix(spawn): establish Claude task channel authority (#4464) * fix(spawn): establish Claude task channel authority * no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference * fix(bin): refuse fm-control.sh exit when the composer holds unproven or pending text (#4458) * fix: guard relaunch exit against pending input * no-mistakes(review): Verifying test run in progress * no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard * no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now * fix(spawn): establish crewmate identity first (#4481) * fix(bin): reconcile redundant secondmate divergence during updates (#4460) * fix: reconcile diverged secondmate updates * no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md * no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md * feat: enable gpt-5.6-luna max reasoning for crew dispatch (#4497) * fix(dispatch): support Codex Luna max effort * no-mistakes(review): use portable CODEX_HOME path in codex effort reference * feat(calm): render smooth Unicode swell with asymmetric two-color sail (#4498) * feat(calm): render smooth Unicode swell * feat(calm): make sails asymmetric * feat(calm): use quarter sail glyph * no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer * no-mistakes(document): docs: sync calm wave phase doc comment * no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。 * fix(bin): supersede stale scout delivery text in brief.md on promotion (#4491) * fix: supersede scout delivery brief on promotion * fix: preserve ship safety contract after promotion * no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch * fix(bin): make captain holds work on hosts with an older JSON::PP, and stop cleanup dropping accents from a held body (#4471) * fix(bin): let captain holds work on hosts with an older JSON::PP Holding a task for the captain, and the cleanup that keeps a captain-held row open, both fail outright on any host whose JSON::PP defaults allow_nonref off - 2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`, but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older library rejects that whole value with "must be object or array". The consequence is fleet-wide on such a host, not one broken command: a worker there cannot formally record a decision for the captain at all. It can only mention the decision in passing in a status line, where it can be missed - which is how a real decision goes unrecorded. The hold reports that the task lost its hold-set stamp; the cleanup cannot return the row to Queued. Both call sites now ask for allow_nonref explicitly rather than inheriting whatever the installed library defaults to. The second one is worth naming: its `/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string case that fails, so the guard selects for the failing input rather than protecting against it. The regression case forces the older default back off for every perl the commands spawn, then drives both paths - holding a task that carries a body, and tearing down a captain-held row whose deliverable must still be appended. It also probes that the simulation genuinely rejects a bare scalar, so the case cannot pass vacuously on a lenient host. Each half was verified failing on its own unfixed call site with that site's real error message. Suites: fm-captain-hold-lifecycle 51 cases, fm-backlog-atomicity 99 cases, 0 failures. Verification limit: the mechanism is reproduced and tested, but neither fix is verified against a real JSON::PP 2.27202 host, because none is in the loop. This laptop runs 4.06, where the bug does not manifest. `bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a brace-delimited object before decoding, so allow_nonref never applies. * fix(bin): stop cleanup silently dropping accented characters from a held body Cleanup rewrites a captain-held row's body to append the finished work's deliverable, and the decoder it reads that body with printed decoded characters to a stream with no `:raw` layer. A character at or below U+00FF then came out as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the accent. `fm_backlog_retain` writes that body straight back through `--body-file`, and nothing reported an error - the character was simply gone from a row still waiting on the captain. The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus `utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already used. Review of the parent commit found this on one of the lines that commit already changed. It predates that change. The test asserts bytes rather than decoded strings, because comparing strings cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any character above U+00FF makes perl print the whole string as UTF-8, so one body carrying both an accent and an em dash passes even unfixed and proves nothing. Verified failing before the fix on the accented row, passing after. Suites: fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures. * no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc * no-mistakes(review): drop whole-file UTF-8 check from retained-body test * no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc * no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc * fix(bin): read codex 0.154's idle braille starfield rows as composer furniture (#4532) * fix(composer): read codex 0.154's idle starfield and status footer as furniture codex-cli 0.154.0 animates a braille "starfield" around its idle composer: on the row above the bold `›` prompt row, on the `›` row behind the SGR-2 dim `Ask Codex to do anything` placeholder, and on the row below it, then draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`). The cells are truecolor greys on both sides of the ghost luminance ceiling, so the brighter ones survive ghost stripping, and the rows below the glyph carry no structural edge. The shared classifier selected the bare `›` shape, extended its wrap region over the two rows beneath the glyph, read the survivors and the footer as wrapped typed input, and answered `pending`; the steering doorbell defers on exactly that verdict, so no doorbell ever reached an idle codex 0.154 pane. bin/fm-composer-lib.sh now recognises that furniture by shape, declared once next to the idle placeholders and reached from the two wrap-region boundary points: - a row whose non-whitespace content is entirely braille cells (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it never counts as wrapped typed content and bounds a bare composer's wrap region; braille behind the glyph row's content is stripped before the emptiness decision when nothing else follows the glyph; a row mixing braille with other text stays typed content; - the codex status footer bounds the wrap region exactly as omp's status row does, anchored on the effort token, a spaced middle dot, and a `~` or `/` path cell, so a typed `fix · tests` stays composer input; - `^Ask Codex to do anything$` joins the verified idle-placeholder set; the ghost strip remains what proves that row empty, and the bare-row rule that bright placeholder text is real input is unchanged. Unchanged: the strict blank-row rule, the styled=0 degradation (a plain cmux/orca capture of this screen still reads `unknown`, never `pending`), FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape. tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte with the divergence (letters in place of the starfield read `pending`) and the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh is the default-on live guard (token-free, skips explicitly without codex or tmux) that launches the installed codex idle and asserts `empty` through both the tmux and the cursorless styled reads, naming codex --version on failure. docs/verification/runtime-backends.md records the dated Herdr evidence: `pending` before, `empty` after, on the captured screen. * no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry --------- Co-authored-by: Todd Billings <todd@usdvcapital.com> * fix(bin): refuse empty text steers in fm-send (#4259) * fix(bin): refuse empty text steers in fm-send A marked secondmate request sent with an empty message delivered only marker and correlation bytes and minted a pending-reply expectation the parent could never see resolved, stalling the fleet with no loud error (#4255). Fail closed on an empty or whitespace-only message on the text path, mirroring the existing --resolve-key refusal. * chore: retain ambient Pi-lens autoformat as its own commit Formatting-only edits produced by ambient Pi-lens autoformat during the msg-loss investigation, kept separate from the behavioural change in c23acba6 so the fix stays reviewable on its own. AGENTS.md is deliberately excluded: its only autoformat edit stripped the trailing space from the documented FM_OPERATIONAL_PREFIX value, which bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11 records as permanent compatibility. Documenting that constant without its trailing space makes the doc wrong about the contract, so that one line was restored rather than retained. * fix(calm): paint the working ship one yellow over all-blue water (#4554) On rose-pine-moon the two-color water (cyan crests over blue troughs) read as a pink stripe over aqua, the yellow left sail and mast clashed with the red right sail, and the hull carried a blue interior run. Every water cell is now blue so the swell reads through glyph height alone, and both sail halves, the mast, and the whole hull are one yellow run. Geometry, cadence, animation, direction flip, resize clamping, and the narrow fallback are unchanged. Update the unit and real-TUI color assertions to the new palette and the Calm docs that described the old one. * fix(bin): stop aging a second mate's active turn from its launch (#4270) * fix(watch): stop aging a second mate's active turn from its launch The parent watcher's second-mate wake-loop stall check exempts a mate that is demonstrably inside an active turn, but secondmate_in_active_turn asked busy_turn_over_age first and returned "not in a turn" whenever that said the bound was crossed. busy_turn_over_age ages from state/<task>.turn-ended, falling back to state/<task>.meta. A second mate's turns end in its own home, so the parent never gets a turn-ended mark for it and the fallback ages the mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago was therefore permanently "over age", the busy pane was never consulted, and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false wake-loop stall. The gate now bounds the busy exemption by <idle> - how long the queue's drain position has not moved - which is evidence this home actually holds. A busy mate stays exempt while the queue has been frozen for less than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so the bound that stops a busy pane from proving liveness forever is kept rather than removed. busy_turn_over_age is untouched; its remaining callers are the ordinary crew busy-pane bound. The regression pins the case that actually broke: a mate whose launch record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn must not escalate, while the same mate with its queue frozen past the bound still publishes exactly one notification. The existing coverage only exercised a freshly launched mate, which passes either way. Reaching that alert now costs a pane capture inside the gate, so the three checkpoints in this suite that assert an alert move from a 1s to a 4s bound - the value the neighbouring active-turn cases already use. The bound is a ceiling, not a wait: the checkpoint returns on the first actionable wake. On a loaded machine a 1s bound missed the alert repeatedly; at 4s it did not miss in 20 runs under the same load. * no-mistakes(review): scope the second-mate active-turn regression test's coverage claim * no-mistakes(document): fix stale second-mate active-turn comments in fm-watch * feat(bin): add read-only PR blocker and reviewer discovery commands (#4278) * feat(bin): add read-only PR blocker and reviewer-discovery commands Two focused, opt-in commands that read GitHub and never write to it. fm-pr-state.sh reports what still blocks one pull request from the author's side: a closed or merged state, draft state, unknown or conflicting mergeability, absent or failing required checks, and a blocking CHANGES_REQUESTED decision explained by each reviewer's latest verdict, marked STALE when it was left at a superseded head. A pull request that only awaits an approval is not reported as blocked, and advisory checks are omitted. Every reading is taken against one exact head; a push that lands mid-read invalidates the whole result rather than mixing two snapshots. fm-pr-reviewers.sh suggests reviewers from the most recent commits to the pull request's exact changed paths, counting each commit once, resolving handles through GitHub's own commit author.login mapping, and excluding the author and Bot accounts. Both stay read-only: no review request, no approval, no merge. Unresolved review-thread state is left unreported because the REST API does not expose it and unattended commands may not use GraphQL. Closes #3731 * no-mistakes(review): accept only PR URLs and stop at terminal state * no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate * no-mistakes(review): stop attributing readings to unverified heads * no-mistakes(review): narrow readiness contract to checks that have reported * no-mistakes(review): read the pull request once, drop the head guard * no-mistakes(document): scope pr-forge isolation proof to its measured members * no-mistakes(document): record uncovered pr-forge members and their pending proof * docs(isolation-proof): re-prove pr-forge at its full membership tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the pr-forge family in this branch, and script_allows_concurrency grants four workers by family membership alone, so both ran concurrently on a proof measured before they existed. Re-proved the family at all eight members: two consecutive runs, 0 failures, each begun with the one-minute load average below 6.0 so the result measures isolation rather than contention. A third run taken between them is disclosed rather than recorded, because it started while the previous run's workers were still decaying. The new durations are not comparable with the six-member measurement above them, so they are not presented as evidence about the two new members, and that record's 1.72x four-worker figure is left as a statement about its own run rather than restated as current. * no-mistakes(review): disclose gh error-text coupling at its matching site and tests * fix(bin): teach validation-round pauses in generated briefs (#2752) * fix(bin): teach validation-round pauses in briefs * no-mistakes(document): Point classifier comments to authoritative pause examples * docs(readme): add star history chart (#4558) * fix(bin): refuse teardown when a task's endpoint close fails (#4510) * fix(teardown): refuse a cleanup whose endpoint close failed bin/fm-teardown.sh discarded both the exit status and the stderr of every fm_backend_kill call, so a close that genuinely failed was indistinguishable from one that succeeded. Teardown continued past it, deleted the task's durable records, returned its worktree, and reported the cleanup as completed. The deleted metadata is the only record of which endpoint belongs to the task, so such a close did not merely leave a stray session behind, it stranded one: nothing was left on disk naming it. The adapters could not carry that signal either. Driven against the real code, every backend arm returned 0 for a genuine failure exactly as it did for an already-exited endpoint, so there was nothing for the four call sites to propagate even once they stopped swallowing it. The tmux arm now resolves a close that did not succeed against the window's exact recorded identity, since kill-window fails the same way for a window that is gone and one that is still there. The Orca arm reports a close its missing CLI never attempted. Both stay silent for an endpoint that is already legitimately gone, and the remaining arms are unchanged: their close-command timing cannot be established without the real Zellij, Orca, and cmux binaries, and a gate that refused ordinary cleanup of an already-exited session would be worse than the defect. docs/verification/runtime-backends.md records what each backend can prove. A reported close failure now reaches teardown's existing retain-and-stop refusal before the records naming the endpoint are removed, matching where the Herdr confirmed-gone gates already sit for the same hazard, and the retained records let a rerun finish once the close works. * no-mistakes(review): refuse unreadable tmux close re-read; honor --force override * no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close * no-mistakes(document): document endpoint-close refusal in its backend and retirement owners * no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree * feat(calm): add flag-gated Claude Code Calm mode (#4565) * feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that brings Calm to Claude Code: the sailboat replaces the stock working row through a Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration, and canonically classified operational user rows draw at zero height. /calm is registered by the hooks module itself and toggles the same per-home config/calm preference the Pi extension uses, so one choice applies on either harness; rows redraw retroactively on toggle and stay hidden across claude --continue. The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS flag is on. Nothing sets that flag in any settings file, and the plugin carries no command file, skill, agent, or classic hook, so it is a complete no-op while the flag is off. The trusted project auto-loads it through an .agents/skills symlink, the only path Claude Code scans for project plugins. Extract the working-ship geometry, bounce track, cadences, and freeze/resume state into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module imports from outside the plugin folder) and have the Pi widget paint that core's frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify operational rows through a port of bin/fm-operational-input.sh's classify command guarded by a corpus parity test against the shell owner. Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering, Raster packing, policy, classifier parity), the mod's own claude plugin test suites behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op, the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272. Docs: record the version-scoped Claude Code evidence and the three bounded gaps in docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md, and make the shared preference, layout, and contributor notes harness-neutral. * no-mistakes(review): Preserve colliding final replies and strengthen parser parity * no-mistakes(review): Preserve final replies and strengthen canonical parity checks * no-mistakes(review): Require exact function-hooks opt-in before Calm activation * no-mistakes(review): Clarify Calm module loading and activation boundaries * no-mistakes(review): Reset Calm presentation state across session starts * no-mistakes(document): Refresh Calm session lifecycle documentation * feat(calm): paint the Claude Code working ship in Claude's own theme colors The captain picked the "Claude native" palette for the Claude Code mod's Raster: every water cell takes the spinner blue of the active theme family (#93a5ff dark, #5769f7 light) and the whole boat takes the Claude orange of the stock spinner (#d77757), one water color and one boat color. The family follows the `theme` setting's prefix, read at load through $.config.list and re-read on a config.set of that row, with `auto` and custom themes falling back to the dark set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte. Rename the shared sprite's color classes from hue names to `water` and `boat`, since each harness now maps them to its own colors; geometry, motion, cadence, and the activation gate are untouched. Tests cover both palettes' packing and the family rule under Node, and the plugin kit drives every theme value, a theme change mid-session, the Calm-off pass-through, and inertness of the menu read while the flag is off. The docs describe the Claude Code colors and record the guard passing on 2.1.273. * no-mistakes(review): Use light palette for unresolved Claude themes * no-mistakes(document): Refresh Claude Calm verification evidence * fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586) * fix(watch): honour a declared wait before wedge-escalating a quiet pane wedge_timer_check escalated on elapsed idle time alone. Nothing asked whether the worker had already said why its pane was quiet, so a lane that declared a bounded external wait climbed the escalation ladder for as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every repeat carried demand-deep-inspection - which by its own wording forbids re-absorbing on the run-step or pane state, so the supervisor could not use the evidence that was there either. The generated brief promises that declaring `paused:` buys the long recheck cadence instead of a wedge, but the timer was still reachable while that declaration stood: a crew that declares a wait and then has an active run or busy pane attributed to it is handed to the timer as provably-working. The declaration is what the worker said about its own silence, so it now outranks a liveness verdict that only says something is running. The consult runs in the at-threshold branch that was about to escalate, beside the worktree walk already there, and costs one status-line read. Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS recheck the declared-wait absorber already uses, so the wait is still rechecked and cannot rot invisibly. Which verb declared it decides the wording, because the two block on different people: a `paused:` wait is owed by an external dependency and asks the reader to confirm it still holds, while a `captain-held:` transfer is owed by the captain reading the recheck and asks them to answer or release the hold. A hold is not rechecked at all while the away-posture record exists, as on every other captain-held path, and that absorb arms no throttle so the recheck is owed in full on return. A declared clearing time that has already passed stops counting, and a lane that never declared one keeps the identical escalation schedule, reason, count and demand-deep-inspection wording, so detection and its worst-case time are unchanged. The deferral restarts the idle timer rather than cancelling it, so a lane that stops waiting escalates again within one threshold. A lane quiet because its own validation run is parked at a gate awaiting a human decision is deliberately out of scope: reading that state needs a signal carrying who the wait is on and what clears it, rather than one inferred from a parked verdict that also covers gates awaiting the crewmate itself. Tests pin both directions for each case and were each confirmed to fail with the consult removed. * no-mistakes(document): docs: honour declared waits in stale-escalation docs * fix(bin): report verified PR state for passed runs (#4624) * fix(bin): derive passed PR state from PR record A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists. For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim. Fixes #4607 * no-mistakes(review): Add bounded GitLab merge-request state reads * no-mistakes(review): Preserve network-free inactive crew-state scans * no-mistakes(document): Document PR record readers in shared library * fix: restore published contribution follow-up (Fixes #4469) (#4627) * fix: restore published contribution follow-up (Fixes #4469) * fix(review): Fix contribution freshness and merge actor routing * fix(review): Restore issue triage and scope contribution follow-up * fix(test): test: assert one wake per contribution signal * fix(document): Document contribution follow-up * fix: restore truthful terminal delivery evidence * fix(review): Disclose unsupported contributions and deduplicate watcher wakes * fix(review): Preserve unmeasured unsupported contributions across Bearings * fix(review): Deduplicate shared contribution wakes and isolate diagnostics * fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower * fix(bin): make remote report transfers explicit and fail-open (#4658) * fix(bin): make a remote-reply document gap self-clearing and re-attemptable A remote mate's undelivered document raised a keyed `blocked` decision that nothing could ever resolve, and any `data/*.md` substring in any mirrored line was an unconditional fetch instruction. A mate announcing a report it had not written yet therefore manufactured a permanent, factually false blocker, and its own explanation of the false alarm manufactured more. The reader has no permanence vocabulary: a report still being written refuses exactly like a path that will never exist. So an undelivered document is now a durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`, re-attempted on the next delta and on the channel's own quiet poll, and retired with a matching `resolved` line naming the local copy once it arrives. The cursor still advances and no delta stalls on one bad pointer. Only a structured `report=data/....md` pointer now offers a document, so a path merely mentioned in prose - including one under another home's mirror tree, which is provably not that mate's to serve - is never fetched. Offers are deduplicated across the whole delta, the escalation names each missing document once and carries the reader's own reason instead of discarding it, and a strictly increasing notice ordinal keeps a later escalation from being swallowed as duplicate bytes. A mirrored line still lands once whichever pointer form it was first written under. * no-mistakes(review): Require structured pointer token boundaries * no-mistakes(review): Unify boundary-safe pointer extraction and rewriting * fix(bin): identify a mirrored line independently of its delivery state Two defects in the boundary-safe pointer work. The at-most-once check compared only the all-remote and all-local renderings of a line, so it could not recognize a mixed one. A line offering two documents where only the first was deliverable mirrored as local-plus-remote; once the second arrived, a cursor-loss whole-log recapture rendered the same line all-local, matched neither alternate, and mirrored a second time. A line's identity is now the canonical form every boundary-valid pointer would take once delivered, derived by the same parser that does extraction and rewriting, so it no longer depends on which documents happened to be deliverable at the time. The pointer map was passed to awk through the process environment. A delta may carry up to the configured 1 MiB bound, and an expanded map of delivered pointers can exceed the platform's exec argument limit, so awk would fail to start; because no caller checked, the empty result would have been appended as blank lines while the cursor advanced past dropped status content. The map now travels in a file, and every call site checks the exit status and stops the ingest rather than committing a delta it could not render. Both passes now run once per stream instead of twice per line. * no-mistakes(review): Abort ingest when document pointer extraction fails * no-mistakes(review): Exclude structured cross-home pointers from document transfer * fix(bin): fail open on an undeliverable remote document instead of tracking it Narrow the remote-reply document fix to the scope the diagnosis actually requires, as decided after measuring a simpler alternative. A document the reader cannot deliver now fails open. The mate's line is mirrored with its own pointer, the cursor advances, and one unkeyed note carries the reader's reason. A note never enters the open-decision fold, so it cannot stand open the way the original keyed block did - which removes the never-clearing false blocker by construction rather than by resolving it. That makes the durable self-clearing obligation unnecessary, so it goes: the per-mate pending-documents record, its notice ordinal and resolved announcements, and the poll-side retry. Canonical line identity goes too, and with it a way to silently drop a genuine status line; mirroring is back to at-most-once on exact bytes. The cross-home exclusion goes as well: under fail-open a cross-home report= either fails harmlessly or is a nested remote report this mate genuinely holds, which is now relayed again. Kept: fetching only on a structured report= pointer, the boundary-correct parser, the file-based rewrite map, and checked extraction and rewrite exit status. The parser now scans behind a sentinel byte so a rejected candidate can no longer give the text right after it a false leading boundary. The reported incident is covered end to end: a report path announced in prose before it exists raises no decision, and the report still arrives through the ledger publisher's structured offer once written. * no-mistakes(review): Preserve source-line identity across remote reply replays * no-mistakes(document): Document remote reply transfer and replay semantics * no-mistakes(lint): Fix staging truncation lint checks * fix(calm): preserve substantive mid-turn responses (#4655) * Preserve substantive Calm mid-turn text * no-mistakes(review): Distinguish newline-preserved replies from short narration * no-mistakes(document): Document Calm mid-turn preservation boundaries * no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check` * fix(bin): preserve PR merge polls across volume remounts (#4656) * fix(bin): re-record PR poll identity after a volume device renumber (Fixes #4260) A volume remount can renumber the state filesystem's st_dev while every inode and byte stays the same; APFS does this across a reboot. A poll registration records its sidecar and check as device:inode, so every poll armed before the remount failed strict validation and the watcher refused all of them as unauthenticated state checks until each was re-armed by hand. There are two device comparisons. fm_pr_private_file_valid compares a live file's device with the state directory's device read in the same invocation: it refuses a file that is not on the state directory's own filesystem and already survives a renumber, so it is unchanged. The registration's recorded identity versus the live identity (from #556, reused by the #932 retirement receipt) binds the registration to the exact files published in its own transaction; its device part is what breaks. When strict capture fails, the watcher now proves the device is the only difference: every other artifact check passes (template bytes, both hashes, private mode, single link, live device, metadata), both recorded identities name one device, and each recorded inode equals its live inode. Only then, under the task's control lock, does it rewrite the two identity lines, repeating the whole proof and comparing the registration's file identity and bytes just before the rename, and then capture strictly again. A swapped, altered, re-moded, relinked, split-device, or foreign-device artifact still fails a proof and is still refused, and a pending retirement receipt blocks the rewrite. Reproduction: on macOS a poll armed on an APFS disk image that was detached and re-attached behind another image moved st_dev 16777239 -> 16777243 with inodes, bytes, mode, and link count unchanged; the real watcher refused it on main and reports its merge with this change. The portable regression test rewrites a real registration's recorded device and drives the watcher. Not changed here: the status presentation cursor keys rows by its own device:inode identity in bin/fm-classify-lib.sh, a different helper that needs its own fix; a retirement receipt left by a reboot between its publication and removal still names the old device and stays refused; custom check trust binds only a content hash and is unaffected. * fix(review): Serialize PR poll publication writers * fix(review): Bound PR poll publication lock scope * fix(bin): keep contribution records when the poll budget runs out (follow-up to #4627) (#4661) A budget that expires partway through an observation no longer records an error or prints the unavailable wake; the URL keeps its prior record and is observed first next poll. forge() flags budget exhaustion at the point it refuses, or when a read is killed at the budget's own deadline, so a genuine forge failure still records the error and wakes. Each distinct URL is now observed once per poll and applied to every owning task. * fix(bin): clear parent pending-replies on local secondmate retirement (#4680) * fix(bin): clear parent pending-replies on local secondmate retirement Local secondmate teardown left resolved parent pending-reply records behind after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced retirement while any reply for that id is still unresolved, and delete every matching record plus its delivery confirmation after a successful local or remote retirement, matching the remote cleanup path. * no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup * no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung * no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern * no-mistakes(review): Pending-replies Basename und corr_id abgleichen * no-mistakes(document): Clarify forced retirement pending-reply cleanup --------- Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc> * fix(bin): accept Orca's composite worktree id when tearing down a task (#4677) * fix(bin): accept Orca's composite worktree id at teardown Teardown refused every Orca-backed task because the endpoint validator checked orca_worktree_id with the simple-atom rule meant for tmux-style window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca returns that id as `<orca id>::<absolute worktree path>`, so the colon and slashes in every real value made validation fail and finished Orca tasks could never be cleaned up. Validate the field as the composite it is: both halves of the first `::` split present, the path half absolute, and no embedded newline, carriage return, or tab. The terminal field keeps the atom check, which is correct for it, and no other backend's validation changes. The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca never returns, which is why the suite passed a check the real value fails. They now carry the composite form, so the tests exercise the real value. * no-mistakes(document): name Orca's repo id in the composite worktree id * no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points * feat(bin): add opt-in typed dispatch resolution (#4692) * feat(bin): add opt-in typed dispatch resolution through typesafe.ai Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or scout profile from a written brief with typesafe.ai's System One model: one Choice question over the rules' `when` texts, then the confidence floor, the rule's `approval` and `floor`, each profile's `provider` and `floor`, one quota-axi snapshot, and the spendPriority argmax all in code. It is off unless TYPESAFE_API_KEY is in the environment or the home's gitignored .env; off means one stderr line, exit 0, and no network call, so firstmate dispatches exactly as before. The key reaches curl on a file descriptor, never argv. Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates the four new optional dispatch fields. Document the schema, the operator contract, the AGENTS.md intake step, and the live and benchmark evidence. * no-mistakes(review): Harden typed dispatch resolution and quota bounds * no-mistakes(review): Validate dispatch floors and ranking evidence * no-mistakes(review): Tighten dispatch response and floor evidence * no-mistakes(review): Neutralize none matching and resolve defaults locally * no-mistakes(review): Preserve providerless profiles outside typed resolution * no-mistakes(review): Validate response usage and reject duplicate profiles * no-mistakes(review): Escalate unverifiable floors and validate probabilities * no-mistakes(review): Validate probability mass and unknown profile floors * no-mistakes(review): Simplify resolver interface and preserve fallback routing * no-mistakes(review): Fix constants and rank partial quota evidence * no-mistakes(review): Add authoritative provider mapping and enforce explicit providers * no-mistakes(review): Declare provider for documented Pi profile * no-mistakes(review): Validate provider identifiers and support Gemini dispatch * no-mistakes(review): Strictly anchor provider identifiers * no-mistakes(review): Validate selectors and preserve fallback candidate evidence * no-mistakes(review): Gate typed validation and harden resolver evidence * no-mistakes(review): Preserve opt-in routing and harden candidate evidence * no-mistakes(review): Prioritize known exhaustion over quota uncertainty * no-mistakes(review): Isolate API secrets and preserve no-key diagnostics * no-mistakes(review): Fallback safely when dispatch rules are absent * no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets * no-mistakes(document): Document typed dispatch safety and fallback behavior * fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753) * test: reproduce buried status declarations in shared readers * fix: share status event reads and preserve open blockers * fix: retain terminal scout and ship status declarations * no-mistakes(review): Fix status chronology, legacy completions, and reader performance * no-mistakes(review): Share terminal decision reconciliation across fleet snapshots * no-mistakes(review): Unify terminal supersession across cached folds and consumers * no-mistakes(review): Filter per-key status history while preserving terminal chronology * no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells * no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses * no-mistakes(document): Document latest-event status read and kind-scoped fold cursor * no-mistakes(lint): Quote literal done in test for-lists for SC1010 * ci: expect 19 snapshot/fleet-view tests This branch adds a fleet-snapshot regression, so the stock macOS Bash lane's hardcoded guard of 18 'ok - ' lines fails on the new count. Bump the guard and its message to 19. * no-mistakes(review): Restore multiline child outcome reporting * no-mistakes(review): Select ledger terminal events through bounded shared reader * no-mistakes(review): Report newest open decision instead of preferring blocked * no-mistakes(review): Require colon before ship/scout terminal supersession in fold * no-mistakes(review): Gate socket-down override on latest event; drop lock matrix * no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions * no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold * no-mistakes(test): Update fleet-view expectations to newest-open-decision rule * no-mistakes(document): Align status-read docs with fold-resolved crew state * no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers * no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree * test: fold terminal-cleanup snapshot coverage into the completed-scout case Keep the ship/scout/secondmate supersession assertions without adding a nineteenth top-level fleet-view test, so CI can stay at the upstream suite count. * no-mistakes(document): Clarify socket-down override expiry in architecture doc * ci: retrigger flaky contribution check * fix(bin): launch codex crewmates with codex's hook layer disabled (#4689) * fix(spawn): launch codex crewmates with codex's hook layer disabled A freshly launched Codex worker never reached its instructions. Codex stopped it on an interactive "Hooks need review" modal whose selection sits on "Review hooks", which is neither trusting nor declining. Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow navigation, so the selection cannot be moved, and pre-accepting the prompt by writing Codex's own trust store would record an operator consent that was never given. The hooks are the machine's own ~/.codex/hooks.json plus any project's .codex/hooks.json. A crewmate needs neither: its turn-end signal is the -c notify= program on the same launch, and Firstmate's project hooks are primary-session infrastructure that stands down in a child worktree. Crewmate and scout launches now pass --disable hooks. That is the opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted hooks; disabling the feature runs none of them and leaves the operator's ~/.codex untouched. An unknown feature name is a hard Codex error, so a release that drops the flag fails the launch loudly instead of silently restoring the modal. A secondmate is a primary in its own home and keeps the project hooks its turn-end guard and session-start digest ride on. Verified on codex-cli 0.151.0: the modal is gone and the turn-end notification still lands. This unblocks the second review that every finished pull request is supposed to get. Fixes kunchenguid/firstmate#4673 * no-mistakes(review): Fix contradictory hook count in Codex verification record * fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710) * fix(bin): settle terminal contributions and wake once per read-failure episode A contribution whose last good observation is merged or closed is final: poll no longer re-reads it, projection keeps it fresh, and a stale error recorded beside it is cleared once. A genuine forge-read failure on an open contribution still records its error on every cycle but prints the unavailable wake only when it starts a failure episode; a successful read ends the episode. Open PRs linked from done tasks keep being observed. The false unavailable beside a complete observation was budget exhaustion mid-observation, already fixed by #4661. * fix(review): Settle terminal contribution owners * fix(review): Deduplicate shared contribution failure episodes * fix(test): Preserve settled terminal contribution records * fix: select authoritative no-mistakes runs (#4476) * fix(crew-state): select authoritative validation runs by identity Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row. Refs: https://github.com/kunchenguid/firstmate/issues/3215 * fix(review): Resolve same-branch run identities beyond capped history * fix(review): Fix run-selection compatibility, races, and worker-state fallbacks * fix(review): Limit run validation to the requested branch * fix(test): Anchor AXI fixtures and document remaining live evidence gaps * fix(document): Clarify run selection documentation and capture ownership * fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation * fix: distinguish captain outcomes from no-op updates (#4738) * fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work MAIN answered a supervision-branch outcome for completed captain-requested work (implementation done, PR ready for review and merge approval) with "Captain, shipshape.", reading section 9's no-action reply as covering it and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no captain-facing response is owed". Section 9 now limits the shipshape reply to true no-ops (idle re-read, empty heartbeat, consequence-free acknowledgement) and requires a short outcome response naming what finished and what word is needed whenever requested work finishes or a result needs the captain's word, even when a transcript entry already shows the substance. The Pi protocol's re-emit rule now says it bounds repetition only, and carries a worked example of the ready-for-review outcome whose correct processing turn a shipshape reply fails. No executable contract evaluates the content of MAIN's captain-facing reply, so the regression is the protocol example in the owner doc rather than a text-match test. * no-mistakes(document): Clarify captain-facing outcomes versus no-ops * docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line The document step condensed the Pi protocol's re-emit rule and dropped the worked example of a finished, ready-for-review outcome whose correct processing turn a "Captain, shipshape." reply fails. That example is the contract's regression: no executable contract evaluates the content of MAIN's captain-facing reply, so the owner doc's example is the test case. Restore it directly under the re-emit rule, prefixed as a regression example that is kept verbatim and never condensed or summarized away. * no-mistakes(review): Clarify captain outcome and decision-word requirements * no-mistakes(document): Clarify captain-facing completion outcomes * docs(pi): require the PR URL in the visible captain-facing outcome reply Captain review on the regression example: drop the sample reply string and say only that the ready-for-review outcome requires relaying a captain-facing outcome response, not just "Captain, shipshape.". Fold in the visible-PR-handoff failure seen this session: after the branch outcome reporting this fix green, MAIN's visible reply was only "Awaiting your merge call." with no PR URL, leaning on the dim anchor. Section 9's URL rule now also covers a review or merge ask and names the visible reply as where the URL goes, sourced from the ready status, pr= metadata, or the supervision branch's summary and never left to a transcript entry. The Pi protocol adds the same-way failure and places the captain-facing text in the final visible assistant reply after the fm_branch_processed call, because Calm hides assistant text emitted in the same step as a tool call as a working note. Investigation verdict, evidence in the PR comment: no recent PR caused the handoff failure; Pi has hidden same-step pre-tool assistant text since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and #4658 touched only remote report transfer. * no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs * no-mistakes(document): Clarify captain-facing supervision outcomes * docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule The consolidated section 9 URL rule narrowed its trigger to a review or merge ask, dropping the "whenever a PR is mentioned" catch-all from #3648 that keeps every PR URL copied from a durable record and never assembled from memory. Restore that trigger as a union with the review or merge ask so the one consolidated rule covers both. * fix(bin): let non-owner Claude Stops exit safely (#4777) * Fix foreign-owner turn-end supervision loop * no-mistakes(review): Scope foreign-owner safe exit to Claude guard * no-mistakes(document): Document Claude foreign-owner safe exit * fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778) Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty indexed array as an unbound variable and aborts the shell. In signal_turnend_panes_churned() the missing_keys loop was reachable with an empty array whenever every churned key already held a fresh .churn-since-* marker (a second churning turn-end inside an open deferral window), so each watcher cycle died about half a minute in and supervision restarted endlessly. The created_keys rollback loops had the same latent crash on their error paths. Audit of bin/ for the same pattern found one more confirmed-reachable case: remote_handoff's noncanonical-body scan iterates to_move, which is empty when a retried remote handoff finds every key already staged in the outbox. All other "${arr[@]}" sites are either count-guarded, guaranteed non-empty by construction, or unreachable while empty. Guard the three reachable expansions with the repo's existing "${arr[@]+...}" idiom. New regression test drives a real watcher through the all-marked churn path; the macos-stock-bash CI lane runs it under real /bin/bash 3.2 via FM_TEST_ONLY. * Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783) The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: require complete captain-facing final responses (#4779) * docs: require complete final responses across harnesses * no-mistakes(document): Document complete final replies for Grok Bot * docs: point Grok replies to the shared contract owner * no-mistakes(review): Clarify final recap without batching decision asks * fix: preserve substantive mid-turn text in Pi Calm (#4788) * fix(calm): preserve substantive Pi mid-turn text * no-mistakes(review): Preserve substantive Pi Calm text per block * no-mistakes(test): Cover shared Calm preservation boundaries behaviorally * no-mistakes(document): Consolidate Calm preservation documentation * fix: harden mail checks and rebalance full-coverage CI (#4800) * Improve CI reliability and rebalance full-coverage validation * no-mistakes(document): Clarify lint partition documentation * fix(bin): answer Kimi 2.0.0 folder-trust dialog during spawn (#4799) * Handle Kimi workspace trust dialog * no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers * no-mistakes(review): Gate Kimi ready on any trust marker and clean captures * no-mistakes(review): Read visible pane for Kimi trust and ready gates * no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate * no-mistakes(review): Harden Kimi viewport capture and trust dialog detection * no-mistakes(document): Document Kimi spawn refusal on cmux and Orca * fix(bin): report a dead-agent record once instead of escalating forever (#4775) * fix(bin): report a record whose agent is gone once instead of escalating forever The wedge escalation path never asked whether there was still an agent to be wedged. A wedge is something stuck that might recover, so re-alarming it earns its cost; an agent that is gone never moves again, its pane never churns, the idle timer never resets, and the escalate path clears its own timer and re-arms with nothing bounding the count. Observed on a live fleet: two finished lanes reached 226 and 203 consecutive escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400 notifications a day from two lanes with no agent running at all. On one, fm-control.sh exit answered already-stopped and fm-crew-state.sh read "failed - run failed". Closing the Herdr pane did not stop it either: with the pane genuinely gone and herdr pane read returning pane_not_found, the count kept climbing, because the poll is driven by the record's window= line rather than by the pane. The cost is not the repetition but that it drowns the alarms that matter. fm_backend_agent_state already separates a thinking agent from a gone one at process level. In t…
doitdigital0495
added a commit
to doitdigital0495/firstmate
that referenced
this pull request
Sep 21, 2026
…ent checks (#20) * fix(bin): treat Claude Code's default external-imports flags as never asked, not declined (#4387) * fix(bin): read Claude Code's default external-imports flags as never asked, not declined (#4378) fm-claude-trust.sh refused the whole trust registration whenever the project-root entry carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code writes that value only on an explicit "No, disable". Claude Code's default project entry carries Approved and WarningShown both false before the dialog is ever shown, so every such project refused every spawn. Only Approved === false with WarningShown === true — the pair the dialog writes on a decline — now counts as a decline. false/false behaves like an absent flag: trust is registered and no import consent is manufactured. New case test_project_root_entry_default_import_flags_are_not_a_decline fails on b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31, bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * no-mistakes(review): Correct harness doc's external-imports decline predicate --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(bin): keep operator-address labels out of no-mistakes intent (#4445) * fix(brief): keep operator address out of composed intent Teach raw-word authoring for intent sections and mid-task relays, with a neutral [captain] provenance marker for legacy mixed tasks. Keep headings and contract prose outside the serialized intent body. The legacy selector already excluded the old speaker labels from its output; preserve that read compatibility. The reproduced leak comes from adding labels inside a modern intent body, not from the legacy selector. Do not scrub actual request content. Add exact serialized-input and generated-contract regressions, retaining refusal of unmarked legacy tasks and coverage of scout promotion. Fixes https://github.com/kunchenguid/firstmate/issues/3882 * no-mistakes(review): Refuse operator-address lines in Captain's intent body * no-mistakes(document): Document operator-address refusal in intent contract comments * fix: classify OpenCode ellipsis hint as idle (#4451) * fix(composer): recognize Grok 1.0.5's oversized titled bottom border as a proven empty composer (#4455) * fix(composer): accept Grok title overhang * no-mistakes(review): summary: named Grok overhang constant, doc caveat, restored tmux typed-title coverage * fix(bin): translate Stop hook timeout signals into durable auto-arm failure (#4474) * fix(bin): recover Claude auto-arm after timeout * no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list * fix(spawn): establish Claude task channel authority (#4464) * fix(spawn): establish Claude task channel authority * no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference * fix(bin): refuse fm-control.sh exit when the composer holds unproven or pending text (#4458) * fix: guard relaunch exit against pending input * no-mistakes(review): Verifying test run in progress * no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard * no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now * fix(spawn): establish crewmate identity first (#4481) * fix(bin): reconcile redundant secondmate divergence during updates (#4460) * fix: reconcile diverged secondmate updates * no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md * no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md * feat: enable gpt-5.6-luna max reasoning for crew dispatch (#4497) * fix(dispatch): support Codex Luna max effort * no-mistakes(review): use portable CODEX_HOME path in codex effort reference * feat(calm): render smooth Unicode swell with asymmetric two-color sail (#4498) * feat(calm): render smooth Unicode swell * feat(calm): make sails asymmetric * feat(calm): use quarter sail glyph * no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer * no-mistakes(document): docs: sync calm wave phase doc comment * no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。 * fix(bin): supersede stale scout delivery text in brief.md on promotion (#4491) * fix: supersede scout delivery brief on promotion * fix: preserve ship safety contract after promotion * no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch * fix(bin): make captain holds work on hosts with an older JSON::PP, and stop cleanup dropping accents from a held body (#4471) * fix(bin): let captain holds work on hosts with an older JSON::PP Holding a task for the captain, and the cleanup that keeps a captain-held row open, both fail outright on any host whose JSON::PP defaults allow_nonref off - 2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`, but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older library rejects that whole value with "must be object or array". The consequence is fleet-wide on such a host, not one broken command: a worker there cannot formally record a decision for the captain at all. It can only mention the decision in passing in a status line, where it can be missed - which is how a real decision goes unrecorded. The hold reports that the task lost its hold-set stamp; the cleanup cannot return the row to Queued. Both call sites now ask for allow_nonref explicitly rather than inheriting whatever the installed library defaults to. The second one is worth naming: its `/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string case that fails, so the guard selects for the failing input rather than protecting against it. The regression case forces the older default back off for every perl the commands spawn, then drives both paths - holding a task that carries a body, and tearing down a captain-held row whose deliverable must still be appended. It also probes that the simulation genuinely rejects a bare scalar, so the case cannot pass vacuously on a lenient host. Each half was verified failing on its own unfixed call site with that site's real error message. Suites: fm-captain-hold-lifecycle 51 cases, fm-backlog-atomicity 99 cases, 0 failures. Verification limit: the mechanism is reproduced and tested, but neither fix is verified against a real JSON::PP 2.27202 host, because none is in the loop. This laptop runs 4.06, where the bug does not manifest. `bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a brace-delimited object before decoding, so allow_nonref never applies. * fix(bin): stop cleanup silently dropping accented characters from a held body Cleanup rewrites a captain-held row's body to append the finished work's deliverable, and the decoder it reads that body with printed decoded characters to a stream with no `:raw` layer. A character at or below U+00FF then came out as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the accent. `fm_backlog_retain` writes that body straight back through `--body-file`, and nothing reported an error - the character was simply gone from a row still waiting on the captain. The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus `utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already used. Review of the parent commit found this on one of the lines that commit already changed. It predates that change. The test asserts bytes rather than decoded strings, because comparing strings cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any character above U+00FF makes perl print the whole string as UTF-8, so one body carrying both an accent and an em dash passes even unfixed and proves nothing. Verified failing before the fix on the accented row, passing after. Suites: fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures. * no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc * no-mistakes(review): drop whole-file UTF-8 check from retained-body test * no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc * no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc * fix(bin): read codex 0.154's idle braille starfield rows as composer furniture (#4532) * fix(composer): read codex 0.154's idle starfield and status footer as furniture codex-cli 0.154.0 animates a braille "starfield" around its idle composer: on the row above the bold `›` prompt row, on the `›` row behind the SGR-2 dim `Ask Codex to do anything` placeholder, and on the row below it, then draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`). The cells are truecolor greys on both sides of the ghost luminance ceiling, so the brighter ones survive ghost stripping, and the rows below the glyph carry no structural edge. The shared classifier selected the bare `›` shape, extended its wrap region over the two rows beneath the glyph, read the survivors and the footer as wrapped typed input, and answered `pending`; the steering doorbell defers on exactly that verdict, so no doorbell ever reached an idle codex 0.154 pane. bin/fm-composer-lib.sh now recognises that furniture by shape, declared once next to the idle placeholders and reached from the two wrap-region boundary points: - a row whose non-whitespace content is entirely braille cells (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it never counts as wrapped typed content and bounds a bare composer's wrap region; braille behind the glyph row's content is stripped before the emptiness decision when nothing else follows the glyph; a row mixing braille with other text stays typed content; - the codex status footer bounds the wrap region exactly as omp's status row does, anchored on the effort token, a spaced middle dot, and a `~` or `/` path cell, so a typed `fix · tests` stays composer input; - `^Ask Codex to do anything$` joins the verified idle-placeholder set; the ghost strip remains what proves that row empty, and the bare-row rule that bright placeholder text is real input is unchanged. Unchanged: the strict blank-row rule, the styled=0 degradation (a plain cmux/orca capture of this screen still reads `unknown`, never `pending`), FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape. tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte with the divergence (letters in place of the starfield read `pending`) and the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh is the default-on live guard (token-free, skips explicitly without codex or tmux) that launches the installed codex idle and asserts `empty` through both the tmux and the cursorless styled reads, naming codex --version on failure. docs/verification/runtime-backends.md records the dated Herdr evidence: `pending` before, `empty` after, on the captured screen. * no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry --------- Co-authored-by: Todd Billings <todd@usdvcapital.com> * fix(bin): refuse empty text steers in fm-send (#4259) * fix(bin): refuse empty text steers in fm-send A marked secondmate request sent with an empty message delivered only marker and correlation bytes and minted a pending-reply expectation the parent could never see resolved, stalling the fleet with no loud error (#4255). Fail closed on an empty or whitespace-only message on the text path, mirroring the existing --resolve-key refusal. * chore: retain ambient Pi-lens autoformat as its own commit Formatting-only edits produced by ambient Pi-lens autoformat during the msg-loss investigation, kept separate from the behavioural change in c23acba6 so the fix stays reviewable on its own. AGENTS.md is deliberately excluded: its only autoformat edit stripped the trailing space from the documented FM_OPERATIONAL_PREFIX value, which bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11 records as permanent compatibility. Documenting that constant without its trailing space makes the doc wrong about the contract, so that one line was restored rather than retained. * fix(calm): paint the working ship one yellow over all-blue water (#4554) On rose-pine-moon the two-color water (cyan crests over blue troughs) read as a pink stripe over aqua, the yellow left sail and mast clashed with the red right sail, and the hull carried a blue interior run. Every water cell is now blue so the swell reads through glyph height alone, and both sail halves, the mast, and the whole hull are one yellow run. Geometry, cadence, animation, direction flip, resize clamping, and the narrow fallback are unchanged. Update the unit and real-TUI color assertions to the new palette and the Calm docs that described the old one. * fix(bin): stop aging a second mate's active turn from its launch (#4270) * fix(watch): stop aging a second mate's active turn from its launch The parent watcher's second-mate wake-loop stall check exempts a mate that is demonstrably inside an active turn, but secondmate_in_active_turn asked busy_turn_over_age first and returned "not in a turn" whenever that said the bound was crossed. busy_turn_over_age ages from state/<task>.turn-ended, falling back to state/<task>.meta. A second mate's turns end in its own home, so the parent never gets a turn-ended mark for it and the fallback ages the mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago was therefore permanently "over age", the busy pane was never consulted, and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false wake-loop stall. The gate now bounds the busy exemption by <idle> - how long the queue's drain position has not moved - which is evidence this home actually holds. A busy mate stays exempt while the queue has been frozen for less than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so the bound that stops a busy pane from proving liveness forever is kept rather than removed. busy_turn_over_age is untouched; its remaining callers are the ordinary crew busy-pane bound. The regression pins the case that actually broke: a mate whose launch record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn must not escalate, while the same mate with its queue frozen past the bound still publishes exactly one notification. The existing coverage only exercised a freshly launched mate, which passes either way. Reaching that alert now costs a pane capture inside the gate, so the three checkpoints in this suite that assert an alert move from a 1s to a 4s bound - the value the neighbouring active-turn cases already use. The bound is a ceiling, not a wait: the checkpoint returns on the first actionable wake. On a loaded machine a 1s bound missed the alert repeatedly; at 4s it did not miss in 20 runs under the same load. * no-mistakes(review): scope the second-mate active-turn regression test's coverage claim * no-mistakes(document): fix stale second-mate active-turn comments in fm-watch * feat(bin): add read-only PR blocker and reviewer discovery commands (#4278) * feat(bin): add read-only PR blocker and reviewer-discovery commands Two focused, opt-in commands that read GitHub and never write to it. fm-pr-state.sh reports what still blocks one pull request from the author's side: a closed or merged state, draft state, unknown or conflicting mergeability, absent or failing required checks, and a blocking CHANGES_REQUESTED decision explained by each reviewer's latest verdict, marked STALE when it was left at a superseded head. A pull request that only awaits an approval is not reported as blocked, and advisory checks are omitted. Every reading is taken against one exact head; a push that lands mid-read invalidates the whole result rather than mixing two snapshots. fm-pr-reviewers.sh suggests reviewers from the most recent commits to the pull request's exact changed paths, counting each commit once, resolving handles through GitHub's own commit author.login mapping, and excluding the author and Bot accounts. Both stay read-only: no review request, no approval, no merge. Unresolved review-thread state is left unreported because the REST API does not expose it and unattended commands may not use GraphQL. Closes #3731 * no-mistakes(review): accept only PR URLs and stop at terminal state * no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate * no-mistakes(review): stop attributing readings to unverified heads * no-mistakes(review): narrow readiness contract to checks that have reported * no-mistakes(review): read the pull request once, drop the head guard * no-mistakes(document): scope pr-forge isolation proof to its measured members * no-mistakes(document): record uncovered pr-forge members and their pending proof * docs(isolation-proof): re-prove pr-forge at its full membership tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the pr-forge family in this branch, and script_allows_concurrency grants four workers by family membership alone, so both ran concurrently on a proof measured before they existed. Re-proved the family at all eight members: two consecutive runs, 0 failures, each begun with the one-minute load average below 6.0 so the result measures isolation rather than contention. A third run taken between them is disclosed rather than recorded, because it started while the previous run's workers were still decaying. The new durations are not comparable with the six-member measurement above them, so they are not presented as evidence about the two new members, and that record's 1.72x four-worker figure is left as a statement about its own run rather than restated as current. * no-mistakes(review): disclose gh error-text coupling at its matching site and tests * fix(bin): teach validation-round pauses in generated briefs (#2752) * fix(bin): teach validation-round pauses in briefs * no-mistakes(document): Point classifier comments to authoritative pause examples * docs(readme): add star history chart (#4558) * fix(bin): refuse teardown when a task's endpoint close fails (#4510) * fix(teardown): refuse a cleanup whose endpoint close failed bin/fm-teardown.sh discarded both the exit status and the stderr of every fm_backend_kill call, so a close that genuinely failed was indistinguishable from one that succeeded. Teardown continued past it, deleted the task's durable records, returned its worktree, and reported the cleanup as completed. The deleted metadata is the only record of which endpoint belongs to the task, so such a close did not merely leave a stray session behind, it stranded one: nothing was left on disk naming it. The adapters could not carry that signal either. Driven against the real code, every backend arm returned 0 for a genuine failure exactly as it did for an already-exited endpoint, so there was nothing for the four call sites to propagate even once they stopped swallowing it. The tmux arm now resolves a close that did not succeed against the window's exact recorded identity, since kill-window fails the same way for a window that is gone and one that is still there. The Orca arm reports a close its missing CLI never attempted. Both stay silent for an endpoint that is already legitimately gone, and the remaining arms are unchanged: their close-command timing cannot be established without the real Zellij, Orca, and cmux binaries, and a gate that refused ordinary cleanup of an already-exited session would be worse than the defect. docs/verification/runtime-backends.md records what each backend can prove. A reported close failure now reaches teardown's existing retain-and-stop refusal before the records naming the endpoint are removed, matching where the Herdr confirmed-gone gates already sit for the same hazard, and the retained records let a rerun finish once the close works. * no-mistakes(review): refuse unreadable tmux close re-read; honor --force override * no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close * no-mistakes(document): document endpoint-close refusal in its backend and retirement owners * no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree * feat(calm): add flag-gated Claude Code Calm mode (#4565) * feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that brings Calm to Claude Code: the sailboat replaces the stock working row through a Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration, and canonically classified operational user rows draw at zero height. /calm is registered by the hooks module itself and toggles the same per-home config/calm preference the Pi extension uses, so one choice applies on either harness; rows redraw retroactively on toggle and stay hidden across claude --continue. The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS flag is on. Nothing sets that flag in any settings file, and the plugin carries no command file, skill, agent, or classic hook, so it is a complete no-op while the flag is off. The trusted project auto-loads it through an .agents/skills symlink, the only path Claude Code scans for project plugins. Extract the working-ship geometry, bounce track, cadences, and freeze/resume state into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module imports from outside the plugin folder) and have the Pi widget paint that core's frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify operational rows through a port of bin/fm-operational-input.sh's classify command guarded by a corpus parity test against the shell owner. Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering, Raster packing, policy, classifier parity), the mod's own claude plugin test suites behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op, the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272. Docs: record the version-scoped Claude Code evidence and the three bounded gaps in docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md, and make the shared preference, layout, and contributor notes harness-neutral. * no-mistakes(review): Preserve colliding final replies and strengthen parser parity * no-mistakes(review): Preserve final replies and strengthen canonical parity checks * no-mistakes(review): Require exact function-hooks opt-in before Calm activation * no-mistakes(review): Clarify Calm module loading and activation boundaries * no-mistakes(review): Reset Calm presentation state across session starts * no-mistakes(document): Refresh Calm session lifecycle documentation * feat(calm): paint the Claude Code working ship in Claude's own theme colors The captain picked the "Claude native" palette for the Claude Code mod's Raster: every water cell takes the spinner blue of the active theme family (#93a5ff dark, #5769f7 light) and the whole boat takes the Claude orange of the stock spinner (#d77757), one water color and one boat color. The family follows the `theme` setting's prefix, read at load through $.config.list and re-read on a config.set of that row, with `auto` and custom themes falling back to the dark set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte. Rename the shared sprite's color classes from hue names to `water` and `boat`, since each harness now maps them to its own colors; geometry, motion, cadence, and the activation gate are untouched. Tests cover both palettes' packing and the family rule under Node, and the plugin kit drives every theme value, a theme change mid-session, the Calm-off pass-through, and inertness of the menu read while the flag is off. The docs describe the Claude Code colors and record the guard passing on 2.1.273. * no-mistakes(review): Use light palette for unresolved Claude themes * no-mistakes(document): Refresh Claude Calm verification evidence * fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586) * fix(watch): honour a declared wait before wedge-escalating a quiet pane wedge_timer_check escalated on elapsed idle time alone. Nothing asked whether the worker had already said why its pane was quiet, so a lane that declared a bounded external wait climbed the escalation ladder for as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every repeat carried demand-deep-inspection - which by its own wording forbids re-absorbing on the run-step or pane state, so the supervisor could not use the evidence that was there either. The generated brief promises that declaring `paused:` buys the long recheck cadence instead of a wedge, but the timer was still reachable while that declaration stood: a crew that declares a wait and then has an active run or busy pane attributed to it is handed to the timer as provably-working. The declaration is what the worker said about its own silence, so it now outranks a liveness verdict that only says something is running. The consult runs in the at-threshold branch that was about to escalate, beside the worktree walk already there, and costs one status-line read. Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS recheck the declared-wait absorber already uses, so the wait is still rechecked and cannot rot invisibly. Which verb declared it decides the wording, because the two block on different people: a `paused:` wait is owed by an external dependency and asks the reader to confirm it still holds, while a `captain-held:` transfer is owed by the captain reading the recheck and asks them to answer or release the hold. A hold is not rechecked at all while the away-posture record exists, as on every other captain-held path, and that absorb arms no throttle so the recheck is owed in full on return. A declared clearing time that has already passed stops counting, and a lane that never declared one keeps the identical escalation schedule, reason, count and demand-deep-inspection wording, so detection and its worst-case time are unchanged. The deferral restarts the idle timer rather than cancelling it, so a lane that stops waiting escalates again within one threshold. A lane quiet because its own validation run is parked at a gate awaiting a human decision is deliberately out of scope: reading that state needs a signal carrying who the wait is on and what clears it, rather than one inferred from a parked verdict that also covers gates awaiting the crewmate itself. Tests pin both directions for each case and were each confirmed to fail with the consult removed. * no-mistakes(document): docs: honour declared waits in stale-escalation docs * fix(bin): report verified PR state for passed runs (#4624) * fix(bin): derive passed PR state from PR record A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists. For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim. Fixes #4607 * no-mistakes(review): Add bounded GitLab merge-request state reads * no-mistakes(review): Preserve network-free inactive crew-state scans * no-mistakes(document): Document PR record readers in shared library * fix: restore published contribution follow-up (Fixes #4469) (#4627) * fix: restore published contribution follow-up (Fixes #4469) * fix(review): Fix contribution freshness and merge actor routing * fix(review): Restore issue triage and scope contribution follow-up * fix(test): test: assert one wake per contribution signal * fix(document): Document contribution follow-up * fix: restore truthful terminal delivery evidence * fix(review): Disclose unsupported contributions and deduplicate watcher wakes * fix(review): Preserve unmeasured unsupported contributions across Bearings * fix(review): Deduplicate shared contribution wakes and isolate diagnostics * fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower * fix(bin): make remote report transfers explicit and fail-open (#4658) * fix(bin): make a remote-reply document gap self-clearing and re-attemptable A remote mate's undelivered document raised a keyed `blocked` decision that nothing could ever resolve, and any `data/*.md` substring in any mirrored line was an unconditional fetch instruction. A mate announcing a report it had not written yet therefore manufactured a permanent, factually false blocker, and its own explanation of the false alarm manufactured more. The reader has no permanence vocabulary: a report still being written refuses exactly like a path that will never exist. So an undelivered document is now a durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`, re-attempted on the next delta and on the channel's own quiet poll, and retired with a matching `resolved` line naming the local copy once it arrives. The cursor still advances and no delta stalls on one bad pointer. Only a structured `report=data/....md` pointer now offers a document, so a path merely mentioned in prose - including one under another home's mirror tree, which is provably not that mate's to serve - is never fetched. Offers are deduplicated across the whole delta, the escalation names each missing document once and carries the reader's own reason instead of discarding it, and a strictly increasing notice ordinal keeps a later escalation from being swallowed as duplicate bytes. A mirrored line still lands once whichever pointer form it was first written under. * no-mistakes(review): Require structured pointer token boundaries * no-mistakes(review): Unify boundary-safe pointer extraction and rewriting * fix(bin): identify a mirrored line independently of its delivery state Two defects in the boundary-safe pointer work. The at-most-once check compared only the all-remote and all-local renderings of a line, so it could not recognize a mixed one. A line offering two documents where only the first was deliverable mirrored as local-plus-remote; once the second arrived, a cursor-loss whole-log recapture rendered the same line all-local, matched neither alternate, and mirrored a second time. A line's identity is now the canonical form every boundary-valid pointer would take once delivered, derived by the same parser that does extraction and rewriting, so it no longer depends on which documents happened to be deliverable at the time. The pointer map was passed to awk through the process environment. A delta may carry up to the configured 1 MiB bound, and an expanded map of delivered pointers can exceed the platform's exec argument limit, so awk would fail to start; because no caller checked, the empty result would have been appended as blank lines while the cursor advanced past dropped status content. The map now travels in a file, and every call site checks the exit status and stops the ingest rather than committing a delta it could not render. Both passes now run once per stream instead of twice per line. * no-mistakes(review): Abort ingest when document pointer extraction fails * no-mistakes(review): Exclude structured cross-home pointers from document transfer * fix(bin): fail open on an undeliverable remote document instead of tracking it Narrow the remote-reply document fix to the scope the diagnosis actually requires, as decided after measuring a simpler alternative. A document the reader cannot deliver now fails open. The mate's line is mirrored with its own pointer, the cursor advances, and one unkeyed note carries the reader's reason. A note never enters the open-decision fold, so it cannot stand open the way the original keyed block did - which removes the never-clearing false blocker by construction rather than by resolving it. That makes the durable self-clearing obligation unnecessary, so it goes: the per-mate pending-documents record, its notice ordinal and resolved announcements, and the poll-side retry. Canonical line identity goes too, and with it a way to silently drop a genuine status line; mirroring is back to at-most-once on exact bytes. The cross-home exclusion goes as well: under fail-open a cross-home report= either fails harmlessly or is a nested remote report this mate genuinely holds, which is now relayed again. Kept: fetching only on a structured report= pointer, the boundary-correct parser, the file-based rewrite map, and checked extraction and rewrite exit status. The parser now scans behind a sentinel byte so a rejected candidate can no longer give the text right after it a false leading boundary. The reported incident is covered end to end: a report path announced in prose before it exists raises no decision, and the report still arrives through the ledger publisher's structured offer once written. * no-mistakes(review): Preserve source-line identity across remote reply replays * no-mistakes(document): Document remote reply transfer and replay semantics * no-mistakes(lint): Fix staging truncation lint checks * fix(calm): preserve substantive mid-turn responses (#4655) * Preserve substantive Calm mid-turn text * no-mistakes(review): Distinguish newline-preserved replies from short narration * no-mistakes(document): Document Calm mid-turn preservation boundaries * no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check` * fix(bin): preserve PR merge polls across volume remounts (#4656) * fix(bin): re-record PR poll identity after a volume device renumber (Fixes #4260) A volume remount can renumber the state filesystem's st_dev while every inode and byte stays the same; APFS does this across a reboot. A poll registration records its sidecar and check as device:inode, so every poll armed before the remount failed strict validation and the watcher refused all of them as unauthenticated state checks until each was re-armed by hand. There are two device comparisons. fm_pr_private_file_valid compares a live file's device with the state directory's device read in the same invocation: it refuses a file that is not on the state directory's own filesystem and already survives a renumber, so it is unchanged. The registration's recorded identity versus the live identity (from #556, reused by the #932 retirement receipt) binds the registration to the exact files published in its own transaction; its device part is what breaks. When strict capture fails, the watcher now proves the device is the only difference: every other artifact check passes (template bytes, both hashes, private mode, single link, live device, metadata), both recorded identities name one device, and each recorded inode equals its live inode. Only then, under the task's control lock, does it rewrite the two identity lines, repeating the whole proof and comparing the registration's file identity and bytes just before the rename, and then capture strictly again. A swapped, altered, re-moded, relinked, split-device, or foreign-device artifact still fails a proof and is still refused, and a pending retirement receipt blocks the rewrite. Reproduction: on macOS a poll armed on an APFS disk image that was detached and re-attached behind another image moved st_dev 16777239 -> 16777243 with inodes, bytes, mode, and link count unchanged; the real watcher refused it on main and reports its merge with this change. The portable regression test rewrites a real registration's recorded device and drives the watcher. Not changed here: the status presentation cursor keys rows by its own device:inode identity in bin/fm-classify-lib.sh, a different helper that needs its own fix; a retirement receipt left by a reboot between its publication and removal still names the old device and stays refused; custom check trust binds only a content hash and is unaffected. * fix(review): Serialize PR poll publication writers * fix(review): Bound PR poll publication lock scope * fix(bin): keep contribution records when the poll budget runs out (follow-up to #4627) (#4661) A budget that expires partway through an observation no longer records an error or prints the unavailable wake; the URL keeps its prior record and is observed first next poll. forge() flags budget exhaustion at the point it refuses, or when a read is killed at the budget's own deadline, so a genuine forge failure still records the error and wakes. Each distinct URL is now observed once per poll and applied to every owning task. * fix(bin): clear parent pending-replies on local secondmate retirement (#4680) * fix(bin): clear parent pending-replies on local secondmate retirement Local secondmate teardown left resolved parent pending-reply records behind after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced retirement while any reply for that id is still unresolved, and delete every matching record plus its delivery confirmation after a successful local or remote retirement, matching the remote cleanup path. * no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup * no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung * no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern * no-mistakes(review): Pending-replies Basename und corr_id abgleichen * no-mistakes(document): Clarify forced retirement pending-reply cleanup --------- Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc> * fix(bin): accept Orca's composite worktree id when tearing down a task (#4677) * fix(bin): accept Orca's composite worktree id at teardown Teardown refused every Orca-backed task because the endpoint validator checked orca_worktree_id with the simple-atom rule meant for tmux-style window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca returns that id as `<orca id>::<absolute worktree path>`, so the colon and slashes in every real value made validation fail and finished Orca tasks could never be cleaned up. Validate the field as the composite it is: both halves of the first `::` split present, the path half absolute, and no embedded newline, carriage return, or tab. The terminal field keeps the atom check, which is correct for it, and no other backend's validation changes. The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca never returns, which is why the suite passed a check the real value fails. They now carry the composite form, so the tests exercise the real value. * no-mistakes(document): name Orca's repo id in the composite worktree id * no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points * feat(bin): add opt-in typed dispatch resolution (#4692) * feat(bin): add opt-in typed dispatch resolution through typesafe.ai Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or scout profile from a written brief with typesafe.ai's System One model: one Choice question over the rules' `when` texts, then the confidence floor, the rule's `approval` and `floor`, each profile's `provider` and `floor`, one quota-axi snapshot, and the spendPriority argmax all in code. It is off unless TYPESAFE_API_KEY is in the environment or the home's gitignored .env; off means one stderr line, exit 0, and no network call, so firstmate dispatches exactly as before. The key reaches curl on a file descriptor, never argv. Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates the four new optional dispatch fields. Document the schema, the operator contract, the AGENTS.md intake step, and the live and benchmark evidence. * no-mistakes(review): Harden typed dispatch resolution and quota bounds * no-mistakes(review): Validate dispatch floors and ranking evidence * no-mistakes(review): Tighten dispatch response and floor evidence * no-mistakes(review): Neutralize none matching and resolve defaults locally * no-mistakes(review): Preserve providerless profiles outside typed resolution * no-mistakes(review): Validate response usage and reject duplicate profiles * no-mistakes(review): Escalate unverifiable floors and validate probabilities * no-mistakes(review): Validate probability mass and unknown profile floors * no-mistakes(review): Simplify resolver interface and preserve fallback routing * no-mistakes(review): Fix constants and rank partial quota evidence * no-mistakes(review): Add authoritative provider mapping and enforce explicit providers * no-mistakes(review): Declare provider for documented Pi profile * no-mistakes(review): Validate provider identifiers and support Gemini dispatch * no-mistakes(review): Strictly anchor provider identifiers * no-mistakes(review): Validate selectors and preserve fallback candidate evidence * no-mistakes(review): Gate typed validation and harden resolver evidence * no-mistakes(review): Preserve opt-in routing and harden candidate evidence * no-mistakes(review): Prioritize known exhaustion over quota uncertainty * no-mistakes(review): Isolate API secrets and preserve no-key diagnostics * no-mistakes(review): Fallback safely when dispatch rules are absent * no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets * no-mistakes(document): Document typed dispatch safety and fallback behavior * fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753) * test: reproduce buried status declarations in shared readers * fix: share status event reads and preserve open blockers * fix: retain terminal scout and ship status declarations * no-mistakes(review): Fix status chronology, legacy completions, and reader performance * no-mistakes(review): Share terminal decision reconciliation across fleet snapshots * no-mistakes(review): Unify terminal supersession across cached folds and consumers * no-mistakes(review): Filter per-key status history while preserving terminal chronology * no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells * no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses * no-mistakes(document): Document latest-event status read and kind-scoped fold cursor * no-mistakes(lint): Quote literal done in test for-lists for SC1010 * ci: expect 19 snapshot/fleet-view tests This branch adds a fleet-snapshot regression, so the stock macOS Bash lane's hardcoded guard of 18 'ok - ' lines fails on the new count. Bump the guard and its message to 19. * no-mistakes(review): Restore multiline child outcome reporting * no-mistakes(review): Select ledger terminal events through bounded shared reader * no-mistakes(review): Report newest open decision instead of preferring blocked * no-mistakes(review): Require colon before ship/scout terminal supersession in fold * no-mistakes(review): Gate socket-down override on latest event; drop lock matrix * no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions * no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold * no-mistakes(test): Update fleet-view expectations to newest-open-decision rule * no-mistakes(document): Align status-read docs with fold-resolved crew state * no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers * no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree * test: fold terminal-cleanup snapshot coverage into the completed-scout case Keep the ship/scout/secondmate supersession assertions without adding a nineteenth top-level fleet-view test, so CI can stay at the upstream suite count. * no-mistakes(document): Clarify socket-down override expiry in architecture doc * ci: retrigger flaky contribution check * fix(bin): launch codex crewmates with codex's hook layer disabled (#4689) * fix(spawn): launch codex crewmates with codex's hook layer disabled A freshly launched Codex worker never reached its instructions. Codex stopped it on an interactive "Hooks need review" modal whose selection sits on "Review hooks", which is neither trusting nor declining. Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow navigation, so the selection cannot be moved, and pre-accepting the prompt by writing Codex's own trust store would record an operator consent that was never given. The hooks are the machine's own ~/.codex/hooks.json plus any project's .codex/hooks.json. A crewmate needs neither: its turn-end signal is the -c notify= program on the same launch, and Firstmate's project hooks are primary-session infrastructure that stands down in a child worktree. Crewmate and scout launches now pass --disable hooks. That is the opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted hooks; disabling the feature runs none of them and leaves the operator's ~/.codex untouched. An unknown feature name is a hard Codex error, so a release that drops the flag fails the launch loudly instead of silently restoring the modal. A secondmate is a primary in its own home and keeps the project hooks its turn-end guard and session-start digest ride on. Verified on codex-cli 0.151.0: the modal is gone and the turn-end notification still lands. This unblocks the second review that every finished pull request is supposed to get. Fixes kunchenguid/firstmate#4673 * no-mistakes(review): Fix contradictory hook count in Codex verification record * fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710) * fix(bin): settle terminal contributions and wake once per read-failure episode A contribution whose last good observation is merged or closed is final: poll no longer re-reads it, projection keeps it fresh, and a stale error recorded beside it is cleared once. A genuine forge-read failure on an open contribution still records its error on every cycle but prints the unavailable wake only when it starts a failure episode; a successful read ends the episode. Open PRs linked from done tasks keep being observed. The false unavailable beside a complete observation was budget exhaustion mid-observation, already fixed by #4661. * fix(review): Settle terminal contribution owners * fix(review): Deduplicate shared contribution failure episodes * fix(test): Preserve settled terminal contribution records * fix: select authoritative no-mistakes runs (#4476) * fix(crew-state): select authoritative validation runs by identity Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row. Refs: https://github.com/kunchenguid/firstmate/issues/3215 * fix(review): Resolve same-branch run identities beyond capped history * fix(review): Fix run-selection compatibility, races, and worker-state fallbacks * fix(review): Limit run validation to the requested branch * fix(test): Anchor AXI fixtures and document remaining live evidence gaps * fix(document): Clarify run selection documentation and capture ownership * fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation * fix: distinguish captain outcomes from no-op updates (#4738) * fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work MAIN answered a supervision-branch outcome for completed captain-requested work (implementation done, PR ready for review and merge approval) with "Captain, shipshape.", reading section 9's no-action reply as covering it and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no captain-facing response is owed". Section 9 now limits the shipshape reply to true no-ops (idle re-read, empty heartbeat, consequence-free acknowledgement) and requires a short outcome response naming what finished and what word is needed whenever requested work finishes or a result needs the captain's word, even when a transcript entry already shows the substance. The Pi protocol's re-emit rule now says it bounds repetition only, and carries a worked example of the ready-for-review outcome whose correct processing turn a shipshape reply fails. No executable contract evaluates the content of MAIN's captain-facing reply, so the regression is the protocol example in the owner doc rather than a text-match test. * no-mistakes(document): Clarify captain-facing outcomes versus no-ops * docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line The document step condensed the Pi protocol's re-emit rule and dropped the worked example of a finished, ready-for-review outcome whose correct processing turn a "Captain, shipshape." reply fails. That example is the contract's regression: no executable contract evaluates the content of MAIN's captain-facing reply, so the owner doc's example is the test case. Restore it directly under the re-emit rule, prefixed as a regression example that is kept verbatim and never condensed or summarized away. * no-mistakes(review): Clarify captain outcome and decision-word requirements * no-mistakes(document): Clarify captain-facing completion outcomes * docs(pi): require the PR URL in the visible captain-facing outcome reply Captain review on the regression example: drop the sample reply string and say only that the ready-for-review outcome requires relaying a captain-facing outcome response, not just "Captain, shipshape.". Fold in the visible-PR-handoff failure seen this session: after the branch outcome reporting this fix green, MAIN's visible reply was only "Awaiting your merge call." with no PR URL, leaning on the dim anchor. Section 9's URL rule now also covers a review or merge ask and names the visible reply as where the URL goes, sourced from the ready status, pr= metadata, or the supervision branch's summary and never left to a transcript entry. The Pi protocol adds the same-way failure and places the captain-facing text in the final visible assistant reply after the fm_branch_processed call, because Calm hides assistant text emitted in the same step as a tool call as a working note. Investigation verdict, evidence in the PR comment: no recent PR caused the handoff failure; Pi has hidden same-step pre-tool assistant text since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and #4658 touched only remote report transfer. * no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs * no-mistakes(document): Clarify captain-facing supervision outcomes * docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule The consolidated section 9 URL rule narrowed its trigger to a review or merge ask, dropping the "whenever a PR is mentioned" catch-all from #3648 that keeps every PR URL copied from a durable record and never assembled from memory. Restore that trigger as a union with the review or merge ask so the one consolidated rule covers both. * fix(bin): let non-owner Claude Stops exit safely (#4777) * Fix foreign-owner turn-end supervision loop * no-mistakes(review): Scope foreign-owner safe exit to Claude guard * no-mistakes(document): Document Claude foreign-owner safe exit * fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778) Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty indexed array as an unbound variable and aborts the shell. In signal_turnend_panes_churned() the missing_keys loop was reachable with an empty array whenever every churned key already held a fresh .churn-since-* marker (a second churning turn-end inside an open deferral window), so each watcher cycle died about half a minute in and supervision restarted endlessly. The created_keys rollback loops had the same latent crash on their error paths. Audit of bin/ for the same pattern found one more confirmed-reachable case: remote_handoff's noncanonical-body scan iterates to_move, which is empty when a retried remote handoff finds every key already staged in the outbox. All other "${arr[@]}" sites are either count-guarded, guaranteed non-empty by construction, or unreachable while empty. Guard the three reachable expansions with the repo's existing "${arr[@]+...}" idiom. New regression test drives a real watcher through the all-marked churn path; the macos-stock-bash CI lane runs it under real /bin/bash 3.2 via FM_TEST_ONLY. * Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783) The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: require complete captain-facing final responses (#4779) * docs: require complete final responses across harnesses * no-mistakes(document): Document complete final replies for Grok Bot * docs: point Grok replies to the shared contract owner * no-mistakes(review): Clarify final recap without batching decision asks * fix: preserve substantive mid-turn text in Pi Calm (#4788) * fix(calm): preserve substantive Pi mid-turn text * no-mistakes(review): Preserve substantive Pi Calm text per block * no-mistakes(test): Cover shared Calm preservation boundaries behaviorally * no-mistakes(document): Consolidate Calm preservation documentation * fix: harden mail checks and rebalance full-coverage CI (#4800) * Improve CI reliability and rebalance full-coverage validation * no-mistakes(document): Clarify lint partition documentation * fix(bin): answer Kimi 2.0.0 folder-trust dialog during spawn (#4799) * Handle Kimi workspace trust dialog * no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers * no-mistakes(review): Gate Kimi ready on any trust marker and clean captures * no-mistakes(review): Read visible pane for Kimi trust and ready gates * no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate * no-mistakes(review): Harden Kimi viewport capture and trust dialog detection * no-mistakes(document): Document Kimi spawn refusal on cmux and Orca * fix(bin): report a dead-agent record once instead of escalating forever (#4775) * fix(bin): report a record whose agent is gone once instead of escalating forever The wedge escalation path never asked whether there was still an agent to be wedged. A wedge is something stuck that might recover, so re-alarming it earns its cost; an agent that is gone never moves again, its pane never churns, the idle timer never resets, and the escalate path clears its own timer and re-arms with nothing bounding the count. Observed on a live fleet: two finished lanes reached 226 and 203 consecutive escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400 notifications a day from two lanes with no agent running at all. On one, fm-control.sh exit answered already-stopped and fm-crew-state.sh read "failed - run failed". Closing the Herdr pane did not stop it either: with the pane genuinely gone and herdr pane read returning pane_not_found, the count kept climbing, because the poll is driven by the record's window= line rather than by the pane. The cost is not the repetition but that it drowns the alarms that matter. fm_backend_agent_state already separates a thinking agent from a gone one at process level. In the branch that was about to escalate, read it once and treat only its two recovery-grade verdicts - dead (endpoint present, no agent in it) and missing (endpoint authoritatively absent) - as proof, reporting that record once and not re-escalating it while it stays that way. Every other verdict, including alive, ambiguous, unreadable, unverified, and a read that failed outright, keeps the identical schedule, reason, and escalation count, so a genuinely wedged live agent is unaffected. The probe costs at most one backend read per window per threshold, the same budget the declared-wait consult and the worktree write probe already take. The report decides nothing about the record's fate: both lanes still held unlanded work and teardown refusing them was correct, so retiring, relaunching, or cleaning up stays with the supervisor. The once-only marker is owned entirely by that function and is dropped by the same read the moment the endpoint stops reading gone, so a replacement launched into the same window escalates normally and its own later death is reported again. Related, and not closed by this: #4412, #4482, #4316. Tests drive the real watcher against a record whose endpoint does not exist and pin both directions: dead and missing report once and never advance the count across later thresholds, while alive, ambiguous, and unreadable endpoints keep escalating with the identical reason and a climbing count. * fix(bin): bind the once-only dead report to the pane it reported Review of the parent commit found a reachable sequence where a later death in the same window lost its promised report. The marker was keyed on the verdict string alone and dropped only when a threshold probe read a non-gone verdict, but probes run only at thresholds: a replacement launched into the same window that dies without ever being probed alive - it crashes at startup, or works and then crashes - was absorbed by the previous death's marker. The pane's first sight yielded only the generic stale wake and every later threshold matched the stale marker, so the second death never got the detailed once-report that both the function's own comment and docs/architecture.md promise. Record the verdict together with the pane hash it was reported for, and absorb a repeat only while both still match. A replacement churns the pane, which resets the stale suppressor, wedge timer, and escalation count while no reset site touches this marker, so the pane half is what tells the second death apart from the first. The live-probe drop stays as it was. Clearing the marker at those reset sites instead would re-open unbounded re-alarming for a dead pane whose display ever ticks, which is the exact defect th…
friesentius
pushed a commit
to friesentius/firstmate
that referenced
this pull request
Sep 21, 2026
…#4656) * fix(bin): re-record PR poll identity after a volume device renumber (Fixes kunchenguid#4260) A volume remount can renumber the state filesystem's st_dev while every inode and byte stays the same; APFS does this across a reboot. A poll registration records its sidecar and check as device:inode, so every poll armed before the remount failed strict validation and the watcher refused all of them as unauthenticated state checks until each was re-armed by hand. There are two device comparisons. fm_pr_private_file_valid compares a live file's device with the state directory's device read in the same invocation: it refuses a file that is not on the state directory's own filesystem and already survives a renumber, so it is unchanged. The registration's recorded identity versus the live identity (from kunchenguid#556, reused by the kunchenguid#932 retirement receipt) binds the registration to the exact files published in its own transaction; its device part is what breaks. When strict capture fails, the watcher now proves the device is the only difference: every other artifact check passes (template bytes, both hashes, private mode, single link, live device, metadata), both recorded identities name one device, and each recorded inode equals its live inode. Only then, under the task's control lock, does it rewrite the two identity lines, repeating the whole proof and comparing the registration's file identity and bytes just before the rename, and then capture strictly again. A swapped, altered, re-moded, relinked, split-device, or foreign-device artifact still fails a proof and is still refused, and a pending retirement receipt blocks the rewrite. Reproduction: on macOS a poll armed on an APFS disk image that was detached and re-attached behind another image moved st_dev 16777239 -> 16777243 with inodes, bytes, mode, and link count unchanged; the real watcher refused it on main and reports its merge with this change. The portable regression test rewrites a real registration's recorded device and drives the watcher. Not changed here: the status presentation cursor keys rows by its own device:inode identity in bin/fm-classify-lib.sh, a different helper that needs its own fix; a retirement receipt left by a reboot between its publication and removal still names the old device and stays refused; custom check trust binds only a content hash and is unaffected. * fix(review): Serialize PR poll publication writers * fix(review): Bound PR poll publication lock scope
friesentius
pushed a commit
to friesentius/firstmate
that referenced
this pull request
Sep 21, 2026
…#4656) * fix(bin): re-record PR poll identity after a volume device renumber (Fixes kunchenguid#4260) A volume remount can renumber the state filesystem's st_dev while every inode and byte stays the same; APFS does this across a reboot. A poll registration records its sidecar and check as device:inode, so every poll armed before the remount failed strict validation and the watcher refused all of them as unauthenticated state checks until each was re-armed by hand. There are two device comparisons. fm_pr_private_file_valid compares a live file's device with the state directory's device read in the same invocation: it refuses a file that is not on the state directory's own filesystem and already survives a renumber, so it is unchanged. The registration's recorded identity versus the live identity (from kunchenguid#556, reused by the kunchenguid#932 retirement receipt) binds the registration to the exact files published in its own transaction; its device part is what breaks. When strict capture fails, the watcher now proves the device is the only difference: every other artifact check passes (template bytes, both hashes, private mode, single link, live device, metadata), both recorded identities name one device, and each recorded inode equals its live inode. Only then, under the task's control lock, does it rewrite the two identity lines, repeating the whole proof and comparing the registration's file identity and bytes just before the rename, and then capture strictly again. A swapped, altered, re-moded, relinked, split-device, or foreign-device artifact still fails a proof and is still refused, and a pending retirement receipt blocks the rewrite. Reproduction: on macOS a poll armed on an APFS disk image that was detached and re-attached behind another image moved st_dev 16777239 -> 16777243 with inodes, bytes, mode, and link count unchanged; the real watcher refused it on main and reports its merge with this change. The portable regression test rewrites a real registration's recorded device and drives the watcher. Not changed here: the status presentation cursor keys rows by its own device:inode identity in bin/fm-classify-lib.sh, a different helper that needs its own fix; a retirement receipt left by a reboot between its publication and removal still names the old device and stays refused; custom check trust binds only a content hash and is unaffected. * fix(review): Serialize PR poll publication writers * fix(review): Bound PR poll publication lock scope
ionnich
added a commit
to ionnich/firstmate
that referenced
this pull request
Sep 22, 2026
* feat(calm): render smooth Unicode swell with asymmetric two-color sail (#4498)
* feat(calm): render smooth Unicode swell
* feat(calm): make sails asymmetric
* feat(calm): use quarter sail glyph
* no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer
* no-mistakes(document): docs: sync calm wave phase doc comment
* no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。
* fix(bin): supersede stale scout delivery text in brief.md on promotion (#4491)
* fix: supersede scout delivery brief on promotion
* fix: preserve ship safety contract after promotion
* no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch
* fix(bin): make captain holds work on hosts with an older JSON::PP, and stop cleanup dropping accents from a held body (#4471)
* fix(bin): let captain holds work on hosts with an older JSON::PP
Holding a task for the captain, and the cleanup that keeps a captain-held row
open, both fail outright on any host whose JSON::PP defaults allow_nonref off -
2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`,
but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older
library rejects that whole value with "must be object or array".
The consequence is fleet-wide on such a host, not one broken command: a worker
there cannot formally record a decision for the captain at all. It can only
mention the decision in passing in a status line, where it can be missed - which
is how a real decision goes unrecorded. The hold reports that the task lost its
hold-set stamp; the cleanup cannot return the row to Queued.
Both call sites now ask for allow_nonref explicitly rather than inheriting
whatever the installed library defaults to. The second one is worth naming: its
`/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string
case that fails, so the guard selects for the failing input rather than
protecting against it.
The regression case forces the older default back off for every perl the commands
spawn, then drives both paths - holding a task that carries a body, and tearing
down a captain-held row whose deliverable must still be appended. It also probes
that the simulation genuinely rejects a bare scalar, so the case cannot pass
vacuously on a lenient host. Each half was verified failing on its own unfixed
call site with that site's real error message. Suites: fm-captain-hold-lifecycle
51 cases, fm-backlog-atomicity 99 cases, 0 failures.
Verification limit: the mechanism is reproduced and tested, but neither fix is
verified against a real JSON::PP 2.27202 host, because none is in the loop. This
laptop runs 4.06, where the bug does not manifest.
`bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a
brace-delimited object before decoding, so allow_nonref never applies.
* fix(bin): stop cleanup silently dropping accented characters from a held body
Cleanup rewrites a captain-held row's body to append the finished work's
deliverable, and the decoder it reads that body with printed decoded characters
to a stream with no `:raw` layer. A character at or below U+00FF then came out
as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the
accent. `fm_backlog_retain` writes that body straight back through
`--body-file`, and nothing reported an error - the character was simply gone
from a row still waiting on the captain.
The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus
`utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already
used.
Review of the parent commit found this on one of the lines that commit already
changed. It predates that change.
The test asserts bytes rather than decoded strings, because comparing strings
cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any
character above U+00FF makes perl print the whole string as UTF-8, so one body
carrying both an accent and an em dash passes even unfixed and proves nothing.
Verified failing before the fix on the accented row, passing after. Suites:
fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures.
* no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc
* no-mistakes(review): drop whole-file UTF-8 check from retained-body test
* no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc
* no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc
* fix(bin): read codex 0.154's idle braille starfield rows as composer furniture (#4532)
* fix(composer): read codex 0.154's idle starfield and status footer as furniture
codex-cli 0.154.0 animates a braille "starfield" around its idle composer:
on the row above the bold `›` prompt row, on the `›` row behind the SGR-2
dim `Ask Codex to do anything` placeholder, and on the row below it, then
draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`).
The cells are truecolor greys on both sides of the ghost luminance ceiling,
so the brighter ones survive ghost stripping, and the rows below the glyph
carry no structural edge. The shared classifier selected the bare `›` shape,
extended its wrap region over the two rows beneath the glyph, read the
survivors and the footer as wrapped typed input, and answered `pending`;
the steering doorbell defers on exactly that verdict, so no doorbell ever
reached an idle codex 0.154 pane.
bin/fm-composer-lib.sh now recognises that furniture by shape, declared
once next to the idle placeholders and reached from the two wrap-region
boundary points:
- a row whose non-whitespace content is entirely braille cells
(U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it
never counts as wrapped typed content and bounds a bare composer's wrap
region; braille behind the glyph row's content is stripped before the
emptiness decision when nothing else follows the glyph; a row mixing
braille with other text stays typed content;
- the codex status footer bounds the wrap region exactly as omp's status
row does, anchored on the effort token, a spaced middle dot, and a `~` or
`/` path cell, so a typed `fix · tests` stays composer input;
- `^Ask Codex to do anything$` joins the verified idle-placeholder set; the
ghost strip remains what proves that row empty, and the bare-row rule that
bright placeholder text is real input is unchanged.
Unchanged: the strict blank-row rule, the styled=0 degradation (a plain
cmux/orca capture of this screen still reads `unknown`, never `pending`),
FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape.
tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte
with the divergence (letters in place of the starfield read `pending`) and
the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh
is the default-on live guard (token-free, skips explicitly without codex or
tmux) that launches the installed codex idle and asserts `empty` through
both the tmux and the cursorless styled reads, naming codex --version on
failure. docs/verification/runtime-backends.md records the dated Herdr
evidence: `pending` before, `empty` after, on the captured screen.
* no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry
---------
Co-authored-by: Todd Billings <todd@usdvcapital.com>
* fix(bin): refuse empty text steers in fm-send (#4259)
* fix(bin): refuse empty text steers in fm-send
A marked secondmate request sent with an empty message delivered only
marker and correlation bytes and minted a pending-reply expectation the
parent could never see resolved, stalling the fleet with no loud error
(#4255). Fail closed on an empty or whitespace-only message on the text
path, mirroring the existing --resolve-key refusal.
* chore: retain ambient Pi-lens autoformat as its own commit
Formatting-only edits produced by ambient Pi-lens autoformat during the
msg-loss investigation, kept separate from the behavioural change in
c23acba6 so the fix stays reviewable on its own.
AGENTS.md is deliberately excluded: its only autoformat edit stripped the
trailing space from the documented FM_OPERATIONAL_PREFIX value, which
bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11
records as permanent compatibility. Documenting that constant without its
trailing space makes the doc wrong about the contract, so that one line was
restored rather than retained.
* fix(calm): paint the working ship one yellow over all-blue water (#4554)
On rose-pine-moon the two-color water (cyan crests over blue troughs) read as
a pink stripe over aqua, the yellow left sail and mast clashed with the red
right sail, and the hull carried a blue interior run. Every water cell is now
blue so the swell reads through glyph height alone, and both sail halves, the
mast, and the whole hull are one yellow run. Geometry, cadence, animation,
direction flip, resize clamping, and the narrow fallback are unchanged.
Update the unit and real-TUI color assertions to the new palette and the Calm
docs that described the old one.
* fix(bin): stop aging a second mate's active turn from its launch (#4270)
* fix(watch): stop aging a second mate's active turn from its launch
The parent watcher's second-mate wake-loop stall check exempts a mate that
is demonstrably inside an active turn, but secondmate_in_active_turn asked
busy_turn_over_age first and returned "not in a turn" whenever that said
the bound was crossed.
busy_turn_over_age ages from state/<task>.turn-ended, falling back to
state/<task>.meta. A second mate's turns end in its own home, so the
parent never gets a turn-ended mark for it and the fallback ages the
mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago
was therefore permanently "over age", the busy pane was never consulted,
and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false
wake-loop stall.
The gate now bounds the busy exemption by <idle> - how long the queue's
drain position has not moved - which is evidence this home actually
holds. A busy mate stays exempt while the queue has been frozen for less
than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so
the bound that stops a busy pane from proving liveness forever is kept
rather than removed. busy_turn_over_age is untouched; its remaining
callers are the ordinary crew busy-pane bound.
The regression pins the case that actually broke: a mate whose launch
record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn
must not escalate, while the same mate with its queue frozen past the
bound still publishes exactly one notification. The existing coverage
only exercised a freshly launched mate, which passes either way.
Reaching that alert now costs a pane capture inside the gate, so the
three checkpoints in this suite that assert an alert move from a 1s to a
4s bound - the value the neighbouring active-turn cases already use. The
bound is a ceiling, not a wait: the checkpoint returns on the first
actionable wake. On a loaded machine a 1s bound missed the alert
repeatedly; at 4s it did not miss in 20 runs under the same load.
* no-mistakes(review): scope the second-mate active-turn regression test's coverage claim
* no-mistakes(document): fix stale second-mate active-turn comments in fm-watch
* feat(bin): add read-only PR blocker and reviewer discovery commands (#4278)
* feat(bin): add read-only PR blocker and reviewer-discovery commands
Two focused, opt-in commands that read GitHub and never write to it.
fm-pr-state.sh reports what still blocks one pull request from the
author's side: a closed or merged state, draft state, unknown or
conflicting mergeability, absent or failing required checks, and a
blocking CHANGES_REQUESTED decision explained by each reviewer's latest
verdict, marked STALE when it was left at a superseded head. A pull
request that only awaits an approval is not reported as blocked, and
advisory checks are omitted. Every reading is taken against one exact
head; a push that lands mid-read invalidates the whole result rather
than mixing two snapshots.
fm-pr-reviewers.sh suggests reviewers from the most recent commits to
the pull request's exact changed paths, counting each commit once,
resolving handles through GitHub's own commit author.login mapping, and
excluding the author and Bot accounts.
Both stay read-only: no review request, no approval, no merge.
Unresolved review-thread state is left unreported because the REST API
does not expose it and unattended commands may not use GraphQL.
Closes #3731
* no-mistakes(review): accept only PR URLs and stop at terminal state
* no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate
* no-mistakes(review): stop attributing readings to unverified heads
* no-mistakes(review): narrow readiness contract to checks that have reported
* no-mistakes(review): read the pull request once, drop the head guard
* no-mistakes(document): scope pr-forge isolation proof to its measured members
* no-mistakes(document): record uncovered pr-forge members and their pending proof
* docs(isolation-proof): re-prove pr-forge at its full membership
tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the
pr-forge family in this branch, and script_allows_concurrency grants
four workers by family membership alone, so both ran concurrently on a
proof measured before they existed.
Re-proved the family at all eight members: two consecutive runs, 0
failures, each begun with the one-minute load average below 6.0 so the
result measures isolation rather than contention. A third run taken
between them is disclosed rather than recorded, because it started
while the previous run's workers were still decaying.
The new durations are not comparable with the six-member measurement
above them, so they are not presented as evidence about the two new
members, and that record's 1.72x four-worker figure is left as a
statement about its own run rather than restated as current.
* no-mistakes(review): disclose gh error-text coupling at its matching site and tests
* fix(bin): teach validation-round pauses in generated briefs (#2752)
* fix(bin): teach validation-round pauses in briefs
* no-mistakes(document): Point classifier comments to authoritative pause examples
* docs(readme): add star history chart (#4558)
* fix(bin): refuse teardown when a task's endpoint close fails (#4510)
* fix(teardown): refuse a cleanup whose endpoint close failed
bin/fm-teardown.sh discarded both the exit status and the stderr of every
fm_backend_kill call, so a close that genuinely failed was indistinguishable
from one that succeeded. Teardown continued past it, deleted the task's durable
records, returned its worktree, and reported the cleanup as completed. The
deleted metadata is the only record of which endpoint belongs to the task, so
such a close did not merely leave a stray session behind, it stranded one:
nothing was left on disk naming it.
The adapters could not carry that signal either. Driven against the real code,
every backend arm returned 0 for a genuine failure exactly as it did for an
already-exited endpoint, so there was nothing for the four call sites to
propagate even once they stopped swallowing it.
The tmux arm now resolves a close that did not succeed against the window's
exact recorded identity, since kill-window fails the same way for a window that
is gone and one that is still there. The Orca arm reports a close its missing
CLI never attempted. Both stay silent for an endpoint that is already
legitimately gone, and the remaining arms are unchanged: their close-command
timing cannot be established without the real Zellij, Orca, and cmux binaries,
and a gate that refused ordinary cleanup of an already-exited session would be
worse than the defect. docs/verification/runtime-backends.md records what each
backend can prove.
A reported close failure now reaches teardown's existing retain-and-stop
refusal before the records naming the endpoint are removed, matching where the
Herdr confirmed-gone gates already sit for the same hazard, and the retained
records let a rerun finish once the close works.
* no-mistakes(review): refuse unreadable tmux close re-read; honor --force override
* no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close
* no-mistakes(document): document endpoint-close refusal in its backend and retirement owners
* no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree
* feat(calm): add flag-gated Claude Code Calm mode (#4565)
* feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag
Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that
brings Calm to Claude Code: the sailboat replaces the stock working row through a
Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration,
and canonically classified operational user rows draw at zero height. /calm is
registered by the hooks module itself and toggles the same per-home config/calm
preference the Pi extension uses, so one choice applies on either harness; rows
redraw retroactively on toggle and stay hidden across claude --continue.
The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS
flag is on. Nothing sets that flag in any settings file, and the plugin carries no
command file, skill, agent, or classic hook, so it is a complete no-op while the
flag is off. The trusted project auto-loads it through an .agents/skills symlink,
the only path Claude Code scans for project plugins.
Extract the working-ship geometry, bounce track, cadences, and freeze/resume state
into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module
imports from outside the plugin folder) and have the Pi widget paint that core's
frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify
operational rows through a port of bin/fm-operational-input.sh's classify command
guarded by a corpus parity test against the shell owner.
Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering,
Raster packing, policy, classifier parity), the mod's own claude plugin test suites
behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op,
the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272.
Docs: record the version-scoped Claude Code evidence and the three bounded gaps in
docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md,
and make the shared preference, layout, and contributor notes harness-neutral.
* no-mistakes(review): Preserve colliding final replies and strengthen parser parity
* no-mistakes(review): Preserve final replies and strengthen canonical parity checks
* no-mistakes(review): Require exact function-hooks opt-in before Calm activation
* no-mistakes(review): Clarify Calm module loading and activation boundaries
* no-mistakes(review): Reset Calm presentation state across session starts
* no-mistakes(document): Refresh Calm session lifecycle documentation
* feat(calm): paint the Claude Code working ship in Claude's own theme colors
The captain picked the "Claude native" palette for the Claude Code mod's Raster:
every water cell takes the spinner blue of the active theme family (#93a5ff dark,
#5769f7 light) and the whole boat takes the Claude orange of the stock spinner
(#d77757), one water color and one boat color. The family follows the `theme`
setting's prefix, read at load through $.config.list and re-read on a
config.set of that row, with `auto` and custom themes falling back to the dark
set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte.
Rename the shared sprite's color classes from hue names to `water` and `boat`,
since each harness now maps them to its own colors; geometry, motion, cadence,
and the activation gate are untouched.
Tests cover both palettes' packing and the family rule under Node, and the
plugin kit drives every theme value, a theme change mid-session, the Calm-off
pass-through, and inertness of the menu read while the flag is off. The docs
describe the Claude Code colors and record the guard passing on 2.1.273.
* no-mistakes(review): Use light palette for unresolved Claude themes
* no-mistakes(document): Refresh Claude Calm verification evidence
* fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586)
* fix(watch): honour a declared wait before wedge-escalating a quiet pane
wedge_timer_check escalated on elapsed idle time alone. Nothing asked
whether the worker had already said why its pane was quiet, so a lane
that declared a bounded external wait climbed the escalation ladder for
as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every
repeat carried demand-deep-inspection - which by its own wording forbids
re-absorbing on the run-step or pane state, so the supervisor could not
use the evidence that was there either.
The generated brief promises that declaring `paused:` buys the long
recheck cadence instead of a wedge, but the timer was still reachable
while that declaration stood: a crew that declares a wait and then has an
active run or busy pane attributed to it is handed to the timer as
provably-working. The declaration is what the worker said about its own
silence, so it now outranks a liveness verdict that only says something
is running.
The consult runs in the at-threshold branch that was about to escalate,
beside the worktree walk already there, and costs one status-line read.
Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS
recheck the declared-wait absorber already uses, so the wait is still
rechecked and cannot rot invisibly. Which verb declared it decides the
wording, because the two block on different people: a `paused:` wait is
owed by an external dependency and asks the reader to confirm it still
holds, while a `captain-held:` transfer is owed by the captain reading
the recheck and asks them to answer or release the hold. A hold is not
rechecked at all while the away-posture record exists, as on every other
captain-held path, and that absorb arms no throttle so the recheck is
owed in full on return.
A declared clearing time that has already passed stops counting, and a
lane that never declared one keeps the identical escalation schedule,
reason, count and demand-deep-inspection wording, so detection and its
worst-case time are unchanged. The deferral restarts the idle timer
rather than cancelling it, so a lane that stops waiting escalates again
within one threshold.
A lane quiet because its own validation run is parked at a gate awaiting
a human decision is deliberately out of scope: reading that state needs a
signal carrying who the wait is on and what clears it, rather than one
inferred from a parked verdict that also covers gates awaiting the
crewmate itself.
Tests pin both directions for each case and were each confirmed to fail
with the consult removed.
* no-mistakes(document): docs: honour declared waits in stale-escalation docs
* fix(bin): report verified PR state for passed runs (#4624)
* fix(bin): derive passed PR state from PR record
A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists.
For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim.
Fixes #4607
* no-mistakes(review): Add bounded GitLab merge-request state reads
* no-mistakes(review): Preserve network-free inactive crew-state scans
* no-mistakes(document): Document PR record readers in shared library
* fix: restore published contribution follow-up (Fixes #4469) (#4627)
* fix: restore published contribution follow-up (Fixes #4469)
* fix(review): Fix contribution freshness and merge actor routing
* fix(review): Restore issue triage and scope contribution follow-up
* fix(test): test: assert one wake per contribution signal
* fix(document): Document contribution follow-up
* fix: restore truthful terminal delivery evidence
* fix(review): Disclose unsupported contributions and deduplicate watcher wakes
* fix(review): Preserve unmeasured unsupported contributions across Bearings
* fix(review): Deduplicate shared contribution wakes and isolate diagnostics
* fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower
* fix(bin): make remote report transfers explicit and fail-open (#4658)
* fix(bin): make a remote-reply document gap self-clearing and re-attemptable
A remote mate's undelivered document raised a keyed `blocked` decision that
nothing could ever resolve, and any `data/*.md` substring in any mirrored line
was an unconditional fetch instruction. A mate announcing a report it had not
written yet therefore manufactured a permanent, factually false blocker, and
its own explanation of the false alarm manufactured more.
The reader has no permanence vocabulary: a report still being written refuses
exactly like a path that will never exist. So an undelivered document is now a
durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`,
re-attempted on the next delta and on the channel's own quiet poll, and retired
with a matching `resolved` line naming the local copy once it arrives. The
cursor still advances and no delta stalls on one bad pointer.
Only a structured `report=data/....md` pointer now offers a document, so a path
merely mentioned in prose - including one under another home's mirror tree,
which is provably not that mate's to serve - is never fetched. Offers are
deduplicated across the whole delta, the escalation names each missing document
once and carries the reader's own reason instead of discarding it, and a
strictly increasing notice ordinal keeps a later escalation from being
swallowed as duplicate bytes. A mirrored line still lands once whichever
pointer form it was first written under.
* no-mistakes(review): Require structured pointer token boundaries
* no-mistakes(review): Unify boundary-safe pointer extraction and rewriting
* fix(bin): identify a mirrored line independently of its delivery state
Two defects in the boundary-safe pointer work.
The at-most-once check compared only the all-remote and all-local renderings
of a line, so it could not recognize a mixed one. A line offering two documents
where only the first was deliverable mirrored as local-plus-remote; once the
second arrived, a cursor-loss whole-log recapture rendered the same line
all-local, matched neither alternate, and mirrored a second time. A line's
identity is now the canonical form every boundary-valid pointer would take once
delivered, derived by the same parser that does extraction and rewriting, so it
no longer depends on which documents happened to be deliverable at the time.
The pointer map was passed to awk through the process environment. A delta may
carry up to the configured 1 MiB bound, and an expanded map of delivered
pointers can exceed the platform's exec argument limit, so awk would fail to
start; because no caller checked, the empty result would have been appended as
blank lines while the cursor advanced past dropped status content. The map now
travels in a file, and every call site checks the exit status and stops the
ingest rather than committing a delta it could not render.
Both passes now run once per stream instead of twice per line.
* no-mistakes(review): Abort ingest when document pointer extraction fails
* no-mistakes(review): Exclude structured cross-home pointers from document transfer
* fix(bin): fail open on an undeliverable remote document instead of tracking it
Narrow the remote-reply document fix to the scope the diagnosis actually
requires, as decided after measuring a simpler alternative.
A document the reader cannot deliver now fails open. The mate's line is
mirrored with its own pointer, the cursor advances, and one unkeyed note
carries the reader's reason. A note never enters the open-decision fold, so it
cannot stand open the way the original keyed block did - which removes the
never-clearing false blocker by construction rather than by resolving it.
That makes the durable self-clearing obligation unnecessary, so it goes: the
per-mate pending-documents record, its notice ordinal and resolved
announcements, and the poll-side retry. Canonical line identity goes too, and
with it a way to silently drop a genuine status line; mirroring is back to
at-most-once on exact bytes. The cross-home exclusion goes as well: under
fail-open a cross-home report= either fails harmlessly or is a nested remote
report this mate genuinely holds, which is now relayed again.
Kept: fetching only on a structured report= pointer, the boundary-correct
parser, the file-based rewrite map, and checked extraction and rewrite exit
status. The parser now scans behind a sentinel byte so a rejected candidate can
no longer give the text right after it a false leading boundary.
The reported incident is covered end to end: a report path announced in prose
before it exists raises no decision, and the report still arrives through the
ledger publisher's structured offer once written.
* no-mistakes(review): Preserve source-line identity across remote reply replays
* no-mistakes(document): Document remote reply transfer and replay semantics
* no-mistakes(lint): Fix staging truncation lint checks
* fix(calm): preserve substantive mid-turn responses (#4655)
* Preserve substantive Calm mid-turn text
* no-mistakes(review): Distinguish newline-preserved replies from short narration
* no-mistakes(document): Document Calm mid-turn preservation boundaries
* no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`
* fix(bin): preserve PR merge polls across volume remounts (#4656)
* fix(bin): re-record PR poll identity after a volume device renumber (Fixes #4260)
A volume remount can renumber the state filesystem's st_dev while every
inode and byte stays the same; APFS does this across a reboot. A poll
registration records its sidecar and check as device:inode, so every poll
armed before the remount failed strict validation and the watcher refused
all of them as unauthenticated state checks until each was re-armed by hand.
There are two device comparisons. fm_pr_private_file_valid compares a live
file's device with the state directory's device read in the same invocation:
it refuses a file that is not on the state directory's own filesystem and
already survives a renumber, so it is unchanged. The registration's recorded
identity versus the live identity (from #556, reused by the #932 retirement
receipt) binds the registration to the exact files published in its own
transaction; its device part is what breaks.
When strict capture fails, the watcher now proves the device is the only
difference: every other artifact check passes (template bytes, both hashes,
private mode, single link, live device, metadata), both recorded identities
name one device, and each recorded inode equals its live inode. Only then,
under the task's control lock, does it rewrite the two identity lines,
repeating the whole proof and comparing the registration's file identity and
bytes just before the rename, and then capture strictly again. A swapped,
altered, re-moded, relinked, split-device, or foreign-device artifact still
fails a proof and is still refused, and a pending retirement receipt blocks
the rewrite.
Reproduction: on macOS a poll armed on an APFS disk image that was detached
and re-attached behind another image moved st_dev 16777239 -> 16777243 with
inodes, bytes, mode, and link count unchanged; the real watcher refused it on
main and reports its merge with this change. The portable regression test
rewrites a real registration's recorded device and drives the watcher.
Not changed here: the status presentation cursor keys rows by its own
device:inode identity in bin/fm-classify-lib.sh, a different helper that
needs its own fix; a retirement receipt left by a reboot between its
publication and removal still names the old device and stays refused; custom
check trust binds only a content hash and is unaffected.
* fix(review): Serialize PR poll publication writers
* fix(review): Bound PR poll publication lock scope
* fix(bin): keep contribution records when the poll budget runs out (follow-up to #4627) (#4661)
A budget that expires partway through an observation no longer records an
error or prints the unavailable wake; the URL keeps its prior record and is
observed first next poll. forge() flags budget exhaustion at the point it
refuses, or when a read is killed at the budget's own deadline, so a genuine
forge failure still records the error and wakes. Each distinct URL is now
observed once per poll and applied to every owning task.
* fix(bin): clear parent pending-replies on local secondmate retirement (#4680)
* fix(bin): clear parent pending-replies on local secondmate retirement
Local secondmate teardown left resolved parent pending-reply records behind
after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced
retirement while any reply for that id is still unresolved, and delete every
matching record plus its delivery confirmation after a successful local or
remote retirement, matching the remote cleanup path.
* no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup
* no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung
* no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern
* no-mistakes(review): Pending-replies Basename und corr_id abgleichen
* no-mistakes(document): Clarify forced retirement pending-reply cleanup
---------
Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>
* fix(bin): accept Orca's composite worktree id when tearing down a task (#4677)
* fix(bin): accept Orca's composite worktree id at teardown
Teardown refused every Orca-backed task because the endpoint validator
checked orca_worktree_id with the simple-atom rule meant for tmux-style
window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca
returns that id as `<orca id>::<absolute worktree path>`, so the colon and
slashes in every real value made validation fail and finished Orca tasks
could never be cleaned up.
Validate the field as the composite it is: both halves of the first `::`
split present, the path half absolute, and no embedded newline, carriage
return, or tab. The terminal field keeps the atom check, which is correct
for it, and no other backend's validation changes.
The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca
never returns, which is why the suite passed a check the real value fails.
They now carry the composite form, so the tests exercise the real value.
* no-mistakes(document): name Orca's repo id in the composite worktree id
* no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points
* feat(bin): add opt-in typed dispatch resolution (#4692)
* feat(bin): add opt-in typed dispatch resolution through typesafe.ai
Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or
scout profile from a written brief with typesafe.ai's System One model:
one Choice question over the rules' `when` texts, then the confidence
floor, the rule's `approval` and `floor`, each profile's `provider` and
`floor`, one quota-axi snapshot, and the spendPriority argmax all in code.
It is off unless TYPESAFE_API_KEY is in the environment or the home's
gitignored .env; off means one stderr line, exit 0, and no network call,
so firstmate dispatches exactly as before. The key reaches curl on a file
descriptor, never argv.
Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and
the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new
tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates
the four new optional dispatch fields. Document the schema, the operator
contract, the AGENTS.md intake step, and the live and benchmark evidence.
* no-mistakes(review): Harden typed dispatch resolution and quota bounds
* no-mistakes(review): Validate dispatch floors and ranking evidence
* no-mistakes(review): Tighten dispatch response and floor evidence
* no-mistakes(review): Neutralize none matching and resolve defaults locally
* no-mistakes(review): Preserve providerless profiles outside typed resolution
* no-mistakes(review): Validate response usage and reject duplicate profiles
* no-mistakes(review): Escalate unverifiable floors and validate probabilities
* no-mistakes(review): Validate probability mass and unknown profile floors
* no-mistakes(review): Simplify resolver interface and preserve fallback routing
* no-mistakes(review): Fix constants and rank partial quota evidence
* no-mistakes(review): Add authoritative provider mapping and enforce explicit providers
* no-mistakes(review): Declare provider for documented Pi profile
* no-mistakes(review): Validate provider identifiers and support Gemini dispatch
* no-mistakes(review): Strictly anchor provider identifiers
* no-mistakes(review): Validate selectors and preserve fallback candidate evidence
* no-mistakes(review): Gate typed validation and harden resolver evidence
* no-mistakes(review): Preserve opt-in routing and harden candidate evidence
* no-mistakes(review): Prioritize known exhaustion over quota uncertainty
* no-mistakes(review): Isolate API secrets and preserve no-key diagnostics
* no-mistakes(review): Fallback safely when dispatch rules are absent
* no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets
* no-mistakes(document): Document typed dispatch safety and fallback behavior
* fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753)
* test: reproduce buried status declarations in shared readers
* fix: share status event reads and preserve open blockers
* fix: retain terminal scout and ship status declarations
* no-mistakes(review): Fix status chronology, legacy completions, and reader performance
* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots
* no-mistakes(review): Unify terminal supersession across cached folds and consumers
* no-mistakes(review): Filter per-key status history while preserving terminal chronology
* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells
* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses
* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor
* no-mistakes(lint): Quote literal done in test for-lists for SC1010
* ci: expect 19 snapshot/fleet-view tests
This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.
* no-mistakes(review): Restore multiline child outcome reporting
* no-mistakes(review): Select ledger terminal events through bounded shared reader
* no-mistakes(review): Report newest open decision instead of preferring blocked
* no-mistakes(review): Require colon before ship/scout terminal supersession in fold
* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix
* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions
* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold
* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule
* no-mistakes(document): Align status-read docs with fold-resolved crew state
* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers
* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree
* test: fold terminal-cleanup snapshot coverage into the completed-scout case
Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.
* no-mistakes(document): Clarify socket-down override expiry in architecture doc
* ci: retrigger flaky contribution check
* fix(bin): launch codex crewmates with codex's hook layer disabled (#4689)
* fix(spawn): launch codex crewmates with codex's hook layer disabled
A freshly launched Codex worker never reached its instructions. Codex
stopped it on an interactive "Hooks need review" modal whose selection
sits on "Review hooks", which is neither trusting nor declining.
Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow
navigation, so the selection cannot be moved, and pre-accepting the
prompt by writing Codex's own trust store would record an operator
consent that was never given.
The hooks are the machine's own ~/.codex/hooks.json plus any project's
.codex/hooks.json. A crewmate needs neither: its turn-end signal is the
-c notify= program on the same launch, and Firstmate's project hooks are
primary-session infrastructure that stands down in a child worktree.
Crewmate and scout launches now pass --disable hooks. That is the
opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted
hooks; disabling the feature runs none of them and leaves the operator's
~/.codex untouched. An unknown feature name is a hard Codex error, so a
release that drops the flag fails the launch loudly instead of silently
restoring the modal. A secondmate is a primary in its own home and keeps
the project hooks its turn-end guard and session-start digest ride on.
Verified on codex-cli 0.151.0: the modal is gone and the turn-end
notification still lands.
This unblocks the second review that every finished pull request is supposed to get.
Fixes kunchenguid/firstmate#4673
* no-mistakes(review): Fix contradictory hook count in Codex verification record
* fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710)
* fix(bin): settle terminal contributions and wake once per read-failure episode
A contribution whose last good observation is merged or closed is final:
poll no longer re-reads it, projection keeps it fresh, and a stale error
recorded beside it is cleared once. A genuine forge-read failure on an open
contribution still records its error on every cycle but prints the
unavailable wake only when it starts a failure episode; a successful read
ends the episode. Open PRs linked from done tasks keep being observed.
The false unavailable beside a complete observation was budget exhaustion
mid-observation, already fixed by #4661.
* fix(review): Settle terminal contribution owners
* fix(review): Deduplicate shared contribution failure episodes
* fix(test): Preserve settled terminal contribution records
* fix: select authoritative no-mistakes runs (#4476)
* fix(crew-state): select authoritative validation runs by identity
Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row.
Refs: https://github.com/kunchenguid/firstmate/issues/3215
* fix(review): Resolve same-branch run identities beyond capped history
* fix(review): Fix run-selection compatibility, races, and worker-state fallbacks
* fix(review): Limit run validation to the requested branch
* fix(test): Anchor AXI fixtures and document remaining live evidence gaps
* fix(document): Clarify run selection documentation and capture ownership
* fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation
* fix: distinguish captain outcomes from no-op updates (#4738)
* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work
MAIN answered a supervision-branch outcome for completed captain-requested
work (implementation done, PR ready for review and merge approval) with
"Captain, shipshape.", reading section 9's no-action reply as covering it
and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no
captain-facing response is owed".
Section 9 now limits the shipshape reply to true no-ops (idle re-read,
empty heartbeat, consequence-free acknowledgement) and requires a short
outcome response naming what finished and what word is needed whenever
requested work finishes or a result needs the captain's word, even when a
transcript entry already shows the substance. The Pi protocol's re-emit
rule now says it bounds repetition only, and carries a worked example of
the ready-for-review outcome whose correct processing turn a shipshape
reply fails.
No executable contract evaluates the content of MAIN's captain-facing
reply, so the regression is the protocol example in the owner doc rather
than a text-match test.
* no-mistakes(document): Clarify captain-facing outcomes versus no-ops
* docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line
The document step condensed the Pi protocol's re-emit rule and dropped the
worked example of a finished, ready-for-review outcome whose correct
processing turn a "Captain, shipshape." reply fails. That example is the
contract's regression: no executable contract evaluates the content of
MAIN's captain-facing reply, so the owner doc's example is the test case.
Restore it directly under the re-emit rule, prefixed as a regression
example that is kept verbatim and never condensed or summarized away.
* no-mistakes(review): Clarify captain outcome and decision-word requirements
* no-mistakes(document): Clarify captain-facing completion outcomes
* docs(pi): require the PR URL in the visible captain-facing outcome reply
Captain review on the regression example: drop the sample reply string
and say only that the ready-for-review outcome requires relaying a
captain-facing outcome response, not just "Captain, shipshape.".
Fold in the visible-PR-handoff failure seen this session: after the
branch outcome reporting this fix green, MAIN's visible reply was only
"Awaiting your merge call." with no PR URL, leaning on the dim anchor.
Section 9's URL rule now also covers a review or merge ask and names the
visible reply as where the URL goes, sourced from the ready status, pr=
metadata, or the supervision branch's summary and never left to a
transcript entry. The Pi protocol adds the same-way failure and places
the captain-facing text in the final visible assistant reply after the
fm_branch_processed call, because Calm hides assistant text emitted in
the same step as a tool call as a working note.
Investigation verdict, evidence in the PR comment: no recent PR caused
the handoff failure; Pi has hidden same-step pre-tool assistant text
since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and
#4658 touched only remote report transfer.
* no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs
* no-mistakes(document): Clarify captain-facing supervision outcomes
* docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule
The consolidated section 9 URL rule narrowed its trigger to a review or
merge ask, dropping the "whenever a PR is mentioned" catch-all from
#3648 that keeps every PR URL copied from a durable record and never
assembled from memory. Restore that trigger as a union with the review
or merge ask so the one consolidated rule covers both.
* fix(bin): let non-owner Claude Stops exit safely (#4777)
* Fix foreign-owner turn-end supervision loop
* no-mistakes(review): Scope foreign-owner safe exit to Claude guard
* no-mistakes(document): Document Claude foreign-owner safe exit
* fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778)
Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty
indexed array as an unbound variable and aborts the shell. In
signal_turnend_panes_churned() the missing_keys loop was reachable with
an empty array whenever every churned key already held a fresh
.churn-since-* marker (a second churning turn-end inside an open
deferral window), so each watcher cycle died about half a minute in and
supervision restarted endlessly. The created_keys rollback loops had the
same latent crash on their error paths.
Audit of bin/ for the same pattern found one more confirmed-reachable
case: remote_handoff's noncanonical-body scan iterates to_move, which is
empty when a retried remote handoff finds every key already staged in
the outbox. All other "${arr[@]}" sites are either count-guarded,
guaranteed non-empty by construction, or unreachable while empty.
Guard the three reachable expansions with the repo's existing
"${arr[@]+...}" idiom. New regression test drives a real watcher
through the all-marked churn path; the macos-stock-bash CI lane runs it
under real /bin/bash 3.2 via FM_TEST_ONLY.
* Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783)
The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: require complete captain-facing final responses (#4779)
* docs: require complete final responses across harnesses
* no-mistakes(document): Document complete final replies for Grok Bot
* docs: point Grok replies to the shared contract owner
* no-mistakes(review): Clarify final recap without batching decision asks
* fix: preserve substantive mid-turn text in Pi Calm (#4788)
* fix(calm): preserve substantive Pi mid-turn text
* no-mistakes(review): Preserve substantive Pi Calm text per block
* no-mistakes(test): Cover shared Calm preservation boundaries behaviorally
* no-mistakes(document): Consolidate Calm preservation documentation
* fix: harden mail checks and rebalance full-coverage CI (#4800)
* Improve CI reliability and rebalance full-coverage validation
* no-mistakes(document): Clarify lint partition documentation
* fix(bin): answer Kimi 2.0.0 folder-trust dialog during spawn (#4799)
* Handle Kimi workspace trust dialog
* no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers
* no-mistakes(review): Gate Kimi ready on any trust marker and clean captures
* no-mistakes(review): Read visible pane for Kimi trust and ready gates
* no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate
* no-mistakes(review): Harden Kimi viewport capture and trust dialog detection
* no-mistakes(document): Document Kimi spawn refusal on cmux and Orca
* fix(bin): report a dead-agent record once instead of escalating forever (#4775)
* fix(bin): report a record whose agent is gone once instead of escalating forever
The wedge escalation path never asked whether there was still an agent to be
wedged. A wedge is something stuck that might recover, so re-alarming it earns
its cost; an agent that is gone never moves again, its pane never churns, the
idle timer never resets, and the escalate path clears its own timer and re-arms
with nothing bounding the count.
Observed on a live fleet: two finished lanes reached 226 and 203 consecutive
escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400
notifications a day from two lanes with no agent running at all. On one,
fm-control.sh exit answered already-stopped and fm-crew-state.sh read
"failed - run failed". Closing the Herdr pane did not stop it either: with the
pane genuinely gone and herdr pane read returning pane_not_found, the count kept
climbing, because the poll is driven by the record's window= line rather than by
the pane. The cost is not the repetition but that it drowns the alarms that
matter.
fm_backend_agent_state already separates a thinking agent from a gone one at
process level. In the branch that was about to escalate, read it once and treat
only its two recovery-grade verdicts - dead (endpoint present, no agent in it)
and missing (endpoint authoritatively absent) - as proof, reporting that record
once and not re-escalating it while it stays that way. Every other verdict,
including alive, ambiguous, unreadable, unverified, and a read that failed
outright, keeps the identical schedule, reason, and escalation count, so a
genuinely wedged live agent is unaffected. The probe costs at most one backend
read per window per threshold, the same budget the declared-wait consult and the
worktree write probe already take.
The report decides nothing about the record's fate: both lanes still held
unlanded work and teardown refusing them was correct, so retiring, relaunching,
or cleaning up stays with the supervisor. The once-only marker is owned entirely
by that function and is dropped by the same read the moment the endpoint stops
reading gone, so a replacement launched into the same window escalates normally
and its own later death is reported again.
Related, and not closed by this: #4412, #4482, #4316.
Tests drive the real watcher against a record whose endpoint does not exist and
pin both directions: dead and missing report once and never advance the count
across later thresholds, while alive, ambiguous, and unreadable endpoints keep
escalating with the identical reason and a climbing count.
* fix(bin): bind the once-only dead report to the pane it reported
Review of the parent commit found a reachable sequence where a later death in
the same window lost its promised report. The marker was keyed on the verdict
string alone and dropped only when a threshold probe read a non-gone verdict,
but probes run only at thresholds: a replacement launched into the same window
that dies without ever being probed alive - it crashes at startup, or works and
then crashes - was absorbed by the previous death's marker. The pane's first
sight yielded only the generic stale wake and every later threshold matched the
stale marker, so the second death never got the detailed once-report that both
the function's own comment and docs/architecture.md promise.
Record the verdict together with the pane hash it was reported for, and absorb a
repeat only while both still match. A replacement churns the pane, which resets
the stale suppressor, wedge timer, and escalation count while no reset site
touches this marker, so the pane half is what tells the second death apart from
the first. The live-probe drop stays as it was.
Clearing the marker at those reset sites instead would re-open unbounded
re-alarming for a dead pane whose display ever ticks, which is the exact defect
the parent commit exists to close.
The noise bound is unchanged: an unchanged dead pane still absorbs on every
later threshold and never advances the escalation count, and every verdict short
of proof still escalates exactly as before.
* no-mistakes(review): Key the dead-record once-marker on the busy incarnation token
* no-mistakes(document): Document dead-record escalation cap in stale-pane config entry
* no-mistakes(document): Add busy-state inventory line to AGENTS.md
* no-mistakes(document): Document dead-record probe on busy-turn-bound wedge path
* fix(bin): create captain-hold rows when Beads requires due (#4854)
Captain holds have no due semantics and are a hold kind, not a Beads issue
type. The create path now waives due.required and maps to native type task.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: disable compact adviser for spawned agents (#4877)
* feat(bin): launch every spawned agent with the compact adviser disabled
Every crewmate, scout, and secondmate Firstmate launches now starts with
COMPACT_ADVISER_DISABLE=1, on a fresh spawn and on a relaunch alike, so an
unattended session never activates the compact adviser.
The value is unconditional: no configuration file gates it and there is no
override, unlike the trace carrier beside it.
Three carriers deliver it, because no single one covers every launch shape.
The pane shell receives an export beside GOTMPDIR, so the agent's own children
inherit it too.
The launch command carries an explicit assignment, prepended outermost so it
wins over any ambient value the pane already held.
The cleared launch environment sets it again at the `env -i` boundary and keeps
COMPACT_ADVISER_DISABLE in the fixed operational floor, which is what preserves
the switch when config/launch-env-allowlist empties the environment, and what
delivers it on a remote host that never had the value.
bin/fm-control.sh relaunch, the bootstrap secondmate relaunch, and the remote
secondmate transport all rebuild their launch through bin/fm-spawn.sh, so they
inherit the same floor.
The captain's own primary session is untouched.
The two new suites drive the real spawn and then execute the launch command the
pane actually received, with the harness replaced by a probe that prints its own
environment, rather than matching script text.
They cover ship and secondmate launches with the allowlist absent and enabled,
the pane export and its ordering, fm-control.sh relaunch, and the full parent to
remote-host chain.
* no-mistakes(review): Export compact-adviser disable across compound launches
* no-mistakes(document): Document spawned-agent compact-adviser environment guarantee
* fix(bin): preserve Claude lock ownership after helper recycling (#4894)
* fix(bin): let a background Claude session keep owning its session lock
Session-lock ownership was decided by process ancestry alone. Under an
unattended Claude session the model loop runs in a transient bg-spare
bridged to the front-end by a shared daemon; when that bridge is
recycled the contiguous claude-named ancestry from a hook to the
recorded owner breaks while the owner pid stays alive, so the Stop
auto-arm stood down as a foreign live owner, the turn-end guard ended
every turn with its read-only diagnostic, and fm-lock.sh refused - a
self-sustaining outage until restart.
Ownership is now ancestry membership OR a trusted same-session id,
never id-first:
- fm-session-lock-lib.sh accepts CLAUDE_CODE_SESSION_ID only when
CLAUDE_PID is a Claude-shaped member of the current contiguous run,
compares it against the id recorded in state/.lock-session, and
requires the recorded pid to still be a live harness. No id, no
sidecar, an untrusted id, a different id, or a dead recorded pid
leaves the ancestry verdict unchanged. Ids are never read from ps
argv.
- fm-lock.sh accepts a same-session holder at both refusal sites,
writes, refreshes, and clears the sidecar only under its claim lock
(including the early already-mine exit, skipped only while the
deferred startup sweep leases that lock), keeps it byte-identical
across a same-session confirmation, records CLAUDE_PID on lock line 1
for a session with a trusted id so a shared daemon or front-end that
outlives the session never keeps a dead session's lock alive, never
rewrites a live line 1 on a same-session confirmation, and names t…
KaranSantra
added a commit
to KaranSantra/firstmate
that referenced
this pull request
Sep 23, 2026
) * fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586) * fix(watch): honour a declared wait before wedge-escalating a quiet pane wedge_timer_check escalated on elapsed idle time alone. Nothing asked whether the worker had already said why its pane was quiet, so a lane that declared a bounded external wait climbed the escalation ladder for as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every repeat carried demand-deep-inspection - which by its own wording forbids re-absorbing on the run-step or pane state, so the supervisor could not use the evidence that was there either. The generated brief promises that declaring `paused:` buys the long recheck cadence instead of a wedge, but the timer was still reachable while that declaration stood: a crew that declares a wait and then has an active run or busy pane attributed to it is handed to the timer as provably-working. The declaration is what the worker said about its own silence, so it now outranks a liveness verdict that only says something is running. The consult runs in the at-threshold branch that was about to escalate, beside the worktree walk already there, and costs one status-line read. Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS recheck the declared-wait absorber already uses, so the wait is still rechecked and cannot rot invisibly. Which verb declared it decides the wording, because the two block on different people: a `paused:` wait is owed by an external dependency and asks the reader to confirm it still holds, while a `captain-held:` transfer is owed by the captain reading the recheck and asks them to answer or release the hold. A hold is not rechecked at all while the away-posture record exists, as on every other captain-held path, and that absorb arms no throttle so the recheck is owed in full on return. A declared clearing time that has already passed stops counting, and a lane that never declared one keeps the identical escalation schedule, reason, count and demand-deep-inspection wording, so detection and its worst-case time are unchanged. The deferral restarts the idle timer rather than cancelling it, so a lane that stops waiting escalates again within one threshold. A lane quiet because its own validation run is parked at a gate awaiting a human decision is deliberately out of scope: reading that state needs a signal carrying who the wait is on and what clears it, rather than one inferred from a parked verdict that also covers gates awaiting the crewmate itself. Tests pin both directions for each case and were each confirmed to fail with the consult removed. * no-mistakes(document): docs: honour declared waits in stale-escalation docs * fix(bin): report verified PR state for passed runs (#4624) * fix(bin): derive passed PR state from PR record A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists. For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim. Fixes #4607 * no-mistakes(review): Add bounded GitLab merge-request state reads * no-mistakes(review): Preserve network-free inactive crew-state scans * no-mistakes(document): Document PR record readers in shared library * fix: restore published contribution follow-up (Fixes #4469) (#4627) * fix: restore published contribution follow-up (Fixes #4469) * fix(review): Fix contribution freshness and merge actor routing * fix(review): Restore issue triage and scope contribution follow-up * fix(test): test: assert one wake per contribution signal * fix(document): Document contribution follow-up * fix: restore truthful terminal delivery evidence * fix(review): Disclose unsupported contributions and deduplicate watcher wakes * fix(review): Preserve unmeasured unsupported contributions across Bearings * fix(review): Deduplicate shared contribution wakes and isolate diagnostics * fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower * fix(bin): make remote report transfers explicit and fail-open (#4658) * fix(bin): make a remote-reply document gap self-clearing and re-attemptable A remote mate's undelivered document raised a keyed `blocked` decision that nothing could ever resolve, and any `data/*.md` substring in any mirrored line was an unconditional fetch instruction. A mate announcing a report it had not written yet therefore manufactured a permanent, factually false blocker, and its own explanation of the false alarm manufactured more. The reader has no permanence vocabulary: a report still being written refuses exactly like a path that will never exist. So an undelivered document is now a durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`, re-attempted on the next delta and on the channel's own quiet poll, and retired with a matching `resolved` line naming the local copy once it arrives. The cursor still advances and no delta stalls on one bad pointer. Only a structured `report=data/....md` pointer now offers a document, so a path merely mentioned in prose - including one under another home's mirror tree, which is provably not that mate's to serve - is never fetched. Offers are deduplicated across the whole delta, the escalation names each missing document once and carries the reader's own reason instead of discarding it, and a strictly increasing notice ordinal keeps a later escalation from being swallowed as duplicate bytes. A mirrored line still lands once whichever pointer form it was first written under. * no-mistakes(review): Require structured pointer token boundaries * no-mistakes(review): Unify boundary-safe pointer extraction and rewriting * fix(bin): identify a mirrored line independently of its delivery state Two defects in the boundary-safe pointer work. The at-most-once check compared only the all-remote and all-local renderings of a line, so it could not recognize a mixed one. A line offering two documents where only the first was deliverable mirrored as local-plus-remote; once the second arrived, a cursor-loss whole-log recapture rendered the same line all-local, matched neither alternate, and mirrored a second time. A line's identity is now the canonical form every boundary-valid pointer would take once delivered, derived by the same parser that does extraction and rewriting, so it no longer depends on which documents happened to be deliverable at the time. The pointer map was passed to awk through the process environment. A delta may carry up to the configured 1 MiB bound, and an expanded map of delivered pointers can exceed the platform's exec argument limit, so awk would fail to start; because no caller checked, the empty result would have been appended as blank lines while the cursor advanced past dropped status content. The map now travels in a file, and every call site checks the exit status and stops the ingest rather than committing a delta it could not render. Both passes now run once per stream instead of twice per line. * no-mistakes(review): Abort ingest when document pointer extraction fails * no-mistakes(review): Exclude structured cross-home pointers from document transfer * fix(bin): fail open on an undeliverable remote document instead of tracking it Narrow the remote-reply document fix to the scope the diagnosis actually requires, as decided after measuring a simpler alternative. A document the reader cannot deliver now fails open. The mate's line is mirrored with its own pointer, the cursor advances, and one unkeyed note carries the reader's reason. A note never enters the open-decision fold, so it cannot stand open the way the original keyed block did - which removes the never-clearing false blocker by construction rather than by resolving it. That makes the durable self-clearing obligation unnecessary, so it goes: the per-mate pending-documents record, its notice ordinal and resolved announcements, and the poll-side retry. Canonical line identity goes too, and with it a way to silently drop a genuine status line; mirroring is back to at-most-once on exact bytes. The cross-home exclusion goes as well: under fail-open a cross-home report= either fails harmlessly or is a nested remote report this mate genuinely holds, which is now relayed again. Kept: fetching only on a structured report= pointer, the boundary-correct parser, the file-based rewrite map, and checked extraction and rewrite exit status. The parser now scans behind a sentinel byte so a rejected candidate can no longer give the text right after it a false leading boundary. The reported incident is covered end to end: a report path announced in prose before it exists raises no decision, and the report still arrives through the ledger publisher's structured offer once written. * no-mistakes(review): Preserve source-line identity across remote reply replays * no-mistakes(document): Document remote reply transfer and replay semantics * no-mistakes(lint): Fix staging truncation lint checks * fix(calm): preserve substantive mid-turn responses (#4655) * Preserve substantive Calm mid-turn text * no-mistakes(review): Distinguish newline-preserved replies from short narration * no-mistakes(document): Document Calm mid-turn preservation boundaries * no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check` * fix(bin): preserve PR merge polls across volume remounts (#4656) * fix(bin): re-record PR poll identity after a volume device renumber (Fixes #4260) A volume remount can renumber the state filesystem's st_dev while every inode and byte stays the same; APFS does this across a reboot. A poll registration records its sidecar and check as device:inode, so every poll armed before the remount failed strict validation and the watcher refused all of them as unauthenticated state checks until each was re-armed by hand. There are two device comparisons. fm_pr_private_file_valid compares a live file's device with the state directory's device read in the same invocation: it refuses a file that is not on the state directory's own filesystem and already survives a renumber, so it is unchanged. The registration's recorded identity versus the live identity (from #556, reused by the #932 retirement receipt) binds the registration to the exact files published in its own transaction; its device part is what breaks. When strict capture fails, the watcher now proves the device is the only difference: every other artifact check passes (template bytes, both hashes, private mode, single link, live device, metadata), both recorded identities name one device, and each recorded inode equals its live inode. Only then, under the task's control lock, does it rewrite the two identity lines, repeating the whole proof and comparing the registration's file identity and bytes just before the rename, and then capture strictly again. A swapped, altered, re-moded, relinked, split-device, or foreign-device artifact still fails a proof and is still refused, and a pending retirement receipt blocks the rewrite. Reproduction: on macOS a poll armed on an APFS disk image that was detached and re-attached behind another image moved st_dev 16777239 -> 16777243 with inodes, bytes, mode, and link count unchanged; the real watcher refused it on main and reports its merge with this change. The portable regression test rewrites a real registration's recorded device and drives the watcher. Not changed here: the status presentation cursor keys rows by its own device:inode identity in bin/fm-classify-lib.sh, a different helper that needs its own fix; a retirement receipt left by a reboot between its publication and removal still names the old device and stays refused; custom check trust binds only a content hash and is unaffected. * fix(review): Serialize PR poll publication writers * fix(review): Bound PR poll publication lock scope * fix(bin): keep contribution records when the poll budget runs out (follow-up to #4627) (#4661) A budget that expires partway through an observation no longer records an error or prints the unavailable wake; the URL keeps its prior record and is observed first next poll. forge() flags budget exhaustion at the point it refuses, or when a read is killed at the budget's own deadline, so a genuine forge failure still records the error and wakes. Each distinct URL is now observed once per poll and applied to every owning task. * fix(bin): clear parent pending-replies on local secondmate retirement (#4680) * fix(bin): clear parent pending-replies on local secondmate retirement Local secondmate teardown left resolved parent pending-reply records behind after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced retirement while any reply for that id is still unresolved, and delete every matching record plus its delivery confirmation after a successful local or remote retirement, matching the remote cleanup path. * no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup * no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung * no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern * no-mistakes(review): Pending-replies Basename und corr_id abgleichen * no-mistakes(document): Clarify forced retirement pending-reply cleanup --------- Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc> * fix(bin): accept Orca's composite worktree id when tearing down a task (#4677) * fix(bin): accept Orca's composite worktree id at teardown Teardown refused every Orca-backed task because the endpoint validator checked orca_worktree_id with the simple-atom rule meant for tmux-style window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca returns that id as `<orca id>::<absolute worktree path>`, so the colon and slashes in every real value made validation fail and finished Orca tasks could never be cleaned up. Validate the field as the composite it is: both halves of the first `::` split present, the path half absolute, and no embedded newline, carriage return, or tab. The terminal field keeps the atom check, which is correct for it, and no other backend's validation changes. The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca never returns, which is why the suite passed a check the real value fails. They now carry the composite form, so the tests exercise the real value. * no-mistakes(document): name Orca's repo id in the composite worktree id * no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points * feat(bin): add opt-in typed dispatch resolution (#4692) * feat(bin): add opt-in typed dispatch resolution through typesafe.ai Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or scout profile from a written brief with typesafe.ai's System One model: one Choice question over the rules' `when` texts, then the confidence floor, the rule's `approval` and `floor`, each profile's `provider` and `floor`, one quota-axi snapshot, and the spendPriority argmax all in code. It is off unless TYPESAFE_API_KEY is in the environment or the home's gitignored .env; off means one stderr line, exit 0, and no network call, so firstmate dispatches exactly as before. The key reaches curl on a file descriptor, never argv. Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates the four new optional dispatch fields. Document the schema, the operator contract, the AGENTS.md intake step, and the live and benchmark evidence. * no-mistakes(review): Harden typed dispatch resolution and quota bounds * no-mistakes(review): Validate dispatch floors and ranking evidence * no-mistakes(review): Tighten dispatch response and floor evidence * no-mistakes(review): Neutralize none matching and resolve defaults locally * no-mistakes(review): Preserve providerless profiles outside typed resolution * no-mistakes(review): Validate response usage and reject duplicate profiles * no-mistakes(review): Escalate unverifiable floors and validate probabilities * no-mistakes(review): Validate probability mass and unknown profile floors * no-mistakes(review): Simplify resolver interface and preserve fallback routing * no-mistakes(review): Fix constants and rank partial quota evidence * no-mistakes(review): Add authoritative provider mapping and enforce explicit providers * no-mistakes(review): Declare provider for documented Pi profile * no-mistakes(review): Validate provider identifiers and support Gemini dispatch * no-mistakes(review): Strictly anchor provider identifiers * no-mistakes(review): Validate selectors and preserve fallback candidate evidence * no-mistakes(review): Gate typed validation and harden resolver evidence * no-mistakes(review): Preserve opt-in routing and harden candidate evidence * no-mistakes(review): Prioritize known exhaustion over quota uncertainty * no-mistakes(review): Isolate API secrets and preserve no-key diagnostics * no-mistakes(review): Fallback safely when dispatch rules are absent * no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets * no-mistakes(document): Document typed dispatch safety and fallback behavior * fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753) * test: reproduce buried status declarations in shared readers * fix: share status event reads and preserve open blockers * fix: retain terminal scout and ship status declarations * no-mistakes(review): Fix status chronology, legacy completions, and reader performance * no-mistakes(review): Share terminal decision reconciliation across fleet snapshots * no-mistakes(review): Unify terminal supersession across cached folds and consumers * no-mistakes(review): Filter per-key status history while preserving terminal chronology * no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells * no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses * no-mistakes(document): Document latest-event status read and kind-scoped fold cursor * no-mistakes(lint): Quote literal done in test for-lists for SC1010 * ci: expect 19 snapshot/fleet-view tests This branch adds a fleet-snapshot regression, so the stock macOS Bash lane's hardcoded guard of 18 'ok - ' lines fails on the new count. Bump the guard and its message to 19. * no-mistakes(review): Restore multiline child outcome reporting * no-mistakes(review): Select ledger terminal events through bounded shared reader * no-mistakes(review): Report newest open decision instead of preferring blocked * no-mistakes(review): Require colon before ship/scout terminal supersession in fold * no-mistakes(review): Gate socket-down override on latest event; drop lock matrix * no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions * no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold * no-mistakes(test): Update fleet-view expectations to newest-open-decision rule * no-mistakes(document): Align status-read docs with fold-resolved crew state * no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers * no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree * test: fold terminal-cleanup snapshot coverage into the completed-scout case Keep the ship/scout/secondmate supersession assertions without adding a nineteenth top-level fleet-view test, so CI can stay at the upstream suite count. * no-mistakes(document): Clarify socket-down override expiry in architecture doc * ci: retrigger flaky contribution check * fix(bin): launch codex crewmates with codex's hook layer disabled (#4689) * fix(spawn): launch codex crewmates with codex's hook layer disabled A freshly launched Codex worker never reached its instructions. Codex stopped it on an interactive "Hooks need review" modal whose selection sits on "Review hooks", which is neither trusting nor declining. Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow navigation, so the selection cannot be moved, and pre-accepting the prompt by writing Codex's own trust store would record an operator consent that was never given. The hooks are the machine's own ~/.codex/hooks.json plus any project's .codex/hooks.json. A crewmate needs neither: its turn-end signal is the -c notify= program on the same launch, and Firstmate's project hooks are primary-session infrastructure that stands down in a child worktree. Crewmate and scout launches now pass --disable hooks. That is the opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted hooks; disabling the feature runs none of them and leaves the operator's ~/.codex untouched. An unknown feature name is a hard Codex error, so a release that drops the flag fails the launch loudly instead of silently restoring the modal. A secondmate is a primary in its own home and keeps the project hooks its turn-end guard and session-start digest ride on. Verified on codex-cli 0.151.0: the modal is gone and the turn-end notification still lands. This unblocks the second review that every finished pull request is supposed to get. Fixes kunchenguid/firstmate#4673 * no-mistakes(review): Fix contradictory hook count in Codex verification record * fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710) * fix(bin): settle terminal contributions and wake once per read-failure episode A contribution whose last good observation is merged or closed is final: poll no longer re-reads it, projection keeps it fresh, and a stale error recorded beside it is cleared once. A genuine forge-read failure on an open contribution still records its error on every cycle but prints the unavailable wake only when it starts a failure episode; a successful read ends the episode. Open PRs linked from done tasks keep being observed. The false unavailable beside a complete observation was budget exhaustion mid-observation, already fixed by #4661. * fix(review): Settle terminal contribution owners * fix(review): Deduplicate shared contribution failure episodes * fix(test): Preserve settled terminal contribution records * fix: select authoritative no-mistakes runs (#4476) * fix(crew-state): select authoritative validation runs by identity Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row. Refs: https://github.com/kunchenguid/firstmate/issues/3215 * fix(review): Resolve same-branch run identities beyond capped history * fix(review): Fix run-selection compatibility, races, and worker-state fallbacks * fix(review): Limit run validation to the requested branch * fix(test): Anchor AXI fixtures and document remaining live evidence gaps * fix(document): Clarify run selection documentation and capture ownership * fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation * fix: distinguish captain outcomes from no-op updates (#4738) * fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work MAIN answered a supervision-branch outcome for completed captain-requested work (implementation done, PR ready for review and merge approval) with "Captain, shipshape.", reading section 9's no-action reply as covering it and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no captain-facing response is owed". Section 9 now limits the shipshape reply to true no-ops (idle re-read, empty heartbeat, consequence-free acknowledgement) and requires a short outcome response naming what finished and what word is needed whenever requested work finishes or a result needs the captain's word, even when a transcript entry already shows the substance. The Pi protocol's re-emit rule now says it bounds repetition only, and carries a worked example of the ready-for-review outcome whose correct processing turn a shipshape reply fails. No executable contract evaluates the content of MAIN's captain-facing reply, so the regression is the protocol example in the owner doc rather than a text-match test. * no-mistakes(document): Clarify captain-facing outcomes versus no-ops * docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line The document step condensed the Pi protocol's re-emit rule and dropped the worked example of a finished, ready-for-review outcome whose correct processing turn a "Captain, shipshape." reply fails. That example is the contract's regression: no executable contract evaluates the content of MAIN's captain-facing reply, so the owner doc's example is the test case. Restore it directly under the re-emit rule, prefixed as a regression example that is kept verbatim and never condensed or summarized away. * no-mistakes(review): Clarify captain outcome and decision-word requirements * no-mistakes(document): Clarify captain-facing completion outcomes * docs(pi): require the PR URL in the visible captain-facing outcome reply Captain review on the regression example: drop the sample reply string and say only that the ready-for-review outcome requires relaying a captain-facing outcome response, not just "Captain, shipshape.". Fold in the visible-PR-handoff failure seen this session: after the branch outcome reporting this fix green, MAIN's visible reply was only "Awaiting your merge call." with no PR URL, leaning on the dim anchor. Section 9's URL rule now also covers a review or merge ask and names the visible reply as where the URL goes, sourced from the ready status, pr= metadata, or the supervision branch's summary and never left to a transcript entry. The Pi protocol adds the same-way failure and places the captain-facing text in the final visible assistant reply after the fm_branch_processed call, because Calm hides assistant text emitted in the same step as a tool call as a working note. Investigation verdict, evidence in the PR comment: no recent PR caused the handoff failure; Pi has hidden same-step pre-tool assistant text since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and #4658 touched only remote report transfer. * no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs * no-mistakes(document): Clarify captain-facing supervision outcomes * docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule The consolidated section 9 URL rule narrowed its trigger to a review or merge ask, dropping the "whenever a PR is mentioned" catch-all from #3648 that keeps every PR URL copied from a durable record and never assembled from memory. Restore that trigger as a union with the review or merge ask so the one consolidated rule covers both. * fix(bin): let non-owner Claude Stops exit safely (#4777) * Fix foreign-owner turn-end supervision loop * no-mistakes(review): Scope foreign-owner safe exit to Claude guard * no-mistakes(document): Document Claude foreign-owner safe exit * fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778) Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty indexed array as an unbound variable and aborts the shell. In signal_turnend_panes_churned() the missing_keys loop was reachable with an empty array whenever every churned key already held a fresh .churn-since-* marker (a second churning turn-end inside an open deferral window), so each watcher cycle died about half a minute in and supervision restarted endlessly. The created_keys rollback loops had the same latent crash on their error paths. Audit of bin/ for the same pattern found one more confirmed-reachable case: remote_handoff's noncanonical-body scan iterates to_move, which is empty when a retried remote handoff finds every key already staged in the outbox. All other "${arr[@]}" sites are either count-guarded, guaranteed non-empty by construction, or unreachable while empty. Guard the three reachable expansions with the repo's existing "${arr[@]+...}" idiom. New regression test drives a real watcher through the all-marked churn path; the macos-stock-bash CI lane runs it under real /bin/bash 3.2 via FM_TEST_ONLY. * Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783) The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: require complete captain-facing final responses (#4779) * docs: require complete final responses across harnesses * no-mistakes(document): Document complete final replies for Grok Bot * docs: point Grok replies to the shared contract owner * no-mistakes(review): Clarify final recap without batching decision asks * fix: preserve substantive mid-turn text in Pi Calm (#4788) * fix(calm): preserve substantive Pi mid-turn text * no-mistakes(review): Preserve substantive Pi Calm text per block * no-mistakes(test): Cover shared Calm preservation boundaries behaviorally * no-mistakes(document): Consolidate Calm preservation documentation * fix: harden mail checks and rebalance full-coverage CI (#4800) * Improve CI reliability and rebalance full-coverage validation * no-mistakes(document): Clarify lint partition documentation * fix(bin): answer Kimi 2.0.0 folder-trust dialog during spawn (#4799) * Handle Kimi workspace trust dialog * no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers * no-mistakes(review): Gate Kimi ready on any trust marker and clean captures * no-mistakes(review): Read visible pane for Kimi trust and ready gates * no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate * no-mistakes(review): Harden Kimi viewport capture and trust dialog detection * no-mistakes(document): Document Kimi spawn refusal on cmux and Orca * fix(bin): report a dead-agent record once instead of escalating forever (#4775) * fix(bin): report a record whose agent is gone once instead of escalating forever The wedge escalation path never asked whether there was still an agent to be wedged. A wedge is something stuck that might recover, so re-alarming it earns its cost; an agent that is gone never moves again, its pane never churns, the idle timer never resets, and the escalate path clears its own timer and re-arms with nothing bounding the count. Observed on a live fleet: two finished lanes reached 226 and 203 consecutive escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400 notifications a day from two lanes with no agent running at all. On one, fm-control.sh exit answered already-stopped and fm-crew-state.sh read "failed - run failed". Closing the Herdr pane did not stop it either: with the pane genuinely gone and herdr pane read returning pane_not_found, the count kept climbing, because the poll is driven by the record's window= line rather than by the pane. The cost is not the repetition but that it drowns the alarms that matter. fm_backend_agent_state already separates a thinking agent from a gone one at process level. In the branch that was about to escalate, read it once and treat only its two recovery-grade verdicts - dead (endpoint present, no agent in it) and missing (endpoint authoritatively absent) - as proof, reporting that record once and not re-escalating it while it stays that way. Every other verdict, including alive, ambiguous, unreadable, unverified, and a read that failed outright, keeps the identical schedule, reason, and escalation count, so a genuinely wedged live agent is unaffected. The probe costs at most one backend read per window per threshold, the same budget the declared-wait consult and the worktree write probe already take. The report decides nothing about the record's fate: both lanes still held unlanded work and teardown refusing them was correct, so retiring, relaunching, or cleaning up stays with the supervisor. The once-only marker is owned entirely by that function and is dropped by the same read the moment the endpoint stops reading gone, so a replacement launched into the same window escalates normally and its own later death is reported again. Related, and not closed by this: #4412, #4482, #4316. Tests drive the real watcher against a record whose endpoint does not exist and pin both directions: dead and missing report once and never advance the count across later thresholds, while alive, ambiguous, and unreadable endpoints keep escalating with the identical reason and a climbing count. * fix(bin): bind the once-only dead report to the pane it reported Review of the parent commit found a reachable sequence where a later death in the same window lost its promised report. The marker was keyed on the verdict string alone and dropped only when a threshold probe read a non-gone verdict, but probes run only at thresholds: a replacement launched into the same window that dies without ever being probed alive - it crashes at startup, or works and then crashes - was absorbed by the previous death's marker. The pane's first sight yielded only the generic stale wake and every later threshold matched the stale marker, so the second death never got the detailed once-report that both the function's own comment and docs/architecture.md promise. Record the verdict together with the pane hash it was reported for, and absorb a repeat only while both still match. A replacement churns the pane, which resets the stale suppressor, wedge timer, and escalation count while no reset site touches this marker, so the pane half is what tells the second death apart from the first. The live-probe drop stays as it was. Clearing the marker at those reset sites instead would re-open unbounded re-alarming for a dead pane whose display ever ticks, which is the exact defect the parent commit exists to close. The noise bound is unchanged: an unchanged dead pane still absorbs on every later threshold and never advances the escalation count, and every verdict short of proof still escalates exactly as before. * no-mistakes(review): Key the dead-record once-marker on the busy incarnation token * no-mistakes(document): Document dead-record escalation cap in stale-pane config entry * no-mistakes(document): Add busy-state inventory line to AGENTS.md * no-mistakes(document): Document dead-record probe on busy-turn-bound wedge path * fix(bin): create captain-hold rows when Beads requires due (#4854) Captain holds have no due semantics and are a hold kind, not a Beads issue type. The create path now waives due.required and maps to native type task. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: disable compact adviser for spawned agents (#4877) * feat(bin): launch every spawned agent with the compact adviser disabled Every crewmate, scout, and secondmate Firstmate launches now starts with COMPACT_ADVISER_DISABLE=1, on a fresh spawn and on a relaunch alike, so an unattended session never activates the compact adviser. The value is unconditional: no configuration file gates it and there is no override, unlike the trace carrier beside it. Three carriers deliver it, because no single one covers every launch shape. The pane shell receives an export beside GOTMPDIR, so the agent's own children inherit it too. The launch command carries an explicit assignment, prepended outermost so it wins over any ambient value the pane already held. The cleared launch environment sets it again at the `env -i` boundary and keeps COMPACT_ADVISER_DISABLE in the fixed operational floor, which is what preserves the switch when config/launch-env-allowlist empties the environment, and what delivers it on a remote host that never had the value. bin/fm-control.sh relaunch, the bootstrap secondmate relaunch, and the remote secondmate transport all rebuild their launch through bin/fm-spawn.sh, so they inherit the same floor. The captain's own primary session is untouched. The two new suites drive the real spawn and then execute the launch command the pane actually received, with the harness replaced by a probe that prints its own environment, rather than matching script text. They cover ship and secondmate launches with the allowlist absent and enabled, the pane export and its ordering, fm-control.sh relaunch, and the full parent to remote-host chain. * no-mistakes(review): Export compact-adviser disable across compound launches * no-mistakes(document): Document spawned-agent compact-adviser environment guarantee * fix(bin): preserve Claude lock ownership after helper recycling (#4894) * fix(bin): let a background Claude session keep owning its session lock Session-lock ownership was decided by process ancestry alone. Under an unattended Claude session the model loop runs in a transient bg-spare bridged to the front-end by a shared daemon; when that bridge is recycled the contiguous claude-named ancestry from a hook to the recorded owner breaks while the owner pid stays alive, so the Stop auto-arm stood down as a foreign live owner, the turn-end guard ended every turn with its read-only diagnostic, and fm-lock.sh refused - a self-sustaining outage until restart. Ownership is now ancestry membership OR a trusted same-session id, never id-first: - fm-session-lock-lib.sh accepts CLAUDE_CODE_SESSION_ID only when CLAUDE_PID is a Claude-shaped member of the current contiguous run, compares it against the id recorded in state/.lock-session, and requires the recorded pid to still be a live harness. No id, no sidecar, an untrusted id, a different id, or a dead recorded pid leaves the ancestry verdict unchanged. Ids are never read from ps argv. - fm-lock.sh accepts a same-session holder at both refusal sites, writes, refreshes, and clears the sidecar only under its claim lock (including the early already-mine exit, skipped only while the deferred startup sweep leases that lock), keeps it byte-identical across a same-session confirmation, records CLAUDE_PID on lock line 1 for a session with a trusted id so a shared daemon or front-end that outlives the session never keeps a dead session's lock alive, never rewrites a live line 1 on a same-session confirmation, and names the recorded id in the live-owner refusal. - The .lock line-1 format is unchanged, so every reader that takes the whole first line as the pid keeps working; the guard's foreign-owner exit is unchanged and inherits the fix through the shared predicate. Tests: the ancestry suite drives the ancestry and id signals apart in a deterministic process table (asserting the divergence) and runs a real orphaned front-end/daemon/pty-host/spare tree through six phases with the real lock, auto-arm, and guard scripts; the foreign-owner repro keeps its negative control and adds a same-id positive control. Disclosure: no live unattended Claude background session ran on the verifying machine. The topology is documented by the real process listings in #3902, #2314, #3398, and #4066; coverage is the structural predicate plus the executable fixtures, not a live pass. Residual: bin/fm-sessionstart-nudge.sh keeps its own private ancestry walk (it only decides whether to print a nudge) and may nudge on a resume in the recycled case. Out of scope, deliberately: no structured lock format, no guard budget changes, no daemon-identity rejection, no fork lineage. * no-mistakes(review): Wait for claim lock; revert failed sidecars * no-mistakes(review): Revalidate ownership after wait; restore sidecars * no-mistakes(review): Roll back sidecar by publication phase * no-mistakes(review): Restore sidecar only if lock line is unchanged * no-mistakes(review): Trust session ids without a spelling allowlist * no-mistakes(review): Disarm sidecar rollback before backup cleanup * no-mistakes(document): Updated session-lock ownership documentation * feat: park main under the away posture on Pi (#4889) * feat: park main under the away posture on Pi While the away-posture record exists on a Pi primary, the supervision branch takes every actionable wake, no processing turn opens on main, captain rows accumulate for the return brief, and main's standing authority relocates to the branch through the existing guarded scripts. - lib/fm-branch-dispatch.ts: read the record at every routing decision; while it exists claim check, decision-owned, and heartbeat rows too, keeping the two broken-queue vetoes; expose checkSeqs so a claimed check row lifts task scoping. - fm-primary-pi-watch.ts: offer every actionable row under the record; a declined wake and every watcher-failure alarm still reach main. - fm-branch-supervision.ts: drop the legacy .afk decline; append a fixed POSTURE: AWAY tail carrying the record's read-back verbatim per wake; open no processing request while the record exists, re-checked immediately before a request would open and at every run boundary; present the accumulated rows at the first run boundary after archive. - fm-lease-lib.sh: fm_lease_forbid_branch passes the branch for opted-in actions only while fm-afk-contract.sh validate succeeds on a confirmed live record; PR merge, fresh spawn, and decision answer opt in, local landing never does. - fm-send.sh: a --resolve-key naming an open needs-decision or captain-held task is a decision answer and meets the partition; blocked: keys stay steering. - fm-spawn.sh: enforce the record's spend cap for a fresh ordinary spawn by either actor; relaunches and secondmates exempt. - fm-branch-prompt.sh: fixed Postures section and the verbatim ask-user-authority policy; the prefix stays byte-stable. - fm-afk-return.sh: count what the away session handled from the store. - docs, afk skill, AGENTS.md stub: main parked on Pi, green merge gate absolute while away. - tests: watcher and branch extension suites, fleet-record, merge, and decision-answer suites cover the relocation, the vetoes, the tail, the parked processing turn, the cancellation, the re-presentation, and the spend cap; dated live-guard evidence recorded. * no-mistakes(review): Refuse branch merge after preflight archive race * no-mistakes(review): Fix away wake, spawn, and processing races * no-mistakes(review): Suppress parked processing; narrow away-only rejection * no-mistakes(review): Abort dedicated processing; gate branch spawn once * no-mistakes(review): Stamp away-only on the dispatch offer * no-mistakes(review): Treat invalid away records as spend-cap absence * no-mistakes(review): Drop spawn test hook; abort processing-opened runs * no-mistakes(review): Bind abort to opening prompt; cap-read absence * no-mistakes(review): Limit away branch spawn to queued work only * no-mistakes(document): Correct AFK posture documentation * ci: standardize workflow timeouts into three tiers (#4910) * ci: simplify CI job timeouts to a three-tier policy Replace the scattered per-job timeout values (10m parallel, 25m lint, 30m serial, 10m macOS) with three readable tiers, each a hang tripwire with headroom rather than a packing estimate: - fast (5m): coverage guard, repo invariants, timing aggregate - normal (30m, one shared budget): lint partitions, portable parallel shards, portable serial shards, macOS stock Bash - heavy (Herdr only): 20m step tripwire on the family run so always() cleanup still runs, under a 75m job-level last-resort backstop The workflow's header comment states the policy and points at docs/fm-test-portable-shards.md "Timeouts", which now owns it, and each job names its tier beside timeout-minutes. tests/fm-ci-workflow.test.sh asserts the policy against the parsed workflow instead of the old per-job minute values: every job joins exactly one tier, exactly three distinct job-level values exist, the fast tier stays within 5-10 minutes, the normal budget stays at least double the modeled parallel lane sum reported by fm-test-run.sh --check-coverage, and the Herdr step tripwire stays below its job backstop with an always() cleanup after it. Concurrency supersession, shard counts, lane membership, and fail-fast settings are unchanged. * no-mistakes(review): Decouple the normal timeout from packing estimates * no-mistakes(review): Assert Herdr teardown follows the family run * no-mistakes(review): Pin Herdr family-run timeout to 20 minutes * no-mistakes(review): Ignore comments when identifying Herdr steps * no-mistakes(review): Identify Herdr steps by declarative ids * no-mistakes(document): Clarify authoritative three-tier timeout policy * fix(bin): keep supervisor status closes from waking the same home (#4895) * fix(bin): keep supervisor status closes from waking the same home A drain that already folded OPEN DECISIONS has presented those bytes even when the watcher has no matching seen marker. Treat that fold, and the presentation cursor, as known so the bookkeeping close stays quiet while later worker lines still signal. * no-mistakes(review): Keep folded worker failures waking past supervisor closes * no-mistakes(review): Wake on unlisted folded worker lines; batch multi-key closes * no-mistakes(review): Stop folded worker resolved lines from counting as already read * no-mistakes(document): Correct self-announced close marker contract in docs * fix(bin): stop labeling Herdr as experimental (#4972) * Stop steering operators away from Herdr * no-mistakes(review): Neutralize remaining Herdr opt-out documentation wording * fix(bin): treat a live no-mistakes run as current after rebase (#4973) * fix(bin): treat a live no-mistakes run as current after rebase A running run on the task's branch is authoritative regardless of head. Matching only the local head made a rebased in-flight run look failed. * no-mistakes(review): restrict coarse live-any-head to foreign-branch answers * no-mistakes(review): reject gate-parked runs from the executing predicate * no-mistakes(review): hoist gate-marker patterns into single run-lib owner * no-mistakes(review): require live daemon for head-free run binding * no-mistakes(review): require answered daemon-down before unbinding live runs * no-mistakes(review): extend daemon guard to anchored continuation routes * no-mistakes(review): delete live-any-head; restore dead-daemon verdict * no-mistakes(review): keep parked gates parked; name dead daemon everywhere * no-mistakes(review): set dead-daemon verdict instead of emitting early * no-mistakes(review): align selected route with legacy dead-daemon handling * no-mistakes(review): drop unproven-record binds; narrow coarse gate reading * no-mistakes(review): narrow header, drop vestigial guard, retarget tests * no-mistakes(review): revert coarse gate override; require answered-down probe * no-mistakes(review): cache one daemon probe; stop duplicating run id * no-mistakes(review): restrict coarse dead-daemon verdict to moved-off rows * no-mistakes(review): delete coarse dead-daemon extension and gate note * no-mistakes(review): delete remaining coarse dead-daemon block and stale docs * no-mistakes(document): document rebase-safe live-run bind and unverified-record verdict * fix(bin): prevent long worker launch command truncation (#4994) * fix(bin): stage the launch command in a private file and type a short source line A long launch line typed while the fresh pane shell is still busy waits in the terminal's canonical line buffer, which drops input past about 1,024 bytes on macOS, so the pane was left at an unfinished command with no agent running. fm-spawn now writes the assembled command to the task's own temp root under umask 077 and types only a short line that sources it. Refs #4559 * fix(bin): keep the per-task temp root private before staging the launch command The root lives at a predictable path under /tmp and now holds the whole launch command. Create it with mode 0700, refuse one that already exists as anything but a directory owned by this user that nobody else can write, and tighten an owned one, so no other local user can plant or swap the staged file. Refs #4559 * fix(bin): enforce private staged launch file mode * test(spawn): cover long staged Claude launches * no-mistakes(review): Namespace launch files and prove truncation staging * no-mistakes(review): Use immutable per-spawn launch filenames * no-mistakes(document): Document staged launch delivery safeguards * no-mistakes(ci): Updated eight behavior tests/fakes to execute or inspect immutable staged launch files instead of expecting inline launch commands. This restores Muse, secondmate lifecycle/restart, remote trace/parent binding, compact-adviser, and Orca coverage. All affected tests, dispatch-profile regression, fixture tests, syntax checks, ShellCheck, and git diff checks pass --------- Co-authored-by: Vytautas Stankus <svycka@gmail.com> * test: authorize isolated Herdr lab validation (#4998) * Add isolated Herdr runbook to test instructions * no-mistakes(review): Drop substring matching from test.instructions contract * no-mistakes(review): Assert commands.test key absence in YAML * Drop unit-first sentence and instructions contract test Captain-scoped follow-up on the Herdr-lab test.instructions ship: keep the lab safety runbook only, and leave the no-mistakes contract test focused on commands.test absence. * docs(vision): accept vendor-semantics and 9k AGENTS ceiling (#4873) (#5001) * docs(vision): accept vendor-semantics and 9k contract-ceiling amendments (#4873) Replace the pixels-of-today's-UI rule with a quarantined, version-pinned surface-adapter exception recorded as standing debt. Cap the always-loaded contract at 9,000 words and require prune-or-trigger before a crossing change lands. Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com> * docs(vision): restore accepted three-sentence vendor-semantics form (#4873) Replace the compressed paraphrase with the issue's accepted wording: a named quarantined version-pinned adapter, expected to break, recorded as standing debt that never hardens into a shared contract. Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com> --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com> * feat(bin): defer the wedge escalation for a lane parked at a supervisor-owed gate (#4974) * fix(watch): recheck a gate awaiting a human instead of wedge-escalating it A lane whose validation run is parked at a gate waiting on a human decision is correctly quiet, but nothing in its status line says so: the evidence is the pipeline's own gate state rather than anything the worker wrote. The wedge timer read that silence as a suspected wedge and climbed the escalation ladder for as long as the wait lasted, and each escalation cost a supervising turn. The landed declared-wait consult does not reach it, because a live ordinary crewmate never reports a declared pause, and raising FM_STALE_ESCALATE_SECS would delay genuine wedge detection for every lane by the same amount. The threshold now reads a second, independent record when the status line accounts for nothing: whether the crew's current state is a gate whose answer is owed by a human. That is minted only from the gate's own findings table, by a row whose `action` column is exactly `ask-user`, located by position out of the table header the way nm_gate_step_row already reads its row - never searched for over the run payload, where a finding's free-text description or a branch name satisfies a search just as well. A gate awaiting the CREWMATE's own answer keeps the unchanged escalation schedule, reason and demand-deep-inspection wording, because a crewmate that goes quiet before answering its own gate is exactly the wedge the ladder exists to catch. Each kind of wait now carries the human it is on, the action that clears it, and whether that human is the captain as data alongside the verdict, rather than as wording chosen per branch where the recheck is written, so the deferral cannot word one kind of wait as another and a new kind cannot ship without deciding all of them. A parked gate has no written record of when its wait began, so its recheck publishes no wait age at all rather than one read from the quiet window this deferral resets on every pass, which would report the same small number for a gate of any age. Like every other captain-facing recheck here it is absorbed in silence while the away-posture record exists, arming no throttle, so the recheck is owed in full the moment the record is archived. The consult runs only in the at-threshold branch that was about to escalate, beside the worktree walk already there, and only for lanes whose status line explained nothing. Closes #3055 * no-mistakes(review): require an unanswered decision before deferring a parked gate * no-mistakes(review): reset the away-silenced timer, fail-safe findings parse, US-joined wait records * test(watch): pass the pane hash wedge_timer_check now takes Upstream gave wedge_timer_check a sixth <pane-hash> argument for its dead-record probe. The malformed-wait-record rounds drive the real function directly, so they pass one, and stub fm_backend_agent_state to a live agent so the probe that runs after a refused deferral keeps the unchanged ladder rather than reading a backend the child shell has none of. * no-mistakes(review): Bind parked-gate wait to its run, owe it firstmate * no-mistakes(document): correct wait-kind count, crew-state reader scope, gate-key coupling * feat(watch): make the parked-gate wait deferral opt-in The wedge timer deferring a lane parked at a validation gate is new supervision behaviour rather than a restored one, and it decides which lanes give up the escalation ladder, so it now ships as a default-off per-home option instead of changing every home on upgrade. config/wedge-defer-parked-gate arms it. The flag is read before the decision fold, so an unconfigured home spends no fold or current-state read, writes no record, and keeps the unchanged escalation schedule, reasons and demand-deep-inspection wording; a test counts the reader calls in both directions to pin that. It is not inherited by secondmate homes: each home supervises its own crew and owns that trade separately, the same reason config/turnend-churn-absorb is home-local. The away-posture absorb returns to leaving the idle timer alone, which it had restarted only because the costly consult could reach it. A parked-gate wait is owed to the supervisor rather than the captain, so it never enters that branch, and the recheck owed on return is again owed in full the moment the record is archived. * test(watch): pin that the away-silenced hold leaves the idle timer alone The absorb no longer restarts the timer, so the recheck owed on return is owed in full rather than a cadence into the return. Nothing asserted that, so a restart could be reintroduced silently. * no-mistakes(review): document away-silence rationale, pin captured gate component * no-mistakes(test): anchor gate row scan to the braced findings header * no-mistakes(document): pin same-block gate row invariant in crew-state comment * fix(bin): reclaim a task whose herdr endpoint was destroyed (#5007) * fix(control): let the owning seat reclaim a task whose endpoint is gone A destroyed pane or workspace made `missing` a terminal state. Relaunch accepted only `dead` and said to stop the agent first; exit refused `missing` and said to reconcile the task first; there is no reconcile verb. Each command named the other as its prerequisite, so a task whose terminal went away could not be reclaimed by anything, and a no-mistakes approval it was parked on had no seat left to answer it. `missing` is agent-free a fortiori: there is no endpoint, so there is no agent in it. Widen the existing guards rather than add a verb. - fm-spawn --relaunch accepts a positively proven `missing` and creates one fresh endpoint in the recorded worktree; the record it already republishes rebinds the task to it. A `dead` endpoint is still adopted in place. - fm-control exit reports `endpoint-gone` instead of dying, so the relaunch transaction's stop step no longer dead-ends, and re-resolves the endpoint from the record before verifying the replacement. The duplicate-agent refusal is untouched: both verdicts come from the same recovery-grade classifier, which claims `missing` only from positive absence, so `alive`, `ambiguous`, and `unreadable` all still refuse. The backends' own create paths refuse a live same-labeled endpoint as a second independent guard. The worktree, its branch, commits, uncommitted changes, armed poll and registration, record rows, and status log are all untouched - a reclaim is a recovery, never a teardown. A secondmate is excluded: its gone-endpoint recovery already has one owner in the session-start liveness sweep, so relaunch refuses and names it rather than becoming a second path to the same outcome. Tests reproduce both halves of the deadlock, the reclaim succeeding, unlanded work surviving it, and the refusals that still hold. * no-mistakes(review): prove endpoint absence per backend before reclaim rebinds * no-mistakes(review): give exit and relaunch one absence proof; pin herdr rebind session * no-mistakes(review): narrow endpoint reclaim to herdr; tmux refuses honestly * no-mistakes(review): stop refusals and docs asserting unestablished causes * no-mistakes(review): stop herdr fixture helper losing tmp-root registration * no-mistakes(review): document workspace drift and absence-probe server residue * no-mistakes(review): correct rebind limitation to its one reachable case * no-mistakes(review): stop claiming reclaim leaves instructions untouched * no-mistakes(document): scope fm-control-lib purity claim, note reclaim coverage * no-mistakes(rebase): read the staged launch file in the herdr fixture Rebasing onto main picked up #4994, which stages a long worker launch command into a script and delivers the short `. '<path>'` line instead of the literal command. The tmux fake and tests/fixtures.sh were updated for that; the herdr fake this branch adds was written before it and still keyed "an agent now exists on this pane" off the literal `encode launch-brief` text, so after the rebase it never marked the rebound pane live and the reclaim's alive-wait read `dead`. Dereference the staged file first, exactly as the tmux fake above does. Test-fixture only; no production path changes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * no-mistakes(document): note reclaim placement in herdr and scripts inventories --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(bin): stamp status events with their emission time (#3764) * test(status): reproduce missing event emission time * wip(status): preserve optional event emission time * test(status): document indirect clock stub invocation * no-mistakes(review): Preserve historical status bytes during reply recovery * no-mistakes(test): Fix timestamped status assertions and remote fixture dependencies * no-mistakes(review): Preserve captain regex overrides for timestamped status events * no-mistakes(document): Clarify status event timing and publication contracts * no-mistakes(lint): Quote literal done to satisfy ShellCheck * no-mistakes(ci): Captain, updated .github/workflows/ci.yml to expect 19 snapshot tests instead of 18, matching the PR’s added regression. Reproduced the failure before the fix. Stock Bash 3.2.57 verification passed: parse sweep, 19 snapshot tests, 53 Bearings tests, and the public-followup regression. Workflow lint and diff checks passed * no-mistakes(test): Preserve terminal notifications with malformed timestamp tags * no-mistakes(test): Stamp Rovo spawn failures with emission time * no-mistakes(document): Verify status event documentation * no-mistakes(lint): Fix ShellCheck quoting in status emission-time tests * no-mistakes(ci): Captain, fixed four lifecycle assertions to accept emission timestamps while preserving publication and retry checks. Reproduced the CI failure before the fix. The lifecycle suite now passes with six Beads capability skips; syntax, targeted ShellCheck, and diff checks passed * no-mistakes(ci): Captain, fixed malformed timestamp colons hiding actionable events using shared normalization. Original bytes and unknown ages are preserved. Regression reproduced before the fix; classifier and remote-reply suites, targeted lint, syntax, and diff checks passed * no-mistakes(review): Stamp remote escalations at call sites, drop new flag * no-mistakes(review): Accept stamped escalation and close lines in test assertions * no-mistakes(review): Restore reserved-key answered-note guard for stamped closes * test(status): accept optional emission time in PR-provenance assertions The #4148 provenance test landed on main with exact unstamped greps. Parent-channel lines from this branch carry [at=<epoch>], so strip only that tag before the same exact match. No production change. * no-mistakes(review): Accept stamped ready signal in PR fallback scrape * no-mistakes(review): Drop relay flag, stamp parent events at call sites * no-mistakes(review): Stamp worker terminal-signal instructions, revert fm-on fixture * no-mistakes(review): Accept optional stamp in live cmux drift guard * no-mistakes(review): Restore original test invocation order in two suites * no-mistakes(review): Strip only well-formed numeric status time tags * no-mistakes(document): Drop stale unstamped PR-ready line spelling from channel doc * no-mistakes(review): Stamp agy spawn-failure status lines with event time * fix(bin): normalize status event times in-shell and freeze the budget test clock Two paths made a status event's emission time cost more than it should. The captain-relevance fallback piped every line through awk to drop a well-formed `[at=<epoch>]` tag before matching, so a supervisor sweep paid a fork per line just to prepare a regex match. Shell parameter expansion does the same strip with no fork, and the retry-dedup scan now reuses that one helper instead o…
dsantosg1103
added a commit
to dsantosg1103/firstmate
that referenced
this pull request
Sep 24, 2026
* fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586)
* fix(watch): honour a declared wait before wedge-escalating a quiet pane
wedge_timer_check escalated on elapsed idle time alone. Nothing asked
whether the worker had already said why its pane was quiet, so a lane
that declared a bounded external wait climbed the escalation ladder for
as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every
repeat carried demand-deep-inspection - which by its own wording forbids
re-absorbing on the run-step or pane state, so the supervisor could not
use the evidence that was there either.
The generated brief promises that declaring `paused:` buys the long
recheck cadence instead of a wedge, but the timer was still reachable
while that declaration stood: a crew that declares a wait and then has an
active run or busy pane attributed to it is handed to the timer as
provably-working. The declaration is what the worker said about its own
silence, so it now outranks a liveness verdict that only says something
is running.
The consult runs in the at-threshold branch that was about to escalate,
beside the worktree walk already there, and costs one status-line read.
Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS
recheck the declared-wait absorber already uses, so the wait is still
rechecked and cannot rot invisibly. Which verb declared it decides the
wording, because the two block on different people: a `paused:` wait is
owed by an external dependency and asks the reader to confirm it still
holds, while a `captain-held:` transfer is owed by the captain reading
the recheck and asks them to answer or release the hold. A hold is not
rechecked at all while the away-posture record exists, as on every other
captain-held path, and that absorb arms no throttle so the recheck is
owed in full on return.
A declared clearing time that has already passed stops counting, and a
lane that never declared one keeps the identical escalation schedule,
reason, count and demand-deep-inspection wording, so detection and its
worst-case time are unchanged. The deferral restarts the idle timer
rather than cancelling it, so a lane that stops waiting escalates again
within one threshold.
A lane quiet because its own validation run is parked at a gate awaiting
a human decision is deliberately out of scope: reading that state needs a
signal carrying who the wait is on and what clears it, rather than one
inferred from a parked verdict that also covers gates awaiting the
crewmate itself.
Tests pin both directions for each case and were each confirmed to fail
with the consult removed.
* no-mistakes(document): docs: honour declared waits in stale-escalation docs
* fix(bin): report verified PR state for passed runs (#4624)
* fix(bin): derive passed PR state from PR record
A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists.
For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim.
Fixes #4607
* no-mistakes(review): Add bounded GitLab merge-request state reads
* no-mistakes(review): Preserve network-free inactive crew-state scans
* no-mistakes(document): Document PR record readers in shared library
* fix: restore published contribution follow-up (Fixes #4469) (#4627)
* fix: restore published contribution follow-up (Fixes #4469)
* fix(review): Fix contribution freshness and merge actor routing
* fix(review): Restore issue triage and scope contribution follow-up
* fix(test): test: assert one wake per contribution signal
* fix(document): Document contribution follow-up
* fix: restore truthful terminal delivery evidence
* fix(review): Disclose unsupported contributions and deduplicate watcher wakes
* fix(review): Preserve unmeasured unsupported contributions across Bearings
* fix(review): Deduplicate shared contribution wakes and isolate diagnostics
* fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower
* fix(bin): make remote report transfers explicit and fail-open (#4658)
* fix(bin): make a remote-reply document gap self-clearing and re-attemptable
A remote mate's undelivered document raised a keyed `blocked` decision that
nothing could ever resolve, and any `data/*.md` substring in any mirrored line
was an unconditional fetch instruction. A mate announcing a report it had not
written yet therefore manufactured a permanent, factually false blocker, and
its own explanation of the false alarm manufactured more.
The reader has no permanence vocabulary: a report still being written refuses
exactly like a path that will never exist. So an undelivered document is now a
durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`,
re-attempted on the next delta and on the channel's own quiet poll, and retired
with a matching `resolved` line naming the local copy once it arrives. The
cursor still advances and no delta stalls on one bad pointer.
Only a structured `report=data/....md` pointer now offers a document, so a path
merely mentioned in prose - including one under another home's mirror tree,
which is provably not that mate's to serve - is never fetched. Offers are
deduplicated across the whole delta, the escalation names each missing document
once and carries the reader's own reason instead of discarding it, and a
strictly increasing notice ordinal keeps a later escalation from being
swallowed as duplicate bytes. A mirrored line still lands once whichever
pointer form it was first written under.
* no-mistakes(review): Require structured pointer token boundaries
* no-mistakes(review): Unify boundary-safe pointer extraction and rewriting
* fix(bin): identify a mirrored line independently of its delivery state
Two defects in the boundary-safe pointer work.
The at-most-once check compared only the all-remote and all-local renderings
of a line, so it could not recognize a mixed one. A line offering two documents
where only the first was deliverable mirrored as local-plus-remote; once the
second arrived, a cursor-loss whole-log recapture rendered the same line
all-local, matched neither alternate, and mirrored a second time. A line's
identity is now the canonical form every boundary-valid pointer would take once
delivered, derived by the same parser that does extraction and rewriting, so it
no longer depends on which documents happened to be deliverable at the time.
The pointer map was passed to awk through the process environment. A delta may
carry up to the configured 1 MiB bound, and an expanded map of delivered
pointers can exceed the platform's exec argument limit, so awk would fail to
start; because no caller checked, the empty result would have been appended as
blank lines while the cursor advanced past dropped status content. The map now
travels in a file, and every call site checks the exit status and stops the
ingest rather than committing a delta it could not render.
Both passes now run once per stream instead of twice per line.
* no-mistakes(review): Abort ingest when document pointer extraction fails
* no-mistakes(review): Exclude structured cross-home pointers from document transfer
* fix(bin): fail open on an undeliverable remote document instead of tracking it
Narrow the remote-reply document fix to the scope the diagnosis actually
requires, as decided after measuring a simpler alternative.
A document the reader cannot deliver now fails open. The mate's line is
mirrored with its own pointer, the cursor advances, and one unkeyed note
carries the reader's reason. A note never enters the open-decision fold, so it
cannot stand open the way the original keyed block did - which removes the
never-clearing false blocker by construction rather than by resolving it.
That makes the durable self-clearing obligation unnecessary, so it goes: the
per-mate pending-documents record, its notice ordinal and resolved
announcements, and the poll-side retry. Canonical line identity goes too, and
with it a way to silently drop a genuine status line; mirroring is back to
at-most-once on exact bytes. The cross-home exclusion goes as well: under
fail-open a cross-home report= either fails harmlessly or is a nested remote
report this mate genuinely holds, which is now relayed again.
Kept: fetching only on a structured report= pointer, the boundary-correct
parser, the file-based rewrite map, and checked extraction and rewrite exit
status. The parser now scans behind a sentinel byte so a rejected candidate can
no longer give the text right after it a false leading boundary.
The reported incident is covered end to end: a report path announced in prose
before it exists raises no decision, and the report still arrives through the
ledger publisher's structured offer once written.
* no-mistakes(review): Preserve source-line identity across remote reply replays
* no-mistakes(document): Document remote reply transfer and replay semantics
* no-mistakes(lint): Fix staging truncation lint checks
* fix(calm): preserve substantive mid-turn responses (#4655)
* Preserve substantive Calm mid-turn text
* no-mistakes(review): Distinguish newline-preserved replies from short narration
* no-mistakes(document): Document Calm mid-turn preservation boundaries
* no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`
* fix(bin): preserve PR merge polls across volume remounts (#4656)
* fix(bin): re-record PR poll identity after a volume device renumber (Fixes #4260)
A volume remount can renumber the state filesystem's st_dev while every
inode and byte stays the same; APFS does this across a reboot. A poll
registration records its sidecar and check as device:inode, so every poll
armed before the remount failed strict validation and the watcher refused
all of them as unauthenticated state checks until each was re-armed by hand.
There are two device comparisons. fm_pr_private_file_valid compares a live
file's device with the state directory's device read in the same invocation:
it refuses a file that is not on the state directory's own filesystem and
already survives a renumber, so it is unchanged. The registration's recorded
identity versus the live identity (from #556, reused by the #932 retirement
receipt) binds the registration to the exact files published in its own
transaction; its device part is what breaks.
When strict capture fails, the watcher now proves the device is the only
difference: every other artifact check passes (template bytes, both hashes,
private mode, single link, live device, metadata), both recorded identities
name one device, and each recorded inode equals its live inode. Only then,
under the task's control lock, does it rewrite the two identity lines,
repeating the whole proof and comparing the registration's file identity and
bytes just before the rename, and then capture strictly again. A swapped,
altered, re-moded, relinked, split-device, or foreign-device artifact still
fails a proof and is still refused, and a pending retirement receipt blocks
the rewrite.
Reproduction: on macOS a poll armed on an APFS disk image that was detached
and re-attached behind another image moved st_dev 16777239 -> 16777243 with
inodes, bytes, mode, and link count unchanged; the real watcher refused it on
main and reports its merge with this change. The portable regression test
rewrites a real registration's recorded device and drives the watcher.
Not changed here: the status presentation cursor keys rows by its own
device:inode identity in bin/fm-classify-lib.sh, a different helper that
needs its own fix; a retirement receipt left by a reboot between its
publication and removal still names the old device and stays refused; custom
check trust binds only a content hash and is unaffected.
* fix(review): Serialize PR poll publication writers
* fix(review): Bound PR poll publication lock scope
* fix(bin): keep contribution records when the poll budget runs out (follow-up to #4627) (#4661)
A budget that expires partway through an observation no longer records an
error or prints the unavailable wake; the URL keeps its prior record and is
observed first next poll. forge() flags budget exhaustion at the point it
refuses, or when a read is killed at the budget's own deadline, so a genuine
forge failure still records the error and wakes. Each distinct URL is now
observed once per poll and applied to every owning task.
* fix(bin): clear parent pending-replies on local secondmate retirement (#4680)
* fix(bin): clear parent pending-replies on local secondmate retirement
Local secondmate teardown left resolved parent pending-reply records behind
after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced
retirement while any reply for that id is still unresolved, and delete every
matching record plus its delivery confirmation after a successful local or
remote retirement, matching the remote cleanup path.
* no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup
* no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung
* no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern
* no-mistakes(review): Pending-replies Basename und corr_id abgleichen
* no-mistakes(document): Clarify forced retirement pending-reply cleanup
---------
Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>
* fix(bin): accept Orca's composite worktree id when tearing down a task (#4677)
* fix(bin): accept Orca's composite worktree id at teardown
Teardown refused every Orca-backed task because the endpoint validator
checked orca_worktree_id with the simple-atom rule meant for tmux-style
window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca
returns that id as `<orca id>::<absolute worktree path>`, so the colon and
slashes in every real value made validation fail and finished Orca tasks
could never be cleaned up.
Validate the field as the composite it is: both halves of the first `::`
split present, the path half absolute, and no embedded newline, carriage
return, or tab. The terminal field keeps the atom check, which is correct
for it, and no other backend's validation changes.
The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca
never returns, which is why the suite passed a check the real value fails.
They now carry the composite form, so the tests exercise the real value.
* no-mistakes(document): name Orca's repo id in the composite worktree id
* no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points
* feat(bin): add opt-in typed dispatch resolution (#4692)
* feat(bin): add opt-in typed dispatch resolution through typesafe.ai
Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or
scout profile from a written brief with typesafe.ai's System One model:
one Choice question over the rules' `when` texts, then the confidence
floor, the rule's `approval` and `floor`, each profile's `provider` and
`floor`, one quota-axi snapshot, and the spendPriority argmax all in code.
It is off unless TYPESAFE_API_KEY is in the environment or the home's
gitignored .env; off means one stderr line, exit 0, and no network call,
so firstmate dispatches exactly as before. The key reaches curl on a file
descriptor, never argv.
Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and
the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new
tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates
the four new optional dispatch fields. Document the schema, the operator
contract, the AGENTS.md intake step, and the live and benchmark evidence.
* no-mistakes(review): Harden typed dispatch resolution and quota bounds
* no-mistakes(review): Validate dispatch floors and ranking evidence
* no-mistakes(review): Tighten dispatch response and floor evidence
* no-mistakes(review): Neutralize none matching and resolve defaults locally
* no-mistakes(review): Preserve providerless profiles outside typed resolution
* no-mistakes(review): Validate response usage and reject duplicate profiles
* no-mistakes(review): Escalate unverifiable floors and validate probabilities
* no-mistakes(review): Validate probability mass and unknown profile floors
* no-mistakes(review): Simplify resolver interface and preserve fallback routing
* no-mistakes(review): Fix constants and rank partial quota evidence
* no-mistakes(review): Add authoritative provider mapping and enforce explicit providers
* no-mistakes(review): Declare provider for documented Pi profile
* no-mistakes(review): Validate provider identifiers and support Gemini dispatch
* no-mistakes(review): Strictly anchor provider identifiers
* no-mistakes(review): Validate selectors and preserve fallback candidate evidence
* no-mistakes(review): Gate typed validation and harden resolver evidence
* no-mistakes(review): Preserve opt-in routing and harden candidate evidence
* no-mistakes(review): Prioritize known exhaustion over quota uncertainty
* no-mistakes(review): Isolate API secrets and preserve no-key diagnostics
* no-mistakes(review): Fallback safely when dispatch rules are absent
* no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets
* no-mistakes(document): Document typed dispatch safety and fallback behavior
* fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753)
* test: reproduce buried status declarations in shared readers
* fix: share status event reads and preserve open blockers
* fix: retain terminal scout and ship status declarations
* no-mistakes(review): Fix status chronology, legacy completions, and reader performance
* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots
* no-mistakes(review): Unify terminal supersession across cached folds and consumers
* no-mistakes(review): Filter per-key status history while preserving terminal chronology
* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells
* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses
* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor
* no-mistakes(lint): Quote literal done in test for-lists for SC1010
* ci: expect 19 snapshot/fleet-view tests
This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.
* no-mistakes(review): Restore multiline child outcome reporting
* no-mistakes(review): Select ledger terminal events through bounded shared reader
* no-mistakes(review): Report newest open decision instead of preferring blocked
* no-mistakes(review): Require colon before ship/scout terminal supersession in fold
* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix
* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions
* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold
* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule
* no-mistakes(document): Align status-read docs with fold-resolved crew state
* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers
* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree
* test: fold terminal-cleanup snapshot coverage into the completed-scout case
Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.
* no-mistakes(document): Clarify socket-down override expiry in architecture doc
* ci: retrigger flaky contribution check
* fix(bin): launch codex crewmates with codex's hook layer disabled (#4689)
* fix(spawn): launch codex crewmates with codex's hook layer disabled
A freshly launched Codex worker never reached its instructions. Codex
stopped it on an interactive "Hooks need review" modal whose selection
sits on "Review hooks", which is neither trusting nor declining.
Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow
navigation, so the selection cannot be moved, and pre-accepting the
prompt by writing Codex's own trust store would record an operator
consent that was never given.
The hooks are the machine's own ~/.codex/hooks.json plus any project's
.codex/hooks.json. A crewmate needs neither: its turn-end signal is the
-c notify= program on the same launch, and Firstmate's project hooks are
primary-session infrastructure that stands down in a child worktree.
Crewmate and scout launches now pass --disable hooks. That is the
opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted
hooks; disabling the feature runs none of them and leaves the operator's
~/.codex untouched. An unknown feature name is a hard Codex error, so a
release that drops the flag fails the launch loudly instead of silently
restoring the modal. A secondmate is a primary in its own home and keeps
the project hooks its turn-end guard and session-start digest ride on.
Verified on codex-cli 0.151.0: the modal is gone and the turn-end
notification still lands.
This unblocks the second review that every finished pull request is supposed to get.
Fixes kunchenguid/firstmate#4673
* no-mistakes(review): Fix contradictory hook count in Codex verification record
* fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710)
* fix(bin): settle terminal contributions and wake once per read-failure episode
A contribution whose last good observation is merged or closed is final:
poll no longer re-reads it, projection keeps it fresh, and a stale error
recorded beside it is cleared once. A genuine forge-read failure on an open
contribution still records its error on every cycle but prints the
unavailable wake only when it starts a failure episode; a successful read
ends the episode. Open PRs linked from done tasks keep being observed.
The false unavailable beside a complete observation was budget exhaustion
mid-observation, already fixed by #4661.
* fix(review): Settle terminal contribution owners
* fix(review): Deduplicate shared contribution failure episodes
* fix(test): Preserve settled terminal contribution records
* fix: select authoritative no-mistakes runs (#4476)
* fix(crew-state): select authoritative validation runs by identity
Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row.
Refs: https://github.com/kunchenguid/firstmate/issues/3215
* fix(review): Resolve same-branch run identities beyond capped history
* fix(review): Fix run-selection compatibility, races, and worker-state fallbacks
* fix(review): Limit run validation to the requested branch
* fix(test): Anchor AXI fixtures and document remaining live evidence gaps
* fix(document): Clarify run selection documentation and capture ownership
* fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation
* fix: distinguish captain outcomes from no-op updates (#4738)
* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work
MAIN answered a supervision-branch outcome for completed captain-requested
work (implementation done, PR ready for review and merge approval) with
"Captain, shipshape.", reading section 9's no-action reply as covering it
and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no
captain-facing response is owed".
Section 9 now limits the shipshape reply to true no-ops (idle re-read,
empty heartbeat, consequence-free acknowledgement) and requires a short
outcome response naming what finished and what word is needed whenever
requested work finishes or a result needs the captain's word, even when a
transcript entry already shows the substance. The Pi protocol's re-emit
rule now says it bounds repetition only, and carries a worked example of
the ready-for-review outcome whose correct processing turn a shipshape
reply fails.
No executable contract evaluates the content of MAIN's captain-facing
reply, so the regression is the protocol example in the owner doc rather
than a text-match test.
* no-mistakes(document): Clarify captain-facing outcomes versus no-ops
* docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line
The document step condensed the Pi protocol's re-emit rule and dropped the
worked example of a finished, ready-for-review outcome whose correct
processing turn a "Captain, shipshape." reply fails. That example is the
contract's regression: no executable contract evaluates the content of
MAIN's captain-facing reply, so the owner doc's example is the test case.
Restore it directly under the re-emit rule, prefixed as a regression
example that is kept verbatim and never condensed or summarized away.
* no-mistakes(review): Clarify captain outcome and decision-word requirements
* no-mistakes(document): Clarify captain-facing completion outcomes
* docs(pi): require the PR URL in the visible captain-facing outcome reply
Captain review on the regression example: drop the sample reply string
and say only that the ready-for-review outcome requires relaying a
captain-facing outcome response, not just "Captain, shipshape.".
Fold in the visible-PR-handoff failure seen this session: after the
branch outcome reporting this fix green, MAIN's visible reply was only
"Awaiting your merge call." with no PR URL, leaning on the dim anchor.
Section 9's URL rule now also covers a review or merge ask and names the
visible reply as where the URL goes, sourced from the ready status, pr=
metadata, or the supervision branch's summary and never left to a
transcript entry. The Pi protocol adds the same-way failure and places
the captain-facing text in the final visible assistant reply after the
fm_branch_processed call, because Calm hides assistant text emitted in
the same step as a tool call as a working note.
Investigation verdict, evidence in the PR comment: no recent PR caused
the handoff failure; Pi has hidden same-step pre-tool assistant text
since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and
#4658 touched only remote report transfer.
* no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs
* no-mistakes(document): Clarify captain-facing supervision outcomes
* docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule
The consolidated section 9 URL rule narrowed its trigger to a review or
merge ask, dropping the "whenever a PR is mentioned" catch-all from
#3648 that keeps every PR URL copied from a durable record and never
assembled from memory. Restore that trigger as a union with the review
or merge ask so the one consolidated rule covers both.
* fix(bin): let non-owner Claude Stops exit safely (#4777)
* Fix foreign-owner turn-end supervision loop
* no-mistakes(review): Scope foreign-owner safe exit to Claude guard
* no-mistakes(document): Document Claude foreign-owner safe exit
* fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778)
Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty
indexed array as an unbound variable and aborts the shell. In
signal_turnend_panes_churned() the missing_keys loop was reachable with
an empty array whenever every churned key already held a fresh
.churn-since-* marker (a second churning turn-end inside an open
deferral window), so each watcher cycle died about half a minute in and
supervision restarted endlessly. The created_keys rollback loops had the
same latent crash on their error paths.
Audit of bin/ for the same pattern found one more confirmed-reachable
case: remote_handoff's noncanonical-body scan iterates to_move, which is
empty when a retried remote handoff finds every key already staged in
the outbox. All other "${arr[@]}" sites are either count-guarded,
guaranteed non-empty by construction, or unreachable while empty.
Guard the three reachable expansions with the repo's existing
"${arr[@]+...}" idiom. New regression test drives a real watcher
through the all-marked churn path; the macos-stock-bash CI lane runs it
under real /bin/bash 3.2 via FM_TEST_ONLY.
* Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783)
The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: require complete captain-facing final responses (#4779)
* docs: require complete final responses across harnesses
* no-mistakes(document): Document complete final replies for Grok Bot
* docs: point Grok replies to the shared contract owner
* no-mistakes(review): Clarify final recap without batching decision asks
* fix: preserve substantive mid-turn text in Pi Calm (#4788)
* fix(calm): preserve substantive Pi mid-turn text
* no-mistakes(review): Preserve substantive Pi Calm text per block
* no-mistakes(test): Cover shared Calm preservation boundaries behaviorally
* no-mistakes(document): Consolidate Calm preservation documentation
* fix: harden mail checks and rebalance full-coverage CI (#4800)
* Improve CI reliability and rebalance full-coverage validation
* no-mistakes(document): Clarify lint partition documentation
* fix(bin): answer Kimi 2.0.0 folder-trust dialog during spawn (#4799)
* Handle Kimi workspace trust dialog
* no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers
* no-mistakes(review): Gate Kimi ready on any trust marker and clean captures
* no-mistakes(review): Read visible pane for Kimi trust and ready gates
* no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate
* no-mistakes(review): Harden Kimi viewport capture and trust dialog detection
* no-mistakes(document): Document Kimi spawn refusal on cmux and Orca
* fix(bin): report a dead-agent record once instead of escalating forever (#4775)
* fix(bin): report a record whose agent is gone once instead of escalating forever
The wedge escalation path never asked whether there was still an agent to be
wedged. A wedge is something stuck that might recover, so re-alarming it earns
its cost; an agent that is gone never moves again, its pane never churns, the
idle timer never resets, and the escalate path clears its own timer and re-arms
with nothing bounding the count.
Observed on a live fleet: two finished lanes reached 226 and 203 consecutive
escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400
notifications a day from two lanes with no agent running at all. On one,
fm-control.sh exit answered already-stopped and fm-crew-state.sh read
"failed - run failed". Closing the Herdr pane did not stop it either: with the
pane genuinely gone and herdr pane read returning pane_not_found, the count kept
climbing, because the poll is driven by the record's window= line rather than by
the pane. The cost is not the repetition but that it drowns the alarms that
matter.
fm_backend_agent_state already separates a thinking agent from a gone one at
process level. In the branch that was about to escalate, read it once and treat
only its two recovery-grade verdicts - dead (endpoint present, no agent in it)
and missing (endpoint authoritatively absent) - as proof, reporting that record
once and not re-escalating it while it stays that way. Every other verdict,
including alive, ambiguous, unreadable, unverified, and a read that failed
outright, keeps the identical schedule, reason, and escalation count, so a
genuinely wedged live agent is unaffected. The probe costs at most one backend
read per window per threshold, the same budget the declared-wait consult and the
worktree write probe already take.
The report decides nothing about the record's fate: both lanes still held
unlanded work and teardown refusing them was correct, so retiring, relaunching,
or cleaning up stays with the supervisor. The once-only marker is owned entirely
by that function and is dropped by the same read the moment the endpoint stops
reading gone, so a replacement launched into the same window escalates normally
and its own later death is reported again.
Related, and not closed by this: #4412, #4482, #4316.
Tests drive the real watcher against a record whose endpoint does not exist and
pin both directions: dead and missing report once and never advance the count
across later thresholds, while alive, ambiguous, and unreadable endpoints keep
escalating with the identical reason and a climbing count.
* fix(bin): bind the once-only dead report to the pane it reported
Review of the parent commit found a reachable sequence where a later death in
the same window lost its promised report. The marker was keyed on the verdict
string alone and dropped only when a threshold probe read a non-gone verdict,
but probes run only at thresholds: a replacement launched into the same window
that dies without ever being probed alive - it crashes at startup, or works and
then crashes - was absorbed by the previous death's marker. The pane's first
sight yielded only the generic stale wake and every later threshold matched the
stale marker, so the second death never got the detailed once-report that both
the function's own comment and docs/architecture.md promise.
Record the verdict together with the pane hash it was reported for, and absorb a
repeat only while both still match. A replacement churns the pane, which resets
the stale suppressor, wedge timer, and escalation count while no reset site
touches this marker, so the pane half is what tells the second death apart from
the first. The live-probe drop stays as it was.
Clearing the marker at those reset sites instead would re-open unbounded
re-alarming for a dead pane whose display ever ticks, which is the exact defect
the parent commit exists to close.
The noise bound is unchanged: an unchanged dead pane still absorbs on every
later threshold and never advances the escalation count, and every verdict short
of proof still escalates exactly as before.
* no-mistakes(review): Key the dead-record once-marker on the busy incarnation token
* no-mistakes(document): Document dead-record escalation cap in stale-pane config entry
* no-mistakes(document): Add busy-state inventory line to AGENTS.md
* no-mistakes(document): Document dead-record probe on busy-turn-bound wedge path
* fix(bin): create captain-hold rows when Beads requires due (#4854)
Captain holds have no due semantics and are a hold kind, not a Beads issue
type. The create path now waives due.required and maps to native type task.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: disable compact adviser for spawned agents (#4877)
* feat(bin): launch every spawned agent with the compact adviser disabled
Every crewmate, scout, and secondmate Firstmate launches now starts with
COMPACT_ADVISER_DISABLE=1, on a fresh spawn and on a relaunch alike, so an
unattended session never activates the compact adviser.
The value is unconditional: no configuration file gates it and there is no
override, unlike the trace carrier beside it.
Three carriers deliver it, because no single one covers every launch shape.
The pane shell receives an export beside GOTMPDIR, so the agent's own children
inherit it too.
The launch command carries an explicit assignment, prepended outermost so it
wins over any ambient value the pane already held.
The cleared launch environment sets it again at the `env -i` boundary and keeps
COMPACT_ADVISER_DISABLE in the fixed operational floor, which is what preserves
the switch when config/launch-env-allowlist empties the environment, and what
delivers it on a remote host that never had the value.
bin/fm-control.sh relaunch, the bootstrap secondmate relaunch, and the remote
secondmate transport all rebuild their launch through bin/fm-spawn.sh, so they
inherit the same floor.
The captain's own primary session is untouched.
The two new suites drive the real spawn and then execute the launch command the
pane actually received, with the harness replaced by a probe that prints its own
environment, rather than matching script text.
They cover ship and secondmate launches with the allowlist absent and enabled,
the pane export and its ordering, fm-control.sh relaunch, and the full parent to
remote-host chain.
* no-mistakes(review): Export compact-adviser disable across compound launches
* no-mistakes(document): Document spawned-agent compact-adviser environment guarantee
* fix(bin): preserve Claude lock ownership after helper recycling (#4894)
* fix(bin): let a background Claude session keep owning its session lock
Session-lock ownership was decided by process ancestry alone. Under an
unattended Claude session the model loop runs in a transient bg-spare
bridged to the front-end by a shared daemon; when that bridge is
recycled the contiguous claude-named ancestry from a hook to the
recorded owner breaks while the owner pid stays alive, so the Stop
auto-arm stood down as a foreign live owner, the turn-end guard ended
every turn with its read-only diagnostic, and fm-lock.sh refused - a
self-sustaining outage until restart.
Ownership is now ancestry membership OR a trusted same-session id,
never id-first:
- fm-session-lock-lib.sh accepts CLAUDE_CODE_SESSION_ID only when
CLAUDE_PID is a Claude-shaped member of the current contiguous run,
compares it against the id recorded in state/.lock-session, and
requires the recorded pid to still be a live harness. No id, no
sidecar, an untrusted id, a different id, or a dead recorded pid
leaves the ancestry verdict unchanged. Ids are never read from ps
argv.
- fm-lock.sh accepts a same-session holder at both refusal sites,
writes, refreshes, and clears the sidecar only under its claim lock
(including the early already-mine exit, skipped only while the
deferred startup sweep leases that lock), keeps it byte-identical
across a same-session confirmation, records CLAUDE_PID on lock line 1
for a session with a trusted id so a shared daemon or front-end that
outlives the session never keeps a dead session's lock alive, never
rewrites a live line 1 on a same-session confirmation, and names the
recorded id in the live-owner refusal.
- The .lock line-1 format is unchanged, so every reader that takes the
whole first line as the pid keeps working; the guard's foreign-owner
exit is unchanged and inherits the fix through the shared predicate.
Tests: the ancestry suite drives the ancestry and id signals apart in a
deterministic process table (asserting the divergence) and runs a real
orphaned front-end/daemon/pty-host/spare tree through six phases with
the real lock, auto-arm, and guard scripts; the foreign-owner repro
keeps its negative control and adds a same-id positive control.
Disclosure: no live unattended Claude background session ran on the
verifying machine. The topology is documented by the real process
listings in #3902, #2314, #3398, and #4066; coverage is the structural
predicate plus the executable fixtures, not a live pass.
Residual: bin/fm-sessionstart-nudge.sh keeps its own private ancestry
walk (it only decides whether to print a nudge) and may nudge on a
resume in the recycled case.
Out of scope, deliberately: no structured lock format, no guard budget
changes, no daemon-identity rejection, no fork lineage.
* no-mistakes(review): Wait for claim lock; revert failed sidecars
* no-mistakes(review): Revalidate ownership after wait; restore sidecars
* no-mistakes(review): Roll back sidecar by publication phase
* no-mistakes(review): Restore sidecar only if lock line is unchanged
* no-mistakes(review): Trust session ids without a spelling allowlist
* no-mistakes(review): Disarm sidecar rollback before backup cleanup
* no-mistakes(document): Updated session-lock ownership documentation
* feat: park main under the away posture on Pi (#4889)
* feat: park main under the away posture on Pi
While the away-posture record exists on a Pi primary, the supervision branch
takes every actionable wake, no processing turn opens on main, captain rows
accumulate for the return brief, and main's standing authority relocates to
the branch through the existing guarded scripts.
- lib/fm-branch-dispatch.ts: read the record at every routing decision; while
it exists claim check, decision-owned, and heartbeat rows too, keeping the
two broken-queue vetoes; expose checkSeqs so a claimed check row lifts task
scoping.
- fm-primary-pi-watch.ts: offer every actionable row under the record; a
declined wake and every watcher-failure alarm still reach main.
- fm-branch-supervision.ts: drop the legacy .afk decline; append a fixed
POSTURE: AWAY tail carrying the record's read-back verbatim per wake; open no
processing request while the record exists, re-checked immediately before a
request would open and at every run boundary; present the accumulated rows
at the first run boundary after archive.
- fm-lease-lib.sh: fm_lease_forbid_branch passes the branch for opted-in
actions only while fm-afk-contract.sh validate succeeds on a confirmed live
record; PR merge, fresh spawn, and decision answer opt in, local landing
never does.
- fm-send.sh: a --resolve-key naming an open needs-decision or captain-held
task is a decision answer and meets the partition; blocked: keys stay
steering.
- fm-spawn.sh: enforce the record's spend cap for a fresh ordinary spawn by
either actor; relaunches and secondmates exempt.
- fm-branch-prompt.sh: fixed Postures section and the verbatim
ask-user-authority policy; the prefix stays byte-stable.
- fm-afk-return.sh: count what the away session handled from the store.
- docs, afk skill, AGENTS.md stub: main parked on Pi, green merge gate
absolute while away.
- tests: watcher and branch extension suites, fleet-record, merge, and
decision-answer suites cover the relocation, the vetoes, the tail, the
parked processing turn, the cancellation, the re-presentation, and the
spend cap; dated live-guard evidence recorded.
* no-mistakes(review): Refuse branch merge after preflight archive race
* no-mistakes(review): Fix away wake, spawn, and processing races
* no-mistakes(review): Suppress parked processing; narrow away-only rejection
* no-mistakes(review): Abort dedicated processing; gate branch spawn once
* no-mistakes(review): Stamp away-only on the dispatch offer
* no-mistakes(review): Treat invalid away records as spend-cap absence
* no-mistakes(review): Drop spawn test hook; abort processing-opened runs
* no-mistakes(review): Bind abort to opening prompt; cap-read absence
* no-mistakes(review): Limit away branch spawn to queued work only
* no-mistakes(document): Correct AFK posture documentation
* ci: standardize workflow timeouts into three tiers (#4910)
* ci: simplify CI job timeouts to a three-tier policy
Replace the scattered per-job timeout values (10m parallel, 25m lint, 30m
serial, 10m macOS) with three readable tiers, each a hang tripwire with
headroom rather than a packing estimate:
- fast (5m): coverage guard, repo invariants, timing aggregate
- normal (30m, one shared budget): lint partitions, portable parallel
shards, portable serial shards, macOS stock Bash
- heavy (Herdr only): 20m step tripwire on the family run so always()
cleanup still runs, under a 75m job-level last-resort backstop
The workflow's header comment states the policy and points at
docs/fm-test-portable-shards.md "Timeouts", which now owns it, and each
job names its tier beside timeout-minutes. tests/fm-ci-workflow.test.sh
asserts the policy against the parsed workflow instead of the old
per-job minute values: every job joins exactly one tier, exactly three
distinct job-level values exist, the fast tier stays within 5-10
minutes, the normal budget stays at least double the modeled parallel
lane sum reported by fm-test-run.sh --check-coverage, and the Herdr step
tripwire stays below its job backstop with an always() cleanup after it.
Concurrency supersession, shard counts, lane membership, and fail-fast
settings are unchanged.
* no-mistakes(review): Decouple the normal timeout from packing estimates
* no-mistakes(review): Assert Herdr teardown follows the family run
* no-mistakes(review): Pin Herdr family-run timeout to 20 minutes
* no-mistakes(review): Ignore comments when identifying Herdr steps
* no-mistakes(review): Identify Herdr steps by declarative ids
* no-mistakes(document): Clarify authoritative three-tier timeout policy
* fix(bin): keep supervisor status closes from waking the same home (#4895)
* fix(bin): keep supervisor status closes from waking the same home
A drain that already folded OPEN DECISIONS has presented those bytes even
when the watcher has no matching seen marker. Treat that fold, and the
presentation cursor, as known so the bookkeeping close stays quiet while
later worker lines still signal.
* no-mistakes(review): Keep folded worker failures waking past supervisor closes
* no-mistakes(review): Wake on unlisted folded worker lines; batch multi-key closes
* no-mistakes(review): Stop folded worker resolved lines from counting as already read
* no-mistakes(document): Correct self-announced close marker contract in docs
* fix(bin): stop labeling Herdr as experimental (#4972)
* Stop steering operators away from Herdr
* no-mistakes(review): Neutralize remaining Herdr opt-out documentation wording
* fix(bin): treat a live no-mistakes run as current after rebase (#4973)
* fix(bin): treat a live no-mistakes run as current after rebase
A running run on the task's branch is authoritative regardless of head.
Matching only the local head made a rebased in-flight run look failed.
* no-mistakes(review): restrict coarse live-any-head to foreign-branch answers
* no-mistakes(review): reject gate-parked runs from the executing predicate
* no-mistakes(review): hoist gate-marker patterns into single run-lib owner
* no-mistakes(review): require live daemon for head-free run binding
* no-mistakes(review): require answered daemon-down before unbinding live runs
* no-mistakes(review): extend daemon guard to anchored continuation routes
* no-mistakes(review): delete live-any-head; restore dead-daemon verdict
* no-mistakes(review): keep parked gates parked; name dead daemon everywhere
* no-mistakes(review): set dead-daemon verdict instead of emitting early
* no-mistakes(review): align selected route with legacy dead-daemon handling
* no-mistakes(review): drop unproven-record binds; narrow coarse gate reading
* no-mistakes(review): narrow header, drop vestigial guard, retarget tests
* no-mistakes(review): revert coarse gate override; require answered-down probe
* no-mistakes(review): cache one daemon probe; stop duplicating run id
* no-mistakes(review): restrict coarse dead-daemon verdict to moved-off rows
* no-mistakes(review): delete coarse dead-daemon extension and gate note
* no-mistakes(review): delete remaining coarse dead-daemon block and stale docs
* no-mistakes(document): document rebase-safe live-run bind and unverified-record verdict
* fix(bin): prevent long worker launch command truncation (#4994)
* fix(bin): stage the launch command in a private file and type a short source line
A long launch line typed while the fresh pane shell is still busy waits in the
terminal's canonical line buffer, which drops input past about 1,024 bytes on
macOS, so the pane was left at an unfinished command with no agent running.
fm-spawn now writes the assembled command to the task's own temp root under
umask 077 and types only a short line that sources it.
Refs #4559
* fix(bin): keep the per-task temp root private before staging the launch command
The root lives at a predictable path under /tmp and now holds the whole launch
command. Create it with mode 0700, refuse one that already exists as anything but
a directory owned by this user that nobody else can write, and tighten an owned
one, so no other local user can plant or swap the staged file.
Refs #4559
* fix(bin): enforce private staged launch file mode
* test(spawn): cover long staged Claude launches
* no-mistakes(review): Namespace launch files and prove truncation staging
* no-mistakes(review): Use immutable per-spawn launch filenames
* no-mistakes(document): Document staged launch delivery safeguards
* no-mistakes(ci): Updated eight behavior tests/fakes to execute or inspect immutable staged launch files instead of expecting inline launch commands. This restores Muse, secondmate lifecycle/restart, remote trace/parent binding, compact-adviser, and Orca coverage. All affected tests, dispatch-profile regression, fixture tests, syntax checks, ShellCheck, and git diff checks pass
---------
Co-authored-by: Vytautas Stankus <svycka@gmail.com>
* test: authorize isolated Herdr lab validation (#4998)
* Add isolated Herdr runbook to test instructions
* no-mistakes(review): Drop substring matching from test.instructions contract
* no-mistakes(review): Assert commands.test key absence in YAML
* Drop unit-first sentence and instructions contract test
Captain-scoped follow-up on the Herdr-lab test.instructions ship:
keep the lab safety runbook only, and leave the no-mistakes contract
test focused on commands.test absence.
* docs(vision): accept vendor-semantics and 9k AGENTS ceiling (#4873) (#5001)
* docs(vision): accept vendor-semantics and 9k contract-ceiling amendments (#4873)
Replace the pixels-of-today's-UI rule with a quarantined, version-pinned
surface-adapter exception recorded as standing debt. Cap the always-loaded
contract at 9,000 words and require prune-or-trigger before a crossing change
lands.
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
* docs(vision): restore accepted three-sentence vendor-semantics form (#4873)
Replace the compressed paraphrase with the issue's accepted wording:
a named quarantined version-pinned adapter, expected to break, recorded
as standing debt that never hardens into a shared contract.
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
* feat(bin): defer the wedge escalation for a lane parked at a supervisor-owed gate (#4974)
* fix(watch): recheck a gate awaiting a human instead of wedge-escalating it
A lane whose validation run is parked at a gate waiting on a human
decision is correctly quiet, but nothing in its status line says so: the
evidence is the pipeline's own gate state rather than anything the worker
wrote. The wedge timer read that silence as a suspected wedge and climbed
the escalation ladder for as long as the wait lasted, and each escalation
cost a supervising turn. The landed declared-wait consult does not reach
it, because a live ordinary crewmate never reports a declared pause, and
raising FM_STALE_ESCALATE_SECS would delay genuine wedge detection for
every lane by the same amount.
The threshold now reads a second, independent record when the status line
accounts for nothing: whether the crew's current state is a gate whose
answer is owed by a human. That is minted only from the gate's own
findings table, by a row whose `action` column is exactly `ask-user`,
located by position out of the table header the way nm_gate_step_row
already reads its row - never searched for over the run payload, where a
finding's free-text description or a branch name satisfies a search just
as well. A gate awaiting the CREWMATE's own answer keeps the unchanged
escalation schedule, reason and demand-deep-inspection wording, because a
crewmate that goes quiet before answering its own gate is exactly the
wedge the ladder exists to catch.
Each kind of wait now carries the human it is on, the action that clears
it, and whether that human is the captain as data alongside the verdict,
rather than as wording chosen per branch where the recheck is written, so
the deferral cannot word one kind of wait as another and a new kind
cannot ship without deciding all of them. A parked gate has no written
record of when its wait began, so its recheck publishes no wait age at
all rather than one read from the quiet window this deferral resets on
every pass, which would report the same small number for a gate of any
age. Like every other captain-facing recheck here it is absorbed in
silence while the away-posture record exists, arming no throttle, so the
recheck is owed in full the moment the record is archived.
The consult runs only in the at-threshold branch that was about to
escalate, beside the worktree walk already there, and only for lanes
whose status line explained nothing.
Closes #3055
* no-mistakes(review): require an unanswered decision before deferring a parked gate
* no-mistakes(review): reset the away-silenced timer, fail-safe findings parse, US-joined wait records
* test(watch): pass the pane hash wedge_timer_check now takes
Upstream gave wedge_timer_check a sixth <pane-hash> argument for its
dead-record probe. The malformed-wait-record rounds drive the real function
directly, so they pass one, and stub fm_backend_agent_state to a live agent so
the probe that runs after a refused deferral keeps the unchanged ladder rather
than reading a backend the child shell has none of.
* no-mistakes(review): Bind parked-gate wait to its run, owe it firstmate
* no-mistakes(document): correct wait-kind count, crew-state reader scope, gate-key coupling
* feat(watch): make the parked-gate wait deferral opt-in
The wedge timer deferring a lane parked at a validation gate is new
supervision behaviour rather than a restored one, and it decides which
lanes give up the escalation ladder, so it now ships as a default-off
per-home option instead of changing every home on upgrade.
config/wedge-defer-parked-gate arms it. The flag is read before the
decision fold, so an unconfigured home spends no fold or current-state
read, writes no record, and keeps the unchanged escalation schedule,
reasons and demand-deep-inspection wording; a test counts the reader
calls in both directions to pin that.
It is not inherited by secondmate homes: each home supervises its own
crew and owns that trade separately, the same reason
config/turnend-churn-absorb is home-local.
The away-posture absorb returns to leaving the idle timer alone, which
it had restarted only because the costly consult could reach it. A
parked-gate wait is owed to the supervisor rather than the captain, so
it never enters that branch, and the recheck owed on return is again
owed in full the moment the record is archived.
* test(watch): pin that the away-silenced hold leaves the idle timer alone
The absorb no longer restarts the timer, so the recheck owed on return is
owed in full rather than a cadence into the return. Nothing asserted
that, so a restart could be reintroduced silently.
* no-mistakes(review): document away-silence rationale, pin captured gate component
* no-mistakes(test): anchor gate row scan to the braced findings header
* no-mistakes(document): pin same-block gate row invariant in crew-state comment
* fix(bin): reclaim a task whose herdr endpoint was destroyed (#5007)
* fix(control): let the owning seat reclaim a task whose endpoint is gone
A destroyed pane or workspace made `missing` a terminal state. Relaunch
accepted only `dead` and said to stop the agent first; exit refused
`missing` and said to reconcile the task first; there is no reconcile
verb. Each command named the other as its prerequisite, so a task whose
terminal went away could not be reclaimed by anything, and a no-mistakes
approval it was parked on had no seat left to answer it.
`missing` is agent-free a fortiori: there is no endpoint, so there is no
agent in it. Widen the existing guards rather than add a verb.
- fm-spawn --relaunch accepts a positively proven `missing` and creates
one fresh endpoint in the recorded worktree; the record it already
republishes rebinds the task to it. A `dead` endpoint is still adopted
in place.
- fm-control exit reports `endpoint-gone` instead of dying, so the
relaunch transaction's stop step no longer dead-ends, and re-resolves
the endpoint from the record before verifying the replacement.
The duplicate-agent refusal is untouched: both verdicts come from the
same recovery-grade classifier, which claims `missing` only from positive
absence, so `alive`, `ambiguous`, and `unreadable` all still refuse. The
backends' own create paths refuse a live same-labeled endpoint as a
second independent guard. The worktree, its branch, commits, uncommitted
changes, armed poll and registration, record rows, and status log are all
untouched - a reclaim is a recovery, never a teardown.
A secondmate is excluded: its gone-endpoint recovery already has one
owner in the session-start liveness sweep, so relaunch refuses and names
it rather than becoming a second path to the same outcome.
Tests reproduce both halves of the deadlock, the reclaim succeeding,
unlanded work surviving it, and the refusals that still hold.
* no-mistakes(review): prove endpoint absence per backend before reclaim rebinds
* no-mistakes(review): give exit and relaunch one absence proof; pin herdr rebind session
* no-mistakes(review): narrow endpoint reclaim to herdr; tmux refuses honestly
* no-mistakes(review): stop refusals and docs asserting unestablished causes
* no-mistakes(review): stop herdr fixture helper losing tmp-root registration
* no-mistakes(review): document workspace drift and absence-probe server residue
* no-mistakes(review): correct rebind limitation to its one reachable case
* no-mistakes(review): stop claiming reclaim leaves instructions untouched
* no-mistakes(document): scope fm-control-lib purity claim, note reclaim coverage
* no-mistakes(rebase): read the staged launch file in the herdr fixture
Rebasing onto main picked up #4994, which stages a long worker launch
command into a script and delivers the short `. '<path>'` line instead of
the literal command. The tmux fake and tests/fixtures.sh were updated for
that; the herdr fake this branch adds was written before it and still
keyed "an agent now exists on this pane" off the literal
`encode launch-brief` text, so after the rebase it never marked the
rebound pane live and the reclaim's alive-wait read `dead`.
Dereference the staged file first, exactly as the tmux fake above does.
Test-fixture only; no production path changes.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* no-mistakes(document): note reclaim placement in herdr and scripts inventories
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat(bin): stamp status events with their emission time (#3764)
* test(status): reproduce missing event emission time
* wip(status): preserve optional event emission time
* test(status): document indirect clock stub invocation
* no-mistakes(review): Preserve historical status bytes during reply recovery
* no-mistakes(test): Fix timestamped status assertions and remote fixture dependencies
* no-mistakes(review): Preserve captain regex overrides for timestamped status events
* no-mistakes(document): Clarify status event timing and publication contracts
* no-mistakes(lint): Quote literal done to satisfy ShellCheck
* no-mistakes(ci): Captain, updated .github/workflows/ci.yml to expect 19 snapshot tests instead of 18, matching the PR’s added regression. Reproduced the failure before the fix. Stock Bash 3.2.57 verification passed: parse sweep, 19 snapshot tests, 53 Bearings tests, and the public-followup regression. Workflow lint and diff checks passed
* no-mistakes(test): Preserve terminal notifications with malformed timestamp tags
* no-mistakes(test): Stamp Rovo spawn failures with emission time
* no-mistakes(document): Verify status event documentation
* no-mistakes(lint): Fix ShellCheck quoting in status emission-time tests
* no-mistakes(ci): Captain, fixed four lifecycle assertions to accept emission timestamps while preserving publication and retry checks. Reproduced the CI failure before the fix. The lifecycle suite now passes with six Beads capability skips; syntax, targeted ShellCheck, and diff checks passed
* no-mistakes(ci): Captain, fixed malformed timestamp colons hiding actionable events using shared normalization. Original bytes and unknown ages are preserved. Regression reproduced before the fix; classifier and remote-reply suites, targeted lint, syntax, and diff checks passed
* no-mistakes(review): Stamp remote escalations at call sites, drop new flag
* no-mistakes(review): Accept stamped escalation and close lines in test assertions
* no-mistakes(review): Restore reserved-key answered-note guard for stamped closes
* test(status): accept optional emission time in PR-provenance assertions
The #4148 provenance test landed on main with exact unstamped greps.
Parent-channel lines from this branch carry [at=<epoch>], so strip only
that tag before the same exact match. No production change.
* no-mistakes(review): Accept stamped ready signal in PR fallback scrape
* no-mistakes(review): Drop relay flag, stamp parent events at call sites
* no-mistakes(review): Stamp worker terminal-signal instructions, revert fm-on fixture
* no-mistakes(review): Accept optional stamp in live cmux drift guard
* no-mistakes(review): Restore original test invocation order in two suites
* no-mistakes(review): Strip only well-formed numeric status time tags
* no-mistakes(document): Drop stale unstamped PR-ready line spelling from channel doc
* no-mistakes(review): Stamp agy spawn-failure status lines with event time
* fix(bin): normalize status event times in-shell and freeze the budget test clock
Two paths made a status event's emission time cost more than it should.
The captain-relevance fallback piped every line through awk to drop a
well-formed `[at=<epoch>]` tag before matching, so a supervisor sweep paid a
fork per line just to prepare a regex match. Shell parameter expansion does the
same strip with no fork, and the retry-dedup scan now reuses that one helper
instead of carrying a …
rub-a-dub-dub
added a commit
to rub-a-dub-dub/firstmate
that referenced
this pull request
Sep 27, 2026
…gences (#29) * fix(bin): treat Claude Code's default external-imports flags as never asked, not declined (#4387) * fix(bin): read Claude Code's default external-imports flags as never asked, not declined (#4378) fm-claude-trust.sh refused the whole trust registration whenever the project-root entry carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code writes that value only on an explicit "No, disable". Claude Code's default project entry carries Approved and WarningShown both false before the dialog is ever shown, so every such project refused every spawn. Only Approved === false with WarningShown === true — the pair the dialog writes on a decline — now counts as a decline. false/false behaves like an absent flag: trust is registered and no import consent is manufactured. New case test_project_root_entry_default_import_flags_are_not_a_decline fails on b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31, bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * no-mistakes(review): Correct harness doc's external-imports decline predicate --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(bin): keep operator-address labels out of no-mistakes intent (#4445) * fix(brief): keep operator address out of composed intent Teach raw-word authoring for intent sections and mid-task relays, with a neutral [captain] provenance marker for legacy mixed tasks. Keep headings and contract prose outside the serialized intent body. The legacy selector already excluded the old speaker labels from its output; preserve that read compatibility. The reproduced leak comes from adding labels inside a modern intent body, not from the legacy selector. Do not scrub actual request content. Add exact serialized-input and generated-contract regressions, retaining refusal of unmarked legacy tasks and coverage of scout promotion. Fixes https://github.com/kunchenguid/firstmate/issues/3882 * no-mistakes(review): Refuse operator-address lines in Captain's intent body * no-mistakes(document): Document operator-address refusal in intent contract comments * fix: classify OpenCode ellipsis hint as idle (#4451) * fix(composer): recognize Grok 1.0.5's oversized titled bottom border as a proven empty composer (#4455) * fix(composer): accept Grok title overhang * no-mistakes(review): summary: named Grok overhang constant, doc caveat, restored tmux typed-title coverage * fix(bin): translate Stop hook timeout signals into durable auto-arm failure (#4474) * fix(bin): recover Claude auto-arm after timeout * no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list * fix(spawn): establish Claude task channel authority (#4464) * fix(spawn): establish Claude task channel authority * no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference * fix(bin): refuse fm-control.sh exit when the composer holds unproven or pending text (#4458) * fix: guard relaunch exit against pending input * no-mistakes(review): Verifying test run in progress * no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard * no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now * fix(spawn): establish crewmate identity first (#4481) * fix(bin): reconcile redundant secondmate divergence during updates (#4460) * fix: reconcile diverged secondmate updates * no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md * no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md * feat: enable gpt-5.6-luna max reasoning for crew dispatch (#4497) * fix(dispatch): support Codex Luna max effort * no-mistakes(review): use portable CODEX_HOME path in codex effort reference * feat(calm): render smooth Unicode swell with asymmetric two-color sail (#4498) * feat(calm): render smooth Unicode swell * feat(calm): make sails asymmetric * feat(calm): use quarter sail glyph * no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer * no-mistakes(document): docs: sync calm wave phase doc comment * no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。 * fix(bin): supersede stale scout delivery text in brief.md on promotion (#4491) * fix: supersede scout delivery brief on promotion * fix: preserve ship safety contract after promotion * no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch * fix(bin): make captain holds work on hosts with an older JSON::PP, and stop cleanup dropping accents from a held body (#4471) * fix(bin): let captain holds work on hosts with an older JSON::PP Holding a task for the captain, and the cleanup that keeps a captain-held row open, both fail outright on any host whose JSON::PP defaults allow_nonref off - 2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`, but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older library rejects that whole value with "must be object or array". The consequence is fleet-wide on such a host, not one broken command: a worker there cannot formally record a decision for the captain at all. It can only mention the decision in passing in a status line, where it can be missed - which is how a real decision goes unrecorded. The hold reports that the task lost its hold-set stamp; the cleanup cannot return the row to Queued. Both call sites now ask for allow_nonref explicitly rather than inheriting whatever the installed library defaults to. The second one is worth naming: its `/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string case that fails, so the guard selects for the failing input rather than protecting against it. The regression case forces the older default back off for every perl the commands spawn, then drives both paths - holding a task that carries a body, and tearing down a captain-held row whose deliverable must still be appended. It also probes that the simulation genuinely rejects a bare scalar, so the case cannot pass vacuously on a lenient host. Each half was verified failing on its own unfixed call site with that site's real error message. Suites: fm-captain-hold-lifecycle 51 cases, fm-backlog-atomicity 99 cases, 0 failures. Verification limit: the mechanism is reproduced and tested, but neither fix is verified against a real JSON::PP 2.27202 host, because none is in the loop. This laptop runs 4.06, where the bug does not manifest. `bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a brace-delimited object before decoding, so allow_nonref never applies. * fix(bin): stop cleanup silently dropping accented characters from a held body Cleanup rewrites a captain-held row's body to append the finished work's deliverable, and the decoder it reads that body with printed decoded characters to a stream with no `:raw` layer. A character at or below U+00FF then came out as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the accent. `fm_backlog_retain` writes that body straight back through `--body-file`, and nothing reported an error - the character was simply gone from a row still waiting on the captain. The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus `utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already used. Review of the parent commit found this on one of the lines that commit already changed. It predates that change. The test asserts bytes rather than decoded strings, because comparing strings cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any character above U+00FF makes perl print the whole string as UTF-8, so one body carrying both an accent and an em dash passes even unfixed and proves nothing. Verified failing before the fix on the accented row, passing after. Suites: fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures. * no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc * no-mistakes(review): drop whole-file UTF-8 check from retained-body test * no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc * no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc * fix(bin): read codex 0.154's idle braille starfield rows as composer furniture (#4532) * fix(composer): read codex 0.154's idle starfield and status footer as furniture codex-cli 0.154.0 animates a braille "starfield" around its idle composer: on the row above the bold `›` prompt row, on the `›` row behind the SGR-2 dim `Ask Codex to do anything` placeholder, and on the row below it, then draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`). The cells are truecolor greys on both sides of the ghost luminance ceiling, so the brighter ones survive ghost stripping, and the rows below the glyph carry no structural edge. The shared classifier selected the bare `›` shape, extended its wrap region over the two rows beneath the glyph, read the survivors and the footer as wrapped typed input, and answered `pending`; the steering doorbell defers on exactly that verdict, so no doorbell ever reached an idle codex 0.154 pane. bin/fm-composer-lib.sh now recognises that furniture by shape, declared once next to the idle placeholders and reached from the two wrap-region boundary points: - a row whose non-whitespace content is entirely braille cells (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it never counts as wrapped typed content and bounds a bare composer's wrap region; braille behind the glyph row's content is stripped before the emptiness decision when nothing else follows the glyph; a row mixing braille with other text stays typed content; - the codex status footer bounds the wrap region exactly as omp's status row does, anchored on the effort token, a spaced middle dot, and a `~` or `/` path cell, so a typed `fix · tests` stays composer input; - `^Ask Codex to do anything$` joins the verified idle-placeholder set; the ghost strip remains what proves that row empty, and the bare-row rule that bright placeholder text is real input is unchanged. Unchanged: the strict blank-row rule, the styled=0 degradation (a plain cmux/orca capture of this screen still reads `unknown`, never `pending`), FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape. tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte with the divergence (letters in place of the starfield read `pending`) and the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh is the default-on live guard (token-free, skips explicitly without codex or tmux) that launches the installed codex idle and asserts `empty` through both the tmux and the cursorless styled reads, naming codex --version on failure. docs/verification/runtime-backends.md records the dated Herdr evidence: `pending` before, `empty` after, on the captured screen. * no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry --------- Co-authored-by: Todd Billings <todd@usdvcapital.com> * fix(bin): refuse empty text steers in fm-send (#4259) * fix(bin): refuse empty text steers in fm-send A marked secondmate request sent with an empty message delivered only marker and correlation bytes and minted a pending-reply expectation the parent could never see resolved, stalling the fleet with no loud error (#4255). Fail closed on an empty or whitespace-only message on the text path, mirroring the existing --resolve-key refusal. * chore: retain ambient Pi-lens autoformat as its own commit Formatting-only edits produced by ambient Pi-lens autoformat during the msg-loss investigation, kept separate from the behavioural change in c23acba6 so the fix stays reviewable on its own. AGENTS.md is deliberately excluded: its only autoformat edit stripped the trailing space from the documented FM_OPERATIONAL_PREFIX value, which bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11 records as permanent compatibility. Documenting that constant without its trailing space makes the doc wrong about the contract, so that one line was restored rather than retained. * fix(calm): paint the working ship one yellow over all-blue water (#4554) On rose-pine-moon the two-color water (cyan crests over blue troughs) read as a pink stripe over aqua, the yellow left sail and mast clashed with the red right sail, and the hull carried a blue interior run. Every water cell is now blue so the swell reads through glyph height alone, and both sail halves, the mast, and the whole hull are one yellow run. Geometry, cadence, animation, direction flip, resize clamping, and the narrow fallback are unchanged. Update the unit and real-TUI color assertions to the new palette and the Calm docs that described the old one. * fix(bin): stop aging a second mate's active turn from its launch (#4270) * fix(watch): stop aging a second mate's active turn from its launch The parent watcher's second-mate wake-loop stall check exempts a mate that is demonstrably inside an active turn, but secondmate_in_active_turn asked busy_turn_over_age first and returned "not in a turn" whenever that said the bound was crossed. busy_turn_over_age ages from state/<task>.turn-ended, falling back to state/<task>.meta. A second mate's turns end in its own home, so the parent never gets a turn-ended mark for it and the fallback ages the mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago was therefore permanently "over age", the busy pane was never consulted, and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false wake-loop stall. The gate now bounds the busy exemption by <idle> - how long the queue's drain position has not moved - which is evidence this home actually holds. A busy mate stays exempt while the queue has been frozen for less than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so the bound that stops a busy pane from proving liveness forever is kept rather than removed. busy_turn_over_age is untouched; its remaining callers are the ordinary crew busy-pane bound. The regression pins the case that actually broke: a mate whose launch record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn must not escalate, while the same mate with its queue frozen past the bound still publishes exactly one notification. The existing coverage only exercised a freshly launched mate, which passes either way. Reaching that alert now costs a pane capture inside the gate, so the three checkpoints in this suite that assert an alert move from a 1s to a 4s bound - the value the neighbouring active-turn cases already use. The bound is a ceiling, not a wait: the checkpoint returns on the first actionable wake. On a loaded machine a 1s bound missed the alert repeatedly; at 4s it did not miss in 20 runs under the same load. * no-mistakes(review): scope the second-mate active-turn regression test's coverage claim * no-mistakes(document): fix stale second-mate active-turn comments in fm-watch * feat(bin): add read-only PR blocker and reviewer discovery commands (#4278) * feat(bin): add read-only PR blocker and reviewer-discovery commands Two focused, opt-in commands that read GitHub and never write to it. fm-pr-state.sh reports what still blocks one pull request from the author's side: a closed or merged state, draft state, unknown or conflicting mergeability, absent or failing required checks, and a blocking CHANGES_REQUESTED decision explained by each reviewer's latest verdict, marked STALE when it was left at a superseded head. A pull request that only awaits an approval is not reported as blocked, and advisory checks are omitted. Every reading is taken against one exact head; a push that lands mid-read invalidates the whole result rather than mixing two snapshots. fm-pr-reviewers.sh suggests reviewers from the most recent commits to the pull request's exact changed paths, counting each commit once, resolving handles through GitHub's own commit author.login mapping, and excluding the author and Bot accounts. Both stay read-only: no review request, no approval, no merge. Unresolved review-thread state is left unreported because the REST API does not expose it and unattended commands may not use GraphQL. Closes #3731 * no-mistakes(review): accept only PR URLs and stop at terminal state * no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate * no-mistakes(review): stop attributing readings to unverified heads * no-mistakes(review): narrow readiness contract to checks that have reported * no-mistakes(review): read the pull request once, drop the head guard * no-mistakes(document): scope pr-forge isolation proof to its measured members * no-mistakes(document): record uncovered pr-forge members and their pending proof * docs(isolation-proof): re-prove pr-forge at its full membership tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the pr-forge family in this branch, and script_allows_concurrency grants four workers by family membership alone, so both ran concurrently on a proof measured before they existed. Re-proved the family at all eight members: two consecutive runs, 0 failures, each begun with the one-minute load average below 6.0 so the result measures isolation rather than contention. A third run taken between them is disclosed rather than recorded, because it started while the previous run's workers were still decaying. The new durations are not comparable with the six-member measurement above them, so they are not presented as evidence about the two new members, and that record's 1.72x four-worker figure is left as a statement about its own run rather than restated as current. * no-mistakes(review): disclose gh error-text coupling at its matching site and tests * fix(bin): teach validation-round pauses in generated briefs (#2752) * fix(bin): teach validation-round pauses in briefs * no-mistakes(document): Point classifier comments to authoritative pause examples * docs(readme): add star history chart (#4558) * fix(bin): refuse teardown when a task's endpoint close fails (#4510) * fix(teardown): refuse a cleanup whose endpoint close failed bin/fm-teardown.sh discarded both the exit status and the stderr of every fm_backend_kill call, so a close that genuinely failed was indistinguishable from one that succeeded. Teardown continued past it, deleted the task's durable records, returned its worktree, and reported the cleanup as completed. The deleted metadata is the only record of which endpoint belongs to the task, so such a close did not merely leave a stray session behind, it stranded one: nothing was left on disk naming it. The adapters could not carry that signal either. Driven against the real code, every backend arm returned 0 for a genuine failure exactly as it did for an already-exited endpoint, so there was nothing for the four call sites to propagate even once they stopped swallowing it. The tmux arm now resolves a close that did not succeed against the window's exact recorded identity, since kill-window fails the same way for a window that is gone and one that is still there. The Orca arm reports a close its missing CLI never attempted. Both stay silent for an endpoint that is already legitimately gone, and the remaining arms are unchanged: their close-command timing cannot be established without the real Zellij, Orca, and cmux binaries, and a gate that refused ordinary cleanup of an already-exited session would be worse than the defect. docs/verification/runtime-backends.md records what each backend can prove. A reported close failure now reaches teardown's existing retain-and-stop refusal before the records naming the endpoint are removed, matching where the Herdr confirmed-gone gates already sit for the same hazard, and the retained records let a rerun finish once the close works. * no-mistakes(review): refuse unreadable tmux close re-read; honor --force override * no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close * no-mistakes(document): document endpoint-close refusal in its backend and retirement owners * no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree * feat(calm): add flag-gated Claude Code Calm mode (#4565) * feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that brings Calm to Claude Code: the sailboat replaces the stock working row through a Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration, and canonically classified operational user rows draw at zero height. /calm is registered by the hooks module itself and toggles the same per-home config/calm preference the Pi extension uses, so one choice applies on either harness; rows redraw retroactively on toggle and stay hidden across claude --continue. The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS flag is on. Nothing sets that flag in any settings file, and the plugin carries no command file, skill, agent, or classic hook, so it is a complete no-op while the flag is off. The trusted project auto-loads it through an .agents/skills symlink, the only path Claude Code scans for project plugins. Extract the working-ship geometry, bounce track, cadences, and freeze/resume state into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module imports from outside the plugin folder) and have the Pi widget paint that core's frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify operational rows through a port of bin/fm-operational-input.sh's classify command guarded by a corpus parity test against the shell owner. Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering, Raster packing, policy, classifier parity), the mod's own claude plugin test suites behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op, the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272. Docs: record the version-scoped Claude Code evidence and the three bounded gaps in docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md, and make the shared preference, layout, and contributor notes harness-neutral. * no-mistakes(review): Preserve colliding final replies and strengthen parser parity * no-mistakes(review): Preserve final replies and strengthen canonical parity checks * no-mistakes(review): Require exact function-hooks opt-in before Calm activation * no-mistakes(review): Clarify Calm module loading and activation boundaries * no-mistakes(review): Reset Calm presentation state across session starts * no-mistakes(document): Refresh Calm session lifecycle documentation * feat(calm): paint the Claude Code working ship in Claude's own theme colors The captain picked the "Claude native" palette for the Claude Code mod's Raster: every water cell takes the spinner blue of the active theme family (#93a5ff dark, #5769f7 light) and the whole boat takes the Claude orange of the stock spinner (#d77757), one water color and one boat color. The family follows the `theme` setting's prefix, read at load through $.config.list and re-read on a config.set of that row, with `auto` and custom themes falling back to the dark set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte. Rename the shared sprite's color classes from hue names to `water` and `boat`, since each harness now maps them to its own colors; geometry, motion, cadence, and the activation gate are untouched. Tests cover both palettes' packing and the family rule under Node, and the plugin kit drives every theme value, a theme change mid-session, the Calm-off pass-through, and inertness of the menu read while the flag is off. The docs describe the Claude Code colors and record the guard passing on 2.1.273. * no-mistakes(review): Use light palette for unresolved Claude themes * no-mistakes(document): Refresh Claude Calm verification evidence * fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586) * fix(watch): honour a declared wait before wedge-escalating a quiet pane wedge_timer_check escalated on elapsed idle time alone. Nothing asked whether the worker had already said why its pane was quiet, so a lane that declared a bounded external wait climbed the escalation ladder for as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every repeat carried demand-deep-inspection - which by its own wording forbids re-absorbing on the run-step or pane state, so the supervisor could not use the evidence that was there either. The generated brief promises that declaring `paused:` buys the long recheck cadence instead of a wedge, but the timer was still reachable while that declaration stood: a crew that declares a wait and then has an active run or busy pane attributed to it is handed to the timer as provably-working. The declaration is what the worker said about its own silence, so it now outranks a liveness verdict that only says something is running. The consult runs in the at-threshold branch that was about to escalate, beside the worktree walk already there, and costs one status-line read. Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS recheck the declared-wait absorber already uses, so the wait is still rechecked and cannot rot invisibly. Which verb declared it decides the wording, because the two block on different people: a `paused:` wait is owed by an external dependency and asks the reader to confirm it still holds, while a `captain-held:` transfer is owed by the captain reading the recheck and asks them to answer or release the hold. A hold is not rechecked at all while the away-posture record exists, as on every other captain-held path, and that absorb arms no throttle so the recheck is owed in full on return. A declared clearing time that has already passed stops counting, and a lane that never declared one keeps the identical escalation schedule, reason, count and demand-deep-inspection wording, so detection and its worst-case time are unchanged. The deferral restarts the idle timer rather than cancelling it, so a lane that stops waiting escalates again within one threshold. A lane quiet because its own validation run is parked at a gate awaiting a human decision is deliberately out of scope: reading that state needs a signal carrying who the wait is on and what clears it, rather than one inferred from a parked verdict that also covers gates awaiting the crewmate itself. Tests pin both directions for each case and were each confirmed to fail with the consult removed. * no-mistakes(document): docs: honour declared waits in stale-escalation docs * fix(bin): report verified PR state for passed runs (#4624) * fix(bin): derive passed PR state from PR record A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists. For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim. Fixes #4607 * no-mistakes(review): Add bounded GitLab merge-request state reads * no-mistakes(review): Preserve network-free inactive crew-state scans * no-mistakes(document): Document PR record readers in shared library * fix: restore published contribution follow-up (Fixes #4469) (#4627) * fix: restore published contribution follow-up (Fixes #4469) * fix(review): Fix contribution freshness and merge actor routing * fix(review): Restore issue triage and scope contribution follow-up * fix(test): test: assert one wake per contribution signal * fix(document): Document contribution follow-up * fix: restore truthful terminal delivery evidence * fix(review): Disclose unsupported contributions and deduplicate watcher wakes * fix(review): Preserve unmeasured unsupported contributions across Bearings * fix(review): Deduplicate shared contribution wakes and isolate diagnostics * fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower * fix(bin): make remote report transfers explicit and fail-open (#4658) * fix(bin): make a remote-reply document gap self-clearing and re-attemptable A remote mate's undelivered document raised a keyed `blocked` decision that nothing could ever resolve, and any `data/*.md` substring in any mirrored line was an unconditional fetch instruction. A mate announcing a report it had not written yet therefore manufactured a permanent, factually false blocker, and its own explanation of the false alarm manufactured more. The reader has no permanence vocabulary: a report still being written refuses exactly like a path that will never exist. So an undelivered document is now a durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`, re-attempted on the next delta and on the channel's own quiet poll, and retired with a matching `resolved` line naming the local copy once it arrives. The cursor still advances and no delta stalls on one bad pointer. Only a structured `report=data/....md` pointer now offers a document, so a path merely mentioned in prose - including one under another home's mirror tree, which is provably not that mate's to serve - is never fetched. Offers are deduplicated across the whole delta, the escalation names each missing document once and carries the reader's own reason instead of discarding it, and a strictly increasing notice ordinal keeps a later escalation from being swallowed as duplicate bytes. A mirrored line still lands once whichever pointer form it was first written under. * no-mistakes(review): Require structured pointer token boundaries * no-mistakes(review): Unify boundary-safe pointer extraction and rewriting * fix(bin): identify a mirrored line independently of its delivery state Two defects in the boundary-safe pointer work. The at-most-once check compared only the all-remote and all-local renderings of a line, so it could not recognize a mixed one. A line offering two documents where only the first was deliverable mirrored as local-plus-remote; once the second arrived, a cursor-loss whole-log recapture rendered the same line all-local, matched neither alternate, and mirrored a second time. A line's identity is now the canonical form every boundary-valid pointer would take once delivered, derived by the same parser that does extraction and rewriting, so it no longer depends on which documents happened to be deliverable at the time. The pointer map was passed to awk through the process environment. A delta may carry up to the configured 1 MiB bound, and an expanded map of delivered pointers can exceed the platform's exec argument limit, so awk would fail to start; because no caller checked, the empty result would have been appended as blank lines while the cursor advanced past dropped status content. The map now travels in a file, and every call site checks the exit status and stops the ingest rather than committing a delta it could not render. Both passes now run once per stream instead of twice per line. * no-mistakes(review): Abort ingest when document pointer extraction fails * no-mistakes(review): Exclude structured cross-home pointers from document transfer * fix(bin): fail open on an undeliverable remote document instead of tracking it Narrow the remote-reply document fix to the scope the diagnosis actually requires, as decided after measuring a simpler alternative. A document the reader cannot deliver now fails open. The mate's line is mirrored with its own pointer, the cursor advances, and one unkeyed note carries the reader's reason. A note never enters the open-decision fold, so it cannot stand open the way the original keyed block did - which removes the never-clearing false blocker by construction rather than by resolving it. That makes the durable self-clearing obligation unnecessary, so it goes: the per-mate pending-documents record, its notice ordinal and resolved announcements, and the poll-side retry. Canonical line identity goes too, and with it a way to silently drop a genuine status line; mirroring is back to at-most-once on exact bytes. The cross-home exclusion goes as well: under fail-open a cross-home report= either fails harmlessly or is a nested remote report this mate genuinely holds, which is now relayed again. Kept: fetching only on a structured report= pointer, the boundary-correct parser, the file-based rewrite map, and checked extraction and rewrite exit status. The parser now scans behind a sentinel byte so a rejected candidate can no longer give the text right after it a false leading boundary. The reported incident is covered end to end: a report path announced in prose before it exists raises no decision, and the report still arrives through the ledger publisher's structured offer once written. * no-mistakes(review): Preserve source-line identity across remote reply replays * no-mistakes(document): Document remote reply transfer and replay semantics * no-mistakes(lint): Fix staging truncation lint checks * fix(calm): preserve substantive mid-turn responses (#4655) * Preserve substantive Calm mid-turn text * no-mistakes(review): Distinguish newline-preserved replies from short narration * no-mistakes(document): Document Calm mid-turn preservation boundaries * no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check` * fix(bin): preserve PR merge polls across volume remounts (#4656) * fix(bin): re-record PR poll identity after a volume device renumber (Fixes #4260) A volume remount can renumber the state filesystem's st_dev while every inode and byte stays the same; APFS does this across a reboot. A poll registration records its sidecar and check as device:inode, so every poll armed before the remount failed strict validation and the watcher refused all of them as unauthenticated state checks until each was re-armed by hand. There are two device comparisons. fm_pr_private_file_valid compares a live file's device with the state directory's device read in the same invocation: it refuses a file that is not on the state directory's own filesystem and already survives a renumber, so it is unchanged. The registration's recorded identity versus the live identity (from #556, reused by the #932 retirement receipt) binds the registration to the exact files published in its own transaction; its device part is what breaks. When strict capture fails, the watcher now proves the device is the only difference: every other artifact check passes (template bytes, both hashes, private mode, single link, live device, metadata), both recorded identities name one device, and each recorded inode equals its live inode. Only then, under the task's control lock, does it rewrite the two identity lines, repeating the whole proof and comparing the registration's file identity and bytes just before the rename, and then capture strictly again. A swapped, altered, re-moded, relinked, split-device, or foreign-device artifact still fails a proof and is still refused, and a pending retirement receipt blocks the rewrite. Reproduction: on macOS a poll armed on an APFS disk image that was detached and re-attached behind another image moved st_dev 16777239 -> 16777243 with inodes, bytes, mode, and link count unchanged; the real watcher refused it on main and reports its merge with this change. The portable regression test rewrites a real registration's recorded device and drives the watcher. Not changed here: the status presentation cursor keys rows by its own device:inode identity in bin/fm-classify-lib.sh, a different helper that needs its own fix; a retirement receipt left by a reboot between its publication and removal still names the old device and stays refused; custom check trust binds only a content hash and is unaffected. * fix(review): Serialize PR poll publication writers * fix(review): Bound PR poll publication lock scope * fix(bin): keep contribution records when the poll budget runs out (follow-up to #4627) (#4661) A budget that expires partway through an observation no longer records an error or prints the unavailable wake; the URL keeps its prior record and is observed first next poll. forge() flags budget exhaustion at the point it refuses, or when a read is killed at the budget's own deadline, so a genuine forge failure still records the error and wakes. Each distinct URL is now observed once per poll and applied to every owning task. * fix(bin): clear parent pending-replies on local secondmate retirement (#4680) * fix(bin): clear parent pending-replies on local secondmate retirement Local secondmate teardown left resolved parent pending-reply records behind after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced retirement while any reply for that id is still unresolved, and delete every matching record plus its delivery confirmation after a successful local or remote retirement, matching the remote cleanup path. * no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup * no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung * no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern * no-mistakes(review): Pending-replies Basename und corr_id abgleichen * no-mistakes(document): Clarify forced retirement pending-reply cleanup --------- Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc> * fix(bin): accept Orca's composite worktree id when tearing down a task (#4677) * fix(bin): accept Orca's composite worktree id at teardown Teardown refused every Orca-backed task because the endpoint validator checked orca_worktree_id with the simple-atom rule meant for tmux-style window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca returns that id as `<orca id>::<absolute worktree path>`, so the colon and slashes in every real value made validation fail and finished Orca tasks could never be cleaned up. Validate the field as the composite it is: both halves of the first `::` split present, the path half absolute, and no embedded newline, carriage return, or tab. The terminal field keeps the atom check, which is correct for it, and no other backend's validation changes. The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca never returns, which is why the suite passed a check the real value fails. They now carry the composite form, so the tests exercise the real value. * no-mistakes(document): name Orca's repo id in the composite worktree id * no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points * feat(bin): add opt-in typed dispatch resolution (#4692) * feat(bin): add opt-in typed dispatch resolution through typesafe.ai Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or scout profile from a written brief with typesafe.ai's System One model: one Choice question over the rules' `when` texts, then the confidence floor, the rule's `approval` and `floor`, each profile's `provider` and `floor`, one quota-axi snapshot, and the spendPriority argmax all in code. It is off unless TYPESAFE_API_KEY is in the environment or the home's gitignored .env; off means one stderr line, exit 0, and no network call, so firstmate dispatches exactly as before. The key reaches curl on a file descriptor, never argv. Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates the four new optional dispatch fields. Document the schema, the operator contract, the AGENTS.md intake step, and the live and benchmark evidence. * no-mistakes(review): Harden typed dispatch resolution and quota bounds * no-mistakes(review): Validate dispatch floors and ranking evidence * no-mistakes(review): Tighten dispatch response and floor evidence * no-mistakes(review): Neutralize none matching and resolve defaults locally * no-mistakes(review): Preserve providerless profiles outside typed resolution * no-mistakes(review): Validate response usage and reject duplicate profiles * no-mistakes(review): Escalate unverifiable floors and validate probabilities * no-mistakes(review): Validate probability mass and unknown profile floors * no-mistakes(review): Simplify resolver interface and preserve fallback routing * no-mistakes(review): Fix constants and rank partial quota evidence * no-mistakes(review): Add authoritative provider mapping and enforce explicit providers * no-mistakes(review): Declare provider for documented Pi profile * no-mistakes(review): Validate provider identifiers and support Gemini dispatch * no-mistakes(review): Strictly anchor provider identifiers * no-mistakes(review): Validate selectors and preserve fallback candidate evidence * no-mistakes(review): Gate typed validation and harden resolver evidence * no-mistakes(review): Preserve opt-in routing and harden candidate evidence * no-mistakes(review): Prioritize known exhaustion over quota uncertainty * no-mistakes(review): Isolate API secrets and preserve no-key diagnostics * no-mistakes(review): Fallback safely when dispatch rules are absent * no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets * no-mistakes(document): Document typed dispatch safety and fallback behavior * fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753) * test: reproduce buried status declarations in shared readers * fix: share status event reads and preserve open blockers * fix: retain terminal scout and ship status declarations * no-mistakes(review): Fix status chronology, legacy completions, and reader performance * no-mistakes(review): Share terminal decision reconciliation across fleet snapshots * no-mistakes(review): Unify terminal supersession across cached folds and consumers * no-mistakes(review): Filter per-key status history while preserving terminal chronology * no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells * no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses * no-mistakes(document): Document latest-event status read and kind-scoped fold cursor * no-mistakes(lint): Quote literal done in test for-lists for SC1010 * ci: expect 19 snapshot/fleet-view tests This branch adds a fleet-snapshot regression, so the stock macOS Bash lane's hardcoded guard of 18 'ok - ' lines fails on the new count. Bump the guard and its message to 19. * no-mistakes(review): Restore multiline child outcome reporting * no-mistakes(review): Select ledger terminal events through bounded shared reader * no-mistakes(review): Report newest open decision instead of preferring blocked * no-mistakes(review): Require colon before ship/scout terminal supersession in fold * no-mistakes(review): Gate socket-down override on latest event; drop lock matrix * no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions * no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold * no-mistakes(test): Update fleet-view expectations to newest-open-decision rule * no-mistakes(document): Align status-read docs with fold-resolved crew state * no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers * no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree * test: fold terminal-cleanup snapshot coverage into the completed-scout case Keep the ship/scout/secondmate supersession assertions without adding a nineteenth top-level fleet-view test, so CI can stay at the upstream suite count. * no-mistakes(document): Clarify socket-down override expiry in architecture doc * ci: retrigger flaky contribution check * fix(bin): launch codex crewmates with codex's hook layer disabled (#4689) * fix(spawn): launch codex crewmates with codex's hook layer disabled A freshly launched Codex worker never reached its instructions. Codex stopped it on an interactive "Hooks need review" modal whose selection sits on "Review hooks", which is neither trusting nor declining. Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow navigation, so the selection cannot be moved, and pre-accepting the prompt by writing Codex's own trust store would record an operator consent that was never given. The hooks are the machine's own ~/.codex/hooks.json plus any project's .codex/hooks.json. A crewmate needs neither: its turn-end signal is the -c notify= program on the same launch, and Firstmate's project hooks are primary-session infrastructure that stands down in a child worktree. Crewmate and scout launches now pass --disable hooks. That is the opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted hooks; disabling the feature runs none of them and leaves the operator's ~/.codex untouched. An unknown feature name is a hard Codex error, so a release that drops the flag fails the launch loudly instead of silently restoring the modal. A secondmate is a primary in its own home and keeps the project hooks its turn-end guard and session-start digest ride on. Verified on codex-cli 0.151.0: the modal is gone and the turn-end notification still lands. This unblocks the second review that every finished pull request is supposed to get. Fixes kunchenguid/firstmate#4673 * no-mistakes(review): Fix contradictory hook count in Codex verification record * fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710) * fix(bin): settle terminal contributions and wake once per read-failure episode A contribution whose last good observation is merged or closed is final: poll no longer re-reads it, projection keeps it fresh, and a stale error recorded beside it is cleared once. A genuine forge-read failure on an open contribution still records its error on every cycle but prints the unavailable wake only when it starts a failure episode; a successful read ends the episode. Open PRs linked from done tasks keep being observed. The false unavailable beside a complete observation was budget exhaustion mid-observation, already fixed by #4661. * fix(review): Settle terminal contribution owners * fix(review): Deduplicate shared contribution failure episodes * fix(test): Preserve settled terminal contribution records * fix: select authoritative no-mistakes runs (#4476) * fix(crew-state): select authoritative validation runs by identity Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row. Refs: https://github.com/kunchenguid/firstmate/issues/3215 * fix(review): Resolve same-branch run identities beyond capped history * fix(review): Fix run-selection compatibility, races, and worker-state fallbacks * fix(review): Limit run validation to the requested branch * fix(test): Anchor AXI fixtures and document remaining live evidence gaps * fix(document): Clarify run selection documentation and capture ownership * fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation * fix: distinguish captain outcomes from no-op updates (#4738) * fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work MAIN answered a supervision-branch outcome for completed captain-requested work (implementation done, PR ready for review and merge approval) with "Captain, shipshape.", reading section 9's no-action reply as covering it and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no captain-facing response is owed". Section 9 now limits the shipshape reply to true no-ops (idle re-read, empty heartbeat, consequence-free acknowledgement) and requires a short outcome response naming what finished and what word is needed whenever requested work finishes or a result needs the captain's word, even when a transcript entry already shows the substance. The Pi protocol's re-emit rule now says it bounds repetition only, and carries a worked example of the ready-for-review outcome whose correct processing turn a shipshape reply fails. No executable contract evaluates the content of MAIN's captain-facing reply, so the regression is the protocol example in the owner doc rather than a text-match test. * no-mistakes(document): Clarify captain-facing outcomes versus no-ops * docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line The document step condensed the Pi protocol's re-emit rule and dropped the worked example of a finished, ready-for-review outcome whose correct processing turn a "Captain, shipshape." reply fails. That example is the contract's regression: no executable contract evaluates the content of MAIN's captain-facing reply, so the owner doc's example is the test case. Restore it directly under the re-emit rule, prefixed as a regression example that is kept verbatim and never condensed or summarized away. * no-mistakes(review): Clarify captain outcome and decision-word requirements * no-mistakes(document): Clarify captain-facing completion outcomes * docs(pi): require the PR URL in the visible captain-facing outcome reply Captain review on the regression example: drop the sample reply string and say only that the ready-for-review outcome requires relaying a captain-facing outcome response, not just "Captain, shipshape.". Fold in the visible-PR-handoff failure seen this session: after the branch outcome reporting this fix green, MAIN's visible reply was only "Awaiting your merge call." with no PR URL, leaning on the dim anchor. Section 9's URL rule now also covers a review or merge ask and names the visible reply as where the URL goes, sourced from the ready status, pr= metadata, or the supervision branch's summary and never left to a transcript entry. The Pi protocol adds the same-way failure and places the captain-facing text in the final visible assistant reply after the fm_branch_processed call, because Calm hides assistant text emitted in the same step as a tool call as a working note. Investigation verdict, evidence in the PR comment: no recent PR caused the handoff failure; Pi has hidden same-step pre-tool assistant text since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and #4658 touched only remote report transfer. * no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs * no-mistakes(document): Clarify captain-facing supervision outcomes * docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule The consolidated section 9 URL rule narrowed its trigger to a review or merge ask, dropping the "whenever a PR is mentioned" catch-all from #3648 that keeps every PR URL copied from a durable record and never assembled from memory. Restore that trigger as a union with the review or merge ask so the one consolidated rule covers both. * fix(bin): let non-owner Claude Stops exit safely (#4777) * Fix foreign-owner turn-end supervision loop * no-mistakes(review): Scope foreign-owner safe exit to Claude guard * no-mistakes(document): Document Claude foreign-owner safe exit * fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778) Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty indexed array as an unbound variable and aborts the shell. In signal_turnend_panes_churned() the missing_keys loop was reachable with an empty array whenever every churned key already held a fresh .churn-since-* marker (a second churning turn-end inside an open deferral window), so each watcher cycle died about half a minute in and supervision restarted endlessly. The created_keys rollback loops had the same latent crash on their error paths. Audit of bin/ for the same pattern found one more confirmed-reachable case: remote_handoff's noncanonical-body scan iterates to_move, which is empty when a retried remote handoff finds every key already staged in the outbox. All other "${arr[@]}" sites are either count-guarded, guaranteed non-empty by construction, or unreachable while empty. Guard the three reachable expansions with the repo's existing "${arr[@]+...}" idiom. New regression test drives a real watcher through the all-marked churn path; the macos-stock-bash CI lane runs it under real /bin/bash 3.2 via FM_TEST_ONLY. * Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783) The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: require complete captain-facing final responses (#4779) * docs: require complete final responses across harnesses * no-mistakes(document): Document complete final replies for Grok Bot * docs: point Grok replies to the shared contract owner * no-mistakes(review): Clarify final recap without batching decision asks * fix: preserve substantive mid-turn text in Pi Calm (#4788) * fix(calm): preserve substantive Pi mid-turn text * no-mistakes(review): Preserve substantive Pi Calm text per block * no-mistakes(test): Cover shared Calm preservation boundaries behaviorally * no-mistakes(document): Consolidate Calm preservation documentation * fix: harden mail checks and rebalance full-coverage CI (#4800) * Improve CI reliability and rebalance full-coverage validation * no-mistakes(document): Clarify lint partition documentation * fix(bin): answer Kimi 2.0.0 folder-trust dialog during spawn (#4799) * Handle Kimi workspace trust dialog * no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers * no-mistakes(review): Gate Kimi ready on any trust marker and clean captures * no-mistakes(review): Read visible pane for Kimi trust and ready gates * no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate * no-mistakes(review): Harden Kimi viewport capture and trust dialog detection * no-mistakes(document): Document Kimi spawn refusal on cmux and Orca * fix(bin): report a dead-agent record once instead of escalating forever (#4775) * fix(bin): report a record whose agent is gone once instead of escalating forever The wedge escalation path never asked whether there was still an agent to be wedged. A wedge is something stuck that might recover, so re-alarming it earns its cost; an agent that is gone never moves again, its pane never churns, the idle timer never resets, and the escalate path clears its own timer and re-arms with nothing bounding the count. Observed on a live fleet: two finished lanes reached 226 and 203 consecutive escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400 notifications a day from two lanes with no agent running at all. On one, fm-control.sh exit answered already-stopped and fm-crew-state.sh read "failed - run failed". Closing the Herdr pane did not stop it either: with the pane genuinely gone and herdr pane read returning pane_not_found, the count kept climbing, because the poll is driven by the record's window= line rather than by the pane. The cost is not the repetition but that it drowns the alarms that matter. fm_backend_agent_state already separates a thinking agent from a gone one at process level. In the branch that was about to escalate, read it once and treat only its two recovery-grade verdicts - dead (endpoint present, no agent in it) and missing (endpoint authoritatively absent) - as proof, reporting that record once and not re-escalating it while it stays that way. Every other verdict, including alive, ambiguous, unreadable, unverified, and a read that failed outright, keeps the identical schedule, reason, and escalation count, so a genuinely wedged live agent is unaffected. The probe costs at most one backend read per window per threshold, the same budget the declared-wait consult and the worktree write probe already take. The report decides nothing about the record's fate: both lanes still held unlanded work and teardown refusing them was correct, so retiring, relaunching, or cleaning up stays with the supervisor. The once-only marker is owned entirely by that function and is dropped by the same read the moment the endpoint stops reading gone, so a replacement launched into the same window escalates normally and its own later death is reported again. Related, and not closed by this: #4412, #4482, #4316. Tests drive the real watcher against a record whose endpoint does not exist and pin both directions: dead and missing report once and never advance the count across later thresholds, while alive, ambiguous, and unreadable endpoints keep escalating with the identical reason and a climbing count. * fix(bin): bind the once-only dead report to the pane it reported Review of the parent commit found a reachable sequence where a later death in the same window lost its promised report. The marker was keyed on the verdict string alone and dropped only when a threshold probe read a non-gone verdict, but probes run only at thresholds: a replacement launched into the same window that dies without ever being probed alive - it crashes at startup, or works and then crashes - was absorbed by the previous death's marker. The pane's first sight yielded only the generic stale wake and every later threshold matched the stale marker, so the second death never got the detailed once-report that both the function's own comment and docs/architecture.md promise. Record the verdict together with the pane hash it was reported for, and absorb a repeat only while both still match. A replacement churns the pane, which resets the stale suppressor, wedge timer, and escalation count while no reset site touches this marker, so the pane half is what tells the second death apart from the first. The live-probe drop stays as it was. Clearing the marker at those reset sites instead would re-open unbounded re-alarming for a dead pane whose display ever ticks, which is the exact defect the pa…
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Complete the PR-check security remediation on existing PR 556 without merging or runtime activation. Validate task identifiers and canonical GitHub PR URLs before side effects; keep watcher polls byte-static with data-only private state, canonical provenance, exact metadata binding, transactional publication, and fail-closed non-executing migration; preserve merge, custom-check, X-mode, and unrelated lifecycle behavior. Require private single-link artifacts and safe cleanup, authenticate all checks at execution, migrate only the exact legitimate historical X polling identity under watcher exclusion to the strict current identity without execution, reject linked, symlinked, byte-mismatched, wrong-device, wrong-mode, and nonordinary lookalikes, and publish X artifacts through same-directory private temporaries plus guarded atomic rename without following destination symlinks. Preserve all prior pipeline fixes, keep the complete rendered public PR diff disclosure-safe, run focused and full validation without warning suppression, push only the existing branch to existing PR 556, and finish at one exact checks-green head.
What Changed
Risk Assessment
✅ Low: The latest changes are narrowly scoped to hardening X artifact reads/writes through the shared private single-link publication and validation path, and I did not find remaining material contradictions to the stated intent.
Testing
The preexisting baseline full-suite result was already successful; I reran focused PR-check, X-mode, and PR-merge suites, captured an end-to-end CLI transcript proving invalid-input refusal, private static poll publication, metadata binding rejection, non-executing migration, and guarded X outbox publication, then reran the complete configured
tests/*.test.shsuite withtmux 3.6a; all passed and the worktree remained clean.Evidence: PR-check security E2E transcript
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
🔧 **Rebase** - 1 issue found → auto-fixed ✅
AGENTS.md- merge conflict rebasing onto origin/main🔧 Fix applied.
✅ Re-checked - no issues remain.
🔧 **Review** - 3 issues found → auto-fixed (2) ✅
bin/fm-x-poll.sh:105- Intent requires "publish X artifacts through same-directory private temporaries plus guarded atomic rename without following destination symlinks", but X inbox publication still writes through$INBOX/$REQ.json.tmpand thenmv -fto$INBOX/$REQ.jsonwithout validatingx-inbox, the temp file, or the destination as private single-link ordinary files. A symlinkedstate/x-inboxor temp path can redirect the write outside private state.bin/fm-x-lib.sh:384- Intent requires "Require private single-link artifacts" and guarded X artifact publication, butfmx_context_registry_setcreatesstate/x-contextwith plainmkdir -pand publishes withmv -f "$tmp" "$file"without rejecting a symlinked context directory or validating the final registry file. This leaves the durable X context artifact outside the new private/single-link guarantees.bin/fm-x-reply.sh:301- Intent requires X artifacts be published through private same-directory temporaries and guarded atomic rename, but dry-run outbox still writes directly withprintf ... > "$outbox_file"after plainmkdir -p. A symlinkedstate/x-outboxor destination can receive the would-be reply payload, and readers can observe a partial write.🔧 Fix: Guard X artifact publication
3 errors still open:
bin/fm-x-poll.sh:49- Intent requires "Require private single-link artifacts" and "publish X artifacts through same-directory private temporaries plus guarded atomic rename without following destination symlinks", but the X poll diagnostic still usesprintf ... > "$ERROR_FILE"after plainmkdir -p "$STATE". A preexistingstate/x-poll.errorsymlink or hardlink is still followed/written instead of rejected through the private artifact boundary.bin/fm-x-lib.sh:292- Intent requires private single-link X artifacts, but the reply-context reader still trustsstate/x-inbox/<request_id>.jsondirectly. A symlinked or hardlinked inbox record can be read byfm-x-reply/follow-up context resolution even though the new publisher would reject writing that artifact.bin/fm-x-lib.sh:481- Intent requires private single-link X artifacts, butfmx_context_registry_getreadsstate/x-context/<request_id>.jsonafter only-f; symlinked directories, symlinked files, and hardlinked records can still supply platform/budget data to reply/follow-up resolution. The write path was hardened, but existing unsafe registry artifacts remain trusted.🔧 Fix: Enforce private X artifact reads
✅ Re-checked - no issues remain.
🔧 **Test** - 1 issue found → auto-fixed ✅
command -v tmux >/dev/null || { echo "tmux is required for e2e tests" >&2; exit 1; }; tmux -V; rc=0; for t in tests/*.test.sh; do echo "== $t =="; bash "$t" || rc=1; done; exit "$rc"🔧 Fix: Fix backend compatibility fixture dependencies
✅ Re-checked - no issues remain.
command -v tmux >/dev/null || { echo "tmux is required for e2e tests" >&2; exit 1; }; tmux -V; rc=0; for t in tests/*.test.sh; do echo "== $t =="; bash "$t" || rc=1; done; exit "$rc"Baseline already reported successful before this round:command -v tmux >/dev/null || { echo "tmux is required for e2e tests" >&2; exit 1; }; tmux -V; rc=0; for t in tests/*.test.sh; do echo "== $t =="; bash "$t" || rc=1; done; exit "$rc"bash tests/fm-pr-check-security.test.shbash tests/fm-x-mode.test.shbash tests/fm-pr-merge.test.shManual evidence transcript exercisingfm-pr-check.sh, generated PR poll execution for OPEN and MERGED states,fm-pr-check-migrate.sh, andfm-x-reply.shdry-run symlink guarding, written to/var/folders/5x/4nqprlbx0518k3ybcb1sz6gr0000gn/T/no-mistakes-evidence/01KXMSQ4Z7DX1WE3CKRY1X98J7/pr-check-security-e2e-transcript.txtcommand -v tmux >/dev/null || { echo "tmux is required for e2e tests" >&2; exit 1; }; tmux -V; rc=0; for t in tests/*.test.sh; do echo "== $t =="; bash "$t" || rc=1; done; exit "$rc"git status --short✅ **Document** - passed
✅ No issues found.
🔧 **Lint** - 1 issue found → auto-fixed ✅
🔧 Fix: Remove unused x-mode test locals
✅ Re-checked - no issues remain.
✅ **Push** - passed
✅ No issues found.