fix: copy PR URLs from durable records - #3648
Merged
Merged
Conversation
Supervision reported a plausible but dead PR link three times because its prompt demanded a full https:// URL at a moment when only a PR number was observable, so the model assembled an owner/repository from memory, and the PR check then accepted that URL and wrote it into the task record, after which the model kept defending its own tool-endorsed guess over the worker's real link. Three changes close that chain without any live forge lookup, so private forges are treated exactly like public ones: - bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy or abstain" section requires a URL to be copied verbatim from a durable record (the done: PR <url> status line, pr= metadata, or the backlog note), forbids assembling owner, repository, host, or number from memory, and has the branch report only the identifier it actually holds when no record names the URL yet, leaving the PR check unarmed until the worker's ready line arrives. AGENTS.md section 7 and 9 carry the same copy-or-abstain rule for main in place of the bare full-URL mandate. - Worker briefs (bin/fm-brief.sh, ship and scout rules) require the full https:// URL wherever a PR is mentioned - status line, terminal, or summary - never a bare "PR 108", so the link is in view as early as the number is. - bin/fm-pr-check.sh refuses, offline and before any side effect, a URL that the task's own done lines contradict, printing both spellings; a log naming no URL still records the argument as before. fm_pr_status_ready_urls in bin/fm-pr-lib.sh owns reading those lines. The refusal also reaches bin/fm-pr-merge.sh, so nothing merges under a contradicted URL. Tests cover the offline refusal with zero side effects, the recorded spelling being accepted, markdown-wrapped and punctuated URLs, working lines not counting, the merge wrapper propagation, a self-hosted merge request with no forge call, the prompt carrying the rule, and the brief carrying the worker rule.
Confidence Score: 5/5The PR appears safe to merge because the previously reported backlog-note PR misbinding is no longer authorized and no blocking failure remains. No blocking failure remains. Reviews (2): Last reviewed commit: "no-mistakes(ci): Removed backlog notes a..." | Re-trigger Greptile |
…rce. PR URLs may now be copied only from the task’s `done: PR <url>` status or canonical `pr=` metadata; otherwise supervision reports only the known identifier and leaves PR checking unarmed. Updated related guidance/docs and verified with branch-supervision tests, brief tests, ShellCheck, and `git diff --check`
AgardnerAU
added a commit
to AgardnerAU/firstmate
that referenced
this pull request
Sep 4, 2026
* refactor(quota): extract mid-task polling and candidate selection into dedicated scripts (#3221)
* Add quota exhaustion detection and safe fallback helpers
- bin/fm-procevent-quota.sh: generic procevent adapter that arms a
recurring quota-axi --json poll and wakes firstmate when a tracked
provider's effectivePercentRemaining drops below a threshold or its
runway.status becomes exhausted_now.
- bin/fm-quota-choose.sh: worker-side helper that picks the first ranked
harness:model candidate with positive effectivePercentRemaining.
- AGENTS.md and .agents/skills/quota-array-dispatch/SKILL.md: document
the new helpers and the mid-task quota-exhaustion wake path.
- tests/fm-quota-choose.test.sh: unit tests with a mocked quota-axi JSON
source.
* no-mistakes(review): Fix quota polling and scope bounds
* no-mistakes(review): Enforce safe default quota selection
* no-mistakes(review): Handle decimal quota values safely
* no-mistakes(review): Fail closed on invalid quota inputs
* no-mistakes(review): Reject empty quota candidate segments
* no-mistakes(review): Harden quota parsing and timeout ownership
* no-mistakes(review): Reuse captured quota snapshots consistently
* no-mistakes(review): Match quota using explicit candidate providers
* no-mistakes(review): Centralize fail-closed quota schema validation
* no-mistakes(review): Reject out-of-range quota percentages
* no-mistakes(review): Validate quota runway status enum
* no-mistakes(review): Tighten quota scope and status contracts
* no-mistakes(review): Preserve unknown quota and exact product bounds
* no-mistakes(review): Preserve provider-level unknown quota
* no-mistakes(review): Reuse canonical verified harness validation
* no-mistakes(document): Document mid-task quota handling
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* fix(docs): restore default routing contract, keep quota helper optional
Restore the AGENTS.md section 4 always-loaded routing paragraph the PR
had deleted, so the standing TOON-first intake, spendPriority ranker,
every-candidate accounting, and load-trigger contract stay exactly as
before this PR. The mid-task quota wake is optional and must not alter
default routing.
Restore the quota-array-dispatch skill ownership line to section 4 as
the always-loaded intake boundary owner; keep the worker-side helper
section as an addition only, without rewiring ownership or load
triggers to section 13.
* fix(bin): use harness-keyed quota matching in optional helper
Revert fm-quota-choose.sh from harness:provider:model tuples back to
harness:model candidates with harness-keyed provider matching, per the
resolved ask-user finding. The helper is optional; authoritative
multi-provider routing (provider discovery from the harness catalog and
quota matching by that explicit provider) stays owned by AGENTS.md
section 4 and the quota-array-dispatch skill intake procedure, not the
helper.
Document the multi-provider limitation in the helper header and the
quota-array-dispatch skill: the helper maps each harness to one primary
provider family only, so a candidate whose established provider differs
from that primary family is checked against the wrong quota row. Use it
only when the brief fixed the candidate order and every candidate's
provider is the harness's primary family.
The helper still consumes one already-captured default-TOON or JSON
snapshot via stdin or --snapshot and never calls quota-axi itself, so
it selects from the same quota state as the intake.
* no-mistakes(review): Fix Muse quota mapping and helper contract docs
* no-mistakes(review): Reject known-empty quotas and map quota tests explicitly
* no-mistakes(review): Preserve unmeasured candidates and enforce snapshot reuse
* no-mistakes(review): Fix quota retirement and dependent regression coverage
* no-mistakes(review): Accept zero-row quota TOON snapshots
* no-mistakes(review): Enforce quota semantics status consistency
* no-mistakes(review): Veto dispatch on any exhausted applicable scope
* no-mistakes(review): Record exhausted quota scope in wake details
* no-mistakes(review): Fix quota help and control dependency coverage
* no-mistakes(review): Decode quoted TOON fields and document quota wakes
* no-mistakes(review): Validate zero-row TOON and map timeout coverage
* no-mistakes(review): Reject multi-value JSON and malformed TOON envelopes
* no-mistakes(review): Validate complete nonzero TOON envelopes
* no-mistakes(review): Accept producer-shaped quota TOON envelopes
* no-mistakes(review): Support empty quota arrays and validate counted rows
* no-mistakes(review): Harden TOON completion, scopes, and quoted fields
* no-mistakes(review): Preserve unknown-headroom exhaustion and reject trailing fields
* no-mistakes(review): Allow unknown headroom under known semantics
* no-mistakes(review): Reject noncanonical quota identities
* no-mistakes(review): Preserve empty quota polling and validate attention identities
* no-mistakes(review): Reject noncanonical provider watches
* no-mistakes(review): Validate all candidates before quota selection
* no-mistakes(document): Correct quota helper safety documentation
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* fix: surface comments on Lavish annotations (#3371)
* fix(bin): keep typed Lavish comments when an element is also annotated
read preferred element text over prompt, so an annotate-and-comment
item dropped the captain's words. Surface prompt as its own field.
Co-authored-by: Cursor <cursoragent@cursor.com>
* no-mistakes(review): Filter non-comment prompts from Lavish reader output
* no-mistakes(document): Clarify Lavish comment presentation contract
* no-mistakes(ci): Fixed Lavish reader comment provenance: non-choice prompts are now emitted even when identical to element text. Added observable regression coverage for identical selector+comment input while retaining pure annotation/message coverage. Reader cases, bash syntax, and diff checks pass. Full fm-procevent suite stops earlier at unrelated “reconcile never claimed” setup failure
* no-mistakes(ci): Fixed duplicate pure-annotation prompts by emitting `prompt:` only when it differs from captured element text. Updated behavioral coverage for selector+comment, pure annotation, and pure message cases. Focused reader regressions, syntax checks, and diff checks pass. Full suite remains blocked by the pre-existing “reconcile never claimed the registered source” failure
* fix(bin): always emit Lavish comments and use real annotation fixtures
Stop inferring comment provenance from prompt==text. Real pure
annotations have no prompt, so always-emit does not duplicate.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: support first public-followup registration on Bash 3.2 (#3420)
* Fix public-followup register crashing on empty lock arrays under bash 3.2.
bash 3.2 with set -u treats "${arr[@]}" on an empty array as unbound, so the first register in a fresh home aborted before taking the registry lock.
The empty-lock regression also runs under the existing stock macOS Bash CI lane so pre-fix code would fail there.
* no-mistakes(document): Document stock Bash registration coverage
* no-mistakes(ci): Pinned the stock macOS Bash CI lane to tasks-axi@0.2.5, eliminating dependency drift. Verified workflow YAML parsing, git diff checks, and the focused regression under /bin/bash 3.2.57 with tasks-axi 0.2.5
* no-mistakes(ci): Fixed the flaky portable CI test: it treated exited zombie processes as live because `kill -0` succeeds for zombies. The watcher and descendant assertions now check process state and regard zombies as exited. Verified `tests/fm-pr-check-security.test.sh`, ShellCheck, `git diff --check`, and the focused Bash public-followup regression
* fix(bin): isolate new Herdr server environments (#2792)
* fix(herdr): isolate server launch environment
* no-mistakes(review): Clear inherited supervision model from Herdr launches
* no-mistakes(document): Document Herdr server launch environment isolation
* fix: surface inbound Relay media to responding agents (#3442)
* fix: surface inbound Relay attachments to the responding agent
A Discord support thread's screenshots were never seen by the agent
handling the mention. The relay delivered them and the poll stashed
them: the reporter's images arrived on the `thread_starter` entry of
`in_reply_to_chain` while the mention's own media list was empty. The
gap was in the responder's playbook, which enumerated a fixed field
list (`request_id`, `text`, `in_reply_to`, `in_reply_to_chain`) and so
made every other field, attachments included, invisible.
Fix it where the gap is, in prose:
- Read the complete payload object rather than a fixed field list, so
media and later relay fields are never skipped again.
- Fetch and view attached media with the agent's own tools, on the
mention and on every chain entry, and call out the common shape where
only the thread starter carries the screenshots.
- Restrict those fetches to known-good platform media hosts over https
(Discord: cdn.discordapp.com, media.discordapp.net,
images-ext-1.discordapp.net, images-ext-2.discordapp.net; X:
pbs.twimg.com, video.twimg.com), report a blocked host instead of
working around it, and treat everything fetched as untrusted public
input on the same terms as the surrounding thread text.
The poll stays out of it and downloads nothing, so no third-party bytes
are pulled on the polling path.
The new test pins the contract the playbook depends on: a mention in the
incident's shape, with an empty top-level media list and screenshots on
the thread starter, must reach the inbox with the payload intact and its
media URLs unfetched.
* no-mistakes(review): Preserve media authority and enforce poll-only fetching
* no-mistakes(document): Clarify Relay attachment safety prose
* fix(bin): defer inactive reconciliation during startup (#3480)
* Defer inactive startup reconciliation
* no-mistakes(review): Queue deferred inactive reconciliation diagnostics durably
* no-mistakes(review): Require worker phases to cover startup requests
* no-mistakes(review): Make diagnostic wakes safely acknowledgeable
* no-mistakes(document): Document deferred startup phase coverage
* fix(bin): bound wake drain presentation lock waits (#3475)
* fix: bound status presentation lock waits
* no-mistakes(review): Distinguish malformed presentation locks from live contention
* no-mistakes(review): Bound no-ack drain queue lock acquisition
* no-mistakes(document): Document bounded presentation-lock drain behavior
* no-mistakes(lint): Annotate bounded lock output global
* no-mistakes(ci): Added deterministic regression coverage for successful bounded-lock acquisition after live contention, verifying helper-to-caller PID ownership handoff and caller release. Verified with bash syntax checks, git diff checks, and the full fm-wake-queue test suite
* fix(bin): retire public follow-ups in remote homes (#3479)
* fix(relay): close a public loop whose work lives in a remote secondmate home
A public-followup loop bound to a REMOTE secondmate could never be closed.
`clear_public_followup_link` (bin/fm-public-followup.sh:701) required an
absolute recorded `work_home_path` for a `secondmate:*` work home, but a remote
route has no local path on this machine, so registration records that field
empty (bin/fm-public-followup.sh:291). Every close ran that clear first, so
`retire` died with "could not clear the legacy X link ... retained for
reconciliation" forever, and `deliver` posted the public reply and then stranded
the loop at `posted`. `--force` never covered that step.
The clear now goes to the remote home over that route's SSH transport, running
`fm-x-followup.sh --clear <work-id>` through `bin/fm-on.sh`. The route is decided
from `data/secondmates.md` before any local path is consulted, so a same-named
local directory can never stand in for a remote home, and registrations already
on disk retire without needing a new field. `fm-on.sh` passes ssh's status
through, so 255 stays the established "delivered but completion unknown" result
this codebase already reconciles: the close is refused, the registration and the
remote link are left exactly as they were, and the message names the unknown
completion instead of claiming a definite failure.
Local secondmate and `main` work homes are untouched, and `--force` still
governs only the unresolved-obligation refusal.
Three regression cases drive a remote route end to end, faking only the ssh
binary at the FM_SSH_BIN seam and then running the real remote entrypoint
against a local checkout, so the clear that must reach the remote home actually
happens there.
* no-mistakes(review): Guard remote link clears by request identity
* no-mistakes(review): Fail guarded clears on unreadable remote state
* no-mistakes(review): Reject guarded clears on non-writable remote state
* no-mistakes(review): Allow no-link retirement in non-writable remote state
* no-mistakes(document): Correct public-followup verification guarantee count
* no-mistakes(ci): Fixed the guarded link-clear race by ensuring absence is decided under the metadata lock whenever publication is possible. Added a behavioral concurrency regression test. Verified with fm-x-mode and fm-public-followup suites, Bash syntax checks, diff checks, and bin/fm-lint.sh
* no-mistakes(ci): Fixed the guarded link-clear race by refusing an unlocked absence decision when a publisher already owns the metadata lock in a non-writable directory. Added a behavioral concurrency regression test. Verified with fm-x-mode, fm-public-followup, syntax/diff checks, and fm-lint
* no-mistakes(ci): Fixed the guarded-clear race by refusing all guarded clears when the metadata parent is non-writable, including apparent link absence. Added a behavioral regression with a publisher waiting to create the lock, updated remote-retirement expectations and verification docs. Passed fm-x-mode, fm-public-followup, fm-lint, documentation audience, Bash syntax, and diff checks
* fix(relay): bound the guarded remote link clear so it refuses instead of hanging
The guarded clear checks that the remote state directory is writable before
taking the metadata lock, but that check cannot close the window: the parent can
turn non-writable between the check and lock creation, and a lock held by a live
holder is indistinguishable from that at the acquire. `fm_lock_acquire_wait` is
an unbounded `while ! try; do sleep 0.1; done`, so either case retried forever
and `deliver` or `retire` wedged with nothing reported, instead of returning the
retained-for-reconciliation refusal the guard exists to produce. This path runs
unattended over the secondmate transport, where a wedge is worse than either
outcome the guard defines.
The guarded clear now acquires through `fm_lock_acquire_wait_bounded`
(FMX_LINK_CLEAR_LOCK_TIMEOUT, default 10 seconds) and refuses on timeout through
the existing failure path. Unguarded local callers keep the ordinary unbounded
wait, so local behavior is unchanged.
The bounded primitive's header no longer claims presentation-only scope, since
this is a second authorized caller; nothing else in the shared lock
infrastructure changed.
The regression holds the metadata lock with a genuinely live process while
leaving the state directory writable, so the refusal can only come from the
bound and never from the writability precondition. Against the unbounded wait it
does not terminate at all; with the bound it refuses, retains the registration,
writes no receipt, and leaves the remote link untouched.
* no-mistakes(review): Harden lock-timeout regression with independent deadline
* no-mistakes(review): Restore no-op guarded clears on read-only state
* no-mistakes(document): Clarify remote public-followup cleanup contract
* fix(bin): support process events under symlinked homes (#3484)
* fix(bin): resolve process-event state roots before validating them
The process-event module validated the caller's spelling of a home's state
root instead of the directory it operates on: it required the supplied path
to equal its own lexical normalization, which rejects any path reached
through a symlinked ancestor. On macOS both /tmp and $TMPDIR are symlinks,
so an operator home under either could never claim a source. Reconcile still
reported the runner started, while the detached runner died writing "cannot
claim source" to the discarded stderr, and the source silently never fired.
Resolve the state root to its physical directory once, then apply the
existing private-directory validation to that resolved directory and derive
every path, recorded claim identity, and later confinement check from it.
This keeps the confinement contract for the directory actually operated on
rather than only for callers that already spelled it physically, and removes
the window where an ancestor symlink could be repointed between check and
use. Homes already spelled physically behave identically.
This was the single cause of both deterministic macOS failures in
tests/fm-procevent.test.sh ("reconcile never claimed the registered source")
and tests/fm-procevent-when.test.sh ("the winning concurrent arm did not
produce an outcome"). The new case pins the behavior with an explicit
symlinked-ancestor home, so it fails without the fix on any platform rather
than only where the temp root happens to be a symlink.
* fix(bin): pin the external capture staging boundary to its physical path
The extension capture path pinned its registry staging boundary by comparing
`pwd -P` against the caller-spelled registry directory, so a home reached
through a symlinked ancestor still refused to start an extension-backed
source after the state root itself resolved correctly. That left such a home
half working: built-in sources ran while external ones failed.
The staging preparer now prints the physical registry directory it validated,
matching the inbox and reservation preparers beside it, and the start path
pins on that returned path. The new end-to-end case drives the shipped
file-signal package from a symlinked home spelling.
* no-mistakes(review): Propagate canonical process-event state roots
* no-mistakes(review): Propagate canonical state to process-event adapters
* no-mistakes(document): Document physical process-event state roots
* fix(pi): deliver captain outcomes as deterministic transcript entries (#3312)
* fix(pi): persist captain outcomes visibly
* no-mistakes(review): Recover captain outcomes after cold-start lock acquisition
* no-mistakes(document): Document cold-start captain-outcome recovery
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes(review): Prove immediate Pi captain-outcome transcript delivery
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* fix(pi): process captain outcomes through a sequence-keyed turn
PR #3312 made every captain-facing supervision outcome a durable, exact-once
visible transcript entry with the read cursor advancing only after that entry
exists. That is the display half of the delivery contract. Left alone it turns
a probabilistic silent loss into a deterministic one: the captain sees an
anchor line, and firstmate never acts, because nothing opens a turn and
nothing records whether main ever processed the outcome.
The 2026-08-31 timeline showed the two shapes this must survive on the
previous hidden-turn path: seven delivered decision outcomes each answered by
an empty assistant message (cursor advanced, no retry, unanswered for close
to three hours), and two answered by an unrelated prior reply. Both happened
because delivery advanced the cursor at enqueue and accepted whatever the
next assistant message was.
Add the processing half on top of the persistence half:
- bin/fm-branch-outcome.sh keeps a processed marker separate from the read
cursor (`unprocessed`, `mark-processed --through`, `processed-init`). It
only advances through an explicit sequence-bound acknowledgement, never
past the read cursor and never backwards; an absent marker reads as zero
and `processed-init` migrates delivered history once so an upgraded home
is not re-presented its past.
- After the visible entry for a captain outcome exists, the extension hands
every still-unprocessed captain row to main as one hidden, typed
`fm-branch-process` request listing each `[seq N] task: summary`, opening
exactly one main turn. Main closes it only by calling the new
`fm_branch_processed` tool with the highest sequence listed. An unrelated,
empty, or paraphrased answer leaves the sequence open, and the same request
is presented again at the end of the next main run and at session start.
The first two presentations of a sequence set open a turn of their own;
after that the request rides the captain's next prompt so an ignored
request cannot loop, and a session replacement resets that budget.
Routine outcomes stay turn-free.
- The regressions cover exactly those incident shapes against the real store
scripts: an empty answer and an unrelated prior answer neither advance the
marker nor stop re-presentation, the acknowledgement is refused beyond the
read cursor and outside lock ownership, a partial acknowledgement keeps the
newer sequence open, and #3312's own assertions now forbid an unkeyed turn
rather than any turn. The store suite pins the marker's bounds and the
migration; the real-SDK guard for appendEntry persistence and model
exclusion is unchanged.
Docs move the protocol from "no model turn" to "one sequence-keyed processing
turn closed only by its acknowledgement", and the verification record carries
the dated run against Pi 0.84.4.
* no-mistakes(review): Harden outcome listing and sequence-bound acknowledgements
* no-mistakes(review): Harden outcome state validation and request pacing
* no-mistakes(review): Reject unsafe sidecars and unterminated outcome stores
* no-mistakes(review): Validate canonical mark-read cursor state
* no-mistakes(review): Guard cursor advancement against corrupt processed state
* no-mistakes(review): Bind acknowledgements to active processing requests
* no-mistakes(review): Reset pacing when processing sequence membership changes
* no-mistakes(review): Enforce silent outcome invariants at storage boundary
* no-mistakes(document): Document hardened captain outcome processing contracts
---------
Co-authored-by: kunchenguid <kun@kunchenguid.com>
* feat: add bounded concurrent Bearings ledger collection (#3481)
* feat: bound Bearings remote ledger collection
* no-mistakes(review): Clarify default remote-ledger collection behavior
* no-mistakes(review): Detach reconcile delivery from watcher loop
* no-mistakes(review): Enforce bounded snapshot and request captures
* no-mistakes(review): Bound legacy summary capture before parsing
* no-mistakes(review): Bound primary remote ledger captures
* no-mistakes(document): Correct snapshot and reconcile documentation
* no-mistakes(lint): Fix ShellCheck quoting in bounded collector
* no-mistakes(ci): Fixed all three CI failures: updated the macOS Bearings assertion to 44 tests, made the home-summary test deterministic and aligned with default ledger consumption, and increased the asynchronous reconcile retirement wait for loaded CI. Verified both focused suites, all 44 Bearings tests, ShellCheck, actionlint, Bash parsing, and git diff checks
* test: await reconcile request retirement
* no-mistakes(review): Avoid empty reconcile queue process churn
* no-mistakes(review): Read ledger summaries from immutable snapshots
* no-mistakes(review): Reject multi-document home ledger streams
* no-mistakes(review): Coalesce durable reconcile requests per target
* no-mistakes(review): Unify reconcile keys and reject snapshot streams
* no-mistakes(review): Key reconcile requests by stable target ID
* no-mistakes(document): Document per-target reconcile request coalescing
* no-mistakes(lint): Remove unused snapshot summary file variable
* no-mistakes(ci): Adjusted the concurrent collector regression’s end-to-end timing ceiling to account for stock macOS process/jq overhead outside the three-second remote collection budget, while remaining below the 15-second serial-read floor. Verified with stock /bin/bash 3.2: all 44 Bearings tests pass; bash syntax and git diff checks pass
* no-mistakes(ci): Fixed legacy summary validation to require exactly one top-level JSON document and added behavioral regression coverage. Stabilized CI by conditionally waiting longer for durable reconcile delivery and synchronously stopping the fm-on worker tree before fixture cleanup. Removed a redundant flaky healthy-path timing assertion; the wedged-reader test still proves concurrent bounded collection. Verified fm-bearings-snapshot, fm-secondmate-reconcile, and fm-on tests, plus project ShellCheck, bash syntax, and git diff checks
* ci: rebalance portable serial test shards (#3489)
* fix(ci): rebalance the portable serial shards on measured durations
The "Behavior portable serial 3" shard ran 17-20 minutes against its
20-minute job cap and intermittently timed out seconds after a passing
test, on branches and on main alike.
Shards are packed longest-processing-time from per-script duration hints,
and those hints were last measured on 2026-08-21 at 116 scripts. The lane
has since grown to 139 scripts and from ~42 to ~63 minutes: 17 scripts had
no hint at all and fell back to the 20 s default, and several existing
hints were low by 2-5x (fm-watch-triage 142 s hinted vs 263 s measured,
fm-public-followup 36 s vs 197 s). The partition therefore looked
perfectly balanced in hint space, 734.6 s per shard, while really running
11.5, 13.6, 18.8 and 16.5 minutes. Script-count balance, which is what the
tests asserted, stayed normal throughout and hid it.
Refresh the hints from the timing artifacts of three green runs, taking
the slowest measurement of each script so the balance holds on a slow
runner, and split the lane across five shards instead of four. Replayed
against those runs' real per-script durations the worst shard is now
12.54 minutes, 63% of the unchanged 20-minute cap, and the serial lane's
wall clock drops from ~20 to ~12.5 minutes.
Bound the drift that caused this rather than relying on the hints being
refreshed by hand: the coverage guard now reports the unmeasured share as
serial_unhinted= and refuses past PORTABLE_SERIAL_MAX_UNHINTED_PERCENT,
which leaves room for newly added tests while making a stale table fail
the guard instead of silently pushing one shard into its cap.
No test changes what it asserts and no test stops running; only the
partition across shards changes.
* no-mistakes(document): Clarify conservative shard timing aggregate
* fix(pi): fall back on incomplete supervision branch prompts (#3491)
* fix(pi): fall back after settled branch errors
* no-mistakes(review): Detect provider errors across prompt compaction
* no-mistakes(review): Preserve in-flight branch state across selection changes
* fix(pi): re-probe supervision branch after cooldown (#3497)
* fix(pi): recover supervision branch after cooldown
* no-mistakes(review): Defer branch recovery until prompt settlement
* no-mistakes(document): Clarify supervision cooldown recovery contract
* fix(bin): remove legacy remote snapshot reads (#3501)
* refactor: remove legacy remote summary reads
* no-mistakes(document): Document ledger-only snapshot reads
* no-mistakes(ci): Fixed the snapshot test fixture so ledger refreshes use the same fake executable PATH as the snapshot consumer. This preserves observable endpoint freshness after removing legacy summary computation. Verified stock Bash parsing and all 44 Bearings tests pass under /bin/bash; git diff checks pass
* no-mistakes(ci): Fixed the CI-only snapshot fixture failure by ensuring the bounded-ledger refresh uses its fake tmux backend. This removes host tmux availability as a source of nondeterminism. Verified all 44 Bearings tests pass, Bash syntax passes, and git diff checks are clean
* no-mistakes(ci): Fixed CI nondeterminism in the Bearings fixture: all local ledger refreshes now use the fixture’s fake tmux backend when available, instead of depending on host tmux state. Verified stock /bin/bash syntax, git diff checks, and all 44 Bearings tests with a deliberately failing host tmux
* fix(pi): preserve watcher continuity across session replacement (#3498)
* fix(pi): rearm watcher after session replacement
* no-mistakes(review): Queue actionable closes across Pi session replacement
* no-mistakes(review): Stop replacement arm when handoff persistence fails
* no-mistakes(review): Preserve actionable wakes through branch and late child races
* no-mistakes(review): Surface late handoff failures without crashing Pi
* no-mistakes(review): Coordinate replacement delivery settlement and unique handoff tokens
* no-mistakes(review): Retry stale deliveries and release settled claims
* no-mistakes(review): Distinguish branch settlement and retry handoff cleanup
* no-mistakes(review): Deduplicate persistent handoff cleanup alerts
* no-mistakes(review): Acknowledge watcher follow-ups only when consumed
* no-mistakes(review): Persist idle follow-ups until agent consumption
* no-mistakes(review): Preserve pending outcomes when handoff persistence fails
* no-mistakes(review): Arm replacement before awaiting prior delivery settlement
* no-mistakes(review): Adopt pending handoffs after lock reclamation
* no-mistakes(review): Prevent stale generations from adopting replacement handoffs
* no-mistakes(review): Scope replacement handoffs by watcher state
* no-mistakes(document): Clarify replacement handoff documentation
* no-mistakes(ci): Fixed the failing branch-extension tests to model the new settlement-promise contract. Failure cases now assert that delivery ownership returns to the watcher instead of expecting direct extension fallback. Verified the updated branch suite, Pi watcher suite, shell syntax, and diff checks
* no-mistakes(review): Update branch settlement tests and preserve chunked outcomes
* no-mistakes(document): Document watcher-owned replacement handoffs
* no-mistakes(document): Verify replacement handoff documentation
* test(pi): cover watcher-owned branch fallback
* no-mistakes(document): Refresh watcher-owned fallback documentation
* fix(bin): resurface task statuses missed by wake handling (#3495)
* fix(bin): resurface terminal statuses lost after branch handling
* test(watch): canonicalize process-event fixture homes
* no-mistakes(review): Index branch outcomes by causal status position
* no-mistakes(review): Recover outcome indexes and deduplicate resurfaced statuses
* no-mistakes(review): Handle legacy ambiguity and oversized status diagnostics
* no-mistakes(review): Keep unclassifiable oversized statuses silent
* no-mistakes(document): Document lost-wake outcome backstop
* no-mistakes(document): Update outcome backstop documentation
* no-mistakes(ci): Fixed CI regressions in wake-drain: parseable reserved-key decisions can no longer bypass the durable decision-fold guard, and status output is prepared and receipt-committed before presentation to prevent repeated one-shot outcomes after later failures. Added a behavioral regression for receipt commit failure and retry. Targeted backstop, correlation-token, decision-cursor, open-decision, unread-status, syntax, and diff checks pass locally. Shard-4 failures appeared unrelated/flaky; the network-parallel test passed locally
* no-mistakes(ci): Fixed the Greptile P1 data-loss issue by committing presentation receipts only after prepared output reaches stdout. Added behavioral coverage proving output failure leaves the backstop retryable and receipt failure may duplicate but never lose a presentation. Relevant wake-drain suites and syntax/diff checks pass. The shard-4 Pi extension failure is unrelated to this PR and did not warrant changes
* no-mistakes(ci): Stabilized tests/fm-bootstrap-network-parallel.test.sh by replacing scheduler-sensitive equal-sleep timing with bounded synchronization between mocked fetch and remote probes. This preserves detection of real serialization while avoiding false failures under CI load. Verified with five consecutive test runs, bash syntax validation, ShellCheck, and git diff checks. The separate Pi stock-rendering failure reproduces locally but is unrelated environment/version drift
* no-mistakes(ci): Fixed Behavior portable serial 4 by adding fm-classify-lib.sh and fm-timeout-lib.sh to the broken-root Pi test fixture; fm-branch-outcome.sh now depends on them. Verified the full Pi branch-extension suite with real-Pi checks skipped, the wake-drain outcome-backstop suite, Bash syntax, and git diff checks. Greptile findings are already addressed at HEAD; the no-mistakes attestation failure is external head-SHA state
* fix(bin): collect follow-up results from remote work homes (#3503)
* fix(bin): deliver typed terminal results from remote work homes
A public commitment whose work is bound to a REMOTE secondmate home could
never receive its typed terminal result. `fm-public-followup.sh brief`
printed an emit command carrying this home's own absolute path and this
checkout's own script path, neither of which exists on the machine the
worker runs on, so the worker had nothing it could write to that the
owning home would ever read - and `consume` kept finding nothing while
the promise stayed open.
The brief is now route-aware: for a remote work home it prints that
route's own code root and home with `--stage-in`, so the typed event is
staged in the home where the work actually runs, and the closing
paragraph names the owning home as the one on the other machine instead
of pointing at the path above it. The owning home collects those staged
results over the same SSH route it reaches that secondmate on, because
the transport only runs outbound: `consume` pulls them into its own
inbox and reconciles them exactly as it reconciles a local report.
Collection is non-destructive until the result is durably held, so a
dropped connection cannot lose a terminal result, and a route that could
not be reached is named in `consume`'s output with the promise left open
rather than reported as an empty inbox.
A local work home is untouched: the brief still prints `--home` with this
home and this checkout's script, and the event still lands directly in
this home's typed terminal-result inbox.
This is the emit-side counterpart of the retire/clear fix in #3479 and
reuses the remote-route resolution that landed with it. Reconciling a
loop bound to a remote route now reaches that route, so the existing
remote cases drive `consume` through the same faked transport their
other steps already use.
* no-mistakes(review): Fail loudly on unresolved routes and invalid staging homes
* no-mistakes(review): Fail collection when remote outbox is unreadable
* no-mistakes(review): Surface reassigned remote routes during empty collection
* no-mistakes(review): Fail remote collection on invalid registrations
* no-mistakes(review): Reject unsafe registration entries during remote collection
* no-mistakes(review): Restore healthy empty remote collection behavior
* no-mistakes(review): Skip remote collection for delivered registrations
* no-mistakes(review): Skip delivered registrations before route validation
* no-mistakes(document): Document remote follow-up collection semantics
* fix(bin): exclude secondmates from home-summary validity (#3504)
* fix(bin): exclude secondmates from home-summary child inventory
kind=secondmate meta records never have backlog rows, so counting them in unowned_children or terminal_in_flight made a clean main home look invalid once earlier ledger checks passed.
* no-mistakes(review): Cover terminal secondmate in-flight exclusion
* no-mistakes(ci): Updated the stock macOS Bash CI snapshot expectation from 15 to 16 tests. Verified all 16 snapshot/fleet-view tests pass under Bash 3.2.57 and `git diff --check` succeeds
* fix(bin): self-heal outcome indexes on first drain (#3509)
* fix(bin): self-heal status-outcome indexes on every drain
Missing ready markers were skipping the lost-wake backstop on non-Pi homes because only the Pi branch ran processed-init. Drain now rebuilds those indexes under the outcome lock and fails closed only on a real store fault.
* no-mistakes(review): Guard held-lock initialization and fail marker writes
* no-mistakes(document): Document cross-harness outcome-index self-healing
* fix(bearings): keep active children underway during captain holds (#3505)
* fix(bearings): keep active children underway beside a captain hold
Project each readable home's active children into Underway independently of the home-level captain-decision classification so a hold no longer hides live work.
* no-mistakes(review): Preserve Underway repos and disclose child truncation
* no-mistakes(review): Fall back to task project for Underway repos
* no-mistakes(ci): Updated the stock macOS Bash CI assertion from 44 to 45 Bearings tests, matching the newly added behavioral regression. Verified all 45 tests pass under /bin/bash, Bash syntax checks pass, and git diff validation is clean
* fix(pi): settle watcher delivery on Pi accepting the follow-up (#3513)
* fix(pi): settle watcher delivery on Pi accepting the follow-up
A follow-up queued while main is streaming joins the running run without
ever raising before_agent_start, so waiting on that event before clearing
the successor pipeline (#3498) stalled every later actionable close: no
successor started, no wake was delivered or offered to the branch, and the
turn-end guard woke main to re-arm by hand after every close.
The pipeline now settles once Pi accepts the follow-up. Consumption is
observed at before_agent_start for an idle main and at the user
message_start for a streaming main, and decides only what a replacement
session (/new, /resume, /fork, reload) replays. An exhausted restoration
delivers its typed failure without launching an arm past the retry bound,
which the stall had hidden. The replacement-coordinator map is typed so the
strict no-emit typecheck passes again.
Tests: the doubles no longer raise before_agent_start for a streaming send,
a portable regression drives two actionable closes while main streams and
proves the successor chain plus consumption-scoped replay, and a
credential-free real-SDK probe pins Pi's event contract for both the
streaming and the idle follow-up.
Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a
* fix(pi): retry a verified successor that fails during wake delivery
A verified successor can exit while the wake it was started for is still
being delivered, most plausibly during a branch turn that holds the
settlement for minutes. Its failure close arrived while the pipeline's
single-flight guard was set, so the close handler skipped the retry, and
the pipeline's end no longer launched an arm, which left the live
generation with no watcher and no retry timer.
The close handler now records that failure when the child had reported
readiness and was not retired by the restoration itself, and the pipeline
runs the ordinary bounded, lock-checked retry for it once the delivery
settles. A restoration started for a later pending supersedes it, and an
exhausted restoration still hands repair to main without a further arm.
The regression holds a branch settlement open while the verified
successor exits with a failure and proves one retry watcher starts after
the settlement releases, none while it is held.
Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a
* fix(bin): bound repeat stale wakes for parked workers (#3532)
* fix(bin): bound repeat stale wakes for a parked but live worker
A worker parked on a declared wait - `paused:` for an external or pipeline
wait, or a verified `captain-held` transfer - kept waking firstmate far inside
FM_PAUSE_RESURFACE_SECS. Observed as five consecutive alarms on one
captain-held worker and dozens across a day on a pipeline wait, and reported
upstream as four wakes in 75 minutes against a 3600s window.
pause_state_class deliberately answers `none` for a still-live agent even under
a declared wait, so a worker genuinely waiting on a decision is never silenced.
That classification is correct and is left alone; it routes every parked but
live worker through surface_nonterminal_stale on first sight of each distinct
stale hash, and an idle parked pane still churns its hash on a clock or a token
counter without changing what is being waited on.
Two places let that churn re-alarm:
- surface_nonterminal_stale queued the wake BEFORE consulting whether a wait was
declared, then wrote `.paused-resurfaced-<key>` - the very throttle that should
have suppressed it. The throttle was never read on this path and was advanced
by the wake it should have prevented.
- The hash-change path cleared that throttle through clear_pause_tracking
whenever the classification came back `none`, so each tick also bought the same
declared wait a fresh window. Fixing only the first site changes nothing.
Read the throttle before anything is queued and advance it only on a wake that
really fires, and on the hash-change path reset only the per-hash bookkeeping
while the declaration still stands, via a clear_stale_hash_tracking split so
neither half of clear_pause_tracking is duplicated. The throttle is keyed to the
declaration, not to the pane.
First sight still wakes, so an inconclusive state is still inspected, and the
window's end still re-surfaces once, so a forgotten wait cannot rot invisibly -
noise traded for a bounded cadence, never for silence. The wake identity stays
the plain `stale: <win>` the away-mode handoff depends on.
Tests cover both observed forms and were confirmed to fail against three
deliberate breaks: each site reverted on its own, and a re-surface that never
fires again.
* fix(document): Clarify declared-wait wake cadence documentation
* fix(ci): Captain, fixed the stale-throttle inheritance: cadence markers now bind to the current wait declaration, so replacement paused and captain-held waits each emit their first plain `stale:` wake. Added behavioral coverage for both forms. Bite proof failed as expected when identity matching was removed, then passed after restoration. Full watcher triage suite, `bin/fm-lint.sh`, syntax checks, and diff checks pass. Changes remain uncommitted for the outer executor
* fix(ci): Captain, fixed the confirmed Greptile finding. `resurface_absorbed` now applies a throttle only when its stored declaration scope matches the current wait, so replacement `paused:` and `captain-held` waits surface immediately without changing classification. Added executable coverage for both absorbed forms. Bite proof failed before the fix at the intended assertion; afterward the full watcher triage suite, `bin/fm-lint.sh`, shell syntax checks, and `git diff --check` passed
* fix(bin): accept the away-mode daemon as the turn-end supervision owner (#3567)
* fix(turnend): accept the away-mode daemon as the supervision owner
While state/.afk exists the away-mode daemon owns supervision and runs
bin/fm-watch.sh one-shot: the watcher exits on every wake and the daemon
starts its replacement. The turn-end guard tested for a live watcher
process holding the watch lock at that instant, so a turn boundary that
landed in the hand-off blocked with "TURN WOULD END BLIND" while
supervision was completely healthy, costing a full handling turn each
time.
Reproduced with the real daemon wrapping the real watcher and the real
guard sampling the same home: 6 of 40 samples blocked, every one of them
with the daemon alive and the beacon 2-3 seconds old, and a new watcher
pid on each cycle. After the fix the same reproduction blocks 0 of 40,
and killing the daemon and its watcher (away mode still on, beacon still
fresh) blocks again.
The guard now accepts a live, identity-matched daemon holding this home
as proof of supervision while away mode is active. The identity match is
the same discipline the watcher lock uses, so a recycled pid or a lock
left by a killed daemon proves nothing. The fresh-beacon half of the
predicate is unchanged: a daemon that stops restarting its watcher still
blocks once the beacon passes grace, a home with no supervisor blocks
exactly as before, and with away mode off the strict watcher predicate is
untouched.
The predicate reads only durable state, so it behaves identically for
every primary harness and runtime backend.
* no-mistakes(document): clarify away-mode daemon supervision proof and test coverage
* no-mistakes(document): generalize stale turn-end predicate summary in architecture.md
* fix(backlog): omit --file from row probes for non-markdown backends (#3582)
* fix(backlog): omit markdown file for beads probes
* no-mistakes(document): Narrow backlog addressing doc to mutations for backend-aware probes
* no-mistakes(ci): Fixed the Greptile P2 review comment (the only failing check) on tests/fm-backlog-atomicity.test.sh. The comment correctly noted that an exported TASKS_AXI_BACKEND environment variable would inherit into the spawned scripts and, because fm_tasks_axi_backend gives it top precedence, override each test case's .tasks.toml backend fixture — making the backend-specific argv assertions fail for environmental reasons. Fix: unset TASKS_AXI_BACKEND in the test harness right after sourcing tests/lib.sh, with a comment explaining why, so every case deterministically exercises its declared backend (4 lines added; no production code touched). Verified: reproduced the leak before the fix (TASKS_AXI_BACKEND=beads made the markdown dispatch case fail with 'beads show failed', exactly the reported failure mode); after the fix the full suite passes (0 failures, exit 0) both with and without TASKS_AXI_BACKEND=beads exported. The added lines are shellcheck-clean (the only shellcheck note, SC1091 on the lib.sh source line, pre-exists this change)
* fix(bin): classify progress updates on requested work as routine (#3589)
The supervision branch's verdict rule escalated every outcome that
answered a captain request, so "the work started" and "still working"
notes reached the captain with nothing to look at. The rule now keeps a
finished result of requested work captain-facing, even when healthy, and
treats start or still-working updates that bring no new artifact,
finding, or decision as routine. The captain list for review-ready PRs,
ask-user findings, exhausted blockers, credentials, and destructive or
security-sensitive cases is unchanged, as are the unsolicited-routine,
silent-fleet-review, and doubt-chooses-captain rules.
The fm_branch_report tool description and the two docs that restated the
old unconditional rule now point at the prompt's "Verdict: routine or
captain" section as the one owner instead of carrying a second copy.
* fix(bin): preserve captain calls during teardown (#3595)
* fix(bin): never close a captain call during cleanup
A scout that held its own work item for the captain, which is what
captain-hold-lifecycle prefers ("hold the work item the question gates"),
was closed by bin/fm-teardown.sh's automatic backlog transition. The
completion gate passed, cleanup ran, and the captain's question moved to
Done with no recorded answer: the one thing the policy says must never
happen. `tasks-axi done` closes a held row silently, and nothing in
teardown asked whether the row was the captain's own call.
bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when
the task is still an open captain call, 1 when it is not, 2 when that
cannot be established. It reads the row through the transition library's
backend-aware probe, so it addresses the same backlog teardown does; the
script's other commands now address the configured data directory the
same way instead of FM_HOME, which also fixes captain holds in a home
with a relocated data directory.
Teardown asks `open` before any destructive step and refuses on 2. On 0
only the close changes: after cleanup and still under the task's own
lock, the row gets one "Deliverable of the finished work" line at the end
of its body and returns to Queued through `tasks-axi reopen`, keeping its
hold, so it lands in Captain's Call instead of reading as work under way.
--force does not lift this: it authorizes discarding unlanded work, never
the captain's question. The deliverable goes into the body because
`tasks-axi update --report` rewrites the title of a row that is not Done.
The crash window reuses the pending-close record teardown already stages:
a `mode=retain` line makes the existing replay record the deliverable and
reopen instead of closing, with the same validator, stale-generation
check, cleanup-incomplete marking, and non-blocking bootstrap lock as an
ordinary close. A retained row the captain answered first simply retires
the record. No parallel record type, recovery command, or second bootstrap
loop is introduced.
Regressions run the real executables: the captain-held scout survives
cleanup queued, held, with its deliverable and on the board, only
`answer` closes it, --force keeps it open, and an ordinary scout still
closes with its report; an interrupted cleanup leaves the row untouched
and the next session start retains it; a relocated backlog keeps the
retention in its one configured file; and a ship row whose hold cannot be
read refuses cleanup before anything destructive.
Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np
* no-mistakes(review): Serialize captain holds and fix backend-aware listing
* no-mistakes(document): Update captain-call retention documentation
* no-mistakes(document): Fix relocated captain-hold backlog diagnostics
* fix(bin): deliver secondmate outcomes to the parent channel (#3592)
* fix(bin): deliver every secondmate outcome on the parent channel from the recording scripts
A secondmate's captain-facing outcomes could miss: the mate model addressed
the captain in its own unread chat instead of appending to the parent
channel, and a PR-ready report, a finding, a decision, a blocker, and a
failure all depended on that one remembered append. Make delivery
structural, so the parent channel never depends on the model:
- bin/fm-parent-channel-lib.sh is the one owner of channel resolution and
exact-line append-once; the merge outcome path and the inactive-outcome
scan now publish through it instead of two private copies.
- bin/fm-inactive-reconcile.sh gains a ledger-first path that runs on every
watcher poll in a secondmate home: a direct child's whole terminal done or
failed line is delivered at once with its note, recorded PR, mode, merge
posture, and scout report pointer, keyed and receipted so it is delivered
once, and the inactive path yields to it. `report <task-id>` runs the same
delivery for a caller holding the child's meta lock.
- bin/fm-pr-check.sh publishes the PR-ready line with the canonical URL at
registration.
- bin/fm-captain-hold.sh publishes a hold and its answer, keyed by task id
and resolution-record count, with no new persisted state.
- bin/fm-teardown.sh delivers the child's final line before removing its
record and refuses, retaining every record, while the channel cannot be
written.
- The charter opens with the parent-channel rule and confines the mate's own
appends to judgement; AGENTS.md carries the carve-out at the persona
address rule and the escalation list.
docs/secondmate-parent-channel.md records the design and its coverage, and
docs/verification/secondmate-parent-channel.md records the live run with real
tmux panes and both real watchers delivering every line with no model.
Supersedes #3569.
* no-mistakes(review): Fix parent outcome retries and reconciliation locking
* no-mistakes(review): Prevent busy children from starving ledger delivery
* no-mistakes(review): Correct ledger metadata and hold occurrence handling
* no-mistakes(review): Disambiguate ledger outcomes and normalize hold reasons
* no-mistakes(review): Close ledger races and preserve teardown records
* no-mistakes(document): Correct parent-channel receipt and scanner documentation
* no-mistakes(lint): Quote done arguments for ShellCheck compliance
* no-mistakes(ci): Fixed both CI failures. Updated GOTMP teardown fixtures for the new final-outcome reporter and isolated them from host tmux state. Updated the PR security assertion to distinguish the accepted PR-ready line from duplicate merge outcomes. Verified with both failing test suites, bash syntax checks, and git diff checks
* no-mistakes(ci): Fixed Greptile’s duplicate-delivery race in bin/fm-inactive-reconcile.sh. Ledger events now claim matching already-delivered inactive receipts using the prior status fingerprint, preventing duplicate parent reports while preserving later same-state completions. Added behavioral regression coverage. Verified inactive-reconcile tests, project lint, documentation audience checks, syntax, and diff checks. Teardown tests passed relevant cases before the documented pre-existing herdr-preflight-missing-adapter failure
* fix(bin): sync remote second mates to primary commit (#3599)
* fix(bin): sync remote second-mate homes to the parent primary commit
Session start and remote launch pointed a remote second-mate home at whatever
Firstmate copy its own host kept, so a home that had already advanced past that
copy refused as a non-fast-forward and every other home stopped at the host's
older commit while the primary ran ahead.
The parent now resolves ITS primary default-branch commit with the existing
helper and hands that commit to the host on both paths. Because a remote home
is a standalone clone, the host imports that one commit before advancing -
already present, else from that host's Firstmate copy without moving it, else
from the home's own origin - and then runs the SAME ff_target guards a local
home gets, so dirty, diverged, feature-branch, and unresolvable targets skip
untouched and the ancestry rules keep one owner. An unimportable target now
names /updatefirstmate instead of failing opaquely, and a host still running an
older Firstmate copy is reported the same way rather than echoing a bare
refusal.
The host-local launch leg no longer re-runs its own secondmate sync, so the
spawn it drives cannot re-target that host's copy after the parent has already
converged the home.
/updatefirstmate is unchanged: it still refreshes the remote code root from that
host's origin and then syncs the home to that refreshed copy, which is what the
sync call with no target commit means.
* no-mistakes(document): Document primary-targeted remote secondmate synchronization
* fix(bin): separate captain intent from firstmate specs (#3597)
* fix(bin): split brief task into captain intent and firstmate spec
Keep no-mistakes --intent as the captain's ask plus later captain words, not the build spec or worker tradeoffs.
* fix(bin): stop task-subsection copies at the next heading
Promotion was swallowing the scout Setup contract into Firstmate spec, and pre-subsection briefs lost their # Task body.
* no-mistakes(review): Validate brief content and preserve nested specifications
* no-mistakes(review): Scope placeholder validation to scaffold-only subsection bodies
* no-mistakes(review): Ignore fenced subsection headings during brief validation
* no-mistakes(review): Preserve captain intent across scout promotion
* no-mistakes(review): Enforce safe intent boundaries for legacy promotions
* no-mistakes(review): Allow marked legacy intent and reject empty promotions
* no-mistakes(review): Scope task parsing and overlay legacy intent contracts
* no-mistakes(review): Overlay current intent contract for all no-mistakes spawns
* no-mistakes(review): Preserve later captain clarifications in intent overlays
* no-mistakes(document): Document brief intent enforcement and ownership
* no-mistakes(ci): Updated spawn-related test fixtures to use valid Captain intent and Firstmate spec subsections, corrected launch-path expectations to launch-brief.md, and resolved ShellCheck quoting findings. Verified with fm-lint.sh and 15 affected behavior tests, including real Herdr tests; all passed
* no-mistakes(ci): Updated stale spawn/promotion fixtures in the Muse, Orca, secondmate-harness, and public-followup suites to provide valid Captain's intent and Firstmate spec subsections. Verified full Orca and secondmate-harness suites, targeted public-followup promotion behavior, Bash syntax, diff checks, and fm-lint
* fix: start a fresh supervision branch for every main session (#3600)
* fix(pi): start a new supervision branch conversation per main session
The supervision branch reopened one recorded conversation forever, so
every main session start reloaded the current generated prompt and then
weeks of accumulated thread, where a superseded rule could still outweigh
today's.
The branch conversation is now scoped to one main session: the session
generation owns the recorded conversation, so a cold start, /new,
/resume, /fork, or a reload always builds a new one, while a rebuild
inside one session (a model or effort change) still continues that
session's own conversation.
The dialog mirror re-anchors with it. Its durable cursor records what the
previous branch conversation received, so a /resume or reload - which
keeps main's own session file - would otherwise leave the new branch
blind to dialog main itself still has. The reset is bounded by the
current main session, and the cursor keeps advancing incrementally within
it. The durable outcome store and its processed marker are untouched, so
unacknowledged captain-facing outcomes still re-present on the new main
session.
* no-mistakes(document): Document fresh Pi supervision conversations
* no-mistakes(ci): Fixed the flaky concurrent inbox failure. Lock acquisition now retries when a competing lock disappears between a failed claim and inspection. Added a behavioral regression covering that race. Verified the full inbox test four times, project lint, and git diff checks
* feat: restart second mates after instruction updates (#3614)
* feat(update): restart second mates whose instructions changed
/updatefirstmate pulled new bytes onto disk and then asked each advanced
second mate to re-read them. A running agent holds AGENTS.md and every
loaded skill frozen from launch and no verified harness offers a reload,
so that steer could not reach a loaded skill at all and left the mate
holding two contradictory copies of its own job description.
An eligible mate is now restarted instead, in the same home and endpoint,
through the existing transactional relaunch. The restart is gated on the
mate first writing down the open work it holds only in conversation - the
open-record half of /stow, never its memory sweeps - so an unregistered
captain call is flushed before the conversation is spent. Anything that
leaves the reload unprovable falls back to the old re-read message and is
reported as exactly that, never as a clean reload.
Remote mates take the same path: fm-remote-secondmate-control.sh gains a
relaunch verb whose host-local leg runs that same control plane, since the
mate is an ordinary local secondmate from its host's point of view. The
primary resolves the profile and passes it explicitly, because
config/secondmate-harness is not inherited and the file on that host
belongs to a different home.
fm-update.sh now splits its advanced live mates into a restart set and a
nudge residual, and both sets require a changed instruction surface, which
also closes the over-nudge against the session-start sweep. Restart is
stricter still: a bin/-only advance reloads itself on the next call, so it
never costs a conversation.
Colocated tests cover the gating, the persist-then-restart order, the
task-subset persist request, each unsafe fallback, the remote hop, and the
remote sync's new instruction-surface report.
* no-mistakes(review): Fix restart correlation, concurrent waits, and lifecycle reporting
* no-mistakes(review): Parallelize relaunches and classify replacement incarnations
* no-mistakes(review): Gate restart actions on live agent state
* no-mistakes(review): Handle failed restart workers without hanging
* no-mistakes(review): Nudge legacy remotes and preserve persist recovery
* no-mistakes(review): Document one-time secondmate restart rollout
* no-mistakes(review): Honor arrived replies and refresh remote profiles
* no-mistakes(review): Revert remote parent profile reconciliation
* no-mistakes(review): Reset remote profile defaults and honor published results
* no-mistakes(review): Preserve fallback nudges for unverifiable secondmates
* no-mistakes(document): Document second-mate restart update flow
* no-mistakes(lint): Fix ShellCheck warnings in restart scripts
* perf: accelerate local validation with bounded concurrency (#3644)
* perf(tests): route gate verification through the bounded concurrent runner
Local validation was the pipeline's dominant cost: across 67 recorded
no-mistakes agent sessions on this repo, 99.3% of command execution was
`bash tests/*.test.sh`, run strictly one script at a time, and 2% of those
calls were killed by an agent-guessed timeout and paid for twice.
Three changes, each measured:
- `.no-mistakes.yaml` pins `commands.test` to
`bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner
already owns changed-file selection, bounded concurrency, the refusal of
unproven scripts, and a generous automatic per-script bound, so the gate's
baseline is neither a serial chain nor a guessed timeout. It stays
intent-targeted - the Test step still runs its evidence agent on top - and
excludes the live-Herdr family the required Herdr lane owns.
- `bin/fm-test-run.sh` gives a plain list of script paths the same bounded
automatic scheduler and automatic bound that `--changed` gets. Naming several
subjects is how a verification round asks for exactly those scripts. The
curated selections are untouched: `--lane` still composes CI shards whose
serial lane must stay serial, `--family` is what the required Herdr lane runs,
and `--all` stays a deliberate complete regression.
- `pr-forge` is admitted to the concurrent-safe family registry on two
consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those,
and records `secondmate` and `session-bootstrap` as refused with the exact
script and reason each failed on, so the refusals are actionable rather than
silent.
Measured on this host, 0 failures on both sides:
verification round, 4 scripts 448s chained -> 231s through the runner (-48%)
pr-forge family 409.2s at 1 worker -> 237.9s at 4 (1.72x)
watcher-wake-lock family 1311.1s at 1 worker -> 539.3s at 4 (2.43x)
A fourth lever was implemented and then removed because the measurement
refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made
`fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s
unchanged, back to back. Those sleeps are not overhead added to the clock -
they are how a test waits for a subject moving on fm-watch.sh's own one-second
cadence - so sampling less often only delays detection. It also broke
`fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a
settled condition. CONTRIBUTING.md records that result so the experiment is not
repeated.
* no-mistakes(review): Separate concurrent runs by isolation proof family
* no-mistakes(review): Limit automatic timeouts to changed-file validation
* no-mistakes(document): Clarify validation concurrency documentation
* fix: copy PR URLs from durable records (#3648)
* fix: copy PR URLs from records or abstain, never assemble them
Supervision reported a plausible but dead PR link three times because its
prompt demanded a full https:// URL at a moment when only a PR number was
observable, so the model assembled an owner/repository from memory, and the PR
check then accepted that URL and wrote it into the task record, after which the
model kept defending its own tool-endorsed guess over the worker's real link.
Three changes close that chain without any live forge lookup, so private
forges are treated exactly like public ones:
- bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy
or abstain" section requires a URL to be copied verbatim from a durable
record (the done: PR <url> status line, pr= metadata, or the backlog note),
forbids assembling owner, repository, host, or number from memory, and has
the branch report only the identifier it actually holds when no record names
the URL yet, leaving the PR check unarmed until the worker's ready line
arrives. AGENTS.md section 7 and 9 carry the same copy-or-abstain rule for
main in place of the bare full-URL mandate.
- Worker briefs (bin/fm-brief.sh, ship and scout rules) require …
Valentino-Sole
added a commit
to Valentino-Sole/firstmate
that referenced
this pull request
Sep 4, 2026
* fix: start a fresh supervision branch for every main session (kunchenguid#3600) * fix(pi): start a new supervision branch conversation per main session The supervision branch reopened one recorded conversation forever, so every main session start reloaded the current generated prompt and then weeks of accumulated thread, where a superseded rule could still outweigh today's. The branch conversation is now scoped to one main session: the session generation owns the recorded conversation, so a cold start, /new, /resume, /fork, or a reload always builds a new one, while a rebuild inside one session (a model or effort change) still continues that session's own conversation. The dialog mirror re-anchors with it. Its durable cursor records what the previous branch conversation received, so a /resume or reload - which keeps main's own session file - would otherwise leave the new branch blind to dialog main itself still has. The reset is bounded by the current main session, and the cursor keeps advancing incrementally within it. The durable outcome store and its processed marker are untouched, so unacknowledged captain-facing outcomes still re-present on the new main session. * no-mistakes(document): Document fresh Pi supervision conversations * no-mistakes(ci): Fixed the flaky concurrent inbox failure. Lock acquisition now retries when a competing lock disappears between a failed claim and inspection. Added a behavioral regression covering that race. Verified the full inbox test four times, project lint, and git diff checks * feat: restart second mates after instruction updates (kunchenguid#3614) * feat(update): restart second mates whose instructions changed /updatefirstmate pulled new bytes onto disk and then asked each advanced second mate to re-read them. A running agent holds AGENTS.md and every loaded skill frozen from launch and no verified harness offers a reload, so that steer could not reach a loaded skill at all and left the mate holding two contradictory copies of its own job description. An eligible mate is now restarted instead, in the same home and endpoint, through the existing transactional relaunch. The restart is gated on the mate first writing down the open work it holds only in conversation - the open-record half of /stow, never its memory sweeps - so an unregistered captain call is flushed before the conversation is spent. Anything that leaves the reload unprovable falls back to the old re-read message and is reported as exactly that, never as a clean reload. Remote mates take the same path: fm-remote-secondmate-control.sh gains a relaunch verb whose host-local leg runs that same control plane, since the mate is an ordinary local secondmate from its host's point of view. The primary resolves the profile and passes it explicitly, because config/secondmate-harness is not inherited and the file on that host belongs to a different home. fm-update.sh now splits its advanced live mates into a restart set and a nudge residual, and both sets require a changed instruction surface, which also closes the over-nudge against the session-start sweep. Restart is stricter still: a bin/-only advance reloads itself on the next call, so it never costs a conversation. Colocated tests cover the gating, the persist-then-restart order, the task-subset persist request, each unsafe fallback, the remote hop, and the remote sync's new instruction-surface report. * no-mistakes(review): Fix restart correlation, concurrent waits, and lifecycle reporting * no-mistakes(review): Parallelize relaunches and classify replacement incarnations * no-mistakes(review): Gate restart actions on live agent state * no-mistakes(review): Handle failed restart workers without hanging * no-mistakes(review): Nudge legacy remotes and preserve persist recovery * no-mistakes(review): Document one-time secondmate restart rollout * no-mistakes(review): Honor arrived replies and refresh remote profiles * no-mistakes(review): Revert remote parent profile reconciliation * no-mistakes(review): Reset remote profile defaults and honor published results * no-mistakes(review): Preserve fallback nudges for unverifiable secondmates * no-mistakes(document): Document second-mate restart update flow * no-mistakes(lint): Fix ShellCheck warnings in restart scripts * perf: accelerate local validation with bounded concurrency (kunchenguid#3644) * perf(tests): route gate verification through the bounded concurrent runner Local validation was the pipeline's dominant cost: across 67 recorded no-mistakes agent sessions on this repo, 99.3% of command execution was `bash tests/*.test.sh`, run strictly one script at a time, and 2% of those calls were killed by an agent-guessed timeout and paid for twice. Three changes, each measured: - `.no-mistakes.yaml` pins `commands.test` to `bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner already owns changed-file selection, bounded concurrency, the refusal of unproven scripts, and a generous automatic per-script bound, so the gate's baseline is neither a serial chain nor a guessed timeout. It stays intent-targeted - the Test step still runs its evidence agent on top - and excludes the live-Herdr family the required Herdr lane owns. - `bin/fm-test-run.sh` gives a plain list of script paths the same bounded automatic scheduler and automatic bound that `--changed` gets. Naming several subjects is how a verification round asks for exactly those scripts. The curated selections are untouched: `--lane` still composes CI shards whose serial lane must stay serial, `--family` is what the required Herdr lane runs, and `--all` stays a deliberate complete regression. - `pr-forge` is admitted to the concurrent-safe family registry on two consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those, and records `secondmate` and `session-bootstrap` as refused with the exact script and reason each failed on, so the refusals are actionable rather than silent. Measured on this host, 0 failures on both sides: verification round, 4 scripts 448s chained -> 231s through the runner (-48%) pr-forge family 409.2s at 1 worker -> 237.9s at 4 (1.72x) watcher-wake-lock family 1311.1s at 1 worker -> 539.3s at 4 (2.43x) A fourth lever was implemented and then removed because the measurement refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made `fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s unchanged, back to back. Those sleeps are not overhead added to the clock - they are how a test waits for a subject moving on fm-watch.sh's own one-second cadence - so sampling less often only delays detection. It also broke `fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a settled condition. CONTRIBUTING.md records that result so the experiment is not repeated. * no-mistakes(review): Separate concurrent runs by isolation proof family * no-mistakes(review): Limit automatic timeouts to changed-file validation * no-mistakes(document): Clarify validation concurrency documentation * fix: copy PR URLs from durable records (kunchenguid#3648) * fix: copy PR URLs from records or abstain, never assemble them Supervision reported a plausible but dead PR link three times because its prompt demanded a full https:// URL at a moment when only a PR number was observable, so the model assembled an owner/repository from memory, and the PR check then accepted that URL and wrote it into the task record, after which the model kept defending its own tool-endorsed guess over the worker's real link. Three changes close that chain without any live forge lookup, so private forges are treated exactly like public ones: - bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy or abstain" section requires a URL to be copied verbatim from a durable record (the done: PR <url> status line, pr= metadata, or the backlog note), forbids assembling owner, repository, host, or number from memory, and has the branch report only the identifier it actually holds when no record names the URL yet, leaving the PR check unarmed until the worker's ready line arrives. AGENTS.md section 7 and 9 carry the same copy-or-abstain rule for main in place of the bare full-URL mandate. - Worker briefs (bin/fm-brief.sh, ship and scout rules) require the full https:// URL wherever a PR is mentioned - status line, terminal, or summary - never a bare "PR 108", so the link is in view as early as the number is. - bin/fm-pr-check.sh refuses, offline and before any side effect, a URL that the task's own done lines contradict, printing both spellings; a log naming no URL still records the argument as before. fm_pr_status_ready_urls in bin/fm-pr-lib.sh owns reading those lines. The refusal also reaches bin/fm-pr-merge.sh, so nothing merges under a contradicted URL. Tests cover the offline refusal with zero side effects, the recorded spelling being accepted, markdown-wrapped and punctuated URLs, working lines not counting, the merge wrapper propagation, a self-hosted merge request with no forge call, the prompt carrying the rule, and the brief carrying the worker rule. * no-mistakes(review): Remove stale PR URL enforcement * no-mistakes(ci): Removed backlog notes as an accepted PR identity source. PR URLs may now be copied only from the task’s `done: PR <url>` status or canonical `pr=` metadata; otherwise supervision reports only the known identifier and leaves PR checking unarmed. Updated related guidance/docs and verified with branch-supervision tests, brief tests, ShellCheck, and `git diff --check` * fix(bin): disable Claude feedback drafts for fleet launches (kunchenguid#3661) * fix(bin): disable Claude's feedback-draft flow for fleet-launched agents Scope --settings '{"feedbackDrafts":"off"}' to every Firstmate-launched Claude crewmate and secondmate, so /bug and /feedback never queue or submit a bug report on the captain's behalf. feedbackDrafts is the documented settings key (Claude Code changelog 2.1.247); the per-launch CLI flag never touches the captain's global settings.json. Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3 * no-mistakes(review): Prevent managed settings from re-enabling Claude feedback drafts * no-mistakes(document): Fix Claude feedback documentation formatting * fix(bin): layer both feedback-draft controls for defense in depth The prior --settings-only fix can be overridden by a managed Claude settings policy (feedbackDrafts precedence). Keep CLAUDE_CODE_SEND_FEEDBACK=0 alongside --settings '{"feedbackDrafts":"off"}': either control alone disables the SendFeedback tool, so a managed override of one still leaves the other in force. Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3 * no-mistakes(document): Document Claude feedback-draft suppression ownership * feat(tests): run three more validation families concurrently (kunchenguid#3662) * perf(tests): admit three more families to concurrent validation The three families that `docs/fm-test-isolation-proof.md` recorded as refused were not refused for concurrency. Each blocker was a test that decided a property by wall clock, or a script filed where it cannot run. Fixing those three things admits all three families and recovers 28.6 minutes of local validation with no assertion removed or weakened. - `tests/fm-backlog-handoff.test.sh` injected its pre-move crash by killing the handoff, sleeping a fixed second, then delegating the move to the real binary. Nothing ever killed the fake, so on a host slow enough for the case's next assertions to take longer than a second, the orphan woke and completed the very move the case requires left undone, and recovery then failed with `Task "pre-move-crash" not found in this backlog`. Watching the two backlogs during the injected crash showed exactly that, the item moving one second after the crash. All four crash injections in the file now go through a new `fm_fake_crash_injector` shim that signals the target and returns only once it is observably gone, and the pre-move fake never delegates the move at all. - `tests/fm-session-start.test.sh` proved the startup digest does not block on a slow current-state read by timing the whole digest against a fixed eight-second sleep, which a loaded host exceeds without the property being violated. It now holds that read open until the case releases it and asserts, the moment the digest returns, that the read has not finished. A digest that waited would wait indefinitely rather than for an interval a slow host can out-run, so the assertion is stronger than the bound it replaces. Its scan budget moves to the maximum, because the old value left two seconds of margin over the fixed sleep and measured the host rather than the deadline that `tests/fm-inactive-reconcile.test.sh` owns. - `fm-backend-herdr-focus-flash-e2e` was filed in the family map's catch-all, which put it in the portable serial lane, where Linux CI gate-skips it: that real-Herdr regression was running nowhere. It moves to `real-herdr-gated` and the required Herdr lane. `fm-claude-stop-autoarm-live-e2e` gate-skips on its opt-in variable and moves to `live-harness-optin`. The 28 remaining ungrouped scripts become an enumerated `standalone` family instead of admitting `unclassified` itself. `unclassified` is the family map's `*)` arm, so admitting it would silently grant concurrency to every test added afterwards, which is exactly the population with no proof. A new test still lands in `unclassified` and stays serial, and `tests/fm-test-run.test.sh` covers that split behaviorally. Each family passes two consecutive four-worker proofs with zero failures. On the production runner, `secondmate` goes 1233.1s to 453.4s, `session-bootstrap` 756.4s to 286.4s, and `standalone` 724.6s to 261.1s: 2.71x overall and 1713.2s recovered. The whole suite runs 177 scripts in 52.6 minutes of wall clock against 121 minutes of summed script time. * no-mistakes(document): Refresh concurrent validation and shard documentation * no-mistakes(ci): Fixed the real-Herdr focus-flash E2E race exposed by reclassification. Part C now starts its persistent child atomically via `pane run` and verifies stable child identity through Herdr’s public `process-info` interface, avoiding the racy send-text/send-keys sequence and platform-specific `ps` matching. Verified with bash syntax checking, ShellCheck, git diff checks, and the complete E2E test on Herdr 0.8.2 * feat: structure no-mistakes ask-user escalations (kunchenguid#3670) * feat(brief): structure no-mistakes ask-user escalation as event + snapshot file Crewmates escalating a no-mistakes ask-user gate now report one status event naming every finding id plus a snapshot file holding the gate's axi finding records verbatim (id, severity, file, line, description, authority), using the same shape even for a single finding. The status line never paraphrases. The format is defined once in fm-dod-lib.sh and rendered into both the scout and ship rule 6 in fm-brief.sh, so a promoted scout - whose rule 6 fm-promote.sh preserves unchanged - gets the identical contract as a freshly-spawned no-mistakes ship worker. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PpiWaDerbYavTLPPtEjQei * no-mistakes(review): Preserve ask-user escalation output contract * no-mistakes(review): Align escalation format test expectation * no-mistakes(review): Scope ask-user escalation instructions correctly * no-mistakes(review): Remove ask-user from generic decision rules --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> * fix(bin): require self-sufficient no-mistakes intent (kunchenguid#3671) * fix(bin): require a self-sufficient no-mistakes intent A no-mistakes worker's --intent is only as useful as the string it passes. PR kunchenguid#3604 shipped with an intent that was only "do 1, 2, 3, 7 from the report": the real contract lived in a private scout report and never reached --intent, so nobody holding that string plus the codebase could have derived the specification. This is pure instruction at the contract's one owner; no spawn-side or promotion-side check is added. - bin/fm-dod-lib.sh: the generated no-mistakes Definition of done now states that the --intent string must be self-sufficient (the string plus the codebase reconstructs roughly the same specification) and tells the worker to write the substance of any report, decision, or PR the captain's intent refers to into --intent rather than the pointer, while Firstmate build instructions and the worker's own decisions still stay out. The spawn-time overlay points back at that rule so its "supersedes" wording cannot cancel it, and the header's owner statement carries the rule. - AGENTS.md section 11 and bin/fm-brief.sh's header ask Firstmate to include the substance of referenced material when filling ## Captain's intent, and section 11 points at the owner of the rule. - tests/fm-brief.test.sh and tests/fm-task-delivery.test.sh assert the rendered brief and launch contract carry the rule. Claude-Session: https://claude.ai/code/session_01YMhEe42q7BAAoN6RxNuzim * no-mistakes(document): Replace incident-specific intent test commentary * fix: accelerate local Bearings snapshot composition (kunchenguid#3499) * Speed local fleet snapshot composition * no-mistakes(review): Stabilize task inventory during concurrent snapshot composition * no-mistakes(document): Document local snapshot observation concurrency * no-mistakes(ci): Fixed CI failures by making empty task manifests compatible with stock macOS Bash 3.2, snapshotting task metadata before concurrent observations to prevent generation drift, strengthening the behavioral race regression, and updating the stock-Bash Bearings test count to 45. Verified fleet snapshot tests (15), Bearings tests (45), workflow lint tests, project lint, Bash 3.2 parsing, and diff checks * no-mistakes(ci): Fixed the Linux CI failure caused by passing large backlog/task JSON through jq command-line arguments, which exceeded the per-argument size limit. Both inventory projections now stream large JSON inputs through stdin. Verified with fm-bearings-snapshot.test.sh (45 tests), fm-fleet-snapshot-view.test.sh (15 tests), Bash syntax, and git diff checks * no-mistakes(ci): Fixed concurrent task teardown during metadata capture: vanished metadata is now omitted while genuine copy failures remain fatal. Added a deterministic public Bearings regression test and updated CI’s expected test count. Verified with the full Bearings suite, workflow-lint suite, Bash syntax checks, and git diff checks * no-mistakes(ci): Fixed PR-caused CI and review issues: streamed large fleet JSON through jq stdin to avoid Linux argument limits, kept crew-state reads bound to captured metadata generations, and strengthened the behavioral race test. Bearings (46 tests), fleet snapshot (15 tests), crew-state, backend, lint, Bash syntax, and diff checks pass locally. Serial shard 5’s unrelated task-inbox segmentation fault appears infrastructural/flaky * no-mistakes(ci): Fixed endpoint-state generation crossing by validating captured spawn_gen before and after local endpoint probes, falling back to exact metadata identity for legacy tasks. Stale probe results now become unknown instead of false unhealthy state. Added a behavioral relaunch-race regression test. Verified the full Bearings snapshot suite, shellcheck, bash syntax, and git diff checks * fix(snapshot): keep live observations generation-coherent * no-mistakes(review): Keep secondmate observations generation-bound without copying reports * no-mistakes(document): Document generation-coherent snapshot observations * test(bearings): measure local read overlap instead of wall-clock budget The large-local-snapshot regression asserted that a whole snapshot composed in under five seconds. That bound measures how loaded the host is, not whether the per-task reads actually overlap, so it failed intermittently on a contended machine: one run in six on a box at load 16-20, landing exactly on the five second boundary. Time a serialized run and a concurrent run of the same workload instead and require the concurrent one to save at least two seconds. Both runs pay the same composition overhead, so the difference isolates the overlap this change delivers. Five one-second reads serialize into five seconds and overlap into about one, and re-serializing the reads collapses the saving to roughly zero, so the assertion still fails loudly if the concurrency regresses. Also bump the pinned Bearings test count to 48, since rebasing onto the current default branch picked up its captain-hold test. * no-mistakes(review): Restore JSON-derived decision flags * no-mistakes(review): Unify status-derived snapshot observations * no-mistakes(ci): Updated the stock macOS Bash CI check’s Bearings test count from 48 to 49. Verified the full Bearings suite passes and emits exactly 49 TAP successes; git diff checks pass * fix: prevent stale supervision wake loops (kunchenguid#3672) * fix(bin): stop the supervision branch's stale-ack and ghost-report loops Clean-slate implementation of the four authorized recommendations from the supervision-ghost-retrigger analysis (items 1, 2, 3, and 7), in their minimal form, superseding PR kunchenguid#3604: - fm_branch_report refuses a task the wake being handled never named. The extension fixes the reportable task set from the eligible rows before each prompt (signal and stale rows resolve to their tasks, a heartbeat allows any task with a live record, fleet is always allowed), so a report typed from memory about a task whose records teardown already removed is never stored or delivered. - An acknowledgement that consumes nothing says "nothing was acknowledged through N" and prints the exact --ack-through / --recovery-generation command for the current presented wake, instead of "re-run the drain", which re-fed the same stale acknowledgement in a loop. - bin/fm-guard.sh no longer tells the branch actor to drain queued wakes while it is handling them; it names the granted rows instead. - Teardown removes state/.<task>.branch-outcome-index for ordinary tasks and descendants; the index rebuild and the append-side index write both skip a task with neither a live record nor a status log, so the branch's report of a teardown it just performed is stored without recreating the index. No new locking, no spawn-generation binding, and no retired-task refusal: the branch can still report the outcome of a task it just tore down, and the teardown test now proves that path end to end. * fix(bin): narrow the branch report scope and guard silence to the minimal form Apply the four review decisions on the clean-slate branch: - A signal or stale prompt may report only the tasks its own rows resolve to; fleet is refused there too. A heartbeat review is not scoped by task at all, so the extension no longer tracks live task records and refuses nothing by task id during a fleet review. - The outcome-index rebuild no longer skips retired tasks; the append-side skip alone keeps a torn-down task's index from being recreated. - bin/fm-guard.sh keeps the queued-wakes warning silent for the branch actor instead of printing a replacement note. * no-mistakes(document): Align supervision docs with scoped wake handling * fix(bin): avoid fleet snapshot argument limits (kunchenguid#3677) * Fix fleet snapshot large JSON transport * no-mistakes(review): Captain: file-back fleet snapshot transport safely * no-mistakes(review): Captain: file-back parent summary aggregation * no-mistakes(ci): Rebased the PR's three commits onto f4d7875 and resolved the fleet snapshot conflict while preserving the base's task-observation lifecycle. Fixed Greptile's valid finding by recursively removing the private mktemp transport directory, so future transport files cannot cause cleanup to fail. Verified with tests/fm-home-summary-refresh.test.sh, bin/fm-lint.sh, git diff --check, and ancestry checks. All passed; the fix remains as an uncommitted worktree change for the outer executor * fix(bin): attribute active runs with unfetched pipeline heads (kunchenguid#3681) * fix(bin): recognize active pipeline fix rounds with unfetched run heads A no-mistakes fix round advances the run head beyond the submitted head, and the pipeline commits in its own checkout, so the task copy never receives the new commit object. fm-crew-state's strict head rule rejected the active row, the coarse runs-list scan skipped it and matched the older failed row at the submitted head, and an active validation read as failed (observed on model-routing-benchmark-hardening: active head ac61c64 vs task copy at fb47636d). fm_nm_runs_status_for_worktree in bin/fm-nm-run-lib.sh now owns runs-ledger attribution: the branch's newest row alone decides, and a newest row whose head cannot resolve locally is recognized only as a provable pipeline-owned continuation - active (running) and anchored by the immediately older row for the same branch having ended at exactly this worktree's HEAD. The reader keeps the axi TOON as full detail for that proven same-branch run. Unanchored, ancestor-anchored, and terminal unresolvable rows stay unattributed, so branch-name coincidence and other tasks' runs never match, and fm_nm_head_matches_worktree keeps its exact prior semantics for teardown (verified by the full teardown suite). Tests: reproduction regression for the unfetched active fix head (reads working via full run-step detail), coarse-path continuation when axi answers another branch, and negative controls for the unanchored active row and the unresolvable terminal row with the historical fallback preserved. Ported onto upstream/main f4d7875, where kunchenguid#3194 independently added the branch_sync custody exemption on the full axi-status path: both mechanisms now coexist, each owning one surface (TOON custody on the full path, the runs ledger on the coarse path). The port deletes the superseded coarse scan-and-skip (nm_runs_status_for_branch) and its now caller-less helpers (fm_nm_head_resolvable, nm_coarse_head_matches_worktree), renames the exemption comment's "the one exemption" phrasing now that a second complementary exemption exists, and points the stale FM_CREW_STATE_RUNS_LIMIT comment at fm_nm_runs_status_for_worktree (judge follow-up #1). The parent coarse-guard test's fixture is the ledger-anchored continuation shape, so its expectation flips to the fixed behavior (working via run-step, never the older failed row); a new mismatched-anchor coarse negative control preserves that guard's original no-anchor protection (pane answers, never the older row). * no-mistakes(document): Clarify pipeline attribution documentation * fix(bin): pre-register claude workspace trust at spawn time (kunchenguid#3663) * fix(bin): pre-register claude workspace trust for task worktrees A claude crewmate launched into a fresh task worktree met Claude Code's interactive workspace-trust dialog before it ever read its brief, and firstmate could not answer it: the key plane carries only Enter, Escape, and C-c with no arrow navigation, and the dialog's selection starts on "No, exit", so the documented Enter recipe ended the session instead of accepting it. Two workers wedged this way and were unblocked only by hand-seeding the trust store per path. --dangerously-skip-permissions does not cover that gate. `claude --help` records the dialog as skipped only in non-interactive mode, through -p or a non-TTY stdout, and a crewmate pane is interactive, so there is no launch flag to reach for. fm-spawn now pre-registers the worktree through bin/fm-claude-trust.sh in the existing claude branch, before the project settings that the same gate would otherwise block, and refuses the spawn when that write fails rather than launching a worker that would wedge. The scope test is the safety property and is structural rather than a path policy: the path must be a linked git worktree, sharing the spawning project's common dir, whose top level is exactly the resolved argument. Git is the ground truth, so the argument is never trusted on its own word, and a primary checkout, an unrelated repo, a worktree subdirectory, a plain directory, and a home directory are each refused rather than warned about or skipped. A treehouse or orca path prefix was deliberately avoided because treehouse's root is configurable, which would make a prefix both wrong and a new policy surface. One structural test covers both worktree providers. tests/fm-claude-trust.test.sh pins both halves, including a case where HOME is itself a valid linked worktree so the home guard is proven load-bearing rather than passing vacuously, plus the spawn-level proof that a claude spawn trusts its worktree and launches with the brief pointed at the same store. The adapter reference no longer tells a firstmate to press Enter on that dialog, and the shared trust reference now names every harness surface: which harnesses gate, which suppress at launch, which dodge the gate, which now pre-registers, and that a claude secondmate is excluded by design. The spawn fixture runs each spawn against a throwaway HOME so the suite cannot write the developer's real store, isolating through HOME rather than CLAUDE_CONFIG_DIR because the spawn forwards a set CLAUDE_CONFIG_DIR onto the launch command that launch-shape assertions read. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HNEN2GLnew27HFyfi4ms4v * fix(bin): create the staged trust store exclusively The staged store was written to a predictable pid-based path with a plain write, which follows a symlink. Where the Claude config directory is writable by another local account, that account could pre-create the path as a symlink and redirect the write into another file the launching user owns. The staged name now carries random bytes and is created with an exclusive "wx" open, so an existing path is refused outright instead of followed. The happy-path test also asserts no staged store survives the rename. The durability comment now states the residual window plainly: the readback proves the entry landed, not that it survives, because a vendor session that rewrites the whole store afterwards can still drop it and no lock closes that window when the writer is Claude itself. The worker then meets the dialog and stalls, which reaches firstmate as the ordinary stale wake rather than as silent success. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HNEN2GLnew27HFyfi4ms4v * no-mistakes(review): neutralise CDPATH in claude trust scope guard * no-mistakes(review): sandbox HOME in spawn tests, drop out-of-scope artifacts * no-mistakes(review): refuse unresolvable git dir, compact store, fix secondmate doc * no-mistakes(review): clear git env overrides, resolve symlinked store target * no-mistakes(review): degrade without node, fix Pi gate claim, record trust proof * no-mistakes(review): refuse without node, pin CLAUDE_CONFIG_DIR in spawn tests * no-mistakes(review): refuse relative config dir and concurrent store modification * no-mistakes(review): correct orca worktree claim, clean staged store on failure * no-mistakes(review): restore pretty-printed store, correct trust dialog docs * no-mistakes(review): arm trust gate before busy state to avoid orphans * no-mistakes(document): record claude trust pre-registration in its owner docs * no-mistakes(document): note orca limit for claude trust pre-registration * no-mistakes(ci): Fixed the Greptile P1 on bin/fm-spawn.sh by moving the Claude trust gate earlier rather than adding cleanup machinery. Diagnosis: Greptile reported that when Claude trust registration fails on tmux/Zellij/cmux/non-projected Herdr, the exit runs after the backend endpoint and /tmp/fm-<id> were created, and the abort trap cleans neither. The endpoint half is pre-existing, deliberate architecture — the two refusals immediately above the gate (the 60s `treehouse get` timeout at fm-spawn.sh:2550 and `validate_spawn_worktree` at :2487) also exit with the endpoint live and direct the operator with "inspect window $T"; spawn_abort_cleanup only reclaims orca endpoints (already covered via ORCA_ABORT_CLEANUP) and herdr projections. The temp-root half was genuinely introduced by this PR: the gate was placed beside the busy-state arm, ~30 lines after `mkdir -p "$TASK_TMP/gotmp"`, and fm-teardown can only find that root through `tasktmp=` in a meta record a refused spawn never publishes. Root-cause fix (smallest correct change, no new subsystem): - bin/fm-spawn.sh — moved the `claude*` trust gate from inside the busy-arm block up to the first point $WT is known, immediately after the `freshen_spawn_worktree_base` block and before TASK_TMP creation, the STATE setup, and the relaunch `clear_relaunch_harness_wiring` retirement. A refusal now leaves no temp root, no retired relaunch wiring, and no busy record; only the endpoint remains, in the same class as the two refusals just above it. - bin/fm-spawn.sh — the refusal message now ends with "inspect window $T", matching the existing convention so control/teardown can identify the endpoint. $T is set for every backend on the non-secondmate path. - bin/fm-spawn.sh:196 — header note corrected from "before any state is armed" to "before any per-task state exists". - tests/fm-claude-trust.test.sh — the existing refused-spawn test's own comment claimed "before any task state exists" but only asserted busy state. Renamed to test_refused_spawn_leaves_no_task_state and added an assertion that /tmp/fm-<id> is absent, with the task id suffixed by the test process pid so the assertion reads only this run's path (a stale /tmp/fm-refusedspawn from the fixed-id version was in fact present on this box). No assertions on implementation source bytes. Verification run locally: - The new assertion fails against the pre-fix bin/fm-spawn.sh ("not ok - a refused spawn stranded a temp root no teardown can find") and passes after — a real before/after regression proof. - tests/fm-claude-trust.test.sh: 20/20 ok. - tests/fm-backend.test.sh, fm-backend-orca, fm-control-relaunch, fm-spawn-dispatch-profile, fm-trace-context-spawn, fm-gotmp: all pass. - tests/fm-backlog-atomicity.test.sh: rc=0, 79 assertions ok. - bin/fm-lint.sh (repo's single lint owner, pinned ShellCheck 0.11.0 + actionlint 1.7.12): clean. - No /tmp/fm-refusedspawn* leftovers after the runs. Scope respected: no trust subsystem, no policy layer, no config surface, no endpoint-cleanup mechanism added; the change is an ordering move plus one error-message clause and the test that pins it. Adapter references and docs made no ordering claim, so none needed updating. Changes are left uncommitted in the worktree for the outer executor --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * fix: restart every live second mate after updates (kunchenguid#3690) * feat(update): restart every live second mate after a successful update /updatefirstmate only restarted a second mate when that pass advanced its AGENTS.md or .agents/skills. An already-current home was skipped entirely, a bin/-only advance was steered instead, and a remote host that could not report its instruction diff was downgraded to a re-read. A running agent also freezes its launch-time wiring - turn-end hooks, harness flags, per-harness feature switches - and none of that is derivable from a file diff, so an unchanged tracked surface is not evidence the agent is already on the current behavior. Restart is now unconditional on a successful update of that home. Every live second mate the pass leaves on the target commit is restarted, whether it advanced or was already there. The safety contract is unchanged: open records are persisted before the agent is replaced, nothing is forced, stashed, or discarded, a home the pass had to skip is not restarted at all, and a mate whose runtime cannot prove a restart keeps the honest re-read path and is never reported as reloaded. bin/fm-ff-lib.sh gains a settled-state hook that fires for a home left at the base whether it advanced or was already there, and never for a skipped one; the instruction-gated hook the session-start convergence sweep uses is untouched. Regressions: fm-update pins the already-current mate into the restart set and the unprovable one into the nudge set, and fm-secondmate-restart drives both real commands end to end - an already-current home is named, persisted, and genuinely replaced with its checkout untouched, while the unprovable one keeps its running agent. * no-mistakes(document): Document unconditional secondmate restarts * fix(bin): close pending-reply decisions via resolve-key (kunchenguid#3696) * fix(bin): close reserved pending-reply keys via fm-send --resolve-key fm-send wrote answered: notes that the reserved-key fold ignores, so operator closes exited 0 while OPEN DECISIONS kept the decision open. Speak the owning library's close vocabulary on that path, and refuse when a reserved close cannot take effect. * no-mistakes(review): Safely quote manual decision-close recovery commands * no-mistakes(review): Reject unclosable overlong decision keys before sending * no-mistakes(review): Remove contract suffix from open decisions hint * no-mistakes(document): Document resolve-key line-cap refusal * fix(bin): prevent false missed-reply escalations (kunchenguid#3697) * fix(bin): stop false missed-reply escalations for same-basename self-home answers A healthy secondmate that wrote corr= to its own state/<id>.status never matched the parent channel, so recovery confirmed and the record escalated as pending-reply-missed. Make the report helper resolve the parent channel itself, skip parent-replies.status as wrong-home, put a readable sighting path on the missed line, and restatement-copy only that same-basename self-home file onto the parent channel. * no-mistakes(review): Resolve late replies before recovery escalation * no-mistakes(review): Tighten reply routing and regression coverage * no-mistakes(review): Preserve reply paths and require explicit home * no-mistakes(review): Encode wrong-home paths before persistence * no-mistakes(document): Document corrected secondmate reply routing * no-mistakes(lint): Fix pending-reply ShellCheck warnings * feat: add verified Gemini crewmate runtime (kunchenguid#3695) * feat(harness): verify gemini as a crewmate runtime adapter Adds Gemini CLI as a fourth dispatch target alongside claude, codex, and grok, scoped to crewmate and scout work only. Every axis was proven against gemini-cli 0.58.0 rather than inferred; docs/verification/runtime-backends.md carries the dated evidence and names what stayed unverified. Busy state is semantic, not rendered: BeforeAgent opens a turn and AfterAgent and SessionEnd close it. AfterAgent also fires on a manual interrupt, so a cancelled turn closes its own record. Three findings shaped the wiring rather than a config line: - --skip-trust and GEMINI_CLI_TRUST_WORKSPACE=true are presented by the CLI as equivalents and are not. A controlled A/B showed --skip-trust leaves project configuration unloaded, so workspace skills never load. - The worktree's .gemini/settings.json is the PROJECT's committed settings file, unlike claude's settings.local.json. Firstmate's hooks therefore go to a firstmate-owned state/<id>.gemini-settings.json reached through GEMINI_CLI_SYSTEM_SETTINGS_PATH, which also works untrusted and merges with a project's own hooks instead of replacing them. - The shipped CLI is a node bundle whose live process reports comm as MainThread, so ancestry cannot see it. GEMINI_CLI=1 is load-bearing and is tested before an inherited CLAUDECODE, and pane liveness identifies gemini from the script argument through the new bin/fm-gemini-lib.sh. Gemini is refused for secondmates: it has no primary supervision protocol. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L * test: clear gemini's marker in launch and detection expectations Every non-gemini launch now clears GEMINI_CLI the way it already clears cursor's markers, so the two tests that pin the exact launch prefix are updated to match. The harness-detection tests that scrub foreign markers before probing ancestry scrub GEMINI_CLI too, so running the suite from inside a gemini session cannot produce a false verdict. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L * docs: classify the gemini harness reference The documentation inventory is the single classification owner for maintained prose surfaces, and every surface must appear in it exactly once. The new harness reference is agent-runtime, matching its siblings. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L * no-mistakes(review): Narrow Gemini ancestry detection * no-mistakes(review): Restrict Gemini hooks to canonical launches * no-mistakes(document): Document Gemini adapter support boundaries * no-mistakes(ci): Fixed Gemini process identity when interpreter or script paths contain whitespace. Tmux liveness now uses NUL-delimited /proc argv on Linux, with the existing flattened ps fallback elsewhere. Added a real-process regression test. Verified with the Gemini harness test suite, full fm-lint, ShellCheck, and git diff --check. The CI and Require no-mistakes runs were action_required/attestation outcomes rather than code failures --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(teardown): conclude parked runs advanced past task copy (kunchenguid#3704) * conclude parked runs the pipeline advanced past the task copy A no-mistakes fix round commits in the daemon's own gate-repo clone, so a run parked at a gate can carry a head whose object the task copy never received. Teardown's strict object-local identity rule then declined to conclude the run, and cleanup left it parked forever holding a fleet slot (observed 2026-09-03; the same masking condition PR 3681 fixed on the read path, now closing the teardown half its scope boundary deferred). task_status_is_own_parked_run now falls back - only when the reported head resolves to no local object - to the one shared runs-ledger attribution rule fm_nm_runs_status_for_worktree (bin/fm-nm-run-lib.sh), whose anchored continuation proof binds the branch's newest active row to this worktree's exact submitted head. Foreign branches, stale history, terminal rows, ancestor-only anchors, diverged newer rows, and ambiguous multi-row shapes all still refuse, and runs that are actively running, fixing, or in CI remain untouched: only the parked-at-a-gate determination ever reaches the abort. No sqlite access, no fetches into another task copy, no custody changes, no duplicated matching logic. * tighten the parked-run ledger fallback and pin both judge corrections The teardown ledger fallback now authorizes concluding this task's parked run only when the shared runs-ledger rule's proved answer is the explicitly active word (running): a terminal newest row - even anchored at exactly the worktree's head - is finished history and never an abort authorization. The read path may classify the same owner's answer; teardown's abort must never fire for a run that already ended. Two bounded pre-validation corrections from the implementation review: - a fetched-object counterfactual pins the strict-rule path: a pipeline fix head fetched into the task copy aborts through object-local identity alone, with an empty ledger and a proof the runs query never fired; - a negative fixture pins the tightened boundary: an unresolvable reported head with a terminal newest same-branch row anchored at the worktree head engages the ledger fallback and still refuses, so the refusal is the terminal-word boundary and not an earlier guard. * no-mistakes(review): Bind teardown ledger fallback to validated run heads * no-mistakes(review): Restore validated advanced-head ledger continuation * no-mistakes(review): Reject invalid ledger dates and terminal statuses * no-mistakes(document): Document teardown ledger scan limit * feat(bin): show requested vs effective model in Herdr agent view Track spawn-config requested_model separately from runtime-verified effective_model, probe Claude/Pi transcripts for exact API ids, push compact display metadata to Herdr, and preserve verified models across relaunch/compaction hooks without inferring aliases as truth. * fix(bin): keep re-probing effective model after first exact reading fm-model-sync.sh only probed for the runtime-verified effective model while it was still pending/UNKNOWN, so a session that later switched models (manual switch, provider fallback) kept displaying the first verified model forever and never appended a fallback-history entry. Probe unconditionally instead; fm_model_record_effective already no-ops when the probed value is unchanged, so this stays cheap. Addresses the Greptile P1 finding on PR kunchenguid#3705's fm-model-sync.sh. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BwfFjeYQcz9cZ3vEZohmpm * fix(bin): distinguish Cursor Grok, direct xAI Grok, and Anthropic Claude in the display Kapitänskorrektur: harness alone conflated Cursor-hosted Grok models (cursor-grok-4.6-*) and direct xAI Grok models (xai/grok-4.6) under one generic label, and displayed Anthropic Claude without naming the provider. Add fm_model_source_label, pattern-matched on the verified exact model id, so the compact display always reads Cursor · Grok, xAI · Grok, or Anthropic · Claude with the exact model id appended. Falls back to the existing harness label for every other model. No routing change: this only affects display strings. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BwfFjeYQcz9cZ3vEZohmpm * fix(bin): wire model-sync into the Pi extension's turn lifecycle fm-model-sync.sh was only invoked from Claude's SessionStart/ UserPromptSubmit/Stop hooks; the Pi harness's own extension (state/<id>.pi-ext.ts) never called it, so a Pi-hosted session (e.g. a pi/xai-grok crewmate) never refreshed its effective model after the first probe and Herdr kept showing the stale value with no fallback-history entry. Call fm-model-sync.sh from the same agent_start/turn_end boundaries Pi already uses for busy-state and the turn-end notification touch. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BwfFjeYQcz9cZ3vEZohmpm * fix(bin): serialize fm-model-sync.sh's meta read-probe-write Overlapping lifecycle events (Pi's agent_start/turn_end, Claude's SessionStart/UserPromptSubmit/Stop) can invoke fm-model-sync.sh concurrently for the same task. The unlocked read-probe-write let interleaved runs revert a newer effective model, mismatch its source, or duplicate a model-history entry. Serialize the critical section through the same per-task meta lock fm-spawn.sh already uses (fm_meta_lock_path + fm_lock_acquire_wait/fm_lock_release), released before every exit path. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BwfFjeYQcz9cZ3vEZohmpm --------- Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> Co-authored-by: Arthur Haro <38157909+haroarthur@users.noreply.github.com> Co-authored-by: Nicolas Payette <nicolas.payette@specira.ai> Co-authored-by: Jon Roosevelt <rooseveltadvisors@gmail.com> Co-authored-by: att430 <41454889+att430@users.noreply.github.com> Co-authored-by: Valentino-Sole <171032438+Valentino-Sole@users.noreply.github.com>
Valentino-Sole
added a commit
to Valentino-Sole/firstmate
that referenced
this pull request
Sep 8, 2026
* fix: start a fresh supervision branch for every main session (kunchenguid#3600) * fix(pi): start a new supervision branch conversation per main session The supervision branch reopened one recorded conversation forever, so every main session start reloaded the current generated prompt and then weeks of accumulated thread, where a superseded rule could still outweigh today's. The branch conversation is now scoped to one main session: the session generation owns the recorded conversation, so a cold start, /new, /resume, /fork, or a reload always builds a new one, while a rebuild inside one session (a model or effort change) still continues that session's own conversation. The dialog mirror re-anchors with it. Its durable cursor records what the previous branch conversation received, so a /resume or reload - which keeps main's own session file - would otherwise leave the new branch blind to dialog main itself still has. The reset is bounded by the current main session, and the cursor keeps advancing incrementally within it. The durable outcome store and its processed marker are untouched, so unacknowledged captain-facing outcomes still re-present on the new main session. * no-mistakes(document): Document fresh Pi supervision conversations * no-mistakes(ci): Fixed the flaky concurrent inbox failure. Lock acquisition now retries when a competing lock disappears between a failed claim and inspection. Added a behavioral regression covering that race. Verified the full inbox test four times, project lint, and git diff checks * feat: restart second mates after instruction updates (kunchenguid#3614) * feat(update): restart second mates whose instructions changed /updatefirstmate pulled new bytes onto disk and then asked each advanced second mate to re-read them. A running agent holds AGENTS.md and every loaded skill frozen from launch and no verified harness offers a reload, so that steer could not reach a loaded skill at all and left the mate holding two contradictory copies of its own job description. An eligible mate is now restarted instead, in the same home and endpoint, through the existing transactional relaunch. The restart is gated on the mate first writing down the open work it holds only in conversation - the open-record half of /stow, never its memory sweeps - so an unregistered captain call is flushed before the conversation is spent. Anything that leaves the reload unprovable falls back to the old re-read message and is reported as exactly that, never as a clean reload. Remote mates take the same path: fm-remote-secondmate-control.sh gains a relaunch verb whose host-local leg runs that same control plane, since the mate is an ordinary local secondmate from its host's point of view. The primary resolves the profile and passes it explicitly, because config/secondmate-harness is not inherited and the file on that host belongs to a different home. fm-update.sh now splits its advanced live mates into a restart set and a nudge residual, and both sets require a changed instruction surface, which also closes the over-nudge against the session-start sweep. Restart is stricter still: a bin/-only advance reloads itself on the next call, so it never costs a conversation. Colocated tests cover the gating, the persist-then-restart order, the task-subset persist request, each unsafe fallback, the remote hop, and the remote sync's new instruction-surface report. * no-mistakes(review): Fix restart correlation, concurrent waits, and lifecycle reporting * no-mistakes(review): Parallelize relaunches and classify replacement incarnations * no-mistakes(review): Gate restart actions on live agent state * no-mistakes(review): Handle failed restart workers without hanging * no-mistakes(review): Nudge legacy remotes and preserve persist recovery * no-mistakes(review): Document one-time secondmate restart rollout * no-mistakes(review): Honor arrived replies and refresh remote profiles * no-mistakes(review): Revert remote parent profile reconciliation * no-mistakes(review): Reset remote profile defaults and honor published results * no-mistakes(review): Preserve fallback nudges for unverifiable secondmates * no-mistakes(document): Document second-mate restart update flow * no-mistakes(lint): Fix ShellCheck warnings in restart scripts * perf: accelerate local validation with bounded concurrency (kunchenguid#3644) * perf(tests): route gate verification through the bounded concurrent runner Local validation was the pipeline's dominant cost: across 67 recorded no-mistakes agent sessions on this repo, 99.3% of command execution was `bash tests/*.test.sh`, run strictly one script at a time, and 2% of those calls were killed by an agent-guessed timeout and paid for twice. Three changes, each measured: - `.no-mistakes.yaml` pins `commands.test` to `bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner already owns changed-file selection, bounded concurrency, the refusal of unproven scripts, and a generous automatic per-script bound, so the gate's baseline is neither a serial chain nor a guessed timeout. It stays intent-targeted - the Test step still runs its evidence agent on top - and excludes the live-Herdr family the required Herdr lane owns. - `bin/fm-test-run.sh` gives a plain list of script paths the same bounded automatic scheduler and automatic bound that `--changed` gets. Naming several subjects is how a verification round asks for exactly those scripts. The curated selections are untouched: `--lane` still composes CI shards whose serial lane must stay serial, `--family` is what the required Herdr lane runs, and `--all` stays a deliberate complete regression. - `pr-forge` is admitted to the concurrent-safe family registry on two consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those, and records `secondmate` and `session-bootstrap` as refused with the exact script and reason each failed on, so the refusals are actionable rather than silent. Measured on this host, 0 failures on both sides: verification round, 4 scripts 448s chained -> 231s through the runner (-48%) pr-forge family 409.2s at 1 worker -> 237.9s at 4 (1.72x) watcher-wake-lock family 1311.1s at 1 worker -> 539.3s at 4 (2.43x) A fourth lever was implemented and then removed because the measurement refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made `fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s unchanged, back to back. Those sleeps are not overhead added to the clock - they are how a test waits for a subject moving on fm-watch.sh's own one-second cadence - so sampling less often only delays detection. It also broke `fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a settled condition. CONTRIBUTING.md records that result so the experiment is not repeated. * no-mistakes(review): Separate concurrent runs by isolation proof family * no-mistakes(review): Limit automatic timeouts to changed-file validation * no-mistakes(document): Clarify validation concurrency documentation * fix: copy PR URLs from durable records (kunchenguid#3648) * fix: copy PR URLs from records or abstain, never assemble them Supervision reported a plausible but dead PR link three times because its prompt demanded a full https:// URL at a moment when only a PR number was observable, so the model assembled an owner/repository from memory, and the PR check then accepted that URL and wrote it into the task record, after which the model kept defending its own tool-endorsed guess over the worker's real link. Three changes close that chain without any live forge lookup, so private forges are treated exactly like public ones: - bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy or abstain" section requires a URL to be copied verbatim from a durable record (the done: PR <url> status line, pr= metadata, or the backlog note), forbids assembling owner, repository, host, or number from memory, and has the branch report only the identifier it actually holds when no record names the URL yet, leaving the PR check unarmed until the worker's ready line arrives. AGENTS.md section 7 and 9 carry the same copy-or-abstain rule for main in place of the bare full-URL mandate. - Worker briefs (bin/fm-brief.sh, ship and scout rules) require the full https:// URL wherever a PR is mentioned - status line, terminal, or summary - never a bare "PR 108", so the link is in view as early as the number is. - bin/fm-pr-check.sh refuses, offline and before any side effect, a URL that the task's own done lines contradict, printing both spellings; a log naming no URL still records the argument as before. fm_pr_status_ready_urls in bin/fm-pr-lib.sh owns reading those lines. The refusal also reaches bin/fm-pr-merge.sh, so nothing merges under a contradicted URL. Tests cover the offline refusal with zero side effects, the recorded spelling being accepted, markdown-wrapped and punctuated URLs, working lines not counting, the merge wrapper propagation, a self-hosted merge request with no forge call, the prompt carrying the rule, and the brief carrying the worker rule. * no-mistakes(review): Remove stale PR URL enforcement * no-mistakes(ci): Removed backlog notes as an accepted PR identity source. PR URLs may now be copied only from the task’s `done: PR <url>` status or canonical `pr=` metadata; otherwise supervision reports only the known identifier and leaves PR checking unarmed. Updated related guidance/docs and verified with branch-supervision tests, brief tests, ShellCheck, and `git diff --check` * fix(bin): disable Claude feedback drafts for fleet launches (kunchenguid#3661) * fix(bin): disable Claude's feedback-draft flow for fleet-launched agents Scope --settings '{"feedbackDrafts":"off"}' to every Firstmate-launched Claude crewmate and secondmate, so /bug and /feedback never queue or submit a bug report on the captain's behalf. feedbackDrafts is the documented settings key (Claude Code changelog 2.1.247); the per-launch CLI flag never touches the captain's global settings.json. Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3 * no-mistakes(review): Prevent managed settings from re-enabling Claude feedback drafts * no-mistakes(document): Fix Claude feedback documentation formatting * fix(bin): layer both feedback-draft controls for defense in depth The prior --settings-only fix can be overridden by a managed Claude settings policy (feedbackDrafts precedence). Keep CLAUDE_CODE_SEND_FEEDBACK=0 alongside --settings '{"feedbackDrafts":"off"}': either control alone disables the SendFeedback tool, so a managed override of one still leaves the other in force. Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3 * no-mistakes(document): Document Claude feedback-draft suppression ownership * feat(tests): run three more validation families concurrently (kunchenguid#3662) * perf(tests): admit three more families to concurrent validation The three families that `docs/fm-test-isolation-proof.md` recorded as refused were not refused for concurrency. Each blocker was a test that decided a property by wall clock, or a script filed where it cannot run. Fixing those three things admits all three families and recovers 28.6 minutes of local validation with no assertion removed or weakened. - `tests/fm-backlog-handoff.test.sh` injected its pre-move crash by killing the handoff, sleeping a fixed second, then delegating the move to the real binary. Nothing ever killed the fake, so on a host slow enough for the case's next assertions to take longer than a second, the orphan woke and completed the very move the case requires left undone, and recovery then failed with `Task "pre-move-crash" not found in this backlog`. Watching the two backlogs during the injected crash showed exactly that, the item moving one second after the crash. All four crash injections in the file now go through a new `fm_fake_crash_injector` shim that signals the target and returns only once it is observably gone, and the pre-move fake never delegates the move at all. - `tests/fm-session-start.test.sh` proved the startup digest does not block on a slow current-state read by timing the whole digest against a fixed eight-second sleep, which a loaded host exceeds without the property being violated. It now holds that read open until the case releases it and asserts, the moment the digest returns, that the read has not finished. A digest that waited would wait indefinitely rather than for an interval a slow host can out-run, so the assertion is stronger than the bound it replaces. Its scan budget moves to the maximum, because the old value left two seconds of margin over the fixed sleep and measured the host rather than the deadline that `tests/fm-inactive-reconcile.test.sh` owns. - `fm-backend-herdr-focus-flash-e2e` was filed in the family map's catch-all, which put it in the portable serial lane, where Linux CI gate-skips it: that real-Herdr regression was running nowhere. It moves to `real-herdr-gated` and the required Herdr lane. `fm-claude-stop-autoarm-live-e2e` gate-skips on its opt-in variable and moves to `live-harness-optin`. The 28 remaining ungrouped scripts become an enumerated `standalone` family instead of admitting `unclassified` itself. `unclassified` is the family map's `*)` arm, so admitting it would silently grant concurrency to every test added afterwards, which is exactly the population with no proof. A new test still lands in `unclassified` and stays serial, and `tests/fm-test-run.test.sh` covers that split behaviorally. Each family passes two consecutive four-worker proofs with zero failures. On the production runner, `secondmate` goes 1233.1s to 453.4s, `session-bootstrap` 756.4s to 286.4s, and `standalone` 724.6s to 261.1s: 2.71x overall and 1713.2s recovered. The whole suite runs 177 scripts in 52.6 minutes of wall clock against 121 minutes of summed script time. * no-mistakes(document): Refresh concurrent validation and shard documentation * no-mistakes(ci): Fixed the real-Herdr focus-flash E2E race exposed by reclassification. Part C now starts its persistent child atomically via `pane run` and verifies stable child identity through Herdr’s public `process-info` interface, avoiding the racy send-text/send-keys sequence and platform-specific `ps` matching. Verified with bash syntax checking, ShellCheck, git diff checks, and the complete E2E test on Herdr 0.8.2 * feat: structure no-mistakes ask-user escalations (kunchenguid#3670) * feat(brief): structure no-mistakes ask-user escalation as event + snapshot file Crewmates escalating a no-mistakes ask-user gate now report one status event naming every finding id plus a snapshot file holding the gate's axi finding records verbatim (id, severity, file, line, description, authority), using the same shape even for a single finding. The status line never paraphrases. The format is defined once in fm-dod-lib.sh and rendered into both the scout and ship rule 6 in fm-brief.sh, so a promoted scout - whose rule 6 fm-promote.sh preserves unchanged - gets the identical contract as a freshly-spawned no-mistakes ship worker. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PpiWaDerbYavTLPPtEjQei * no-mistakes(review): Preserve ask-user escalation output contract * no-mistakes(review): Align escalation format test expectation * no-mistakes(review): Scope ask-user escalation instructions correctly * no-mistakes(review): Remove ask-user from generic decision rules --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> * fix(bin): require self-sufficient no-mistakes intent (kunchenguid#3671) * fix(bin): require a self-sufficient no-mistakes intent A no-mistakes worker's --intent is only as useful as the string it passes. PR kunchenguid#3604 shipped with an intent that was only "do 1, 2, 3, 7 from the report": the real contract lived in a private scout report and never reached --intent, so nobody holding that string plus the codebase could have derived the specification. This is pure instruction at the contract's one owner; no spawn-side or promotion-side check is added. - bin/fm-dod-lib.sh: the generated no-mistakes Definition of done now states that the --intent string must be self-sufficient (the string plus the codebase reconstructs roughly the same specification) and tells the worker to write the substance of any report, decision, or PR the captain's intent refers to into --intent rather than the pointer, while Firstmate build instructions and the worker's own decisions still stay out. The spawn-time overlay points back at that rule so its "supersedes" wording cannot cancel it, and the header's owner statement carries the rule. - AGENTS.md section 11 and bin/fm-brief.sh's header ask Firstmate to include the substance of referenced material when filling ## Captain's intent, and section 11 points at the owner of the rule. - tests/fm-brief.test.sh and tests/fm-task-delivery.test.sh assert the rendered brief and launch contract carry the rule. Claude-Session: https://claude.ai/code/session_01YMhEe42q7BAAoN6RxNuzim * no-mistakes(document): Replace incident-specific intent test commentary * fix: accelerate local Bearings snapshot composition (kunchenguid#3499) * Speed local fleet snapshot composition * no-mistakes(review): Stabilize task inventory during concurrent snapshot composition * no-mistakes(document): Document local snapshot observation concurrency * no-mistakes(ci): Fixed CI failures by making empty task manifests compatible with stock macOS Bash 3.2, snapshotting task metadata before concurrent observations to prevent generation drift, strengthening the behavioral race regression, and updating the stock-Bash Bearings test count to 45. Verified fleet snapshot tests (15), Bearings tests (45), workflow lint tests, project lint, Bash 3.2 parsing, and diff checks * no-mistakes(ci): Fixed the Linux CI failure caused by passing large backlog/task JSON through jq command-line arguments, which exceeded the per-argument size limit. Both inventory projections now stream large JSON inputs through stdin. Verified with fm-bearings-snapshot.test.sh (45 tests), fm-fleet-snapshot-view.test.sh (15 tests), Bash syntax, and git diff checks * no-mistakes(ci): Fixed concurrent task teardown during metadata capture: vanished metadata is now omitted while genuine copy failures remain fatal. Added a deterministic public Bearings regression test and updated CI’s expected test count. Verified with the full Bearings suite, workflow-lint suite, Bash syntax checks, and git diff checks * no-mistakes(ci): Fixed PR-caused CI and review issues: streamed large fleet JSON through jq stdin to avoid Linux argument limits, kept crew-state reads bound to captured metadata generations, and strengthened the behavioral race test. Bearings (46 tests), fleet snapshot (15 tests), crew-state, backend, lint, Bash syntax, and diff checks pass locally. Serial shard 5’s unrelated task-inbox segmentation fault appears infrastructural/flaky * no-mistakes(ci): Fixed endpoint-state generation crossing by validating captured spawn_gen before and after local endpoint probes, falling back to exact metadata identity for legacy tasks. Stale probe results now become unknown instead of false unhealthy state. Added a behavioral relaunch-race regression test. Verified the full Bearings snapshot suite, shellcheck, bash syntax, and git diff checks * fix(snapshot): keep live observations generation-coherent * no-mistakes(review): Keep secondmate observations generation-bound without copying reports * no-mistakes(document): Document generation-coherent snapshot observations * test(bearings): measure local read overlap instead of wall-clock budget The large-local-snapshot regression asserted that a whole snapshot composed in under five seconds. That bound measures how loaded the host is, not whether the per-task reads actually overlap, so it failed intermittently on a contended machine: one run in six on a box at load 16-20, landing exactly on the five second boundary. Time a serialized run and a concurrent run of the same workload instead and require the concurrent one to save at least two seconds. Both runs pay the same composition overhead, so the difference isolates the overlap this change delivers. Five one-second reads serialize into five seconds and overlap into about one, and re-serializing the reads collapses the saving to roughly zero, so the assertion still fails loudly if the concurrency regresses. Also bump the pinned Bearings test count to 48, since rebasing onto the current default branch picked up its captain-hold test. * no-mistakes(review): Restore JSON-derived decision flags * no-mistakes(review): Unify status-derived snapshot observations * no-mistakes(ci): Updated the stock macOS Bash CI check’s Bearings test count from 48 to 49. Verified the full Bearings suite passes and emits exactly 49 TAP successes; git diff checks pass * fix: prevent stale supervision wake loops (kunchenguid#3672) * fix(bin): stop the supervision branch's stale-ack and ghost-report loops Clean-slate implementation of the four authorized recommendations from the supervision-ghost-retrigger analysis (items 1, 2, 3, and 7), in their minimal form, superseding PR kunchenguid#3604: - fm_branch_report refuses a task the wake being handled never named. The extension fixes the reportable task set from the eligible rows before each prompt (signal and stale rows resolve to their tasks, a heartbeat allows any task with a live record, fleet is always allowed), so a report typed from memory about a task whose records teardown already removed is never stored or delivered. - An acknowledgement that consumes nothing says "nothing was acknowledged through N" and prints the exact --ack-through / --recovery-generation command for the current presented wake, instead of "re-run the drain", which re-fed the same stale acknowledgement in a loop. - bin/fm-guard.sh no longer tells the branch actor to drain queued wakes while it is handling them; it names the granted rows instead. - Teardown removes state/.<task>.branch-outcome-index for ordinary tasks and descendants; the index rebuild and the append-side index write both skip a task with neither a live record nor a status log, so the branch's report of a teardown it just performed is stored without recreating the index. No new locking, no spawn-generation binding, and no retired-task refusal: the branch can still report the outcome of a task it just tore down, and the teardown test now proves that path end to end. * fix(bin): narrow the branch report scope and guard silence to the minimal form Apply the four review decisions on the clean-slate branch: - A signal or stale prompt may report only the tasks its own rows resolve to; fleet is refused there too. A heartbeat review is not scoped by task at all, so the extension no longer tracks live task records and refuses nothing by task id during a fleet review. - The outcome-index rebuild no longer skips retired tasks; the append-side skip alone keeps a torn-down task's index from being recreated. - bin/fm-guard.sh keeps the queued-wakes warning silent for the branch actor instead of printing a replacement note. * no-mistakes(document): Align supervision docs with scoped wake handling * fix(bin): avoid fleet snapshot argument limits (kunchenguid#3677) * Fix fleet snapshot large JSON transport * no-mistakes(review): Captain: file-back fleet snapshot transport safely * no-mistakes(review): Captain: file-back parent summary aggregation * no-mistakes(ci): Rebased the PR's three commits onto f4d7875 and resolved the fleet snapshot conflict while preserving the base's task-observation lifecycle. Fixed Greptile's valid finding by recursively removing the private mktemp transport directory, so future transport files cannot cause cleanup to fail. Verified with tests/fm-home-summary-refresh.test.sh, bin/fm-lint.sh, git diff --check, and ancestry checks. All passed; the fix remains as an uncommitted worktree change for the outer executor * fix(bin): attribute active runs with unfetched pipeline heads (kunchenguid#3681) * fix(bin): recognize active pipeline fix rounds with unfetched run heads A no-mistakes fix round advances the run head beyond the submitted head, and the pipeline commits in its own checkout, so the task copy never receives the new commit object. fm-crew-state's strict head rule rejected the active row, the coarse runs-list scan skipped it and matched the older failed row at the submitted head, and an active validation read as failed (observed on model-routing-benchmark-hardening: active head ac61c64 vs task copy at fb47636d). fm_nm_runs_status_for_worktree in bin/fm-nm-run-lib.sh now owns runs-ledger attribution: the branch's newest row alone decides, and a newest row whose head cannot resolve locally is recognized only as a provable pipeline-owned continuation - active (running) and anchored by the immediately older row for the same branch having ended at exactly this worktree's HEAD. The reader keeps the axi TOON as full detail for that proven same-branch run. Unanchored, ancestor-anchored, and terminal unresolvable rows stay unattributed, so branch-name coincidence and other tasks' runs never match, and fm_nm_head_matches_worktree keeps its exact prior semantics for teardown (verified by the full teardown suite). Tests: reproduction regression for the unfetched active fix head (reads working via full run-step detail), coarse-path continuation when axi answers another branch, and negative controls for the unanchored active row and the unresolvable terminal row with the historical fallback preserved. Ported onto upstream/main f4d7875, where kunchenguid#3194 independently added the branch_sync custody exemption on the full axi-status path: both mechanisms now coexist, each owning one surface (TOON custody on the full path, the runs ledger on the coarse path). The port deletes the superseded coarse scan-and-skip (nm_runs_status_for_branch) and its now caller-less helpers (fm_nm_head_resolvable, nm_coarse_head_matches_worktree), renames the exemption comment's "the one exemption" phrasing now that a second complementary exemption exists, and points the stale FM_CREW_STATE_RUNS_LIMIT comment at fm_nm_runs_status_for_worktree (judge follow-up #1). The parent coarse-guard test's fixture is the ledger-anchored continuation shape, so its expectation flips to the fixed behavior (working via run-step, never the older failed row); a new mismatched-anchor coarse negative control preserves that guard's original no-anchor protection (pane answers, never the older row). * no-mistakes(document): Clarify pipeline attribution documentation * fix(bin): pre-register claude workspace trust at spawn time (kunchenguid#3663) * fix(bin): pre-register claude workspace trust for task worktrees A claude crewmate launched into a fresh task worktree met Claude Code's interactive workspace-trust dialog before it ever read its brief, and firstmate could not answer it: the key plane carries only Enter, Escape, and C-c with no arrow navigation, and the dialog's selection starts on "No, exit", so the documented Enter recipe ended the session instead of accepting it. Two workers wedged this way and were unblocked only by hand-seeding the trust store per path. --dangerously-skip-permissions does not cover that gate. `claude --help` records the dialog as skipped only in non-interactive mode, through -p or a non-TTY stdout, and a crewmate pane is interactive, so there is no launch flag to reach for. fm-spawn now pre-registers the worktree through bin/fm-claude-trust.sh in the existing claude branch, before the project settings that the same gate would otherwise block, and refuses the spawn when that write fails rather than launching a worker that would wedge. The scope test is the safety property and is structural rather than a path policy: the path must be a linked git worktree, sharing the spawning project's common dir, whose top level is exactly the resolved argument. Git is the ground truth, so the argument is never trusted on its own word, and a primary checkout, an unrelated repo, a worktree subdirectory, a plain directory, and a home directory are each refused rather than warned about or skipped. A treehouse or orca path prefix was deliberately avoided because treehouse's root is configurable, which would make a prefix both wrong and a new policy surface. One structural test covers both worktree providers. tests/fm-claude-trust.test.sh pins both halves, including a case where HOME is itself a valid linked worktree so the home guard is proven load-bearing rather than passing vacuously, plus the spawn-level proof that a claude spawn trusts its worktree and launches with the brief pointed at the same store. The adapter reference no longer tells a firstmate to press Enter on that dialog, and the shared trust reference now names every harness surface: which harnesses gate, which suppress at launch, which dodge the gate, which now pre-registers, and that a claude secondmate is excluded by design. The spawn fixture runs each spawn against a throwaway HOME so the suite cannot write the developer's real store, isolating through HOME rather than CLAUDE_CONFIG_DIR because the spawn forwards a set CLAUDE_CONFIG_DIR onto the launch command that launch-shape assertions read. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HNEN2GLnew27HFyfi4ms4v * fix(bin): create the staged trust store exclusively The staged store was written to a predictable pid-based path with a plain write, which follows a symlink. Where the Claude config directory is writable by another local account, that account could pre-create the path as a symlink and redirect the write into another file the launching user owns. The staged name now carries random bytes and is created with an exclusive "wx" open, so an existing path is refused outright instead of followed. The happy-path test also asserts no staged store survives the rename. The durability comment now states the residual window plainly: the readback proves the entry landed, not that it survives, because a vendor session that rewrites the whole store afterwards can still drop it and no lock closes that window when the writer is Claude itself. The worker then meets the dialog and stalls, which reaches firstmate as the ordinary stale wake rather than as silent success. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HNEN2GLnew27HFyfi4ms4v * no-mistakes(review): neutralise CDPATH in claude trust scope guard * no-mistakes(review): sandbox HOME in spawn tests, drop out-of-scope artifacts * no-mistakes(review): refuse unresolvable git dir, compact store, fix secondmate doc * no-mistakes(review): clear git env overrides, resolve symlinked store target * no-mistakes(review): degrade without node, fix Pi gate claim, record trust proof * no-mistakes(review): refuse without node, pin CLAUDE_CONFIG_DIR in spawn tests * no-mistakes(review): refuse relative config dir and concurrent store modification * no-mistakes(review): correct orca worktree claim, clean staged store on failure * no-mistakes(review): restore pretty-printed store, correct trust dialog docs * no-mistakes(review): arm trust gate before busy state to avoid orphans * no-mistakes(document): record claude trust pre-registration in its owner docs * no-mistakes(document): note orca limit for claude trust pre-registration * no-mistakes(ci): Fixed the Greptile P1 on bin/fm-spawn.sh by moving the Claude trust gate earlier rather than adding cleanup machinery. Diagnosis: Greptile reported that when Claude trust registration fails on tmux/Zellij/cmux/non-projected Herdr, the exit runs after the backend endpoint and /tmp/fm-<id> were created, and the abort trap cleans neither. The endpoint half is pre-existing, deliberate architecture — the two refusals immediately above the gate (the 60s `treehouse get` timeout at fm-spawn.sh:2550 and `validate_spawn_worktree` at :2487) also exit with the endpoint live and direct the operator with "inspect window $T"; spawn_abort_cleanup only reclaims orca endpoints (already covered via ORCA_ABORT_CLEANUP) and herdr projections. The temp-root half was genuinely introduced by this PR: the gate was placed beside the busy-state arm, ~30 lines after `mkdir -p "$TASK_TMP/gotmp"`, and fm-teardown can only find that root through `tasktmp=` in a meta record a refused spawn never publishes. Root-cause fix (smallest correct change, no new subsystem): - bin/fm-spawn.sh — moved the `claude*` trust gate from inside the busy-arm block up to the first point $WT is known, immediately after the `freshen_spawn_worktree_base` block and before TASK_TMP creation, the STATE setup, and the relaunch `clear_relaunch_harness_wiring` retirement. A refusal now leaves no temp root, no retired relaunch wiring, and no busy record; only the endpoint remains, in the same class as the two refusals just above it. - bin/fm-spawn.sh — the refusal message now ends with "inspect window $T", matching the existing convention so control/teardown can identify the endpoint. $T is set for every backend on the non-secondmate path. - bin/fm-spawn.sh:196 — header note corrected from "before any state is armed" to "before any per-task state exists". - tests/fm-claude-trust.test.sh — the existing refused-spawn test's own comment claimed "before any task state exists" but only asserted busy state. Renamed to test_refused_spawn_leaves_no_task_state and added an assertion that /tmp/fm-<id> is absent, with the task id suffixed by the test process pid so the assertion reads only this run's path (a stale /tmp/fm-refusedspawn from the fixed-id version was in fact present on this box). No assertions on implementation source bytes. Verification run locally: - The new assertion fails against the pre-fix bin/fm-spawn.sh ("not ok - a refused spawn stranded a temp root no teardown can find") and passes after — a real before/after regression proof. - tests/fm-claude-trust.test.sh: 20/20 ok. - tests/fm-backend.test.sh, fm-backend-orca, fm-control-relaunch, fm-spawn-dispatch-profile, fm-trace-context-spawn, fm-gotmp: all pass. - tests/fm-backlog-atomicity.test.sh: rc=0, 79 assertions ok. - bin/fm-lint.sh (repo's single lint owner, pinned ShellCheck 0.11.0 + actionlint 1.7.12): clean. - No /tmp/fm-refusedspawn* leftovers after the runs. Scope respected: no trust subsystem, no policy layer, no config surface, no endpoint-cleanup mechanism added; the change is an ordering move plus one error-message clause and the test that pins it. Adapter references and docs made no ordering claim, so none needed updating. Changes are left uncommitted in the worktree for the outer executor --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * fix: restart every live second mate after updates (kunchenguid#3690) * feat(update): restart every live second mate after a successful update /updatefirstmate only restarted a second mate when that pass advanced its AGENTS.md or .agents/skills. An already-current home was skipped entirely, a bin/-only advance was steered instead, and a remote host that could not report its instruction diff was downgraded to a re-read. A running agent also freezes its launch-time wiring - turn-end hooks, harness flags, per-harness feature switches - and none of that is derivable from a file diff, so an unchanged tracked surface is not evidence the agent is already on the current behavior. Restart is now unconditional on a successful update of that home. Every live second mate the pass leaves on the target commit is restarted, whether it advanced or was already there. The safety contract is unchanged: open records are persisted before the agent is replaced, nothing is forced, stashed, or discarded, a home the pass had to skip is not restarted at all, and a mate whose runtime cannot prove a restart keeps the honest re-read path and is never reported as reloaded. bin/fm-ff-lib.sh gains a settled-state hook that fires for a home left at the base whether it advanced or was already there, and never for a skipped one; the instruction-gated hook the session-start convergence sweep uses is untouched. Regressions: fm-update pins the already-current mate into the restart set and the unprovable one into the nudge set, and fm-secondmate-restart drives both real commands end to end - an already-current home is named, persisted, and genuinely replaced with its checkout untouched, while the unprovable one keeps its running agent. * no-mistakes(document): Document unconditional secondmate restarts * fix(bin): close pending-reply decisions via resolve-key (kunchenguid#3696) * fix(bin): close reserved pending-reply keys via fm-send --resolve-key fm-send wrote answered: notes that the reserved-key fold ignores, so operator closes exited 0 while OPEN DECISIONS kept the decision open. Speak the owning library's close vocabulary on that path, and refuse when a reserved close cannot take effect. * no-mistakes(review): Safely quote manual decision-close recovery commands * no-mistakes(review): Reject unclosable overlong decision keys before sending * no-mistakes(review): Remove contract suffix from open decisions hint * no-mistakes(document): Document resolve-key line-cap refusal * fix(bin): prevent false missed-reply escalations (kunchenguid#3697) * fix(bin): stop false missed-reply escalations for same-basename self-home answers A healthy secondmate that wrote corr= to its own state/<id>.status never matched the parent channel, so recovery confirmed and the record escalated as pending-reply-missed. Make the report helper resolve the parent channel itself, skip parent-replies.status as wrong-home, put a readable sighting path on the missed line, and restatement-copy only that same-basename self-home file onto the parent channel. * no-mistakes(review): Resolve late replies before recovery escalation * no-mistakes(review): Tighten reply routing and regression coverage * no-mistakes(review): Preserve reply paths and require explicit home * no-mistakes(review): Encode wrong-home paths before persistence * no-mistakes(document): Document corrected secondmate reply routing * no-mistakes(lint): Fix pending-reply ShellCheck warnings * feat: add verified Gemini crewmate runtime (kunchenguid#3695) * feat(harness): verify gemini as a crewmate runtime adapter Adds Gemini CLI as a fourth dispatch target alongside claude, codex, and grok, scoped to crewmate and scout work only. Every axis was proven against gemini-cli 0.58.0 rather than inferred; docs/verification/runtime-backends.md carries the dated evidence and names what stayed unverified. Busy state is semantic, not rendered: BeforeAgent opens a turn and AfterAgent and SessionEnd close it. AfterAgent also fires on a manual interrupt, so a cancelled turn closes its own record. Three findings shaped the wiring rather than a config line: - --skip-trust and GEMINI_CLI_TRUST_WORKSPACE=true are presented by the CLI as equivalents and are not. A controlled A/B showed --skip-trust leaves project configuration unloaded, so workspace skills never load. - The worktree's .gemini/settings.json is the PROJECT's committed settings file, unlike claude's settings.local.json. Firstmate's hooks therefore go to a firstmate-owned state/<id>.gemini-settings.json reached through GEMINI_CLI_SYSTEM_SETTINGS_PATH, which also works untrusted and merges with a project's own hooks instead of replacing them. - The shipped CLI is a node bundle whose live process reports comm as MainThread, so ancestry cannot see it. GEMINI_CLI=1 is load-bearing and is tested before an inherited CLAUDECODE, and pane liveness identifies gemini from the script argument through the new bin/fm-gemini-lib.sh. Gemini is refused for secondmates: it has no primary supervision protocol. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L * test: clear gemini's marker in launch and detection expectations Every non-gemini launch now clears GEMINI_CLI the way it already clears cursor's markers, so the two tests that pin the exact launch prefix are updated to match. The harness-detection tests that scrub foreign markers before probing ancestry scrub GEMINI_CLI too, so running the suite from inside a gemini session cannot produce a false verdict. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L * docs: classify the gemini harness reference The documentation inventory is the single classification owner for maintained prose surfaces, and every surface must appear in it exactly once. The new harness reference is agent-runtime, matching its siblings. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L * no-mistakes(review): Narrow Gemini ancestry detection * no-mistakes(review): Restrict Gemini hooks to canonical launches * no-mistakes(document): Document Gemini adapter support boundaries * no-mistakes(ci): Fixed Gemini process identity when interpreter or script paths contain whitespace. Tmux liveness now uses NUL-delimited /proc argv on Linux, with the existing flattened ps fallback elsewhere. Added a real-process regression test. Verified with the Gemini harness test suite, full fm-lint, ShellCheck, and git diff --check. The CI and Require no-mistakes runs were action_required/attestation outcomes rather than code failures --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(teardown): conclude parked runs advanced past task copy (kunchenguid#3704) * conclude parked runs the pipeline advanced past the task copy A no-mistakes fix round commits in the daemon's own gate-repo clone, so a run parked at a gate can carry a head whose object the task copy never received. Teardown's strict object-local identity rule then declined to conclude the run, and cleanup left it parked forever holding a fleet slot (observed 2026-09-03; the same masking condition PR 3681 fixed on the read path, now closing the teardown half its scope boundary deferred). task_status_is_own_parked_run now falls back - only when the reported head resolves to no local object - to the one shared runs-ledger attribution rule fm_nm_runs_status_for_worktree (bin/fm-nm-run-lib.sh), whose anchored continuation proof binds the branch's newest active row to this worktree's exact submitted head. Foreign branches, stale history, terminal rows, ancestor-only anchors, diverged newer rows, and ambiguous multi-row shapes all still refuse, and runs that are actively running, fixing, or in CI remain untouched: only the parked-at-a-gate determination ever reaches the abort. No sqlite access, no fetches into another task copy, no custody changes, no duplicated matching logic. * tighten the parked-run ledger fallback and pin both judge corrections The teardown ledger fallback now authorizes concluding this task's parked run only when the shared runs-ledger rule's proved answer is the explicitly active word (running): a terminal newest row - even anchored at exactly the worktree's head - is finished history and never an abort authorization. The read path may classify the same owner's answer; teardown's abort must never fire for a run that already ended. Two bounded pre-validation corrections from the implementation review: - a fetched-object counterfactual pins the strict-rule path: a pipeline fix head fetched into the task copy aborts through object-local identity alone, with an empty ledger and a proof the runs query never fired; - a negative fixture pins the tightened boundary: an unresolvable reported head with a terminal newest same-branch row anchored at the worktree head engages the ledger fallback and still refuses, so the refusal is the terminal-word boundary and not an earlier guard. * no-mistakes(review): Bind teardown ledger fallback to validated run heads * no-mistakes(review): Restore validated advanced-head ledger continuation * no-mistakes(review): Reject invalid ledger dates and terminal statuses * no-mistakes(document): Document teardown ledger scan limit * feat(bin): show requested vs effective model in Herdr agent view Track spawn-config requested_model separately from runtime-verified effective_model, probe Claude/Pi transcripts for exact API ids, push compact display metadata to Herdr, and preserve verified models across relaunch/compaction hooks without inferring aliases as truth. * fix(bin): keep re-probing effective model after first exact reading fm-model-sync.sh only probed for the runtime-verified effective model while it was still pending/UNKNOWN, so a session that later switched models (manual switch, provider fallback) kept displaying the first verified model forever and never appended a fallback-history entry. Probe unconditionally instead; fm_model_record_effective already no-ops when the probed value is unchanged, so this stays cheap. Addresses the Greptile P1 finding on PR kunchenguid#3705's fm-model-sync.sh. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BwfFjeYQcz9cZ3vEZohmpm * fix(bin): distinguish Cursor Grok, direct xAI Grok, and Anthropic Claude in the display Kapitänskorrektur: harness alone conflated Cursor-hosted Grok models (cursor-grok-4.6-*) and direct xAI Grok models (xai/grok-4.6) under one generic label, and displayed Anthropic Claude without naming the provider. Add fm_model_source_label, pattern-matched on the verified exact model id, so the compact display always reads Cursor · Grok, xAI · Grok, or Anthropic · Claude with the exact model id appended. Falls back to the existing harness label for every other model. No routing change: this only affects display strings. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BwfFjeYQcz9cZ3vEZohmpm * fix(bin): wire model-sync into the Pi extension's turn lifecycle fm-model-sync.sh was only invoked from Claude's SessionStart/ UserPromptSubmit/Stop hooks; the Pi harness's own extension (state/<id>.pi-ext.ts) never called it, so a Pi-hosted session (e.g. a pi/xai-grok crewmate) never refreshed its effective model after the first probe and Herdr kept showing the stale value with no fallback-history entry. Call fm-model-sync.sh from the same agent_start/turn_end boundaries Pi already uses for busy-state and the turn-end notification touch. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BwfFjeYQcz9cZ3vEZohmpm * fix(bin): serialize fm-model-sync.sh's meta read-probe-write Overlapping lifecycle events (Pi's agent_start/turn_end, Claude's SessionStart/UserPromptSubmit/Stop) can invoke fm-model-sync.sh concurrently for the same task. The unlocked read-probe-write let interleaved runs revert a newer effective model, mismatch its source, or duplicate a model-history entry. Serialize the critical section through the same per-task meta lock fm-spawn.sh already uses (fm_meta_lock_path + fm_lock_acquire_wait/fm_lock_release), released before every exit path. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BwfFjeYQcz9cZ3vEZohmpm --------- Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> Co-authored-by: Arthur Haro <38157909+haroarthur@users.noreply.github.com> Co-authored-by: Nicolas Payette <nicolas.payette@specira.ai> Co-authored-by: Jon Roosevelt <rooseveltadvisors@gmail.com> Co-authored-by: att430 <41454889+att430@users.noreply.github.com> Co-authored-by: Valentino-Sole <171032438+Valentino-Sole@users.noreply.github.com>
lytv
pushed a commit
to lytv/mymate
that referenced
this pull request
Sep 8, 2026
* fix: copy PR URLs from records or abstain, never assemble them Supervision reported a plausible but dead PR link three times because its prompt demanded a full https:// URL at a moment when only a PR number was observable, so the model assembled an owner/repository from memory, and the PR check then accepted that URL and wrote it into the task record, after which the model kept defending its own tool-endorsed guess over the worker's real link. Three changes close that chain without any live forge lookup, so private forges are treated exactly like public ones: - bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy or abstain" section requires a URL to be copied verbatim from a durable record (the done: PR <url> status line, pr= metadata, or the backlog note), forbids assembling owner, repository, host, or number from memory, and has the branch report only the identifier it actually holds when no record names the URL yet, leaving the PR check unarmed until the worker's ready line arrives. AGENTS.md section 7 and 9 carry the same copy-or-abstain rule for main in place of the bare full-URL mandate. - Worker briefs (bin/fm-brief.sh, ship and scout rules) require the full https:// URL wherever a PR is mentioned - status line, terminal, or summary - never a bare "PR 108", so the link is in view as early as the number is. - bin/fm-pr-check.sh refuses, offline and before any side effect, a URL that the task's own done lines contradict, printing both spellings; a log naming no URL still records the argument as before. fm_pr_status_ready_urls in bin/fm-pr-lib.sh owns reading those lines. The refusal also reaches bin/fm-pr-merge.sh, so nothing merges under a contradicted URL. Tests cover the offline refusal with zero side effects, the recorded spelling being accepted, markdown-wrapped and punctuated URLs, working lines not counting, the merge wrapper propagation, a self-hosted merge request with no forge call, the prompt carrying the rule, and the brief carrying the worker rule. * no-mistakes(review): Remove stale PR URL enforcement * no-mistakes(ci): Removed backlog notes as an accepted PR identity source. PR URLs may now be copied only from the task’s `done: PR <url>` status or canonical `pr=` metadata; otherwise supervision reports only the known identifier and leaves PR checking unarmed. Updated related guidance/docs and verified with branch-supervision tests, brief tests, ShellCheck, and `git diff --check`
BenWilcox8
pushed a commit
to BenWilcox8/firstmate
that referenced
this pull request
Sep 12, 2026
* fix: copy PR URLs from records or abstain, never assemble them Supervision reported a plausible but dead PR link three times because its prompt demanded a full https:// URL at a moment when only a PR number was observable, so the model assembled an owner/repository from memory, and the PR check then accepted that URL and wrote it into the task record, after which the model kept defending its own tool-endorsed guess over the worker's real link. Three changes close that chain without any live forge lookup, so private forges are treated exactly like public ones: - bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy or abstain" section requires a URL to be copied verbatim from a durable record (the done: PR <url> status line, pr= metadata, or the backlog note), forbids assembling owner, repository, host, or number from memory, and has the branch report only the identifier it actually holds when no record names the URL yet, leaving the PR check unarmed until the worker's ready line arrives. AGENTS.md section 7 and 9 carry the same copy-or-abstain rule for main in place of the bare full-URL mandate. - Worker briefs (bin/fm-brief.sh, ship and scout rules) require the full https:// URL wherever a PR is mentioned - status line, terminal, or summary - never a bare "PR 108", so the link is in view as early as the number is. - bin/fm-pr-check.sh refuses, offline and before any side effect, a URL that the task's own done lines contradict, printing both spellings; a log naming no URL still records the argument as before. fm_pr_status_ready_urls in bin/fm-pr-lib.sh owns reading those lines. The refusal also reaches bin/fm-pr-merge.sh, so nothing merges under a contradicted URL. Tests cover the offline refusal with zero side effects, the recorded spelling being accepted, markdown-wrapped and punctuated URLs, working lines not counting, the merge wrapper propagation, a self-hosted merge request with no forge call, the prompt carrying the rule, and the brief carrying the worker rule. * no-mistakes(review): Remove stale PR URL enforcement * no-mistakes(ci): Removed backlog notes as an accepted PR identity source. PR URLs may now be copied only from the task’s `done: PR <url>` status or canonical `pr=` metadata; otherwise supervision reports only the known identifier and leaves PR checking unarmed. Updated related guidance/docs and verified with branch-supervision tests, brief tests, ShellCheck, and `git diff --check`
limitall
added a commit
to limitall/firstmate
that referenced
this pull request
Sep 16, 2026
The contract asked for a PR's complete https://... link and never said where that link comes from, which is the prompt shape upstream found makes a model assemble a plausible-looking dead URL from an owner, a repository and a half-remembered number (kunchenguid/firstmate#3648). AGENTS.md section 9 is now the one owner of the rule: copy the URL verbatim from the worker's `done: PR <url>` line or the task's recorded `pr=`, never assemble one, and when no record holds it say only the identifier actually held. Section 7's landing paragraph points at that rule instead of restating "full URL, never a bare #number". The bearings skill stated the same sentence three times; it now states it once, extended with the source and the abstain path, and its gather step names the ready line as where a URL on this port actually lives. The mechanical half is Get-FmTaskDeliveredPrUrl. Nothing here writes `pr=`, so the field teardown read was always empty and its two consumers degraded silently: the landed-work test fell through to a `gh-axi pr list --head` lookup, which finds whatever PR points at the branch rather than the one the worker said it delivered. The extractor takes the recorded field first, then the LAST line of the status file matching the exact ready-signal shape `^done: PR (<url>)( checks green)?$`, then nothing. Anchored at both ends, using upstream #4148's pattern, so a PR merely mentioned in a status line is never claimed as the delivery; last match wins because the log is append-only and a re-push reports a second PR; a scout returns nothing at all, because it investigates and never delivers. One value feeds both the landed check and the backlog reminder, so both get a copied URL or none. Returning nothing is a real answer: the reminder keeps its visible PR_URL placeholder rather than letting anything build a link out of a repo name and a number.
jorguez96
added a commit
to jorguez96/firstmate
that referenced
this pull request
Sep 17, 2026
…s resolved (#14) * fix(bin): defer inactive reconciliation during startup (#3480) * Defer inactive startup reconciliation * no-mistakes(review): Queue deferred inactive reconciliation diagnostics durably * no-mistakes(review): Require worker phases to cover startup requests * no-mistakes(review): Make diagnostic wakes safely acknowledgeable * no-mistakes(document): Document deferred startup phase coverage * fix(bin): bound wake drain presentation lock waits (#3475) * fix: bound status presentation lock waits * no-mistakes(review): Distinguish malformed presentation locks from live contention * no-mistakes(review): Bound no-ack drain queue lock acquisition * no-mistakes(document): Document bounded presentation-lock drain behavior * no-mistakes(lint): Annotate bounded lock output global * no-mistakes(ci): Added deterministic regression coverage for successful bounded-lock acquisition after live contention, verifying helper-to-caller PID ownership handoff and caller release. Verified with bash syntax checks, git diff checks, and the full fm-wake-queue test suite * fix(bin): retire public follow-ups in remote homes (#3479) * fix(relay): close a public loop whose work lives in a remote secondmate home A public-followup loop bound to a REMOTE secondmate could never be closed. `clear_public_followup_link` (bin/fm-public-followup.sh:701) required an absolute recorded `work_home_path` for a `secondmate:*` work home, but a remote route has no local path on this machine, so registration records that field empty (bin/fm-public-followup.sh:291). Every close ran that clear first, so `retire` died with "could not clear the legacy X link ... retained for reconciliation" forever, and `deliver` posted the public reply and then stranded the loop at `posted`. `--force` never covered that step. The clear now goes to the remote home over that route's SSH transport, running `fm-x-followup.sh --clear <work-id>` through `bin/fm-on.sh`. The route is decided from `data/secondmates.md` before any local path is consulted, so a same-named local directory can never stand in for a remote home, and registrations already on disk retire without needing a new field. `fm-on.sh` passes ssh's status through, so 255 stays the established "delivered but completion unknown" result this codebase already reconciles: the close is refused, the registration and the remote link are left exactly as they were, and the message names the unknown completion instead of claiming a definite failure. Local secondmate and `main` work homes are untouched, and `--force` still governs only the unresolved-obligation refusal. Three regression cases drive a remote route end to end, faking only the ssh binary at the FM_SSH_BIN seam and then running the real remote entrypoint against a local checkout, so the clear that must reach the remote home actually happens there. * no-mistakes(review): Guard remote link clears by request identity * no-mistakes(review): Fail guarded clears on unreadable remote state * no-mistakes(review): Reject guarded clears on non-writable remote state * no-mistakes(review): Allow no-link retirement in non-writable remote state * no-mistakes(document): Correct public-followup verification guarantee count * no-mistakes(ci): Fixed the guarded link-clear race by ensuring absence is decided under the metadata lock whenever publication is possible. Added a behavioral concurrency regression test. Verified with fm-x-mode and fm-public-followup suites, Bash syntax checks, diff checks, and bin/fm-lint.sh * no-mistakes(ci): Fixed the guarded link-clear race by refusing an unlocked absence decision when a publisher already owns the metadata lock in a non-writable directory. Added a behavioral concurrency regression test. Verified with fm-x-mode, fm-public-followup, syntax/diff checks, and fm-lint * no-mistakes(ci): Fixed the guarded-clear race by refusing all guarded clears when the metadata parent is non-writable, including apparent link absence. Added a behavioral regression with a publisher waiting to create the lock, updated remote-retirement expectations and verification docs. Passed fm-x-mode, fm-public-followup, fm-lint, documentation audience, Bash syntax, and diff checks * fix(relay): bound the guarded remote link clear so it refuses instead of hanging The guarded clear checks that the remote state directory is writable before taking the metadata lock, but that check cannot close the window: the parent can turn non-writable between the check and lock creation, and a lock held by a live holder is indistinguishable from that at the acquire. `fm_lock_acquire_wait` is an unbounded `while ! try; do sleep 0.1; done`, so either case retried forever and `deliver` or `retire` wedged with nothing reported, instead of returning the retained-for-reconciliation refusal the guard exists to produce. This path runs unattended over the secondmate transport, where a wedge is worse than either outcome the guard defines. The guarded clear now acquires through `fm_lock_acquire_wait_bounded` (FMX_LINK_CLEAR_LOCK_TIMEOUT, default 10 seconds) and refuses on timeout through the existing failure path. Unguarded local callers keep the ordinary unbounded wait, so local behavior is unchanged. The bounded primitive's header no longer claims presentation-only scope, since this is a second authorized caller; nothing else in the shared lock infrastructure changed. The regression holds the metadata lock with a genuinely live process while leaving the state directory writable, so the refusal can only come from the bound and never from the writability precondition. Against the unbounded wait it does not terminate at all; with the bound it refuses, retains the registration, writes no receipt, and leaves the remote link untouched. * no-mistakes(review): Harden lock-timeout regression with independent deadline * no-mistakes(review): Restore no-op guarded clears on read-only state * no-mistakes(document): Clarify remote public-followup cleanup contract * fix(bin): support process events under symlinked homes (#3484) * fix(bin): resolve process-event state roots before validating them The process-event module validated the caller's spelling of a home's state root instead of the directory it operates on: it required the supplied path to equal its own lexical normalization, which rejects any path reached through a symlinked ancestor. On macOS both /tmp and $TMPDIR are symlinks, so an operator home under either could never claim a source. Reconcile still reported the runner started, while the detached runner died writing "cannot claim source" to the discarded stderr, and the source silently never fired. Resolve the state root to its physical directory once, then apply the existing private-directory validation to that resolved directory and derive every path, recorded claim identity, and later confinement check from it. This keeps the confinement contract for the directory actually operated on rather than only for callers that already spelled it physically, and removes the window where an ancestor symlink could be repointed between check and use. Homes already spelled physically behave identically. This was the single cause of both deterministic macOS failures in tests/fm-procevent.test.sh ("reconcile never claimed the registered source") and tests/fm-procevent-when.test.sh ("the winning concurrent arm did not produce an outcome"). The new case pins the behavior with an explicit symlinked-ancestor home, so it fails without the fix on any platform rather than only where the temp root happens to be a symlink. * fix(bin): pin the external capture staging boundary to its physical path The extension capture path pinned its registry staging boundary by comparing `pwd -P` against the caller-spelled registry directory, so a home reached through a symlinked ancestor still refused to start an extension-backed source after the state root itself resolved correctly. That left such a home half working: built-in sources ran while external ones failed. The staging preparer now prints the physical registry directory it validated, matching the inbox and reservation preparers beside it, and the start path pins on that returned path. The new end-to-end case drives the shipped file-signal package from a symlinked home spelling. * no-mistakes(review): Propagate canonical process-event state roots * no-mistakes(review): Propagate canonical state to process-event adapters * no-mistakes(document): Document physical process-event state roots * fix(pi): deliver captain outcomes as deterministic transcript entries (#3312) * fix(pi): persist captain outcomes visibly * no-mistakes(review): Recover captain outcomes after cold-start lock acquisition * no-mistakes(document): Document cold-start captain-outcome recovery * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes(review): Prove immediate Pi captain-outcome transcript delivery * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * fix(pi): process captain outcomes through a sequence-keyed turn PR #3312 made every captain-facing supervision outcome a durable, exact-once visible transcript entry with the read cursor advancing only after that entry exists. That is the display half of the delivery contract. Left alone it turns a probabilistic silent loss into a deterministic one: the captain sees an anchor line, and firstmate never acts, because nothing opens a turn and nothing records whether main ever processed the outcome. The 2026-08-31 timeline showed the two shapes this must survive on the previous hidden-turn path: seven delivered decision outcomes each answered by an empty assistant message (cursor advanced, no retry, unanswered for close to three hours), and two answered by an unrelated prior reply. Both happened because delivery advanced the cursor at enqueue and accepted whatever the next assistant message was. Add the processing half on top of the persistence half: - bin/fm-branch-outcome.sh keeps a processed marker separate from the read cursor (`unprocessed`, `mark-processed --through`, `processed-init`). It only advances through an explicit sequence-bound acknowledgement, never past the read cursor and never backwards; an absent marker reads as zero and `processed-init` migrates delivered history once so an upgraded home is not re-presented its past. - After the visible entry for a captain outcome exists, the extension hands every still-unprocessed captain row to main as one hidden, typed `fm-branch-process` request listing each `[seq N] task: summary`, opening exactly one main turn. Main closes it only by calling the new `fm_branch_processed` tool with the highest sequence listed. An unrelated, empty, or paraphrased answer leaves the sequence open, and the same request is presented again at the end of the next main run and at session start. The first two presentations of a sequence set open a turn of their own; after that the request rides the captain's next prompt so an ignored request cannot loop, and a session replacement resets that budget. Routine outcomes stay turn-free. - The regressions cover exactly those incident shapes against the real store scripts: an empty answer and an unrelated prior answer neither advance the marker nor stop re-presentation, the acknowledgement is refused beyond the read cursor and outside lock ownership, a partial acknowledgement keeps the newer sequence open, and #3312's own assertions now forbid an unkeyed turn rather than any turn. The store suite pins the marker's bounds and the migration; the real-SDK guard for appendEntry persistence and model exclusion is unchanged. Docs move the protocol from "no model turn" to "one sequence-keyed processing turn closed only by its acknowledgement", and the verification record carries the dated run against Pi 0.84.4. * no-mistakes(review): Harden outcome listing and sequence-bound acknowledgements * no-mistakes(review): Harden outcome state validation and request pacing * no-mistakes(review): Reject unsafe sidecars and unterminated outcome stores * no-mistakes(review): Validate canonical mark-read cursor state * no-mistakes(review): Guard cursor advancement against corrupt processed state * no-mistakes(review): Bind acknowledgements to active processing requests * no-mistakes(review): Reset pacing when processing sequence membership changes * no-mistakes(review): Enforce silent outcome invariants at storage boundary * no-mistakes(document): Document hardened captain outcome processing contracts --------- Co-authored-by: kunchenguid <kun@kunchenguid.com> * feat: add bounded concurrent Bearings ledger collection (#3481) * feat: bound Bearings remote ledger collection * no-mistakes(review): Clarify default remote-ledger collection behavior * no-mistakes(review): Detach reconcile delivery from watcher loop * no-mistakes(review): Enforce bounded snapshot and request captures * no-mistakes(review): Bound legacy summary capture before parsing * no-mistakes(review): Bound primary remote ledger captures * no-mistakes(document): Correct snapshot and reconcile documentation * no-mistakes(lint): Fix ShellCheck quoting in bounded collector * no-mistakes(ci): Fixed all three CI failures: updated the macOS Bearings assertion to 44 tests, made the home-summary test deterministic and aligned with default ledger consumption, and increased the asynchronous reconcile retirement wait for loaded CI. Verified both focused suites, all 44 Bearings tests, ShellCheck, actionlint, Bash parsing, and git diff checks * test: await reconcile request retirement * no-mistakes(review): Avoid empty reconcile queue process churn * no-mistakes(review): Read ledger summaries from immutable snapshots * no-mistakes(review): Reject multi-document home ledger streams * no-mistakes(review): Coalesce durable reconcile requests per target * no-mistakes(review): Unify reconcile keys and reject snapshot streams * no-mistakes(review): Key reconcile requests by stable target ID * no-mistakes(document): Document per-target reconcile request coalescing * no-mistakes(lint): Remove unused snapshot summary file variable * no-mistakes(ci): Adjusted the concurrent collector regression’s end-to-end timing ceiling to account for stock macOS process/jq overhead outside the three-second remote collection budget, while remaining below the 15-second serial-read floor. Verified with stock /bin/bash 3.2: all 44 Bearings tests pass; bash syntax and git diff checks pass * no-mistakes(ci): Fixed legacy summary validation to require exactly one top-level JSON document and added behavioral regression coverage. Stabilized CI by conditionally waiting longer for durable reconcile delivery and synchronously stopping the fm-on worker tree before fixture cleanup. Removed a redundant flaky healthy-path timing assertion; the wedged-reader test still proves concurrent bounded collection. Verified fm-bearings-snapshot, fm-secondmate-reconcile, and fm-on tests, plus project ShellCheck, bash syntax, and git diff checks * ci: rebalance portable serial test shards (#3489) * fix(ci): rebalance the portable serial shards on measured durations The "Behavior portable serial 3" shard ran 17-20 minutes against its 20-minute job cap and intermittently timed out seconds after a passing test, on branches and on main alike. Shards are packed longest-processing-time from per-script duration hints, and those hints were last measured on 2026-08-21 at 116 scripts. The lane has since grown to 139 scripts and from ~42 to ~63 minutes: 17 scripts had no hint at all and fell back to the 20 s default, and several existing hints were low by 2-5x (fm-watch-triage 142 s hinted vs 263 s measured, fm-public-followup 36 s vs 197 s). The partition therefore looked perfectly balanced in hint space, 734.6 s per shard, while really running 11.5, 13.6, 18.8 and 16.5 minutes. Script-count balance, which is what the tests asserted, stayed normal throughout and hid it. Refresh the hints from the timing artifacts of three green runs, taking the slowest measurement of each script so the balance holds on a slow runner, and split the lane across five shards instead of four. Replayed against those runs' real per-script durations the worst shard is now 12.54 minutes, 63% of the unchanged 20-minute cap, and the serial lane's wall clock drops from ~20 to ~12.5 minutes. Bound the drift that caused this rather than relying on the hints being refreshed by hand: the coverage guard now reports the unmeasured share as serial_unhinted= and refuses past PORTABLE_SERIAL_MAX_UNHINTED_PERCENT, which leaves room for newly added tests while making a stale table fail the guard instead of silently pushing one shard into its cap. No test changes what it asserts and no test stops running; only the partition across shards changes. * no-mistakes(document): Clarify conservative shard timing aggregate * fix(pi): fall back on incomplete supervision branch prompts (#3491) * fix(pi): fall back after settled branch errors * no-mistakes(review): Detect provider errors across prompt compaction * no-mistakes(review): Preserve in-flight branch state across selection changes * fix(pi): re-probe supervision branch after cooldown (#3497) * fix(pi): recover supervision branch after cooldown * no-mistakes(review): Defer branch recovery until prompt settlement * no-mistakes(document): Clarify supervision cooldown recovery contract * fix(bin): remove legacy remote snapshot reads (#3501) * refactor: remove legacy remote summary reads * no-mistakes(document): Document ledger-only snapshot reads * no-mistakes(ci): Fixed the snapshot test fixture so ledger refreshes use the same fake executable PATH as the snapshot consumer. This preserves observable endpoint freshness after removing legacy summary computation. Verified stock Bash parsing and all 44 Bearings tests pass under /bin/bash; git diff checks pass * no-mistakes(ci): Fixed the CI-only snapshot fixture failure by ensuring the bounded-ledger refresh uses its fake tmux backend. This removes host tmux availability as a source of nondeterminism. Verified all 44 Bearings tests pass, Bash syntax passes, and git diff checks are clean * no-mistakes(ci): Fixed CI nondeterminism in the Bearings fixture: all local ledger refreshes now use the fixture’s fake tmux backend when available, instead of depending on host tmux state. Verified stock /bin/bash syntax, git diff checks, and all 44 Bearings tests with a deliberately failing host tmux * fix(pi): preserve watcher continuity across session replacement (#3498) * fix(pi): rearm watcher after session replacement * no-mistakes(review): Queue actionable closes across Pi session replacement * no-mistakes(review): Stop replacement arm when handoff persistence fails * no-mistakes(review): Preserve actionable wakes through branch and late child races * no-mistakes(review): Surface late handoff failures without crashing Pi * no-mistakes(review): Coordinate replacement delivery settlement and unique handoff tokens * no-mistakes(review): Retry stale deliveries and release settled claims * no-mistakes(review): Distinguish branch settlement and retry handoff cleanup * no-mistakes(review): Deduplicate persistent handoff cleanup alerts * no-mistakes(review): Acknowledge watcher follow-ups only when consumed * no-mistakes(review): Persist idle follow-ups until agent consumption * no-mistakes(review): Preserve pending outcomes when handoff persistence fails * no-mistakes(review): Arm replacement before awaiting prior delivery settlement * no-mistakes(review): Adopt pending handoffs after lock reclamation * no-mistakes(review): Prevent stale generations from adopting replacement handoffs * no-mistakes(review): Scope replacement handoffs by watcher state * no-mistakes(document): Clarify replacement handoff documentation * no-mistakes(ci): Fixed the failing branch-extension tests to model the new settlement-promise contract. Failure cases now assert that delivery ownership returns to the watcher instead of expecting direct extension fallback. Verified the updated branch suite, Pi watcher suite, shell syntax, and diff checks * no-mistakes(review): Update branch settlement tests and preserve chunked outcomes * no-mistakes(document): Document watcher-owned replacement handoffs * no-mistakes(document): Verify replacement handoff documentation * test(pi): cover watcher-owned branch fallback * no-mistakes(document): Refresh watcher-owned fallback documentation * fix(bin): resurface task statuses missed by wake handling (#3495) * fix(bin): resurface terminal statuses lost after branch handling * test(watch): canonicalize process-event fixture homes * no-mistakes(review): Index branch outcomes by causal status position * no-mistakes(review): Recover outcome indexes and deduplicate resurfaced statuses * no-mistakes(review): Handle legacy ambiguity and oversized status diagnostics * no-mistakes(review): Keep unclassifiable oversized statuses silent * no-mistakes(document): Document lost-wake outcome backstop * no-mistakes(document): Update outcome backstop documentation * no-mistakes(ci): Fixed CI regressions in wake-drain: parseable reserved-key decisions can no longer bypass the durable decision-fold guard, and status output is prepared and receipt-committed before presentation to prevent repeated one-shot outcomes after later failures. Added a behavioral regression for receipt commit failure and retry. Targeted backstop, correlation-token, decision-cursor, open-decision, unread-status, syntax, and diff checks pass locally. Shard-4 failures appeared unrelated/flaky; the network-parallel test passed locally * no-mistakes(ci): Fixed the Greptile P1 data-loss issue by committing presentation receipts only after prepared output reaches stdout. Added behavioral coverage proving output failure leaves the backstop retryable and receipt failure may duplicate but never lose a presentation. Relevant wake-drain suites and syntax/diff checks pass. The shard-4 Pi extension failure is unrelated to this PR and did not warrant changes * no-mistakes(ci): Stabilized tests/fm-bootstrap-network-parallel.test.sh by replacing scheduler-sensitive equal-sleep timing with bounded synchronization between mocked fetch and remote probes. This preserves detection of real serialization while avoiding false failures under CI load. Verified with five consecutive test runs, bash syntax validation, ShellCheck, and git diff checks. The separate Pi stock-rendering failure reproduces locally but is unrelated environment/version drift * no-mistakes(ci): Fixed Behavior portable serial 4 by adding fm-classify-lib.sh and fm-timeout-lib.sh to the broken-root Pi test fixture; fm-branch-outcome.sh now depends on them. Verified the full Pi branch-extension suite with real-Pi checks skipped, the wake-drain outcome-backstop suite, Bash syntax, and git diff checks. Greptile findings are already addressed at HEAD; the no-mistakes attestation failure is external head-SHA state * fix(bin): collect follow-up results from remote work homes (#3503) * fix(bin): deliver typed terminal results from remote work homes A public commitment whose work is bound to a REMOTE secondmate home could never receive its typed terminal result. `fm-public-followup.sh brief` printed an emit command carrying this home's own absolute path and this checkout's own script path, neither of which exists on the machine the worker runs on, so the worker had nothing it could write to that the owning home would ever read - and `consume` kept finding nothing while the promise stayed open. The brief is now route-aware: for a remote work home it prints that route's own code root and home with `--stage-in`, so the typed event is staged in the home where the work actually runs, and the closing paragraph names the owning home as the one on the other machine instead of pointing at the path above it. The owning home collects those staged results over the same SSH route it reaches that secondmate on, because the transport only runs outbound: `consume` pulls them into its own inbox and reconciles them exactly as it reconciles a local report. Collection is non-destructive until the result is durably held, so a dropped connection cannot lose a terminal result, and a route that could not be reached is named in `consume`'s output with the promise left open rather than reported as an empty inbox. A local work home is untouched: the brief still prints `--home` with this home and this checkout's script, and the event still lands directly in this home's typed terminal-result inbox. This is the emit-side counterpart of the retire/clear fix in #3479 and reuses the remote-route resolution that landed with it. Reconciling a loop bound to a remote route now reaches that route, so the existing remote cases drive `consume` through the same faked transport their other steps already use. * no-mistakes(review): Fail loudly on unresolved routes and invalid staging homes * no-mistakes(review): Fail collection when remote outbox is unreadable * no-mistakes(review): Surface reassigned remote routes during empty collection * no-mistakes(review): Fail remote collection on invalid registrations * no-mistakes(review): Reject unsafe registration entries during remote collection * no-mistakes(review): Restore healthy empty remote collection behavior * no-mistakes(review): Skip remote collection for delivered registrations * no-mistakes(review): Skip delivered registrations before route validation * no-mistakes(document): Document remote follow-up collection semantics * fix(bin): exclude secondmates from home-summary validity (#3504) * fix(bin): exclude secondmates from home-summary child inventory kind=secondmate meta records never have backlog rows, so counting them in unowned_children or terminal_in_flight made a clean main home look invalid once earlier ledger checks passed. * no-mistakes(review): Cover terminal secondmate in-flight exclusion * no-mistakes(ci): Updated the stock macOS Bash CI snapshot expectation from 15 to 16 tests. Verified all 16 snapshot/fleet-view tests pass under Bash 3.2.57 and `git diff --check` succeeds * fix(bin): self-heal outcome indexes on first drain (#3509) * fix(bin): self-heal status-outcome indexes on every drain Missing ready markers were skipping the lost-wake backstop on non-Pi homes because only the Pi branch ran processed-init. Drain now rebuilds those indexes under the outcome lock and fails closed only on a real store fault. * no-mistakes(review): Guard held-lock initialization and fail marker writes * no-mistakes(document): Document cross-harness outcome-index self-healing * fix(bearings): keep active children underway during captain holds (#3505) * fix(bearings): keep active children underway beside a captain hold Project each readable home's active children into Underway independently of the home-level captain-decision classification so a hold no longer hides live work. * no-mistakes(review): Preserve Underway repos and disclose child truncation * no-mistakes(review): Fall back to task project for Underway repos * no-mistakes(ci): Updated the stock macOS Bash CI assertion from 44 to 45 Bearings tests, matching the newly added behavioral regression. Verified all 45 tests pass under /bin/bash, Bash syntax checks pass, and git diff validation is clean * fix(pi): settle watcher delivery on Pi accepting the follow-up (#3513) * fix(pi): settle watcher delivery on Pi accepting the follow-up A follow-up queued while main is streaming joins the running run without ever raising before_agent_start, so waiting on that event before clearing the successor pipeline (#3498) stalled every later actionable close: no successor started, no wake was delivered or offered to the branch, and the turn-end guard woke main to re-arm by hand after every close. The pipeline now settles once Pi accepts the follow-up. Consumption is observed at before_agent_start for an idle main and at the user message_start for a streaming main, and decides only what a replacement session (/new, /resume, /fork, reload) replays. An exhausted restoration delivers its typed failure without launching an arm past the retry bound, which the stall had hidden. The replacement-coordinator map is typed so the strict no-emit typecheck passes again. Tests: the doubles no longer raise before_agent_start for a streaming send, a portable regression drives two actionable closes while main streams and proves the successor chain plus consumption-scoped replay, and a credential-free real-SDK probe pins Pi's event contract for both the streaming and the idle follow-up. Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a * fix(pi): retry a verified successor that fails during wake delivery A verified successor can exit while the wake it was started for is still being delivered, most plausibly during a branch turn that holds the settlement for minutes. Its failure close arrived while the pipeline's single-flight guard was set, so the close handler skipped the retry, and the pipeline's end no longer launched an arm, which left the live generation with no watcher and no retry timer. The close handler now records that failure when the child had reported readiness and was not retired by the restoration itself, and the pipeline runs the ordinary bounded, lock-checked retry for it once the delivery settles. A restoration started for a later pending supersedes it, and an exhausted restoration still hands repair to main without a further arm. The regression holds a branch settlement open while the verified successor exits with a failure and proves one retry watcher starts after the settlement releases, none while it is held. Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a * fix(bin): bound repeat stale wakes for parked workers (#3532) * fix(bin): bound repeat stale wakes for a parked but live worker A worker parked on a declared wait - `paused:` for an external or pipeline wait, or a verified `captain-held` transfer - kept waking firstmate far inside FM_PAUSE_RESURFACE_SECS. Observed as five consecutive alarms on one captain-held worker and dozens across a day on a pipeline wait, and reported upstream as four wakes in 75 minutes against a 3600s window. pause_state_class deliberately answers `none` for a still-live agent even under a declared wait, so a worker genuinely waiting on a decision is never silenced. That classification is correct and is left alone; it routes every parked but live worker through surface_nonterminal_stale on first sight of each distinct stale hash, and an idle parked pane still churns its hash on a clock or a token counter without changing what is being waited on. Two places let that churn re-alarm: - surface_nonterminal_stale queued the wake BEFORE consulting whether a wait was declared, then wrote `.paused-resurfaced-<key>` - the very throttle that should have suppressed it. The throttle was never read on this path and was advanced by the wake it should have prevented. - The hash-change path cleared that throttle through clear_pause_tracking whenever the classification came back `none`, so each tick also bought the same declared wait a fresh window. Fixing only the first site changes nothing. Read the throttle before anything is queued and advance it only on a wake that really fires, and on the hash-change path reset only the per-hash bookkeeping while the declaration still stands, via a clear_stale_hash_tracking split so neither half of clear_pause_tracking is duplicated. The throttle is keyed to the declaration, not to the pane. First sight still wakes, so an inconclusive state is still inspected, and the window's end still re-surfaces once, so a forgotten wait cannot rot invisibly - noise traded for a bounded cadence, never for silence. The wake identity stays the plain `stale: <win>` the away-mode handoff depends on. Tests cover both observed forms and were confirmed to fail against three deliberate breaks: each site reverted on its own, and a re-surface that never fires again. * fix(document): Clarify declared-wait wake cadence documentation * fix(ci): Captain, fixed the stale-throttle inheritance: cadence markers now bind to the current wait declaration, so replacement paused and captain-held waits each emit their first plain `stale:` wake. Added behavioral coverage for both forms. Bite proof failed as expected when identity matching was removed, then passed after restoration. Full watcher triage suite, `bin/fm-lint.sh`, syntax checks, and diff checks pass. Changes remain uncommitted for the outer executor * fix(ci): Captain, fixed the confirmed Greptile finding. `resurface_absorbed` now applies a throttle only when its stored declaration scope matches the current wait, so replacement `paused:` and `captain-held` waits surface immediately without changing classification. Added executable coverage for both absorbed forms. Bite proof failed before the fix at the intended assertion; afterward the full watcher triage suite, `bin/fm-lint.sh`, shell syntax checks, and `git diff --check` passed * fix(bin): accept the away-mode daemon as the turn-end supervision owner (#3567) * fix(turnend): accept the away-mode daemon as the supervision owner While state/.afk exists the away-mode daemon owns supervision and runs bin/fm-watch.sh one-shot: the watcher exits on every wake and the daemon starts its replacement. The turn-end guard tested for a live watcher process holding the watch lock at that instant, so a turn boundary that landed in the hand-off blocked with "TURN WOULD END BLIND" while supervision was completely healthy, costing a full handling turn each time. Reproduced with the real daemon wrapping the real watcher and the real guard sampling the same home: 6 of 40 samples blocked, every one of them with the daemon alive and the beacon 2-3 seconds old, and a new watcher pid on each cycle. After the fix the same reproduction blocks 0 of 40, and killing the daemon and its watcher (away mode still on, beacon still fresh) blocks again. The guard now accepts a live, identity-matched daemon holding this home as proof of supervision while away mode is active. The identity match is the same discipline the watcher lock uses, so a recycled pid or a lock left by a killed daemon proves nothing. The fresh-beacon half of the predicate is unchanged: a daemon that stops restarting its watcher still blocks once the beacon passes grace, a home with no supervisor blocks exactly as before, and with away mode off the strict watcher predicate is untouched. The predicate reads only durable state, so it behaves identically for every primary harness and runtime backend. * no-mistakes(document): clarify away-mode daemon supervision proof and test coverage * no-mistakes(document): generalize stale turn-end predicate summary in architecture.md * fix(backlog): omit --file from row probes for non-markdown backends (#3582) * fix(backlog): omit markdown file for beads probes * no-mistakes(document): Narrow backlog addressing doc to mutations for backend-aware probes * no-mistakes(ci): Fixed the Greptile P2 review comment (the only failing check) on tests/fm-backlog-atomicity.test.sh. The comment correctly noted that an exported TASKS_AXI_BACKEND environment variable would inherit into the spawned scripts and, because fm_tasks_axi_backend gives it top precedence, override each test case's .tasks.toml backend fixture — making the backend-specific argv assertions fail for environmental reasons. Fix: unset TASKS_AXI_BACKEND in the test harness right after sourcing tests/lib.sh, with a comment explaining why, so every case deterministically exercises its declared backend (4 lines added; no production code touched). Verified: reproduced the leak before the fix (TASKS_AXI_BACKEND=beads made the markdown dispatch case fail with 'beads show failed', exactly the reported failure mode); after the fix the full suite passes (0 failures, exit 0) both with and without TASKS_AXI_BACKEND=beads exported. The added lines are shellcheck-clean (the only shellcheck note, SC1091 on the lib.sh source line, pre-exists this change) * fix(bin): classify progress updates on requested work as routine (#3589) The supervision branch's verdict rule escalated every outcome that answered a captain request, so "the work started" and "still working" notes reached the captain with nothing to look at. The rule now keeps a finished result of requested work captain-facing, even when healthy, and treats start or still-working updates that bring no new artifact, finding, or decision as routine. The captain list for review-ready PRs, ask-user findings, exhausted blockers, credentials, and destructive or security-sensitive cases is unchanged, as are the unsolicited-routine, silent-fleet-review, and doubt-chooses-captain rules. The fm_branch_report tool description and the two docs that restated the old unconditional rule now point at the prompt's "Verdict: routine or captain" section as the one owner instead of carrying a second copy. * fix(bin): preserve captain calls during teardown (#3595) * fix(bin): never close a captain call during cleanup A scout that held its own work item for the captain, which is what captain-hold-lifecycle prefers ("hold the work item the question gates"), was closed by bin/fm-teardown.sh's automatic backlog transition. The completion gate passed, cleanup ran, and the captain's question moved to Done with no recorded answer: the one thing the policy says must never happen. `tasks-axi done` closes a held row silently, and nothing in teardown asked whether the row was the captain's own call. bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when the task is still an open captain call, 1 when it is not, 2 when that cannot be established. It reads the row through the transition library's backend-aware probe, so it addresses the same backlog teardown does; the script's other commands now address the configured data directory the same way instead of FM_HOME, which also fixes captain holds in a home with a relocated data directory. Teardown asks `open` before any destructive step and refuses on 2. On 0 only the close changes: after cleanup and still under the task's own lock, the row gets one "Deliverable of the finished work" line at the end of its body and returns to Queued through `tasks-axi reopen`, keeping its hold, so it lands in Captain's Call instead of reading as work under way. --force does not lift this: it authorizes discarding unlanded work, never the captain's question. The deliverable goes into the body because `tasks-axi update --report` rewrites the title of a row that is not Done. The crash window reuses the pending-close record teardown already stages: a `mode=retain` line makes the existing replay record the deliverable and reopen instead of closing, with the same validator, stale-generation check, cleanup-incomplete marking, and non-blocking bootstrap lock as an ordinary close. A retained row the captain answered first simply retires the record. No parallel record type, recovery command, or second bootstrap loop is introduced. Regressions run the real executables: the captain-held scout survives cleanup queued, held, with its deliverable and on the board, only `answer` closes it, --force keeps it open, and an ordinary scout still closes with its report; an interrupted cleanup leaves the row untouched and the next session start retains it; a relocated backlog keeps the retention in its one configured file; and a ship row whose hold cannot be read refuses cleanup before anything destructive. Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np * no-mistakes(review): Serialize captain holds and fix backend-aware listing * no-mistakes(document): Update captain-call retention documentation * no-mistakes(document): Fix relocated captain-hold backlog diagnostics * fix(bin): deliver secondmate outcomes to the parent channel (#3592) * fix(bin): deliver every secondmate outcome on the parent channel from the recording scripts A secondmate's captain-facing outcomes could miss: the mate model addressed the captain in its own unread chat instead of appending to the parent channel, and a PR-ready report, a finding, a decision, a blocker, and a failure all depended on that one remembered append. Make delivery structural, so the parent channel never depends on the model: - bin/fm-parent-channel-lib.sh is the one owner of channel resolution and exact-line append-once; the merge outcome path and the inactive-outcome scan now publish through it instead of two private copies. - bin/fm-inactive-reconcile.sh gains a ledger-first path that runs on every watcher poll in a secondmate home: a direct child's whole terminal done or failed line is delivered at once with its note, recorded PR, mode, merge posture, and scout report pointer, keyed and receipted so it is delivered once, and the inactive path yields to it. `report <task-id>` runs the same delivery for a caller holding the child's meta lock. - bin/fm-pr-check.sh publishes the PR-ready line with the canonical URL at registration. - bin/fm-captain-hold.sh publishes a hold and its answer, keyed by task id and resolution-record count, with no new persisted state. - bin/fm-teardown.sh delivers the child's final line before removing its record and refuses, retaining every record, while the channel cannot be written. - The charter opens with the parent-channel rule and confines the mate's own appends to judgement; AGENTS.md carries the carve-out at the persona address rule and the escalation list. docs/secondmate-parent-channel.md records the design and its coverage, and docs/verification/secondmate-parent-channel.md records the live run with real tmux panes and both real watchers delivering every line with no model. Supersedes #3569. * no-mistakes(review): Fix parent outcome retries and reconciliation locking * no-mistakes(review): Prevent busy children from starving ledger delivery * no-mistakes(review): Correct ledger metadata and hold occurrence handling * no-mistakes(review): Disambiguate ledger outcomes and normalize hold reasons * no-mistakes(review): Close ledger races and preserve teardown records * no-mistakes(document): Correct parent-channel receipt and scanner documentation * no-mistakes(lint): Quote done arguments for ShellCheck compliance * no-mistakes(ci): Fixed both CI failures. Updated GOTMP teardown fixtures for the new final-outcome reporter and isolated them from host tmux state. Updated the PR security assertion to distinguish the accepted PR-ready line from duplicate merge outcomes. Verified with both failing test suites, bash syntax checks, and git diff checks * no-mistakes(ci): Fixed Greptile’s duplicate-delivery race in bin/fm-inactive-reconcile.sh. Ledger events now claim matching already-delivered inactive receipts using the prior status fingerprint, preventing duplicate parent reports while preserving later same-state completions. Added behavioral regression coverage. Verified inactive-reconcile tests, project lint, documentation audience checks, syntax, and diff checks. Teardown tests passed relevant cases before the documented pre-existing herdr-preflight-missing-adapter failure * fix(bin): sync remote second mates to primary commit (#3599) * fix(bin): sync remote second-mate homes to the parent primary commit Session start and remote launch pointed a remote second-mate home at whatever Firstmate copy its own host kept, so a home that had already advanced past that copy refused as a non-fast-forward and every other home stopped at the host's older commit while the primary ran ahead. The parent now resolves ITS primary default-branch commit with the existing helper and hands that commit to the host on both paths. Because a remote home is a standalone clone, the host imports that one commit before advancing - already present, else from that host's Firstmate copy without moving it, else from the home's own origin - and then runs the SAME ff_target guards a local home gets, so dirty, diverged, feature-branch, and unresolvable targets skip untouched and the ancestry rules keep one owner. An unimportable target now names /updatefirstmate instead of failing opaquely, and a host still running an older Firstmate copy is reported the same way rather than echoing a bare refusal. The host-local launch leg no longer re-runs its own secondmate sync, so the spawn it drives cannot re-target that host's copy after the parent has already converged the home. /updatefirstmate is unchanged: it still refreshes the remote code root from that host's origin and then syncs the home to that refreshed copy, which is what the sync call with no target commit means. * no-mistakes(document): Document primary-targeted remote secondmate synchronization * fix(bin): separate captain intent from firstmate specs (#3597) * fix(bin): split brief task into captain intent and firstmate spec Keep no-mistakes --intent as the captain's ask plus later captain words, not the build spec or worker tradeoffs. * fix(bin): stop task-subsection copies at the next heading Promotion was swallowing the scout Setup contract into Firstmate spec, and pre-subsection briefs lost their # Task body. * no-mistakes(review): Validate brief content and preserve nested specifications * no-mistakes(review): Scope placeholder validation to scaffold-only subsection bodies * no-mistakes(review): Ignore fenced subsection headings during brief validation * no-mistakes(review): Preserve captain intent across scout promotion * no-mistakes(review): Enforce safe intent boundaries for legacy promotions * no-mistakes(review): Allow marked legacy intent and reject empty promotions * no-mistakes(review): Scope task parsing and overlay legacy intent contracts * no-mistakes(review): Overlay current intent contract for all no-mistakes spawns * no-mistakes(review): Preserve later captain clarifications in intent overlays * no-mistakes(document): Document brief intent enforcement and ownership * no-mistakes(ci): Updated spawn-related test fixtures to use valid Captain intent and Firstmate spec subsections, corrected launch-path expectations to launch-brief.md, and resolved ShellCheck quoting findings. Verified with fm-lint.sh and 15 affected behavior tests, including real Herdr tests; all passed * no-mistakes(ci): Updated stale spawn/promotion fixtures in the Muse, Orca, secondmate-harness, and public-followup suites to provide valid Captain's intent and Firstmate spec subsections. Verified full Orca and secondmate-harness suites, targeted public-followup promotion behavior, Bash syntax, diff checks, and fm-lint * fix: start a fresh supervision branch for every main session (#3600) * fix(pi): start a new supervision branch conversation per main session The supervision branch reopened one recorded conversation forever, so every main session start reloaded the current generated prompt and then weeks of accumulated thread, where a superseded rule could still outweigh today's. The branch conversation is now scoped to one main session: the session generation owns the recorded conversation, so a cold start, /new, /resume, /fork, or a reload always builds a new one, while a rebuild inside one session (a model or effort change) still continues that session's own conversation. The dialog mirror re-anchors with it. Its durable cursor records what the previous branch conversation received, so a /resume or reload - which keeps main's own session file - would otherwise leave the new branch blind to dialog main itself still has. The reset is bounded by the current main session, and the cursor keeps advancing incrementally within it. The durable outcome store and its processed marker are untouched, so unacknowledged captain-facing outcomes still re-present on the new main session. * no-mistakes(document): Document fresh Pi supervision conversations * no-mistakes(ci): Fixed the flaky concurrent inbox failure. Lock acquisition now retries when a competing lock disappears between a failed claim and inspection. Added a behavioral regression covering that race. Verified the full inbox test four times, project lint, and git diff checks * feat: restart second mates after instruction updates (#3614) * feat(update): restart second mates whose instructions changed /updatefirstmate pulled new bytes onto disk and then asked each advanced second mate to re-read them. A running agent holds AGENTS.md and every loaded skill frozen from launch and no verified harness offers a reload, so that steer could not reach a loaded skill at all and left the mate holding two contradictory copies of its own job description. An eligible mate is now restarted instead, in the same home and endpoint, through the existing transactional relaunch. The restart is gated on the mate first writing down the open work it holds only in conversation - the open-record half of /stow, never its memory sweeps - so an unregistered captain call is flushed before the conversation is spent. Anything that leaves the reload unprovable falls back to the old re-read message and is reported as exactly that, never as a clean reload. Remote mates take the same path: fm-remote-secondmate-control.sh gains a relaunch verb whose host-local leg runs that same control plane, since the mate is an ordinary local secondmate from its host's point of view. The primary resolves the profile and passes it explicitly, because config/secondmate-harness is not inherited and the file on that host belongs to a different home. fm-update.sh now splits its advanced live mates into a restart set and a nudge residual, and both sets require a changed instruction surface, which also closes the over-nudge against the session-start sweep. Restart is stricter still: a bin/-only advance reloads itself on the next call, so it never costs a conversation. Colocated tests cover the gating, the persist-then-restart order, the task-subset persist request, each unsafe fallback, the remote hop, and the remote sync's new instruction-surface report. * no-mistakes(review): Fix restart correlation, concurrent waits, and lifecycle reporting * no-mistakes(review): Parallelize relaunches and classify replacement incarnations * no-mistakes(review): Gate restart actions on live agent state * no-mistakes(review): Handle failed restart workers without hanging * no-mistakes(review): Nudge legacy remotes and preserve persist recovery * no-mistakes(review): Document one-time secondmate restart rollout * no-mistakes(review): Honor arrived replies and refresh remote profiles * no-mistakes(review): Revert remote parent profile reconciliation * no-mistakes(review): Reset remote profile defaults and honor published results * no-mistakes(review): Preserve fallback nudges for unverifiable secondmates * no-mistakes(document): Document second-mate restart update flow * no-mistakes(lint): Fix ShellCheck warnings in restart scripts * perf: accelerate local validation with bounded concurrency (#3644) * perf(tests): route gate verification through the bounded concurrent runner Local validation was the pipeline's dominant cost: across 67 recorded no-mistakes agent sessions on this repo, 99.3% of command execution was `bash tests/*.test.sh`, run strictly one script at a time, and 2% of those calls were killed by an agent-guessed timeout and paid for twice. Three changes, each measured: - `.no-mistakes.yaml` pins `commands.test` to `bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner already owns changed-file selection, bounded concurrency, the refusal of unproven scripts, and a generous automatic per-script bound, so the gate's baseline is neither a serial chain nor a guessed timeout. It stays intent-targeted - the Test step still runs its evidence agent on top - and excludes the live-Herdr family the required Herdr lane owns. - `bin/fm-test-run.sh` gives a plain list of script paths the same bounded automatic scheduler and automatic bound that `--changed` gets. Naming several subjects is how a verification round asks for exactly those scripts. The curated selections are untouched: `--lane` still composes CI shards whose serial lane must stay serial, `--family` is what the required Herdr lane runs, and `--all` stays a deliberate complete regression. - `pr-forge` is admitted to the concurrent-safe family registry on two consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those, and records `secondmate` and `session-bootstrap` as refused with the exact script and reason each failed on, so the refusals are actionable rather than silent. Measured on this host, 0 failures on both sides: verification round, 4 scripts 448s chained -> 231s through the runner (-48%) pr-forge family 409.2s at 1 worker -> 237.9s at 4 (1.72x) watcher-wake-lock family 1311.1s at 1 worker -> 539.3s at 4 (2.43x) A fourth lever was implemented and then removed because the measurement refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made `fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s unchanged, back to back. Those sleeps are not overhead added to the clock - they are how a test waits for a subject moving on fm-watch.sh's own one-second cadence - so sampling less often only delays detection. It also broke `fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a settled condition. CONTRIBUTING.md records that result so the experiment is not repeated. * no-mistakes(review): Separate concurrent runs by isolation proof family * no-mistakes(review): Limit automatic timeouts to changed-file validation * no-mistakes(document): Clarify validation concurrency documentation * fix: copy PR URLs from durable records (#3648) * fix: copy PR URLs from records or abstain, never assemble them Supervision reported a plausible but dead PR link three times because its prompt demanded a full https:// URL at a moment when only a PR number was observable, so the model assembled an owner/repository from memory, and the PR check then accepted that URL and wrote it into the task record, after which the model kept defending its own tool-endorsed guess over the worker's real link. Three changes close that chain without any live forge lookup, so private forges are treated exactly like public ones: - bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy or abstain" section requires a URL to be copied verbatim from a durable record (the done: PR <url> status line, pr= metadata, or the backlog note), forbids assembling owner, repository, host, or number from memory, and has the branch report only the identifier it actually holds when no record names the URL yet, leaving the PR check unarmed until the worker's ready line arrives. AGENTS.md section 7 and 9 carry the same copy-or-abstain rule for main in place of the bare full-URL mandate. - Worker briefs (bin/fm-brief.sh, ship and scout rules) require the full https:// URL wherever a PR is mentioned - status line, terminal, or summary - never a bare "PR 108", so the link is in view as early as the number is. - bin/fm-pr-check.sh refuses, offline and before any side effect, a URL that the task's own done lines contradict, printing both spellings; a log naming no URL still records the argument as before. fm_pr_status_ready_urls in bin/fm-pr-lib.sh owns reading those lines. The refusal also reaches bin/fm-pr-merge.sh, so nothing merges under a contradicted URL. Tests cover the offline refusal with zero side effects, the recorded spelling being accepted, markdown-wrapped and punctuated URLs, working lines not counting, the merge wrapper propagation, a self-hosted merge request with no forge call, the prompt carrying the rule, and the brief carrying the worker rule. * no-mistakes(review): Remove stale PR URL enforcement * no-mistakes(ci): Removed backlog notes as an accepted PR identity source. PR URLs may now be copied only from the task’s `done: PR <url>` status or canonical `pr=` metadata; otherwise supervision reports only the known identifier and leaves PR checking unarmed. Updated related guidance/docs and verified with branch-supervision tests, brief tests, ShellCheck, and `git diff --check` * fix(bin): disable Claude feedback drafts for fleet launches (#3661) * fix(bin): disable Claude's feedback-draft flow for fleet-launched agents Scope --settings '{"feedbackDrafts":"off"}' to every Firstmate-launched Claude crewmate and secondmate, so /bug and /feedback never queue or submit a bug report on the captain's behalf. feedbackDrafts is the documented settings key (Claude Code changelog 2.1.247); the per-launch CLI flag never touches the captain's global settings.json. Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3 * no-mistakes(review): Prevent managed settings from re-enabling Claude feedback drafts * no-mistakes(document): Fix Claude feedback documentation formatting * fix(bin): layer both feedback-draft controls for defense in depth The prior --settings-only fix can be overridden by a managed Claude settings policy (feedbackDrafts precedence). Keep CLAUDE_CODE_SEND_FEEDBACK=0 alongside --settings '{"feedbackDrafts":"off"}': either control alone disables the SendFeedback tool, so a managed override of one still leaves the other in force. Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3 * no-mistakes(document): Document Claude feedback-draft suppression ownership * feat(tests): run three more validation families concurrently (#3662) * perf(tests): admit three more families to concurrent validation The three families that `docs/fm-test-isolation-proof.md` recorded as refused were not refused for concurrency. Each blocker was a test that decided a property by wall clock, or a script filed where it cannot run. Fixing those three things admits all three families and recovers 28.6 minutes of local validation with no assertion removed or weakened. - `tests/fm-backlog-handoff.test.sh` injected its pre-move crash by killing the handoff, sleeping a fixed second, then delegating the move to the real binary. Nothing ever killed the fake, so on a host slow enough for the case's next assertions to take longer than a second, the orphan woke and completed the very move the case requires left undone, and recovery then failed with `Task "pre-move-crash" not found in this backlog`. Watching the two backlogs during the injected crash showed exactly that, the item moving one second after the crash. All four crash injections in the file now go through a new `fm_fake_crash_injector` shim that signals the target and returns only once it is observably gone, and the pre-move fake never delegates the move at all. - `tests/fm-session-start.test.sh` proved the startup digest does not block on a slow current-state read by timing the whole digest against a fixed eight-second sleep, which a loaded host exceeds without the property being violated. It now holds that read open until the case releases it and asserts, the moment the digest returns, that the read has not finished. A digest that waited would wait indefinitely rather than for an interval a slow host can out-run, so the assertion is stronger than the bound it replaces. Its scan budget moves to the maximum, because the old value left two seconds of margin over the fixed sleep and measured the host rather than the deadline that `tests/fm-inactive-reconcile.test.sh` owns. - `fm-backend-herdr-focus-flash-e2e` was filed in the family map's catch-all, which put it in the portable serial lane, where Linux CI gate-skips it: that real-Herdr regression was running nowhere. It moves to `real-herdr-gated` and the required Herdr lane. `fm-claude-stop-autoarm-live-e2e` gate-skips on its opt-in variable and moves to `live-harness-optin`. The 28 remaining ungrouped scripts become an enumerated `standalone` family instead of admitting `unclassified` itself. `unclassified` is the family map's `*)` arm, so admitting it would silently grant concurrency to every test added afterwards, which is exactly the population with no proof. A new test still lands in `unclassified` and stays serial, and `tests/fm-test-run.test.sh` covers that split behaviorally. Each family passes two consecutive four-worker proofs with zero failures. On the production runner, `secondmate` goes 1233.1s to 453.4s, `session-bootstrap` 756.4s to 286.4s, and `standalone` 724.6s to 261.1s: 2.71x overall and 1713.2s recovered. The whole suite runs 177 scripts in 52.6 minutes of wall clock against 121 minutes of summed script time. * no-mistakes(document): Refresh concurrent validation and shard documentation * no-mistakes(ci): Fixed the real-Herdr focus-flash E2E race exposed by reclassification. Part C now starts its persistent child atomically via `pane run` and verifies stable child identity through Herdr’s public `process-info` interface, avoiding the racy send-text/send-keys sequence and platform-specific `ps` matching. Verified with bash syntax checking, ShellCheck, git diff checks, and the complete E2E test on Herdr 0.8.2 * feat: structure no-mistakes ask-user escalations (#3670) * feat(brief): structure no-mistakes ask-user escalation as event + snapshot file Crewmates escalating a no-mistakes ask-user gate now report one status event naming every finding id plus a snapshot file holding the gate's axi finding records verbatim (id, severity, file, line, description, authority), using the same shape even for a single finding. The status line never paraphrases. The format is defined once in fm-dod-lib.sh and rendered into both the scout and ship rule 6 in fm-brief.sh, so a promoted scout - whose rule 6 fm-promote.sh preserves unchanged - gets the identical contract as a freshly-spawned no-mistakes ship worker. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PpiWaDerbYavTLPPtEjQei * no-mistakes(review): Preserve ask-user escalation output contract * no-mistakes(review): Align escalation format test expectation * no-mistakes(review): Scope ask-user escalation instructions correctly * no-mistakes(review): Remove ask-user from generic decision rules --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> * fix(bin): require self-sufficient no-mistakes intent (#3671) * fix(bin): require a self-sufficient no-mistakes intent A no-mistakes worker's --intent is only as useful as the string it passes. PR #3604 shipped with an intent that was only "do 1, 2, 3, 7 from the report": the real contract lived in a private scout report and never reached --intent, so nobody holding that string plus the codebase could have derived the specification. This is pure instruction at the contract's one owner; no spawn-side or promotion-side check is added. - bin/fm-dod-lib.sh: the generated no-mistakes Definition of done now states that the --intent string must be self-sufficient (the string plus the codebase reconstructs roughly the same specification) and tells the worker to write the substance of any report, decision, or PR the captain's intent refers to into --intent rather than the pointer, while Firstmate build instructions and the worker's own decisions still stay out. The spawn-time overlay points back at that rule so its "supersedes" wording cannot cancel it, and the header's owner statement carries the rule. - AGENTS.md section 11 and bin/fm-brief.sh's header ask Firstmate to include the substance of referenced material when filling ## Captain's intent, and section 11 points at the owner of the rule. - tests/fm-brief.test.sh and tests/fm-task-delivery.test.sh assert the rendered brief and launch contract carry the rule. Claude-Session: https://claude.ai/code/session_01YMhEe42q7BAAoN6RxNuzim * no-mistakes(document): Replace incident-specific intent test commentary * fix: accelerate local Bearings snapshot composition (#3499) * Speed local fleet snapshot composition * no-mistakes(review): Stabilize task inventory during concurrent snapshot composition * no-mistakes(document): Document local snapshot observation concurrency * no-mistakes(ci): Fixed CI failures by making empty task manifests compatible with stock macOS Bash 3.2, snapshotting task metadata before concurrent observations to prevent generation drift, strengthening the behavioral race regression, and updating the stock-Bash Bearings test count to 45. Verified fleet snapshot tests (15), Bearings tests (45), workflow lint tests, project lint, Bash 3.2 parsing, and diff checks * no-mistakes(ci): Fixed the Linux CI failure caused by passing large backlog/task JSON through jq command-line arguments, which exceeded the per-argument size limit. Both inventory projections now stream large JSON inputs through stdin. Verified with fm-bearings-snapshot.test.sh (45 tests), fm-fleet-snapshot-view.test.sh (15 tests), Bash s…
kunchenguid
added a commit
that referenced
this pull request
Sep 17, 2026
…olidated URL rule The consolidated section 9 URL rule narrowed its trigger to a review or merge ask, dropping the "whenever a PR is mentioned" catch-all from #3648 that keeps every PR URL copied from a durable record and never assembled from memory. Restore that trigger as a union with the review or merge ask so the one consolidated rule covers both.
kunchenguid
added a commit
that referenced
this pull request
Sep 17, 2026
* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work MAIN answered a supervision-branch outcome for completed captain-requested work (implementation done, PR ready for review and merge approval) with "Captain, shipshape.", reading section 9's no-action reply as covering it and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no captain-facing response is owed". Section 9 now limits the shipshape reply to true no-ops (idle re-read, empty heartbeat, consequence-free acknowledgement) and requires a short outcome response naming what finished and what word is needed whenever requested work finishes or a result needs the captain's word, even when a transcript entry already shows the substance. The Pi protocol's re-emit rule now says it bounds repetition only, and carries a worked example of the ready-for-review outcome whose correct processing turn a shipshape reply fails. No executable contract evaluates the content of MAIN's captain-facing reply, so the regression is the protocol example in the owner doc rather than a text-match test. * no-mistakes(document): Clarify captain-facing outcomes versus no-ops * docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line The document step condensed the Pi protocol's re-emit rule and dropped the worked example of a finished, ready-for-review outcome whose correct processing turn a "Captain, shipshape." reply fails. That example is the contract's regression: no executable contract evaluates the content of MAIN's captain-facing reply, so the owner doc's example is the test case. Restore it directly under the re-emit rule, prefixed as a regression example that is kept verbatim and never condensed or summarized away. * no-mistakes(review): Clarify captain outcome and decision-word requirements * no-mistakes(document): Clarify captain-facing completion outcomes * docs(pi): require the PR URL in the visible captain-facing outcome reply Captain review on the regression example: drop the sample reply string and say only that the ready-for-review outcome requires relaying a captain-facing outcome response, not just "Captain, shipshape.". Fold in the visible-PR-handoff failure seen this session: after the branch outcome reporting this fix green, MAIN's visible reply was only "Awaiting your merge call." with no PR URL, leaning on the dim anchor. Section 9's URL rule now also covers a review or merge ask and names the visible reply as where the URL goes, sourced from the ready status, pr= metadata, or the supervision branch's summary and never left to a transcript entry. The Pi protocol adds the same-way failure and places the captain-facing text in the final visible assistant reply after the fm_branch_processed call, because Calm hides assistant text emitted in the same step as a tool call as a working note. Investigation verdict, evidence in the PR comment: no recent PR caused the handoff failure; Pi has hidden same-step pre-tool assistant text since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and #4658 touched only remote report transfer. * no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs * no-mistakes(document): Clarify captain-facing supervision outcomes * docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule The consolidated section 9 URL rule narrowed its trigger to a review or merge ask, dropping the "whenever a PR is mentioned" catch-all from #3648 that keeps every PR URL copied from a durable record and never assembled from memory. Restore that trigger as a union with the review or merge ask so the one consolidated rule covers both.
vifar
added a commit
to vifar/firstmate
that referenced
this pull request
Sep 18, 2026
* fix(bin): translate Stop hook timeout signals into durable auto-arm failure (#4474)
* fix(bin): recover Claude auto-arm after timeout
* no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list
* fix(spawn): establish Claude task channel authority (#4464)
* fix(spawn): establish Claude task channel authority
* no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference
* fix(bin): refuse fm-control.sh exit when the composer holds unproven or pending text (#4458)
* fix: guard relaunch exit against pending input
* no-mistakes(review): Verifying test run in progress
* no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard
* no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now
* fix(spawn): establish crewmate identity first (#4481)
* fix(bin): reconcile redundant secondmate divergence during updates (#4460)
* fix: reconcile diverged secondmate updates
* no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md
* no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md
* feat: enable gpt-5.6-luna max reasoning for crew dispatch (#4497)
* fix(dispatch): support Codex Luna max effort
* no-mistakes(review): use portable CODEX_HOME path in codex effort reference
* feat(calm): render smooth Unicode swell with asymmetric two-color sail (#4498)
* feat(calm): render smooth Unicode swell
* feat(calm): make sails asymmetric
* feat(calm): use quarter sail glyph
* no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer
* no-mistakes(document): docs: sync calm wave phase doc comment
* no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。
* fix(bin): supersede stale scout delivery text in brief.md on promotion (#4491)
* fix: supersede scout delivery brief on promotion
* fix: preserve ship safety contract after promotion
* no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch
* fix(bin): make captain holds work on hosts with an older JSON::PP, and stop cleanup dropping accents from a held body (#4471)
* fix(bin): let captain holds work on hosts with an older JSON::PP
Holding a task for the captain, and the cleanup that keeps a captain-held row
open, both fail outright on any host whose JSON::PP defaults allow_nonref off -
2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`,
but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older
library rejects that whole value with "must be object or array".
The consequence is fleet-wide on such a host, not one broken command: a worker
there cannot formally record a decision for the captain at all. It can only
mention the decision in passing in a status line, where it can be missed - which
is how a real decision goes unrecorded. The hold reports that the task lost its
hold-set stamp; the cleanup cannot return the row to Queued.
Both call sites now ask for allow_nonref explicitly rather than inheriting
whatever the installed library defaults to. The second one is worth naming: its
`/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string
case that fails, so the guard selects for the failing input rather than
protecting against it.
The regression case forces the older default back off for every perl the commands
spawn, then drives both paths - holding a task that carries a body, and tearing
down a captain-held row whose deliverable must still be appended. It also probes
that the simulation genuinely rejects a bare scalar, so the case cannot pass
vacuously on a lenient host. Each half was verified failing on its own unfixed
call site with that site's real error message. Suites: fm-captain-hold-lifecycle
51 cases, fm-backlog-atomicity 99 cases, 0 failures.
Verification limit: the mechanism is reproduced and tested, but neither fix is
verified against a real JSON::PP 2.27202 host, because none is in the loop. This
laptop runs 4.06, where the bug does not manifest.
`bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a
brace-delimited object before decoding, so allow_nonref never applies.
* fix(bin): stop cleanup silently dropping accented characters from a held body
Cleanup rewrites a captain-held row's body to append the finished work's
deliverable, and the decoder it reads that body with printed decoded characters
to a stream with no `:raw` layer. A character at or below U+00FF then came out
as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the
accent. `fm_backlog_retain` writes that body straight back through
`--body-file`, and nothing reported an error - the character was simply gone
from a row still waiting on the captain.
The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus
`utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already
used.
Review of the parent commit found this on one of the lines that commit already
changed. It predates that change.
The test asserts bytes rather than decoded strings, because comparing strings
cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any
character above U+00FF makes perl print the whole string as UTF-8, so one body
carrying both an accent and an em dash passes even unfixed and proves nothing.
Verified failing before the fix on the accented row, passing after. Suites:
fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures.
* no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc
* no-mistakes(review): drop whole-file UTF-8 check from retained-body test
* no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc
* no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc
* fix(bin): read codex 0.154's idle braille starfield rows as composer furniture (#4532)
* fix(composer): read codex 0.154's idle starfield and status footer as furniture
codex-cli 0.154.0 animates a braille "starfield" around its idle composer:
on the row above the bold `›` prompt row, on the `›` row behind the SGR-2
dim `Ask Codex to do anything` placeholder, and on the row below it, then
draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`).
The cells are truecolor greys on both sides of the ghost luminance ceiling,
so the brighter ones survive ghost stripping, and the rows below the glyph
carry no structural edge. The shared classifier selected the bare `›` shape,
extended its wrap region over the two rows beneath the glyph, read the
survivors and the footer as wrapped typed input, and answered `pending`;
the steering doorbell defers on exactly that verdict, so no doorbell ever
reached an idle codex 0.154 pane.
bin/fm-composer-lib.sh now recognises that furniture by shape, declared
once next to the idle placeholders and reached from the two wrap-region
boundary points:
- a row whose non-whitespace content is entirely braille cells
(U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it
never counts as wrapped typed content and bounds a bare composer's wrap
region; braille behind the glyph row's content is stripped before the
emptiness decision when nothing else follows the glyph; a row mixing
braille with other text stays typed content;
- the codex status footer bounds the wrap region exactly as omp's status
row does, anchored on the effort token, a spaced middle dot, and a `~` or
`/` path cell, so a typed `fix · tests` stays composer input;
- `^Ask Codex to do anything$` joins the verified idle-placeholder set; the
ghost strip remains what proves that row empty, and the bare-row rule that
bright placeholder text is real input is unchanged.
Unchanged: the strict blank-row rule, the styled=0 degradation (a plain
cmux/orca capture of this screen still reads `unknown`, never `pending`),
FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape.
tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte
with the divergence (letters in place of the starfield read `pending`) and
the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh
is the default-on live guard (token-free, skips explicitly without codex or
tmux) that launches the installed codex idle and asserts `empty` through
both the tmux and the cursorless styled reads, naming codex --version on
failure. docs/verification/runtime-backends.md records the dated Herdr
evidence: `pending` before, `empty` after, on the captured screen.
* no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry
---------
Co-authored-by: Todd Billings <todd@usdvcapital.com>
* fix(bin): refuse empty text steers in fm-send (#4259)
* fix(bin): refuse empty text steers in fm-send
A marked secondmate request sent with an empty message delivered only
marker and correlation bytes and minted a pending-reply expectation the
parent could never see resolved, stalling the fleet with no loud error
(#4255). Fail closed on an empty or whitespace-only message on the text
path, mirroring the existing --resolve-key refusal.
* chore: retain ambient Pi-lens autoformat as its own commit
Formatting-only edits produced by ambient Pi-lens autoformat during the
msg-loss investigation, kept separate from the behavioural change in
c23acba6 so the fix stays reviewable on its own.
AGENTS.md is deliberately excluded: its only autoformat edit stripped the
trailing space from the documented FM_OPERATIONAL_PREFIX value, which
bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11
records as permanent compatibility. Documenting that constant without its
trailing space makes the doc wrong about the contract, so that one line was
restored rather than retained.
* fix(calm): paint the working ship one yellow over all-blue water (#4554)
On rose-pine-moon the two-color water (cyan crests over blue troughs) read as
a pink stripe over aqua, the yellow left sail and mast clashed with the red
right sail, and the hull carried a blue interior run. Every water cell is now
blue so the swell reads through glyph height alone, and both sail halves, the
mast, and the whole hull are one yellow run. Geometry, cadence, animation,
direction flip, resize clamping, and the narrow fallback are unchanged.
Update the unit and real-TUI color assertions to the new palette and the Calm
docs that described the old one.
* fix(bin): stop aging a second mate's active turn from its launch (#4270)
* fix(watch): stop aging a second mate's active turn from its launch
The parent watcher's second-mate wake-loop stall check exempts a mate that
is demonstrably inside an active turn, but secondmate_in_active_turn asked
busy_turn_over_age first and returned "not in a turn" whenever that said
the bound was crossed.
busy_turn_over_age ages from state/<task>.turn-ended, falling back to
state/<task>.meta. A second mate's turns end in its own home, so the
parent never gets a turn-ended mark for it and the fallback ages the
mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago
was therefore permanently "over age", the busy pane was never consulted,
and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false
wake-loop stall.
The gate now bounds the busy exemption by <idle> - how long the queue's
drain position has not moved - which is evidence this home actually
holds. A busy mate stays exempt while the queue has been frozen for less
than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so
the bound that stops a busy pane from proving liveness forever is kept
rather than removed. busy_turn_over_age is untouched; its remaining
callers are the ordinary crew busy-pane bound.
The regression pins the case that actually broke: a mate whose launch
record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn
must not escalate, while the same mate with its queue frozen past the
bound still publishes exactly one notification. The existing coverage
only exercised a freshly launched mate, which passes either way.
Reaching that alert now costs a pane capture inside the gate, so the
three checkpoints in this suite that assert an alert move from a 1s to a
4s bound - the value the neighbouring active-turn cases already use. The
bound is a ceiling, not a wait: the checkpoint returns on the first
actionable wake. On a loaded machine a 1s bound missed the alert
repeatedly; at 4s it did not miss in 20 runs under the same load.
* no-mistakes(review): scope the second-mate active-turn regression test's coverage claim
* no-mistakes(document): fix stale second-mate active-turn comments in fm-watch
* feat(bin): add read-only PR blocker and reviewer discovery commands (#4278)
* feat(bin): add read-only PR blocker and reviewer-discovery commands
Two focused, opt-in commands that read GitHub and never write to it.
fm-pr-state.sh reports what still blocks one pull request from the
author's side: a closed or merged state, draft state, unknown or
conflicting mergeability, absent or failing required checks, and a
blocking CHANGES_REQUESTED decision explained by each reviewer's latest
verdict, marked STALE when it was left at a superseded head. A pull
request that only awaits an approval is not reported as blocked, and
advisory checks are omitted. Every reading is taken against one exact
head; a push that lands mid-read invalidates the whole result rather
than mixing two snapshots.
fm-pr-reviewers.sh suggests reviewers from the most recent commits to
the pull request's exact changed paths, counting each commit once,
resolving handles through GitHub's own commit author.login mapping, and
excluding the author and Bot accounts.
Both stay read-only: no review request, no approval, no merge.
Unresolved review-thread state is left unreported because the REST API
does not expose it and unattended commands may not use GraphQL.
Closes #3731
* no-mistakes(review): accept only PR URLs and stop at terminal state
* no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate
* no-mistakes(review): stop attributing readings to unverified heads
* no-mistakes(review): narrow readiness contract to checks that have reported
* no-mistakes(review): read the pull request once, drop the head guard
* no-mistakes(document): scope pr-forge isolation proof to its measured members
* no-mistakes(document): record uncovered pr-forge members and their pending proof
* docs(isolation-proof): re-prove pr-forge at its full membership
tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the
pr-forge family in this branch, and script_allows_concurrency grants
four workers by family membership alone, so both ran concurrently on a
proof measured before they existed.
Re-proved the family at all eight members: two consecutive runs, 0
failures, each begun with the one-minute load average below 6.0 so the
result measures isolation rather than contention. A third run taken
between them is disclosed rather than recorded, because it started
while the previous run's workers were still decaying.
The new durations are not comparable with the six-member measurement
above them, so they are not presented as evidence about the two new
members, and that record's 1.72x four-worker figure is left as a
statement about its own run rather than restated as current.
* no-mistakes(review): disclose gh error-text coupling at its matching site and tests
* fix(bin): teach validation-round pauses in generated briefs (#2752)
* fix(bin): teach validation-round pauses in briefs
* no-mistakes(document): Point classifier comments to authoritative pause examples
* docs(readme): add star history chart (#4558)
* fix(bin): refuse teardown when a task's endpoint close fails (#4510)
* fix(teardown): refuse a cleanup whose endpoint close failed
bin/fm-teardown.sh discarded both the exit status and the stderr of every
fm_backend_kill call, so a close that genuinely failed was indistinguishable
from one that succeeded. Teardown continued past it, deleted the task's durable
records, returned its worktree, and reported the cleanup as completed. The
deleted metadata is the only record of which endpoint belongs to the task, so
such a close did not merely leave a stray session behind, it stranded one:
nothing was left on disk naming it.
The adapters could not carry that signal either. Driven against the real code,
every backend arm returned 0 for a genuine failure exactly as it did for an
already-exited endpoint, so there was nothing for the four call sites to
propagate even once they stopped swallowing it.
The tmux arm now resolves a close that did not succeed against the window's
exact recorded identity, since kill-window fails the same way for a window that
is gone and one that is still there. The Orca arm reports a close its missing
CLI never attempted. Both stay silent for an endpoint that is already
legitimately gone, and the remaining arms are unchanged: their close-command
timing cannot be established without the real Zellij, Orca, and cmux binaries,
and a gate that refused ordinary cleanup of an already-exited session would be
worse than the defect. docs/verification/runtime-backends.md records what each
backend can prove.
A reported close failure now reaches teardown's existing retain-and-stop
refusal before the records naming the endpoint are removed, matching where the
Herdr confirmed-gone gates already sit for the same hazard, and the retained
records let a rerun finish once the close works.
* no-mistakes(review): refuse unreadable tmux close re-read; honor --force override
* no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close
* no-mistakes(document): document endpoint-close refusal in its backend and retirement owners
* no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree
* feat(calm): add flag-gated Claude Code Calm mode (#4565)
* feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag
Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that
brings Calm to Claude Code: the sailboat replaces the stock working row through a
Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration,
and canonically classified operational user rows draw at zero height. /calm is
registered by the hooks module itself and toggles the same per-home config/calm
preference the Pi extension uses, so one choice applies on either harness; rows
redraw retroactively on toggle and stay hidden across claude --continue.
The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS
flag is on. Nothing sets that flag in any settings file, and the plugin carries no
command file, skill, agent, or classic hook, so it is a complete no-op while the
flag is off. The trusted project auto-loads it through an .agents/skills symlink,
the only path Claude Code scans for project plugins.
Extract the working-ship geometry, bounce track, cadences, and freeze/resume state
into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module
imports from outside the plugin folder) and have the Pi widget paint that core's
frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify
operational rows through a port of bin/fm-operational-input.sh's classify command
guarded by a corpus parity test against the shell owner.
Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering,
Raster packing, policy, classifier parity), the mod's own claude plugin test suites
behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op,
the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272.
Docs: record the version-scoped Claude Code evidence and the three bounded gaps in
docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md,
and make the shared preference, layout, and contributor notes harness-neutral.
* no-mistakes(review): Preserve colliding final replies and strengthen parser parity
* no-mistakes(review): Preserve final replies and strengthen canonical parity checks
* no-mistakes(review): Require exact function-hooks opt-in before Calm activation
* no-mistakes(review): Clarify Calm module loading and activation boundaries
* no-mistakes(review): Reset Calm presentation state across session starts
* no-mistakes(document): Refresh Calm session lifecycle documentation
* feat(calm): paint the Claude Code working ship in Claude's own theme colors
The captain picked the "Claude native" palette for the Claude Code mod's Raster:
every water cell takes the spinner blue of the active theme family (#93a5ff dark,
#5769f7 light) and the whole boat takes the Claude orange of the stock spinner
(#d77757), one water color and one boat color. The family follows the `theme`
setting's prefix, read at load through $.config.list and re-read on a
config.set of that row, with `auto` and custom themes falling back to the dark
set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte.
Rename the shared sprite's color classes from hue names to `water` and `boat`,
since each harness now maps them to its own colors; geometry, motion, cadence,
and the activation gate are untouched.
Tests cover both palettes' packing and the family rule under Node, and the
plugin kit drives every theme value, a theme change mid-session, the Calm-off
pass-through, and inertness of the menu read while the flag is off. The docs
describe the Claude Code colors and record the guard passing on 2.1.273.
* no-mistakes(review): Use light palette for unresolved Claude themes
* no-mistakes(document): Refresh Claude Calm verification evidence
* fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586)
* fix(watch): honour a declared wait before wedge-escalating a quiet pane
wedge_timer_check escalated on elapsed idle time alone. Nothing asked
whether the worker had already said why its pane was quiet, so a lane
that declared a bounded external wait climbed the escalation ladder for
as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every
repeat carried demand-deep-inspection - which by its own wording forbids
re-absorbing on the run-step or pane state, so the supervisor could not
use the evidence that was there either.
The generated brief promises that declaring `paused:` buys the long
recheck cadence instead of a wedge, but the timer was still reachable
while that declaration stood: a crew that declares a wait and then has an
active run or busy pane attributed to it is handed to the timer as
provably-working. The declaration is what the worker said about its own
silence, so it now outranks a liveness verdict that only says something
is running.
The consult runs in the at-threshold branch that was about to escalate,
beside the worktree walk already there, and costs one status-line read.
Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS
recheck the declared-wait absorber already uses, so the wait is still
rechecked and cannot rot invisibly. Which verb declared it decides the
wording, because the two block on different people: a `paused:` wait is
owed by an external dependency and asks the reader to confirm it still
holds, while a `captain-held:` transfer is owed by the captain reading
the recheck and asks them to answer or release the hold. A hold is not
rechecked at all while the away-posture record exists, as on every other
captain-held path, and that absorb arms no throttle so the recheck is
owed in full on return.
A declared clearing time that has already passed stops counting, and a
lane that never declared one keeps the identical escalation schedule,
reason, count and demand-deep-inspection wording, so detection and its
worst-case time are unchanged. The deferral restarts the idle timer
rather than cancelling it, so a lane that stops waiting escalates again
within one threshold.
A lane quiet because its own validation run is parked at a gate awaiting
a human decision is deliberately out of scope: reading that state needs a
signal carrying who the wait is on and what clears it, rather than one
inferred from a parked verdict that also covers gates awaiting the
crewmate itself.
Tests pin both directions for each case and were each confirmed to fail
with the consult removed.
* no-mistakes(document): docs: honour declared waits in stale-escalation docs
* fix(bin): report verified PR state for passed runs (#4624)
* fix(bin): derive passed PR state from PR record
A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists.
For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim.
Fixes #4607
* no-mistakes(review): Add bounded GitLab merge-request state reads
* no-mistakes(review): Preserve network-free inactive crew-state scans
* no-mistakes(document): Document PR record readers in shared library
* fix: restore published contribution follow-up (Fixes #4469) (#4627)
* fix: restore published contribution follow-up (Fixes #4469)
* fix(review): Fix contribution freshness and merge actor routing
* fix(review): Restore issue triage and scope contribution follow-up
* fix(test): test: assert one wake per contribution signal
* fix(document): Document contribution follow-up
* fix: restore truthful terminal delivery evidence
* fix(review): Disclose unsupported contributions and deduplicate watcher wakes
* fix(review): Preserve unmeasured unsupported contributions across Bearings
* fix(review): Deduplicate shared contribution wakes and isolate diagnostics
* fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower
* fix(bin): make remote report transfers explicit and fail-open (#4658)
* fix(bin): make a remote-reply document gap self-clearing and re-attemptable
A remote mate's undelivered document raised a keyed `blocked` decision that
nothing could ever resolve, and any `data/*.md` substring in any mirrored line
was an unconditional fetch instruction. A mate announcing a report it had not
written yet therefore manufactured a permanent, factually false blocker, and
its own explanation of the false alarm manufactured more.
The reader has no permanence vocabulary: a report still being written refuses
exactly like a path that will never exist. So an undelivered document is now a
durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`,
re-attempted on the next delta and on the channel's own quiet poll, and retired
with a matching `resolved` line naming the local copy once it arrives. The
cursor still advances and no delta stalls on one bad pointer.
Only a structured `report=data/....md` pointer now offers a document, so a path
merely mentioned in prose - including one under another home's mirror tree,
which is provably not that mate's to serve - is never fetched. Offers are
deduplicated across the whole delta, the escalation names each missing document
once and carries the reader's own reason instead of discarding it, and a
strictly increasing notice ordinal keeps a later escalation from being
swallowed as duplicate bytes. A mirrored line still lands once whichever
pointer form it was first written under.
* no-mistakes(review): Require structured pointer token boundaries
* no-mistakes(review): Unify boundary-safe pointer extraction and rewriting
* fix(bin): identify a mirrored line independently of its delivery state
Two defects in the boundary-safe pointer work.
The at-most-once check compared only the all-remote and all-local renderings
of a line, so it could not recognize a mixed one. A line offering two documents
where only the first was deliverable mirrored as local-plus-remote; once the
second arrived, a cursor-loss whole-log recapture rendered the same line
all-local, matched neither alternate, and mirrored a second time. A line's
identity is now the canonical form every boundary-valid pointer would take once
delivered, derived by the same parser that does extraction and rewriting, so it
no longer depends on which documents happened to be deliverable at the time.
The pointer map was passed to awk through the process environment. A delta may
carry up to the configured 1 MiB bound, and an expanded map of delivered
pointers can exceed the platform's exec argument limit, so awk would fail to
start; because no caller checked, the empty result would have been appended as
blank lines while the cursor advanced past dropped status content. The map now
travels in a file, and every call site checks the exit status and stops the
ingest rather than committing a delta it could not render.
Both passes now run once per stream instead of twice per line.
* no-mistakes(review): Abort ingest when document pointer extraction fails
* no-mistakes(review): Exclude structured cross-home pointers from document transfer
* fix(bin): fail open on an undeliverable remote document instead of tracking it
Narrow the remote-reply document fix to the scope the diagnosis actually
requires, as decided after measuring a simpler alternative.
A document the reader cannot deliver now fails open. The mate's line is
mirrored with its own pointer, the cursor advances, and one unkeyed note
carries the reader's reason. A note never enters the open-decision fold, so it
cannot stand open the way the original keyed block did - which removes the
never-clearing false blocker by construction rather than by resolving it.
That makes the durable self-clearing obligation unnecessary, so it goes: the
per-mate pending-documents record, its notice ordinal and resolved
announcements, and the poll-side retry. Canonical line identity goes too, and
with it a way to silently drop a genuine status line; mirroring is back to
at-most-once on exact bytes. The cross-home exclusion goes as well: under
fail-open a cross-home report= either fails harmlessly or is a nested remote
report this mate genuinely holds, which is now relayed again.
Kept: fetching only on a structured report= pointer, the boundary-correct
parser, the file-based rewrite map, and checked extraction and rewrite exit
status. The parser now scans behind a sentinel byte so a rejected candidate can
no longer give the text right after it a false leading boundary.
The reported incident is covered end to end: a report path announced in prose
before it exists raises no decision, and the report still arrives through the
ledger publisher's structured offer once written.
* no-mistakes(review): Preserve source-line identity across remote reply replays
* no-mistakes(document): Document remote reply transfer and replay semantics
* no-mistakes(lint): Fix staging truncation lint checks
* fix(calm): preserve substantive mid-turn responses (#4655)
* Preserve substantive Calm mid-turn text
* no-mistakes(review): Distinguish newline-preserved replies from short narration
* no-mistakes(document): Document Calm mid-turn preservation boundaries
* no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`
* fix(bin): preserve PR merge polls across volume remounts (#4656)
* fix(bin): re-record PR poll identity after a volume device renumber (Fixes #4260)
A volume remount can renumber the state filesystem's st_dev while every
inode and byte stays the same; APFS does this across a reboot. A poll
registration records its sidecar and check as device:inode, so every poll
armed before the remount failed strict validation and the watcher refused
all of them as unauthenticated state checks until each was re-armed by hand.
There are two device comparisons. fm_pr_private_file_valid compares a live
file's device with the state directory's device read in the same invocation:
it refuses a file that is not on the state directory's own filesystem and
already survives a renumber, so it is unchanged. The registration's recorded
identity versus the live identity (from #556, reused by the #932 retirement
receipt) binds the registration to the exact files published in its own
transaction; its device part is what breaks.
When strict capture fails, the watcher now proves the device is the only
difference: every other artifact check passes (template bytes, both hashes,
private mode, single link, live device, metadata), both recorded identities
name one device, and each recorded inode equals its live inode. Only then,
under the task's control lock, does it rewrite the two identity lines,
repeating the whole proof and comparing the registration's file identity and
bytes just before the rename, and then capture strictly again. A swapped,
altered, re-moded, relinked, split-device, or foreign-device artifact still
fails a proof and is still refused, and a pending retirement receipt blocks
the rewrite.
Reproduction: on macOS a poll armed on an APFS disk image that was detached
and re-attached behind another image moved st_dev 16777239 -> 16777243 with
inodes, bytes, mode, and link count unchanged; the real watcher refused it on
main and reports its merge with this change. The portable regression test
rewrites a real registration's recorded device and drives the watcher.
Not changed here: the status presentation cursor keys rows by its own
device:inode identity in bin/fm-classify-lib.sh, a different helper that
needs its own fix; a retirement receipt left by a reboot between its
publication and removal still names the old device and stays refused; custom
check trust binds only a content hash and is unaffected.
* fix(review): Serialize PR poll publication writers
* fix(review): Bound PR poll publication lock scope
* fix(bin): keep contribution records when the poll budget runs out (follow-up to #4627) (#4661)
A budget that expires partway through an observation no longer records an
error or prints the unavailable wake; the URL keeps its prior record and is
observed first next poll. forge() flags budget exhaustion at the point it
refuses, or when a read is killed at the budget's own deadline, so a genuine
forge failure still records the error and wakes. Each distinct URL is now
observed once per poll and applied to every owning task.
* fix(bin): clear parent pending-replies on local secondmate retirement (#4680)
* fix(bin): clear parent pending-replies on local secondmate retirement
Local secondmate teardown left resolved parent pending-reply records behind
after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced
retirement while any reply for that id is still unresolved, and delete every
matching record plus its delivery confirmation after a successful local or
remote retirement, matching the remote cleanup path.
* no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup
* no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung
* no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern
* no-mistakes(review): Pending-replies Basename und corr_id abgleichen
* no-mistakes(document): Clarify forced retirement pending-reply cleanup
---------
Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>
* fix(bin): accept Orca's composite worktree id when tearing down a task (#4677)
* fix(bin): accept Orca's composite worktree id at teardown
Teardown refused every Orca-backed task because the endpoint validator
checked orca_worktree_id with the simple-atom rule meant for tmux-style
window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca
returns that id as `<orca id>::<absolute worktree path>`, so the colon and
slashes in every real value made validation fail and finished Orca tasks
could never be cleaned up.
Validate the field as the composite it is: both halves of the first `::`
split present, the path half absolute, and no embedded newline, carriage
return, or tab. The terminal field keeps the atom check, which is correct
for it, and no other backend's validation changes.
The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca
never returns, which is why the suite passed a check the real value fails.
They now carry the composite form, so the tests exercise the real value.
* no-mistakes(document): name Orca's repo id in the composite worktree id
* no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points
* feat(bin): add opt-in typed dispatch resolution (#4692)
* feat(bin): add opt-in typed dispatch resolution through typesafe.ai
Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or
scout profile from a written brief with typesafe.ai's System One model:
one Choice question over the rules' `when` texts, then the confidence
floor, the rule's `approval` and `floor`, each profile's `provider` and
`floor`, one quota-axi snapshot, and the spendPriority argmax all in code.
It is off unless TYPESAFE_API_KEY is in the environment or the home's
gitignored .env; off means one stderr line, exit 0, and no network call,
so firstmate dispatches exactly as before. The key reaches curl on a file
descriptor, never argv.
Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and
the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new
tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates
the four new optional dispatch fields. Document the schema, the operator
contract, the AGENTS.md intake step, and the live and benchmark evidence.
* no-mistakes(review): Harden typed dispatch resolution and quota bounds
* no-mistakes(review): Validate dispatch floors and ranking evidence
* no-mistakes(review): Tighten dispatch response and floor evidence
* no-mistakes(review): Neutralize none matching and resolve defaults locally
* no-mistakes(review): Preserve providerless profiles outside typed resolution
* no-mistakes(review): Validate response usage and reject duplicate profiles
* no-mistakes(review): Escalate unverifiable floors and validate probabilities
* no-mistakes(review): Validate probability mass and unknown profile floors
* no-mistakes(review): Simplify resolver interface and preserve fallback routing
* no-mistakes(review): Fix constants and rank partial quota evidence
* no-mistakes(review): Add authoritative provider mapping and enforce explicit providers
* no-mistakes(review): Declare provider for documented Pi profile
* no-mistakes(review): Validate provider identifiers and support Gemini dispatch
* no-mistakes(review): Strictly anchor provider identifiers
* no-mistakes(review): Validate selectors and preserve fallback candidate evidence
* no-mistakes(review): Gate typed validation and harden resolver evidence
* no-mistakes(review): Preserve opt-in routing and harden candidate evidence
* no-mistakes(review): Prioritize known exhaustion over quota uncertainty
* no-mistakes(review): Isolate API secrets and preserve no-key diagnostics
* no-mistakes(review): Fallback safely when dispatch rules are absent
* no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets
* no-mistakes(document): Document typed dispatch safety and fallback behavior
* fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753)
* test: reproduce buried status declarations in shared readers
* fix: share status event reads and preserve open blockers
* fix: retain terminal scout and ship status declarations
* no-mistakes(review): Fix status chronology, legacy completions, and reader performance
* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots
* no-mistakes(review): Unify terminal supersession across cached folds and consumers
* no-mistakes(review): Filter per-key status history while preserving terminal chronology
* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells
* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses
* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor
* no-mistakes(lint): Quote literal done in test for-lists for SC1010
* ci: expect 19 snapshot/fleet-view tests
This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.
* no-mistakes(review): Restore multiline child outcome reporting
* no-mistakes(review): Select ledger terminal events through bounded shared reader
* no-mistakes(review): Report newest open decision instead of preferring blocked
* no-mistakes(review): Require colon before ship/scout terminal supersession in fold
* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix
* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions
* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold
* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule
* no-mistakes(document): Align status-read docs with fold-resolved crew state
* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers
* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree
* test: fold terminal-cleanup snapshot coverage into the completed-scout case
Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.
* no-mistakes(document): Clarify socket-down override expiry in architecture doc
* ci: retrigger flaky contribution check
* fix(bin): launch codex crewmates with codex's hook layer disabled (#4689)
* fix(spawn): launch codex crewmates with codex's hook layer disabled
A freshly launched Codex worker never reached its instructions. Codex
stopped it on an interactive "Hooks need review" modal whose selection
sits on "Review hooks", which is neither trusting nor declining.
Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow
navigation, so the selection cannot be moved, and pre-accepting the
prompt by writing Codex's own trust store would record an operator
consent that was never given.
The hooks are the machine's own ~/.codex/hooks.json plus any project's
.codex/hooks.json. A crewmate needs neither: its turn-end signal is the
-c notify= program on the same launch, and Firstmate's project hooks are
primary-session infrastructure that stands down in a child worktree.
Crewmate and scout launches now pass --disable hooks. That is the
opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted
hooks; disabling the feature runs none of them and leaves the operator's
~/.codex untouched. An unknown feature name is a hard Codex error, so a
release that drops the flag fails the launch loudly instead of silently
restoring the modal. A secondmate is a primary in its own home and keeps
the project hooks its turn-end guard and session-start digest ride on.
Verified on codex-cli 0.151.0: the modal is gone and the turn-end
notification still lands.
This unblocks the second review that every finished pull request is supposed to get.
Fixes kunchenguid/firstmate#4673
* no-mistakes(review): Fix contradictory hook count in Codex verification record
* fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710)
* fix(bin): settle terminal contributions and wake once per read-failure episode
A contribution whose last good observation is merged or closed is final:
poll no longer re-reads it, projection keeps it fresh, and a stale error
recorded beside it is cleared once. A genuine forge-read failure on an open
contribution still records its error on every cycle but prints the
unavailable wake only when it starts a failure episode; a successful read
ends the episode. Open PRs linked from done tasks keep being observed.
The false unavailable beside a complete observation was budget exhaustion
mid-observation, already fixed by #4661.
* fix(review): Settle terminal contribution owners
* fix(review): Deduplicate shared contribution failure episodes
* fix(test): Preserve settled terminal contribution records
* fix: select authoritative no-mistakes runs (#4476)
* fix(crew-state): select authoritative validation runs by identity
Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row.
Refs: https://github.com/kunchenguid/firstmate/issues/3215
* fix(review): Resolve same-branch run identities beyond capped history
* fix(review): Fix run-selection compatibility, races, and worker-state fallbacks
* fix(review): Limit run validation to the requested branch
* fix(test): Anchor AXI fixtures and document remaining live evidence gaps
* fix(document): Clarify run selection documentation and capture ownership
* fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation
* fix: distinguish captain outcomes from no-op updates (#4738)
* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work
MAIN answered a supervision-branch outcome for completed captain-requested
work (implementation done, PR ready for review and merge approval) with
"Captain, shipshape.", reading section 9's no-action reply as covering it
and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no
captain-facing response is owed".
Section 9 now limits the shipshape reply to true no-ops (idle re-read,
empty heartbeat, consequence-free acknowledgement) and requires a short
outcome response naming what finished and what word is needed whenever
requested work finishes or a result needs the captain's word, even when a
transcript entry already shows the substance. The Pi protocol's re-emit
rule now says it bounds repetition only, and carries a worked example of
the ready-for-review outcome whose correct processing turn a shipshape
reply fails.
No executable contract evaluates the content of MAIN's captain-facing
reply, so the regression is the protocol example in the owner doc rather
than a text-match test.
* no-mistakes(document): Clarify captain-facing outcomes versus no-ops
* docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line
The document step condensed the Pi protocol's re-emit rule and dropped the
worked example of a finished, ready-for-review outcome whose correct
processing turn a "Captain, shipshape." reply fails. That example is the
contract's regression: no executable contract evaluates the content of
MAIN's captain-facing reply, so the owner doc's example is the test case.
Restore it directly under the re-emit rule, prefixed as a regression
example that is kept verbatim and never condensed or summarized away.
* no-mistakes(review): Clarify captain outcome and decision-word requirements
* no-mistakes(document): Clarify captain-facing completion outcomes
* docs(pi): require the PR URL in the visible captain-facing outcome reply
Captain review on the regression example: drop the sample reply string
and say only that the ready-for-review outcome requires relaying a
captain-facing outcome response, not just "Captain, shipshape.".
Fold in the visible-PR-handoff failure seen this session: after the
branch outcome reporting this fix green, MAIN's visible reply was only
"Awaiting your merge call." with no PR URL, leaning on the dim anchor.
Section 9's URL rule now also covers a review or merge ask and names the
visible reply as where the URL goes, sourced from the ready status, pr=
metadata, or the supervision branch's summary and never left to a
transcript entry. The Pi protocol adds the same-way failure and places
the captain-facing text in the final visible assistant reply after the
fm_branch_processed call, because Calm hides assistant text emitted in
the same step as a tool call as a working note.
Investigation verdict, evidence in the PR comment: no recent PR caused
the handoff failure; Pi has hidden same-step pre-tool assistant text
since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and
#4658 touched only remote report transfer.
* no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs
* no-mistakes(document): Clarify captain-facing supervision outcomes
* docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule
The consolidated section 9 URL rule narrowed its trigger to a review or
merge ask, dropping the "whenever a PR is mentioned" catch-all from
#3648 that keeps every PR URL copied from a durable record and never
assembled from memory. Restore that trigger as a union with the review
or merge ask so the one consolidated rule covers both.
* fix(bin): let non-owner Claude Stops exit safely (#4777)
* Fix foreign-owner turn-end supervision loop
* no-mistakes(review): Scope foreign-owner safe exit to Claude guard
* no-mistakes(document): Document Claude foreign-owner safe exit
* fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778)
Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty
indexed array as an unbound variable and aborts the shell. In
signal_turnend_panes_churned() the missing_keys loop was reachable with
an empty array whenever every churned key already held a fresh
.churn-since-* marker (a second churning turn-end inside an open
deferral window), so each watcher cycle died about half a minute in and
supervision restarted endlessly. The created_keys rollback loops had the
same latent crash on their error paths.
Audit of bin/ for the same pattern found one more confirmed-reachable
case: remote_handoff's noncanonical-body scan iterates to_move, which is
empty when a retried remote handoff finds every key already staged in
the outbox. All other "${arr[@]}" sites are either count-guarded,
guaranteed non-empty by construction, or unreachable while empty.
Guard the three reachable expansions with the repo's existing
"${arr[@]+...}" idiom. New regression test drives a real watcher
through the all-marked churn path; the macos-stock-bash CI lane runs it
under real /bin/bash 3.2 via FM_TEST_ONLY.
* Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783)
The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: require complete captain-facing final responses (#4779)
* docs: require complete final responses across harnesses
* no-mistakes(document): Document complete final replies for Grok Bot
* docs: point Grok replies to the shared contract owner
* no-mistakes(review): Clarify final recap without batching decision asks
* fix: preserve substantive mid-turn text in Pi Calm (#4788)
* fix(calm): preserve substantive Pi mid-turn text
* no-mistakes(review): Preserve substantive Pi Calm text per block
* no-mistakes(test): Cover shared Calm preservation boundaries behaviorally
* no-mistakes(document): Consolidate Calm preservation documentation
* no-mistakes(ci): Fixed Lint SC2034 in tests/fm-inactive-reconcile.test.sh (unused attempt → _) and tests/fm-spawn-pool-base-freshen.test.sh (shared-record LOCAL_ONLY_TIP disable). Fixed Behavior portable serial 4 by aligning tests/fm-wake-daemon-lifecycle-e2e.test.sh with classify_stale's post-c6e3cc1 contract: recorded meta + FM_FAKE_CREW_STATE working evidence so transient stale self-handles. Verified locally: lifecycle e2e passes; shellcheck -x on the three files exits 0. PR must be raised via no-mistakes is an attestation/pipeline check, not a code defect—no code change for it
* no-mistakes(document): Document Claude foreign-owner session-lock consumers
---------
Co-authored-by: Pablo Ontiveros <pablo.ontiveros@gmail.com>
Co-authored-by: Umer <umeranjum17@gmail.com>
Co-authored-by: Yasuhito Takamiya <yasuhito@hey.com>
Co-authored-by: Marsjohn-11 <74795701+Marsjohn-11@users.noreply.github.com>
Co-authored-by: tbillings28 <todd@toddbillings.com>
Co-authored-by: Todd Billings <todd@usdvcapital.com>
Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: Tiago <tiagop@hey.com>
Co-authored-by: Amin Roudaki <roudaky@gmail.com>
Co-authored-by: Joseph Kim <jokim1@gmail.com>
Co-authored-by: Mickaël Rémond <mremond@process-one.net>
Co-authored-by: Sebastian <80847374+thelad-dev@users.noreply.github.com>
Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>
Co-authored-by: Juan José González Giraldo <juanjose.eng@gmail.com>
Co-authored-by: Cody <72239807+codyjohnsontx@users.noreply.github.com>
Co-authored-by: Pedro Guimarães <21346846+0x7067@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
moughamir
added a commit
to moughamir/firstmate
that referenced
this pull request
Sep 18, 2026
* feat(hermes): add Hermes Agent as verified primary harness Integrate Hermes Agent as a firstmate primary harness with a three-layer bridge: session-start digest injection via plugin hooks, turn-end guard to prevent blind stops during fleet work, and a background watcher for zero-token event-driven supervision. Includes the Hermes session-provider backend adapter (bin/backends/hermes.sh), the firstmate Hermes plugin (bin/hermes-plugin/firstmate/), the plugin installer, supervision protocol doc, harness-adapters skill reference, and verification evidence record. The adapter is EXPERIMENTAL with no dedicated real-Hermes CI lane yet; the portable regression pins the PID-file state machine and dispatch wiring. * fix: select authoritative no-mistakes runs (kunchenguid#4476) * fix(crew-state): select authoritative validation runs by identity Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row. Refs: kunchenguid#3215 * fix(review): Resolve same-branch run identities beyond capped history * fix(review): Fix run-selection compatibility, races, and worker-state fallbacks * fix(review): Limit run validation to the requested branch * fix(test): Anchor AXI fixtures and document remaining live evidence gaps * fix(document): Clarify run selection documentation and capture ownership * fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation * fix: distinguish captain outcomes from no-op updates (kunchenguid#4738) * fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work MAIN answered a supervision-branch outcome for completed captain-requested work (implementation done, PR ready for review and merge approval) with "Captain, shipshape.", reading section 9's no-action reply as covering it and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no captain-facing response is owed". Section 9 now limits the shipshape reply to true no-ops (idle re-read, empty heartbeat, consequence-free acknowledgement) and requires a short outcome response naming what finished and what word is needed whenever requested work finishes or a result needs the captain's word, even when a transcript entry already shows the substance. The Pi protocol's re-emit rule now says it bounds repetition only, and carries a worked example of the ready-for-review outcome whose correct processing turn a shipshape reply fails. No executable contract evaluates the content of MAIN's captain-facing reply, so the regression is the protocol example in the owner doc rather than a text-match test. * no-mistakes(document): Clarify captain-facing outcomes versus no-ops * docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line The document step condensed the Pi protocol's re-emit rule and dropped the worked example of a finished, ready-for-review outcome whose correct processing turn a "Captain, shipshape." reply fails. That example is the contract's regression: no executable contract evaluates the content of MAIN's captain-facing reply, so the owner doc's example is the test case. Restore it directly under the re-emit rule, prefixed as a regression example that is kept verbatim and never condensed or summarized away. * no-mistakes(review): Clarify captain outcome and decision-word requirements * no-mistakes(document): Clarify captain-facing completion outcomes * docs(pi): require the PR URL in the visible captain-facing outcome reply Captain review on the regression example: drop the sample reply string and say only that the ready-for-review outcome requires relaying a captain-facing outcome response, not just "Captain, shipshape.". Fold in the visible-PR-handoff failure seen this session: after the branch outcome reporting this fix green, MAIN's visible reply was only "Awaiting your merge call." with no PR URL, leaning on the dim anchor. Section 9's URL rule now also covers a review or merge ask and names the visible reply as where the URL goes, sourced from the ready status, pr= metadata, or the supervision branch's summary and never left to a transcript entry. The Pi protocol adds the same-way failure and places the captain-facing text in the final visible assistant reply after the fm_branch_processed call, because Calm hides assistant text emitted in the same step as a tool call as a working note. Investigation verdict, evidence in the PR comment: no recent PR caused the handoff failure; Pi has hidden same-step pre-tool assistant text since kunchenguid#2339 (2026-08-13), kunchenguid#4655 changed only the Claude Code mod, and kunchenguid#4658 touched only remote report transfer. * no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs * no-mistakes(document): Clarify captain-facing supervision outcomes * docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule The consolidated section 9 URL rule narrowed its trigger to a review or merge ask, dropping the "whenever a PR is mentioned" catch-all from kunchenguid#3648 that keeps every PR URL copied from a durable record and never assembled from memory. Restore that trigger as a union with the review or merge ask so the one consolidated rule covers both. * fix(bin): let non-owner Claude Stops exit safely (kunchenguid#4777) * Fix foreign-owner turn-end supervision loop * no-mistakes(review): Scope foreign-owner safe exit to Claude guard * no-mistakes(document): Document Claude foreign-owner safe exit * fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (kunchenguid#4778) Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty indexed array as an unbound variable and aborts the shell. In signal_turnend_panes_churned() the missing_keys loop was reachable with an empty array whenever every churned key already held a fresh .churn-since-* marker (a second churning turn-end inside an open deferral window), so each watcher cycle died about half a minute in and supervision restarted endlessly. The created_keys rollback loops had the same latent crash on their error paths. Audit of bin/ for the same pattern found one more confirmed-reachable case: remote_handoff's noncanonical-body scan iterates to_move, which is empty when a retried remote handoff finds every key already staged in the outbox. All other "${arr[@]}" sites are either count-guarded, guaranteed non-empty by construction, or unreachable while empty. Guard the three reachable expansions with the repo's existing "${arr[@]+...}" idiom. New regression test drives a real watcher through the all-marked churn path; the macos-stock-bash CI lane runs it under real /bin/bash 3.2 via FM_TEST_ONLY. * Make the foreign-owner turn-end repro create a Linux-readable session lock. (kunchenguid#4783) The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: require complete captain-facing final responses (kunchenguid#4779) * docs: require complete final responses across harnesses * no-mistakes(document): Document complete final replies for Grok Bot * docs: point Grok replies to the shared contract owner * no-mistakes(review): Clarify final recap without batching decision asks * fix: preserve substantive mid-turn text in Pi Calm (kunchenguid#4788) * fix(calm): preserve substantive Pi mid-turn text * no-mistakes(review): Preserve substantive Pi Calm text per block * no-mistakes(test): Cover shared Calm preservation boundaries behaviorally * no-mistakes(document): Consolidate Calm preservation documentation * fix: harden mail checks and rebalance full-coverage CI (kunchenguid#4800) * Improve CI reliability and rebalance full-coverage validation * no-mistakes(document): Clarify lint partition documentation * fix(bin): answer Kimi 2.0.0 folder-trust dialog during spawn (kunchenguid#4799) * Handle Kimi workspace trust dialog * no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers * no-mistakes(review): Gate Kimi ready on any trust marker and clean captures * no-mistakes(review): Read visible pane for Kimi trust and ready gates * no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate * no-mistakes(review): Harden Kimi viewport capture and trust dialog detection * no-mistakes(document): Document Kimi spawn refusal on cmux and Orca * fix(bin): report a dead-agent record once instead of escalating forever (kunchenguid#4775) * fix(bin): report a record whose agent is gone once instead of escalating forever The wedge escalation path never asked whether there was still an agent to be wedged. A wedge is something stuck that might recover, so re-alarming it earns its cost; an agent that is gone never moves again, its pane never churns, the idle timer never resets, and the escalate path clears its own timer and re-arms with nothing bounding the count. Observed on a live fleet: two finished lanes reached 226 and 203 consecutive escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400 notifications a day from two lanes with no agent running at all. On one, fm-control.sh exit answered already-stopped and fm-crew-state.sh read "failed - run failed". Closing the Herdr pane did not stop it either: with the pane genuinely gone and herdr pane read returning pane_not_found, the count kept climbing, because the poll is driven by the record's window= line rather than by the pane. The cost is not the repetition but that it drowns the alarms that matter. fm_backend_agent_state already separates a thinking agent from a gone one at process level. In the branch that was about to escalate, read it once and treat only its two recovery-grade verdicts - dead (endpoint present, no agent in it) and missing (endpoint authoritatively absent) - as proof, reporting that record once and not re-escalating it while it stays that way. Every other verdict, including alive, ambiguous, unreadable, unverified, and a read that failed outright, keeps the identical schedule, reason, and escalation count, so a genuinely wedged live agent is unaffected. The probe costs at most one backend read per window per threshold, the same budget the declared-wait consult and the worktree write probe already take. The report decides nothing about the record's fate: both lanes still held unlanded work and teardown refusing them was correct, so retiring, relaunching, or cleaning up stays with the supervisor. The once-only marker is owned entirely by that function and is dropped by the same read the moment the endpoint stops reading gone, so a replacement launched into the same window escalates normally and its own later death is reported again. Related, and not closed by this: kunchenguid#4412, kunchenguid#4482, kunchenguid#4316. Tests drive the real watcher against a record whose endpoint does not exist and pin both directions: dead and missing report once and never advance the count across later thresholds, while alive, ambiguous, and unreadable endpoints keep escalating with the identical reason and a climbing count. * fix(bin): bind the once-only dead report to the pane it reported Review of the parent commit found a reachable sequence where a later death in the same window lost its promised report. The marker was keyed on the verdict string alone and dropped only when a threshold probe read a non-gone verdict, but probes run only at thresholds: a replacement launched into the same window that dies without ever being probed alive - it crashes at startup, or works and then crashes - was absorbed by the previous death's marker. The pane's first sight yielded only the generic stale wake and every later threshold matched the stale marker, so the second death never got the detailed once-report that both the function's own comment and docs/architecture.md promise. Record the verdict together with the pane hash it was reported for, and absorb a repeat only while both still match. A replacement churns the pane, which resets the stale suppressor, wedge timer, and escalation count while no reset site touches this marker, so the pane half is what tells the second death apart from the first. The live-probe drop stays as it was. Clearing the marker at those reset sites instead would re-open unbounded re-alarming for a dead pane whose display ever ticks, which is the exact defect the parent commit exists to close. The noise bound is unchanged: an unchanged dead pane still absorbs on every later threshold and never advances the escalation count, and every verdict short of proof still escalates exactly as before. * no-mistakes(review): Key the dead-record once-marker on the busy incarnation token * no-mistakes(document): Document dead-record escalation cap in stale-pane config entry * no-mistakes(document): Add busy-state inventory line to AGENTS.md * no-mistakes(document): Document dead-record probe on busy-turn-bound wedge path --------- Co-authored-by: Mickaël Rémond <mremond@process-one.net> Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Co-authored-by: Pedro Guimarães <21346846+0x7067@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: ShaDev <shazellb@gmail.com> Co-authored-by: Umer <umeranjum17@gmail.com>
sanis
added a commit
to sanis/firstmate
that referenced
this pull request
Sep 19, 2026
* fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (#4424)
* fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (#42)
* fix(pr-merge): read a plan-gated 403 on branch rules as no merge queue
github_read_queue_method left status=unreadable for every failed rules
read, including a 403 whose body is GitHub's own "Upgrade to GitHub
Pro or make this repository public" message. A repository whose plan
cannot expose branch rules cannot have a merge_queue rule either, so
that specific 403 now resolves to status=none instead of unreadable -
unblocking the away-merge grant on private repos without GitHub Pro.
Any other failure (auth, rate limit, network, 404, unrelated 403)
still reads as unreadable.
* no-mistakes(document): Update stale away-merge queue-grant comment for plan-gated 403
---------
Co-authored-by: NewAiCoder <claude@theinbtw.com>
* no-mistakes(review): Fix misleading away-queue-grant comment in fm-pr-merge and its test
* no-mistakes(document): Update architecture.md for plan-gated-403 merge queue exception
---------
Co-authored-by: NewAiCoder <claude@theinbtw.com>
* fix(bin): select suites that read a changed top-level test fixture (#4246)
* fix(tests): select readers of a changed top-level test fixture
bin/fm-test-run.sh --changed recognised shared test helpers by an explicit
list, tests/lib.sh|tests/*-helpers.sh|tests/fixtures.sh. A top-level
tests/*-fixture.sh matched none of those, fell through to the tests/*
catch-all, and was marked unmapped, so selection aborted with "no
changed-test mapping for source path" and the run selected nothing at all.
tests/herdr-client-pair-fixture.sh and tests/remote-herdr-fixture.sh are
real shared fixtures with real consumers, so any branch touching one of
them left a validation pipeline driving --changed with a hard abort rather
than a narrowed selection.
Extend the helper arm to tests/*-fixture.sh rather than routing it through
the tests/fixtures/*/* arm. Both arms resolve consumers with the same
reference scan, and that scan is what selects the right suites here: it
finds exactly the tests that read the fixture. The fixtures/ arm adds only
a directory-keying step, which has nothing to key on for a top-level file,
so the helper arm is the same behaviour with no extra machinery. A
tests/ path nothing reads still reaches the catch-all and still refuses
loudly.
Refs https://github.com/kunchenguid/firstmate/issues/4100
* no-mistakes(test): order nested fixtures arm before top-level fixture glob
* no-mistakes(document): document tests/ shared-file mapping contract and arm order
* no-mistakes(review): drop vacuous test phase, correct header claim, restore comment
* fix(bin): treat Claude Code's default external-imports flags as never asked, not declined (#4387)
* fix(bin): read Claude Code's default external-imports flags as never asked, not declined (#4378)
fm-claude-trust.sh refused the whole trust registration whenever the project-root entry
carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code
writes that value only on an explicit "No, disable". Claude Code's default project
entry carries Approved and WarningShown both false before the dialog is ever shown, so
every such project refused every spawn.
Only Approved === false with WarningShown === true — the pair the dialog writes on a
decline — now counts as a decline. false/false behaves like an absent flag: trust is
registered and no import consent is manufactured.
New case test_project_root_entry_default_import_flags_are_not_a_decline fails on
b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31,
bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* no-mistakes(review): Correct harness doc's external-imports decline predicate
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(bin): keep operator-address labels out of no-mistakes intent (#4445)
* fix(brief): keep operator address out of composed intent
Teach raw-word authoring for intent sections and mid-task relays, with a neutral [captain] provenance marker for legacy mixed tasks. Keep headings and contract prose outside the serialized intent body.
The legacy selector already excluded the old speaker labels from its output; preserve that read compatibility. The reproduced leak comes from adding labels inside a modern intent body, not from the legacy selector. Do not scrub actual request content.
Add exact serialized-input and generated-contract regressions, retaining refusal of unmarked legacy tasks and coverage of scout promotion.
Fixes https://github.com/kunchenguid/firstmate/issues/3882
* no-mistakes(review): Refuse operator-address lines in Captain's intent body
* no-mistakes(document): Document operator-address refusal in intent contract comments
* fix: classify OpenCode ellipsis hint as idle (#4451)
* fix(composer): recognize Grok 1.0.5's oversized titled bottom border as a proven empty composer (#4455)
* fix(composer): accept Grok title overhang
* no-mistakes(review): summary: named Grok overhang constant, doc caveat, restored tmux typed-title coverage
* fix(bin): translate Stop hook timeout signals into durable auto-arm failure (#4474)
* fix(bin): recover Claude auto-arm after timeout
* no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list
* fix(spawn): establish Claude task channel authority (#4464)
* fix(spawn): establish Claude task channel authority
* no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference
* fix(bin): refuse fm-control.sh exit when the composer holds unproven or pending text (#4458)
* fix: guard relaunch exit against pending input
* no-mistakes(review): Verifying test run in progress
* no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard
* no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now
* fix(spawn): establish crewmate identity first (#4481)
* fix(bin): reconcile redundant secondmate divergence during updates (#4460)
* fix: reconcile diverged secondmate updates
* no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md
* no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md
* feat: enable gpt-5.6-luna max reasoning for crew dispatch (#4497)
* fix(dispatch): support Codex Luna max effort
* no-mistakes(review): use portable CODEX_HOME path in codex effort reference
* feat(calm): render smooth Unicode swell with asymmetric two-color sail (#4498)
* feat(calm): render smooth Unicode swell
* feat(calm): make sails asymmetric
* feat(calm): use quarter sail glyph
* no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer
* no-mistakes(document): docs: sync calm wave phase doc comment
* no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。
* fix(bin): supersede stale scout delivery text in brief.md on promotion (#4491)
* fix: supersede scout delivery brief on promotion
* fix: preserve ship safety contract after promotion
* no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch
* fix(bin): make captain holds work on hosts with an older JSON::PP, and stop cleanup dropping accents from a held body (#4471)
* fix(bin): let captain holds work on hosts with an older JSON::PP
Holding a task for the captain, and the cleanup that keeps a captain-held row
open, both fail outright on any host whose JSON::PP defaults allow_nonref off -
2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`,
but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older
library rejects that whole value with "must be object or array".
The consequence is fleet-wide on such a host, not one broken command: a worker
there cannot formally record a decision for the captain at all. It can only
mention the decision in passing in a status line, where it can be missed - which
is how a real decision goes unrecorded. The hold reports that the task lost its
hold-set stamp; the cleanup cannot return the row to Queued.
Both call sites now ask for allow_nonref explicitly rather than inheriting
whatever the installed library defaults to. The second one is worth naming: its
`/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string
case that fails, so the guard selects for the failing input rather than
protecting against it.
The regression case forces the older default back off for every perl the commands
spawn, then drives both paths - holding a task that carries a body, and tearing
down a captain-held row whose deliverable must still be appended. It also probes
that the simulation genuinely rejects a bare scalar, so the case cannot pass
vacuously on a lenient host. Each half was verified failing on its own unfixed
call site with that site's real error message. Suites: fm-captain-hold-lifecycle
51 cases, fm-backlog-atomicity 99 cases, 0 failures.
Verification limit: the mechanism is reproduced and tested, but neither fix is
verified against a real JSON::PP 2.27202 host, because none is in the loop. This
laptop runs 4.06, where the bug does not manifest.
`bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a
brace-delimited object before decoding, so allow_nonref never applies.
* fix(bin): stop cleanup silently dropping accented characters from a held body
Cleanup rewrites a captain-held row's body to append the finished work's
deliverable, and the decoder it reads that body with printed decoded characters
to a stream with no `:raw` layer. A character at or below U+00FF then came out
as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the
accent. `fm_backlog_retain` writes that body straight back through
`--body-file`, and nothing reported an error - the character was simply gone
from a row still waiting on the captain.
The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus
`utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already
used.
Review of the parent commit found this on one of the lines that commit already
changed. It predates that change.
The test asserts bytes rather than decoded strings, because comparing strings
cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any
character above U+00FF makes perl print the whole string as UTF-8, so one body
carrying both an accent and an em dash passes even unfixed and proves nothing.
Verified failing before the fix on the accented row, passing after. Suites:
fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures.
* no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc
* no-mistakes(review): drop whole-file UTF-8 check from retained-body test
* no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc
* no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc
* fix(bin): read codex 0.154's idle braille starfield rows as composer furniture (#4532)
* fix(composer): read codex 0.154's idle starfield and status footer as furniture
codex-cli 0.154.0 animates a braille "starfield" around its idle composer:
on the row above the bold `›` prompt row, on the `›` row behind the SGR-2
dim `Ask Codex to do anything` placeholder, and on the row below it, then
draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`).
The cells are truecolor greys on both sides of the ghost luminance ceiling,
so the brighter ones survive ghost stripping, and the rows below the glyph
carry no structural edge. The shared classifier selected the bare `›` shape,
extended its wrap region over the two rows beneath the glyph, read the
survivors and the footer as wrapped typed input, and answered `pending`;
the steering doorbell defers on exactly that verdict, so no doorbell ever
reached an idle codex 0.154 pane.
bin/fm-composer-lib.sh now recognises that furniture by shape, declared
once next to the idle placeholders and reached from the two wrap-region
boundary points:
- a row whose non-whitespace content is entirely braille cells
(U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it
never counts as wrapped typed content and bounds a bare composer's wrap
region; braille behind the glyph row's content is stripped before the
emptiness decision when nothing else follows the glyph; a row mixing
braille with other text stays typed content;
- the codex status footer bounds the wrap region exactly as omp's status
row does, anchored on the effort token, a spaced middle dot, and a `~` or
`/` path cell, so a typed `fix · tests` stays composer input;
- `^Ask Codex to do anything$` joins the verified idle-placeholder set; the
ghost strip remains what proves that row empty, and the bare-row rule that
bright placeholder text is real input is unchanged.
Unchanged: the strict blank-row rule, the styled=0 degradation (a plain
cmux/orca capture of this screen still reads `unknown`, never `pending`),
FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape.
tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte
with the divergence (letters in place of the starfield read `pending`) and
the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh
is the default-on live guard (token-free, skips explicitly without codex or
tmux) that launches the installed codex idle and asserts `empty` through
both the tmux and the cursorless styled reads, naming codex --version on
failure. docs/verification/runtime-backends.md records the dated Herdr
evidence: `pending` before, `empty` after, on the captured screen.
* no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry
---------
Co-authored-by: Todd Billings <todd@usdvcapital.com>
* fix(bin): refuse empty text steers in fm-send (#4259)
* fix(bin): refuse empty text steers in fm-send
A marked secondmate request sent with an empty message delivered only
marker and correlation bytes and minted a pending-reply expectation the
parent could never see resolved, stalling the fleet with no loud error
(#4255). Fail closed on an empty or whitespace-only message on the text
path, mirroring the existing --resolve-key refusal.
* chore: retain ambient Pi-lens autoformat as its own commit
Formatting-only edits produced by ambient Pi-lens autoformat during the
msg-loss investigation, kept separate from the behavioural change in
c23acba6 so the fix stays reviewable on its own.
AGENTS.md is deliberately excluded: its only autoformat edit stripped the
trailing space from the documented FM_OPERATIONAL_PREFIX value, which
bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11
records as permanent compatibility. Documenting that constant without its
trailing space makes the doc wrong about the contract, so that one line was
restored rather than retained.
* fix(calm): paint the working ship one yellow over all-blue water (#4554)
On rose-pine-moon the two-color water (cyan crests over blue troughs) read as
a pink stripe over aqua, the yellow left sail and mast clashed with the red
right sail, and the hull carried a blue interior run. Every water cell is now
blue so the swell reads through glyph height alone, and both sail halves, the
mast, and the whole hull are one yellow run. Geometry, cadence, animation,
direction flip, resize clamping, and the narrow fallback are unchanged.
Update the unit and real-TUI color assertions to the new palette and the Calm
docs that described the old one.
* fix(bin): stop aging a second mate's active turn from its launch (#4270)
* fix(watch): stop aging a second mate's active turn from its launch
The parent watcher's second-mate wake-loop stall check exempts a mate that
is demonstrably inside an active turn, but secondmate_in_active_turn asked
busy_turn_over_age first and returned "not in a turn" whenever that said
the bound was crossed.
busy_turn_over_age ages from state/<task>.turn-ended, falling back to
state/<task>.meta. A second mate's turns end in its own home, so the
parent never gets a turn-ended mark for it and the fallback ages the
mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago
was therefore permanently "over age", the busy pane was never consulted,
and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false
wake-loop stall.
The gate now bounds the busy exemption by <idle> - how long the queue's
drain position has not moved - which is evidence this home actually
holds. A busy mate stays exempt while the queue has been frozen for less
than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so
the bound that stops a busy pane from proving liveness forever is kept
rather than removed. busy_turn_over_age is untouched; its remaining
callers are the ordinary crew busy-pane bound.
The regression pins the case that actually broke: a mate whose launch
record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn
must not escalate, while the same mate with its queue frozen past the
bound still publishes exactly one notification. The existing coverage
only exercised a freshly launched mate, which passes either way.
Reaching that alert now costs a pane capture inside the gate, so the
three checkpoints in this suite that assert an alert move from a 1s to a
4s bound - the value the neighbouring active-turn cases already use. The
bound is a ceiling, not a wait: the checkpoint returns on the first
actionable wake. On a loaded machine a 1s bound missed the alert
repeatedly; at 4s it did not miss in 20 runs under the same load.
* no-mistakes(review): scope the second-mate active-turn regression test's coverage claim
* no-mistakes(document): fix stale second-mate active-turn comments in fm-watch
* feat(bin): add read-only PR blocker and reviewer discovery commands (#4278)
* feat(bin): add read-only PR blocker and reviewer-discovery commands
Two focused, opt-in commands that read GitHub and never write to it.
fm-pr-state.sh reports what still blocks one pull request from the
author's side: a closed or merged state, draft state, unknown or
conflicting mergeability, absent or failing required checks, and a
blocking CHANGES_REQUESTED decision explained by each reviewer's latest
verdict, marked STALE when it was left at a superseded head. A pull
request that only awaits an approval is not reported as blocked, and
advisory checks are omitted. Every reading is taken against one exact
head; a push that lands mid-read invalidates the whole result rather
than mixing two snapshots.
fm-pr-reviewers.sh suggests reviewers from the most recent commits to
the pull request's exact changed paths, counting each commit once,
resolving handles through GitHub's own commit author.login mapping, and
excluding the author and Bot accounts.
Both stay read-only: no review request, no approval, no merge.
Unresolved review-thread state is left unreported because the REST API
does not expose it and unattended commands may not use GraphQL.
Closes #3731
* no-mistakes(review): accept only PR URLs and stop at terminal state
* no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate
* no-mistakes(review): stop attributing readings to unverified heads
* no-mistakes(review): narrow readiness contract to checks that have reported
* no-mistakes(review): read the pull request once, drop the head guard
* no-mistakes(document): scope pr-forge isolation proof to its measured members
* no-mistakes(document): record uncovered pr-forge members and their pending proof
* docs(isolation-proof): re-prove pr-forge at its full membership
tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the
pr-forge family in this branch, and script_allows_concurrency grants
four workers by family membership alone, so both ran concurrently on a
proof measured before they existed.
Re-proved the family at all eight members: two consecutive runs, 0
failures, each begun with the one-minute load average below 6.0 so the
result measures isolation rather than contention. A third run taken
between them is disclosed rather than recorded, because it started
while the previous run's workers were still decaying.
The new durations are not comparable with the six-member measurement
above them, so they are not presented as evidence about the two new
members, and that record's 1.72x four-worker figure is left as a
statement about its own run rather than restated as current.
* no-mistakes(review): disclose gh error-text coupling at its matching site and tests
* fix(bin): teach validation-round pauses in generated briefs (#2752)
* fix(bin): teach validation-round pauses in briefs
* no-mistakes(document): Point classifier comments to authoritative pause examples
* docs(readme): add star history chart (#4558)
* fix(bin): refuse teardown when a task's endpoint close fails (#4510)
* fix(teardown): refuse a cleanup whose endpoint close failed
bin/fm-teardown.sh discarded both the exit status and the stderr of every
fm_backend_kill call, so a close that genuinely failed was indistinguishable
from one that succeeded. Teardown continued past it, deleted the task's durable
records, returned its worktree, and reported the cleanup as completed. The
deleted metadata is the only record of which endpoint belongs to the task, so
such a close did not merely leave a stray session behind, it stranded one:
nothing was left on disk naming it.
The adapters could not carry that signal either. Driven against the real code,
every backend arm returned 0 for a genuine failure exactly as it did for an
already-exited endpoint, so there was nothing for the four call sites to
propagate even once they stopped swallowing it.
The tmux arm now resolves a close that did not succeed against the window's
exact recorded identity, since kill-window fails the same way for a window that
is gone and one that is still there. The Orca arm reports a close its missing
CLI never attempted. Both stay silent for an endpoint that is already
legitimately gone, and the remaining arms are unchanged: their close-command
timing cannot be established without the real Zellij, Orca, and cmux binaries,
and a gate that refused ordinary cleanup of an already-exited session would be
worse than the defect. docs/verification/runtime-backends.md records what each
backend can prove.
A reported close failure now reaches teardown's existing retain-and-stop
refusal before the records naming the endpoint are removed, matching where the
Herdr confirmed-gone gates already sit for the same hazard, and the retained
records let a rerun finish once the close works.
* no-mistakes(review): refuse unreadable tmux close re-read; honor --force override
* no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close
* no-mistakes(document): document endpoint-close refusal in its backend and retirement owners
* no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree
* feat(calm): add flag-gated Claude Code Calm mode (#4565)
* feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag
Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that
brings Calm to Claude Code: the sailboat replaces the stock working row through a
Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration,
and canonically classified operational user rows draw at zero height. /calm is
registered by the hooks module itself and toggles the same per-home config/calm
preference the Pi extension uses, so one choice applies on either harness; rows
redraw retroactively on toggle and stay hidden across claude --continue.
The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS
flag is on. Nothing sets that flag in any settings file, and the plugin carries no
command file, skill, agent, or classic hook, so it is a complete no-op while the
flag is off. The trusted project auto-loads it through an .agents/skills symlink,
the only path Claude Code scans for project plugins.
Extract the working-ship geometry, bounce track, cadences, and freeze/resume state
into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module
imports from outside the plugin folder) and have the Pi widget paint that core's
frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify
operational rows through a port of bin/fm-operational-input.sh's classify command
guarded by a corpus parity test against the shell owner.
Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering,
Raster packing, policy, classifier parity), the mod's own claude plugin test suites
behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op,
the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272.
Docs: record the version-scoped Claude Code evidence and the three bounded gaps in
docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md,
and make the shared preference, layout, and contributor notes harness-neutral.
* no-mistakes(review): Preserve colliding final replies and strengthen parser parity
* no-mistakes(review): Preserve final replies and strengthen canonical parity checks
* no-mistakes(review): Require exact function-hooks opt-in before Calm activation
* no-mistakes(review): Clarify Calm module loading and activation boundaries
* no-mistakes(review): Reset Calm presentation state across session starts
* no-mistakes(document): Refresh Calm session lifecycle documentation
* feat(calm): paint the Claude Code working ship in Claude's own theme colors
The captain picked the "Claude native" palette for the Claude Code mod's Raster:
every water cell takes the spinner blue of the active theme family (#93a5ff dark,
#5769f7 light) and the whole boat takes the Claude orange of the stock spinner
(#d77757), one water color and one boat color. The family follows the `theme`
setting's prefix, read at load through $.config.list and re-read on a
config.set of that row, with `auto` and custom themes falling back to the dark
set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte.
Rename the shared sprite's color classes from hue names to `water` and `boat`,
since each harness now maps them to its own colors; geometry, motion, cadence,
and the activation gate are untouched.
Tests cover both palettes' packing and the family rule under Node, and the
plugin kit drives every theme value, a theme change mid-session, the Calm-off
pass-through, and inertness of the menu read while the flag is off. The docs
describe the Claude Code colors and record the guard passing on 2.1.273.
* no-mistakes(review): Use light palette for unresolved Claude themes
* no-mistakes(document): Refresh Claude Calm verification evidence
* fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586)
* fix(watch): honour a declared wait before wedge-escalating a quiet pane
wedge_timer_check escalated on elapsed idle time alone. Nothing asked
whether the worker had already said why its pane was quiet, so a lane
that declared a bounded external wait climbed the escalation ladder for
as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every
repeat carried demand-deep-inspection - which by its own wording forbids
re-absorbing on the run-step or pane state, so the supervisor could not
use the evidence that was there either.
The generated brief promises that declaring `paused:` buys the long
recheck cadence instead of a wedge, but the timer was still reachable
while that declaration stood: a crew that declares a wait and then has an
active run or busy pane attributed to it is handed to the timer as
provably-working. The declaration is what the worker said about its own
silence, so it now outranks a liveness verdict that only says something
is running.
The consult runs in the at-threshold branch that was about to escalate,
beside the worktree walk already there, and costs one status-line read.
Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS
recheck the declared-wait absorber already uses, so the wait is still
rechecked and cannot rot invisibly. Which verb declared it decides the
wording, because the two block on different people: a `paused:` wait is
owed by an external dependency and asks the reader to confirm it still
holds, while a `captain-held:` transfer is owed by the captain reading
the recheck and asks them to answer or release the hold. A hold is not
rechecked at all while the away-posture record exists, as on every other
captain-held path, and that absorb arms no throttle so the recheck is
owed in full on return.
A declared clearing time that has already passed stops counting, and a
lane that never declared one keeps the identical escalation schedule,
reason, count and demand-deep-inspection wording, so detection and its
worst-case time are unchanged. The deferral restarts the idle timer
rather than cancelling it, so a lane that stops waiting escalates again
within one threshold.
A lane quiet because its own validation run is parked at a gate awaiting
a human decision is deliberately out of scope: reading that state needs a
signal carrying who the wait is on and what clears it, rather than one
inferred from a parked verdict that also covers gates awaiting the
crewmate itself.
Tests pin both directions for each case and were each confirmed to fail
with the consult removed.
* no-mistakes(document): docs: honour declared waits in stale-escalation docs
* fix(bin): report verified PR state for passed runs (#4624)
* fix(bin): derive passed PR state from PR record
A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists.
For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim.
Fixes #4607
* no-mistakes(review): Add bounded GitLab merge-request state reads
* no-mistakes(review): Preserve network-free inactive crew-state scans
* no-mistakes(document): Document PR record readers in shared library
* fix: restore published contribution follow-up (Fixes #4469) (#4627)
* fix: restore published contribution follow-up (Fixes #4469)
* fix(review): Fix contribution freshness and merge actor routing
* fix(review): Restore issue triage and scope contribution follow-up
* fix(test): test: assert one wake per contribution signal
* fix(document): Document contribution follow-up
* fix: restore truthful terminal delivery evidence
* fix(review): Disclose unsupported contributions and deduplicate watcher wakes
* fix(review): Preserve unmeasured unsupported contributions across Bearings
* fix(review): Deduplicate shared contribution wakes and isolate diagnostics
* fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower
* fix(bin): make remote report transfers explicit and fail-open (#4658)
* fix(bin): make a remote-reply document gap self-clearing and re-attemptable
A remote mate's undelivered document raised a keyed `blocked` decision that
nothing could ever resolve, and any `data/*.md` substring in any mirrored line
was an unconditional fetch instruction. A mate announcing a report it had not
written yet therefore manufactured a permanent, factually false blocker, and
its own explanation of the false alarm manufactured more.
The reader has no permanence vocabulary: a report still being written refuses
exactly like a path that will never exist. So an undelivered document is now a
durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`,
re-attempted on the next delta and on the channel's own quiet poll, and retired
with a matching `resolved` line naming the local copy once it arrives. The
cursor still advances and no delta stalls on one bad pointer.
Only a structured `report=data/....md` pointer now offers a document, so a path
merely mentioned in prose - including one under another home's mirror tree,
which is provably not that mate's to serve - is never fetched. Offers are
deduplicated across the whole delta, the escalation names each missing document
once and carries the reader's own reason instead of discarding it, and a
strictly increasing notice ordinal keeps a later escalation from being
swallowed as duplicate bytes. A mirrored line still lands once whichever
pointer form it was first written under.
* no-mistakes(review): Require structured pointer token boundaries
* no-mistakes(review): Unify boundary-safe pointer extraction and rewriting
* fix(bin): identify a mirrored line independently of its delivery state
Two defects in the boundary-safe pointer work.
The at-most-once check compared only the all-remote and all-local renderings
of a line, so it could not recognize a mixed one. A line offering two documents
where only the first was deliverable mirrored as local-plus-remote; once the
second arrived, a cursor-loss whole-log recapture rendered the same line
all-local, matched neither alternate, and mirrored a second time. A line's
identity is now the canonical form every boundary-valid pointer would take once
delivered, derived by the same parser that does extraction and rewriting, so it
no longer depends on which documents happened to be deliverable at the time.
The pointer map was passed to awk through the process environment. A delta may
carry up to the configured 1 MiB bound, and an expanded map of delivered
pointers can exceed the platform's exec argument limit, so awk would fail to
start; because no caller checked, the empty result would have been appended as
blank lines while the cursor advanced past dropped status content. The map now
travels in a file, and every call site checks the exit status and stops the
ingest rather than committing a delta it could not render.
Both passes now run once per stream instead of twice per line.
* no-mistakes(review): Abort ingest when document pointer extraction fails
* no-mistakes(review): Exclude structured cross-home pointers from document transfer
* fix(bin): fail open on an undeliverable remote document instead of tracking it
Narrow the remote-reply document fix to the scope the diagnosis actually
requires, as decided after measuring a simpler alternative.
A document the reader cannot deliver now fails open. The mate's line is
mirrored with its own pointer, the cursor advances, and one unkeyed note
carries the reader's reason. A note never enters the open-decision fold, so it
cannot stand open the way the original keyed block did - which removes the
never-clearing false blocker by construction rather than by resolving it.
That makes the durable self-clearing obligation unnecessary, so it goes: the
per-mate pending-documents record, its notice ordinal and resolved
announcements, and the poll-side retry. Canonical line identity goes too, and
with it a way to silently drop a genuine status line; mirroring is back to
at-most-once on exact bytes. The cross-home exclusion goes as well: under
fail-open a cross-home report= either fails harmlessly or is a nested remote
report this mate genuinely holds, which is now relayed again.
Kept: fetching only on a structured report= pointer, the boundary-correct
parser, the file-based rewrite map, and checked extraction and rewrite exit
status. The parser now scans behind a sentinel byte so a rejected candidate can
no longer give the text right after it a false leading boundary.
The reported incident is covered end to end: a report path announced in prose
before it exists raises no decision, and the report still arrives through the
ledger publisher's structured offer once written.
* no-mistakes(review): Preserve source-line identity across remote reply replays
* no-mistakes(document): Document remote reply transfer and replay semantics
* no-mistakes(lint): Fix staging truncation lint checks
* fix(calm): preserve substantive mid-turn responses (#4655)
* Preserve substantive Calm mid-turn text
* no-mistakes(review): Distinguish newline-preserved replies from short narration
* no-mistakes(document): Document Calm mid-turn preservation boundaries
* no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`
* fix(bin): preserve PR merge polls across volume remounts (#4656)
* fix(bin): re-record PR poll identity after a volume device renumber (Fixes #4260)
A volume remount can renumber the state filesystem's st_dev while every
inode and byte stays the same; APFS does this across a reboot. A poll
registration records its sidecar and check as device:inode, so every poll
armed before the remount failed strict validation and the watcher refused
all of them as unauthenticated state checks until each was re-armed by hand.
There are two device comparisons. fm_pr_private_file_valid compares a live
file's device with the state directory's device read in the same invocation:
it refuses a file that is not on the state directory's own filesystem and
already survives a renumber, so it is unchanged. The registration's recorded
identity versus the live identity (from #556, reused by the #932 retirement
receipt) binds the registration to the exact files published in its own
transaction; its device part is what breaks.
When strict capture fails, the watcher now proves the device is the only
difference: every other artifact check passes (template bytes, both hashes,
private mode, single link, live device, metadata), both recorded identities
name one device, and each recorded inode equals its live inode. Only then,
under the task's control lock, does it rewrite the two identity lines,
repeating the whole proof and comparing the registration's file identity and
bytes just before the rename, and then capture strictly again. A swapped,
altered, re-moded, relinked, split-device, or foreign-device artifact still
fails a proof and is still refused, and a pending retirement receipt blocks
the rewrite.
Reproduction: on macOS a poll armed on an APFS disk image that was detached
and re-attached behind another image moved st_dev 16777239 -> 16777243 with
inodes, bytes, mode, and link count unchanged; the real watcher refused it on
main and reports its merge with this change. The portable regression test
rewrites a real registration's recorded device and drives the watcher.
Not changed here: the status presentation cursor keys rows by its own
device:inode identity in bin/fm-classify-lib.sh, a different helper that
needs its own fix; a retirement receipt left by a reboot between its
publication and removal still names the old device and stays refused; custom
check trust binds only a content hash and is unaffected.
* fix(review): Serialize PR poll publication writers
* fix(review): Bound PR poll publication lock scope
* fix(bin): keep contribution records when the poll budget runs out (follow-up to #4627) (#4661)
A budget that expires partway through an observation no longer records an
error or prints the unavailable wake; the URL keeps its prior record and is
observed first next poll. forge() flags budget exhaustion at the point it
refuses, or when a read is killed at the budget's own deadline, so a genuine
forge failure still records the error and wakes. Each distinct URL is now
observed once per poll and applied to every owning task.
* fix(bin): clear parent pending-replies on local secondmate retirement (#4680)
* fix(bin): clear parent pending-replies on local secondmate retirement
Local secondmate teardown left resolved parent pending-reply records behind
after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced
retirement while any reply for that id is still unresolved, and delete every
matching record plus its delivery confirmation after a successful local or
remote retirement, matching the remote cleanup path.
* no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup
* no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung
* no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern
* no-mistakes(review): Pending-replies Basename und corr_id abgleichen
* no-mistakes(document): Clarify forced retirement pending-reply cleanup
---------
Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>
* fix(bin): accept Orca's composite worktree id when tearing down a task (#4677)
* fix(bin): accept Orca's composite worktree id at teardown
Teardown refused every Orca-backed task because the endpoint validator
checked orca_worktree_id with the simple-atom rule meant for tmux-style
window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca
returns that id as `<orca id>::<absolute worktree path>`, so the colon and
slashes in every real value made validation fail and finished Orca tasks
could never be cleaned up.
Validate the field as the composite it is: both halves of the first `::`
split present, the path half absolute, and no embedded newline, carriage
return, or tab. The terminal field keeps the atom check, which is correct
for it, and no other backend's validation changes.
The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca
never returns, which is why the suite passed a check the real value fails.
They now carry the composite form, so the tests exercise the real value.
* no-mistakes(document): name Orca's repo id in the composite worktree id
* no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points
* feat(bin): add opt-in typed dispatch resolution (#4692)
* feat(bin): add opt-in typed dispatch resolution through typesafe.ai
Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or
scout profile from a written brief with typesafe.ai's System One model:
one Choice question over the rules' `when` texts, then the confidence
floor, the rule's `approval` and `floor`, each profile's `provider` and
`floor`, one quota-axi snapshot, and the spendPriority argmax all in code.
It is off unless TYPESAFE_API_KEY is in the environment or the home's
gitignored .env; off means one stderr line, exit 0, and no network call,
so firstmate dispatches exactly as before. The key reaches curl on a file
descriptor, never argv.
Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and
the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new
tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates
the four new optional dispatch fields. Document the schema, the operator
contract, the AGENTS.md intake step, and the live and benchmark evidence.
* no-mistakes(review): Harden typed dispatch resolution and quota bounds
* no-mistakes(review): Validate dispatch floors and ranking evidence
* no-mistakes(review): Tighten dispatch response and floor evidence
* no-mistakes(review): Neutralize none matching and resolve defaults locally
* no-mistakes(review): Preserve providerless profiles outside typed resolution
* no-mistakes(review): Validate response usage and reject duplicate profiles
* no-mistakes(review): Escalate unverifiable floors and validate probabilities
* no-mistakes(review): Validate probability mass and unknown profile floors
* no-mistakes(review): Simplify resolver interface and preserve fallback routing
* no-mistakes(review): Fix constants and rank partial quota evidence
* no-mistakes(review): Add authoritative provider mapping and enforce explicit providers
* no-mistakes(review): Declare provider for documented Pi profile
* no-mistakes(review): Validate provider identifiers and support Gemini dispatch
* no-mistakes(review): Strictly anchor provider identifiers
* no-mistakes(review): Validate selectors and preserve fallback candidate evidence
* no-mistakes(review): Gate typed validation and harden resolver evidence
* no-mistakes(review): Preserve opt-in routing and harden candidate evidence
* no-mistakes(review): Prioritize known exhaustion over quota uncertainty
* no-mistakes(review): Isolate API secrets and preserve no-key diagnostics
* no-mistakes(review): Fallback safely when dispatch rules are absent
* no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets
* no-mistakes(document): Document typed dispatch safety and fallback behavior
* fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753)
* test: reproduce buried status declarations in shared readers
* fix: share status event reads and preserve open blockers
* fix: retain terminal scout and ship status declarations
* no-mistakes(review): Fix status chronology, legacy completions, and reader performance
* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots
* no-mistakes(review): Unify terminal supersession across cached folds and consumers
* no-mistakes(review): Filter per-key status history while preserving terminal chronology
* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells
* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses
* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor
* no-mistakes(lint): Quote literal done in test for-lists for SC1010
* ci: expect 19 snapshot/fleet-view tests
This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.
* no-mistakes(review): Restore multiline child outcome reporting
* no-mistakes(review): Select ledger terminal events through bounded shared reader
* no-mistakes(review): Report newest open decision instead of preferring blocked
* no-mistakes(review): Require colon before ship/scout terminal supersession in fold
* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix
* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions
* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold
* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule
* no-mistakes(document): Align status-read docs with fold-resolved crew state
* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers
* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree
* test: fold terminal-cleanup snapshot coverage into the completed-scout case
Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.
* no-mistakes(document): Clarify socket-down override expiry in architecture doc
* ci: retrigger flaky contribution check
* fix(bin): launch codex crewmates with codex's hook layer disabled (#4689)
* fix(spawn): launch codex crewmates with codex's hook layer disabled
A freshly launched Codex worker never reached its instructions. Codex
stopped it on an interactive "Hooks need review" modal whose selection
sits on "Review hooks", which is neither trusting nor declining.
Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow
navigation, so the selection cannot be moved, and pre-accepting the
prompt by writing Codex's own trust store would record an operator
consent that was never given.
The hooks are the machine's own ~/.codex/hooks.json plus any project's
.codex/hooks.json. A crewmate needs neither: its turn-end signal is the
-c notify= program on the same launch, and Firstmate's project hooks are
primary-session infrastructure that stands down in a child worktree.
Crewmate and scout launches now pass --disable hooks. That is the
opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted
hooks; disabling the feature runs none of them and leaves the operator's
~/.codex untouched. An unknown feature name is a hard Codex error, so a
release that drops the flag fails the launch loudly instead of silently
restoring the modal. A secondmate is a primary in its own home and keeps
the project hooks its turn-end guard and session-start digest ride on.
Verified on codex-cli 0.151.0: the modal is gone and the turn-end
notification still lands.
This unblocks the second review that every finished pull request is supposed to get.
Fixes kunchenguid/firstmate#4673
* no-mistakes(review): Fix contradictory hook count in Codex verification record
* fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710)
* fix(bin): settle terminal contributions and wake once per read-failure episode
A contribution whose last good observation is merged or closed is final:
poll no longer re-reads it, projection keeps it fresh, and a stale error
recorded beside it is cleared once. A genuine forge-read failure on an open
contribution still records its error on every cycle but prints the
unavailable wake only when it starts a failure episode; a successful read
ends the episode. Open PRs linked from done tasks keep being observed.
The false unavailable beside a complete observation was budget exhaustion
mid-observation, already fixed by #4661.
* fix(review): Settle terminal contribution owners
* fix(review): Deduplicate shared contribution failure episodes
* fix(test): Preserve settled terminal contribution records
* fix: select authoritative no-mistakes runs (#4476)
* fix(crew-state): select authoritative validation runs by identity
Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row.
Refs: https://github.com/kunchenguid/firstmate/issues/3215
* fix(review): Resolve same-branch run identities beyond capped history
* fix(review): Fix run-selection compatibility, races, and worker-state fallbacks
* fix(review): Limit run validation to the requested branch
* fix(test): Anchor AXI fixtures and document remaining live evidence gaps
* fix(document): Clarify run selection documentation and capture ownership
* fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation
* fix: distinguish captain outcomes from no-op updates (#4738)
* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work
MAIN answered a supervision-branch outcome for completed captain-requested
work (implementation done, PR ready for review and merge approval) with
"Captain, shipshape.", reading section 9's no-action reply as covering it
and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no
captain-facing response is owed".
Section 9 now limits the shipshape reply to true no-ops (idle re-read,
empty heartbeat, consequence-free acknowledgement) and requires a short
outcome response naming what finished and what word is needed whenever
requested work finishes or a result needs the captain's word, even when a
transcript entry already shows the substance. The Pi protocol's re-emit
rule now says it bounds repetition only, and carries a worked example of
the ready-for-review outcome whose correct processing turn a shipshape
reply fails.
No executable contract evaluates the content of MAIN's captain-facing
reply, so the regression is the protocol example in the owner doc rather
than a text-match test.
* no-mistakes(document): Clarify captain-facing outcomes versus no-ops
* docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line
The document step condensed the Pi protocol's re-emit rule and dropped the
worked example of a finished, ready-for-review outcome whose correct
processing turn a "Captain, shipshape." reply fails. That example is the
contract's regression: no executable contract evaluates the content of
MAIN's captain-facing reply, so the owner doc's example is the test case.
Restore it directly under the re-emit rule, prefixed as a regression
example that is kept verbatim and never condensed or summarized away.
* no-mistakes(review): Clarify captain outcome and decision-word requirements
* no-mistakes(document): Clarify captain-facing completion outcomes
* docs(pi): require the PR URL in the visible captain-facing outcome reply
Captain review on the regression example: drop the sample reply string
and say only that the ready-for-review outcome requires relaying a
captain-facing outcome response, not just "Captain, shipshape.".
Fold in the visible-PR-handoff failure seen this session: after the
branch outcome reporting this fix green, MAIN's visible reply was only
"Awaiting your merge call." with no PR URL, leaning on the dim anchor.
Section 9's URL rule now also covers a review or merge ask and names the
visible reply as where the URL goes, sourced from the ready status, pr=
metadata, or the supervision branch's summary and never left to a
transcript entry. The Pi protocol adds the same-way failure and places
the captain-facing text in the final visible assistant reply after the
fm_branch_processed call, because Calm hides assistant text emitted in
the same step as a tool call as a working note.
Investigation verdict, evidence in the PR comment: no recent PR caused
the handoff failure; Pi has hidden same-step pre-tool assistant text
since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and
#4658 touched only remote report transfer.
* no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs
* no-mistakes(document): Clarify captain-facing supervision outcomes
* docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule
The consolidated section 9 URL rule narrowed its trigger to a review or
merge ask, dropping the "whenever a PR is mentioned" catch-all from
#3648 that keeps every PR URL copied from a durable record and never
assembled from memory. Restore that trigger as a union with the review
or merge ask so the one consolidated rule covers both.
* fix(bin): let non-owner Claude Stops exit safely (#4777)
* Fix foreign-owner turn-end supervision loop
* no-mistakes(review): Scope foreign-owner safe exit to Claude guard
* no-mistakes(document): Document Claude foreign-owner safe exit
* fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778)
Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty
indexed array as an unbound variable and aborts the shell. In
signal_turnend_panes_churned() the missing_keys loop was reachable with
an empty array whenever every churned key already held a fresh
.churn-since-* marker (a second churning turn-end inside an open
deferral window), so each watcher cycle died about half a minute in and
supervision restarted endlessly. The created_keys rollback loops had the
same latent crash on their error paths.
Audit of bin/ for the same pattern found one more confirmed-reachable
case: remote_handoff's noncanonical-body scan iterates to_move, which is
empty when a retried remote handoff finds every key already staged in
the outbox. All other "${arr[@]}" sites are either count-guarded,
guaranteed non-empty by construction, or unreachable while empty.
Guard the three reachable expansions with the repo's existing
"${arr[@]+...}" idiom. New regression test drives a real watcher
through the all-marked churn path; the macos-stock-bash CI lane runs it
under real /bin/bash 3.2 via FM_TEST_ONLY.
* Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783)
The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: require complete captain-facing final responses (#4779)
* docs: require complete final responses across harnesses
* no-mistakes(document): Document complete final replies for Grok Bot
* docs: point Grok replies to the shared contract owner
* no-mistakes(review): Clarify final recap without batching decision asks
* fix: preserve substantive mid-turn text in Pi Calm (#4788)
* fix(calm): preserve substantive Pi mid-turn text
* no-mistakes(review): Preserve substantive Pi Calm text per block
* no-mistakes(test): Cover shared Calm preservation boundaries behaviorally
* no-mistakes(document): Consolidate Calm preservation documentation
* fix: harden mail checks and rebalance full-coverage CI (#4800)
* Improve CI reliability and rebalance full-coverage validation
* no-mistakes(document): Clarify lint partition documentation
* fix(bin): answer Kimi 2.0.0 folder-trust dialog during spawn (#4799)
* Handle Kimi workspace trust dialog
* no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers
* no-mistakes(review): Gate Kimi ready on any trust marker and clean captures
* no-mistakes(review): Read visible pane for Kimi trust and ready gates
* no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate
* no-mistakes(review): Harden Kimi viewport capture and trust dialog detection
* no-mistakes(document): Document Kimi spawn refusal on cmux and Orca
* fix(bin): report a dead-agent record once instead of escalating forever (#4775)
* fix(bin): report a record whose agent is gone once instead of escalating forever
The wedge escalation path never asked whether there was still an agent to be
wedged. A wedge is something stuck that might recover, so re-alarming it earns
its cost; an agent that is gone never moves again, its pane never churns, the
idle timer never resets, and the escalate path clears its own timer and re-arms
with nothing bounding the count.
Observed on a live fleet: two finished lanes reached 226 and 203 consecutive
escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400
notifications a day from two lanes with no agent running at all. On one,
fm-control.sh exit answered already-stopped and fm-crew-state.sh read
"failed - run failed". Closing the Herdr pane did not stop it either: with the
pane genuinely gone and herdr pane read returning pane_not_found, the count kept
climbing, because the poll is driven by the record's window= line rather than by
the pane. The cost is not the repetition but that it drowns the alarms that
matter.
fm_backend_agent_state already separates a thinking agent from a gone one at
process level. In the branch that wa…
sctru
added a commit
to sctru/firstmate
that referenced
this pull request
Sep 19, 2026
* fix(bin): safely unregister custom checks (#3369)
* fix(bin): add a safe owner for custom-check retirement
Agents were improvising rm of check files with unset STATE/ID, which wedges
headless panes. Unregister validates the id and state directory first.
Co-authored-by: Cursor <cursoragent@cursor.com>
* no-mistakes(review): Refuse explicitly empty custom-check state overrides
* no-mistakes(document): Document custom-check retirement safety contract
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* refactor(quota): extract mid-task polling and candidate selection into dedicated scripts (#3221)
* Add quota exhaustion detection and safe fallback helpers
- bin/fm-procevent-quota.sh: generic procevent adapter that arms a
recurring quota-axi --json poll and wakes firstmate when a tracked
provider's effectivePercentRemaining drops below a threshold or its
runway.status becomes exhausted_now.
- bin/fm-quota-choose.sh: worker-side helper that picks the first ranked
harness:model candidate with positive effectivePercentRemaining.
- AGENTS.md and .agents/skills/quota-array-dispatch/SKILL.md: document
the new helpers and the mid-task quota-exhaustion wake path.
- tests/fm-quota-choose.test.sh: unit tests with a mocked quota-axi JSON
source.
* no-mistakes(review): Fix quota polling and scope bounds
* no-mistakes(review): Enforce safe default quota selection
* no-mistakes(review): Handle decimal quota values safely
* no-mistakes(review): Fail closed on invalid quota inputs
* no-mistakes(review): Reject empty quota candidate segments
* no-mistakes(review): Harden quota parsing and timeout ownership
* no-mistakes(review): Reuse captured quota snapshots consistently
* no-mistakes(review): Match quota using explicit candidate providers
* no-mistakes(review): Centralize fail-closed quota schema validation
* no-mistakes(review): Reject out-of-range quota percentages
* no-mistakes(review): Validate quota runway status enum
* no-mistakes(review): Tighten quota scope and status contracts
* no-mistakes(review): Preserve unknown quota and exact product bounds
* no-mistakes(review): Preserve provider-level unknown quota
* no-mistakes(review): Reuse canonical verified harness validation
* no-mistakes(document): Document mid-task quota handling
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* fix(docs): restore default routing contract, keep quota helper optional
Restore the AGENTS.md section 4 always-loaded routing paragraph the PR
had deleted, so the standing TOON-first intake, spendPriority ranker,
every-candidate accounting, and load-trigger contract stay exactly as
before this PR. The mid-task quota wake is optional and must not alter
default routing.
Restore the quota-array-dispatch skill ownership line to section 4 as
the always-loaded intake boundary owner; keep the worker-side helper
section as an addition only, without rewiring ownership or load
triggers to section 13.
* fix(bin): use harness-keyed quota matching in optional helper
Revert fm-quota-choose.sh from harness:provider:model tuples back to
harness:model candidates with harness-keyed provider matching, per the
resolved ask-user finding. The helper is optional; authoritative
multi-provider routing (provider discovery from the harness catalog and
quota matching by that explicit provider) stays owned by AGENTS.md
section 4 and the quota-array-dispatch skill intake procedure, not the
helper.
Document the multi-provider limitation in the helper header and the
quota-array-dispatch skill: the helper maps each harness to one primary
provider family only, so a candidate whose established provider differs
from that primary family is checked against the wrong quota row. Use it
only when the brief fixed the candidate order and every candidate's
provider is the harness's primary family.
The helper still consumes one already-captured default-TOON or JSON
snapshot via stdin or --snapshot and never calls quota-axi itself, so
it selects from the same quota state as the intake.
* no-mistakes(review): Fix Muse quota mapping and helper contract docs
* no-mistakes(review): Reject known-empty quotas and map quota tests explicitly
* no-mistakes(review): Preserve unmeasured candidates and enforce snapshot reuse
* no-mistakes(review): Fix quota retirement and dependent regression coverage
* no-mistakes(review): Accept zero-row quota TOON snapshots
* no-mistakes(review): Enforce quota semantics status consistency
* no-mistakes(review): Veto dispatch on any exhausted applicable scope
* no-mistakes(review): Record exhausted quota scope in wake details
* no-mistakes(review): Fix quota help and control dependency coverage
* no-mistakes(review): Decode quoted TOON fields and document quota wakes
* no-mistakes(review): Validate zero-row TOON and map timeout coverage
* no-mistakes(review): Reject multi-value JSON and malformed TOON envelopes
* no-mistakes(review): Validate complete nonzero TOON envelopes
* no-mistakes(review): Accept producer-shaped quota TOON envelopes
* no-mistakes(review): Support empty quota arrays and validate counted rows
* no-mistakes(review): Harden TOON completion, scopes, and quoted fields
* no-mistakes(review): Preserve unknown-headroom exhaustion and reject trailing fields
* no-mistakes(review): Allow unknown headroom under known semantics
* no-mistakes(review): Reject noncanonical quota identities
* no-mistakes(review): Preserve empty quota polling and validate attention identities
* no-mistakes(review): Reject noncanonical provider watches
* no-mistakes(review): Validate all candidates before quota selection
* no-mistakes(document): Correct quota helper safety documentation
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* fix: surface comments on Lavish annotations (#3371)
* fix(bin): keep typed Lavish comments when an element is also annotated
read preferred element text over prompt, so an annotate-and-comment
item dropped the captain's words. Surface prompt as its own field.
Co-authored-by: Cursor <cursoragent@cursor.com>
* no-mistakes(review): Filter non-comment prompts from Lavish reader output
* no-mistakes(document): Clarify Lavish comment presentation contract
* no-mistakes(ci): Fixed Lavish reader comment provenance: non-choice prompts are now emitted even when identical to element text. Added observable regression coverage for identical selector+comment input while retaining pure annotation/message coverage. Reader cases, bash syntax, and diff checks pass. Full fm-procevent suite stops earlier at unrelated “reconcile never claimed” setup failure
* no-mistakes(ci): Fixed duplicate pure-annotation prompts by emitting `prompt:` only when it differs from captured element text. Updated behavioral coverage for selector+comment, pure annotation, and pure message cases. Focused reader regressions, syntax checks, and diff checks pass. Full suite remains blocked by the pre-existing “reconcile never claimed the registered source” failure
* fix(bin): always emit Lavish comments and use real annotation fixtures
Stop inferring comment provenance from prompt==text. Real pure
annotations have no prompt, so always-emit does not duplicate.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: support first public-followup registration on Bash 3.2 (#3420)
* Fix public-followup register crashing on empty lock arrays under bash 3.2.
bash 3.2 with set -u treats "${arr[@]}" on an empty array as unbound, so the first register in a fresh home aborted before taking the registry lock.
The empty-lock regression also runs under the existing stock macOS Bash CI lane so pre-fix code would fail there.
* no-mistakes(document): Document stock Bash registration coverage
* no-mistakes(ci): Pinned the stock macOS Bash CI lane to tasks-axi@0.2.5, eliminating dependency drift. Verified workflow YAML parsing, git diff checks, and the focused regression under /bin/bash 3.2.57 with tasks-axi 0.2.5
* no-mistakes(ci): Fixed the flaky portable CI test: it treated exited zombie processes as live because `kill -0` succeeds for zombies. The watcher and descendant assertions now check process state and regard zombies as exited. Verified `tests/fm-pr-check-security.test.sh`, ShellCheck, `git diff --check`, and the focused Bash public-followup regression
* fix(bin): isolate new Herdr server environments (#2792)
* fix(herdr): isolate server launch environment
* no-mistakes(review): Clear inherited supervision model from Herdr launches
* no-mistakes(document): Document Herdr server launch environment isolation
* fix: surface inbound Relay media to responding agents (#3442)
* fix: surface inbound Relay attachments to the responding agent
A Discord support thread's screenshots were never seen by the agent
handling the mention. The relay delivered them and the poll stashed
them: the reporter's images arrived on the `thread_starter` entry of
`in_reply_to_chain` while the mention's own media list was empty. The
gap was in the responder's playbook, which enumerated a fixed field
list (`request_id`, `text`, `in_reply_to`, `in_reply_to_chain`) and so
made every other field, attachments included, invisible.
Fix it where the gap is, in prose:
- Read the complete payload object rather than a fixed field list, so
media and later relay fields are never skipped again.
- Fetch and view attached media with the agent's own tools, on the
mention and on every chain entry, and call out the common shape where
only the thread starter carries the screenshots.
- Restrict those fetches to known-good platform media hosts over https
(Discord: cdn.discordapp.com, media.discordapp.net,
images-ext-1.discordapp.net, images-ext-2.discordapp.net; X:
pbs.twimg.com, video.twimg.com), report a blocked host instead of
working around it, and treat everything fetched as untrusted public
input on the same terms as the surrounding thread text.
The poll stays out of it and downloads nothing, so no third-party bytes
are pulled on the polling path.
The new test pins the contract the playbook depends on: a mention in the
incident's shape, with an empty top-level media list and screenshots on
the thread starter, must reach the inbox with the payload intact and its
media URLs unfetched.
* no-mistakes(review): Preserve media authority and enforce poll-only fetching
* no-mistakes(document): Clarify Relay attachment safety prose
* fix(bin): defer inactive reconciliation during startup (#3480)
* Defer inactive startup reconciliation
* no-mistakes(review): Queue deferred inactive reconciliation diagnostics durably
* no-mistakes(review): Require worker phases to cover startup requests
* no-mistakes(review): Make diagnostic wakes safely acknowledgeable
* no-mistakes(document): Document deferred startup phase coverage
* fix(bin): bound wake drain presentation lock waits (#3475)
* fix: bound status presentation lock waits
* no-mistakes(review): Distinguish malformed presentation locks from live contention
* no-mistakes(review): Bound no-ack drain queue lock acquisition
* no-mistakes(document): Document bounded presentation-lock drain behavior
* no-mistakes(lint): Annotate bounded lock output global
* no-mistakes(ci): Added deterministic regression coverage for successful bounded-lock acquisition after live contention, verifying helper-to-caller PID ownership handoff and caller release. Verified with bash syntax checks, git diff checks, and the full fm-wake-queue test suite
* fix(bin): retire public follow-ups in remote homes (#3479)
* fix(relay): close a public loop whose work lives in a remote secondmate home
A public-followup loop bound to a REMOTE secondmate could never be closed.
`clear_public_followup_link` (bin/fm-public-followup.sh:701) required an
absolute recorded `work_home_path` for a `secondmate:*` work home, but a remote
route has no local path on this machine, so registration records that field
empty (bin/fm-public-followup.sh:291). Every close ran that clear first, so
`retire` died with "could not clear the legacy X link ... retained for
reconciliation" forever, and `deliver` posted the public reply and then stranded
the loop at `posted`. `--force` never covered that step.
The clear now goes to the remote home over that route's SSH transport, running
`fm-x-followup.sh --clear <work-id>` through `bin/fm-on.sh`. The route is decided
from `data/secondmates.md` before any local path is consulted, so a same-named
local directory can never stand in for a remote home, and registrations already
on disk retire without needing a new field. `fm-on.sh` passes ssh's status
through, so 255 stays the established "delivered but completion unknown" result
this codebase already reconciles: the close is refused, the registration and the
remote link are left exactly as they were, and the message names the unknown
completion instead of claiming a definite failure.
Local secondmate and `main` work homes are untouched, and `--force` still
governs only the unresolved-obligation refusal.
Three regression cases drive a remote route end to end, faking only the ssh
binary at the FM_SSH_BIN seam and then running the real remote entrypoint
against a local checkout, so the clear that must reach the remote home actually
happens there.
* no-mistakes(review): Guard remote link clears by request identity
* no-mistakes(review): Fail guarded clears on unreadable remote state
* no-mistakes(review): Reject guarded clears on non-writable remote state
* no-mistakes(review): Allow no-link retirement in non-writable remote state
* no-mistakes(document): Correct public-followup verification guarantee count
* no-mistakes(ci): Fixed the guarded link-clear race by ensuring absence is decided under the metadata lock whenever publication is possible. Added a behavioral concurrency regression test. Verified with fm-x-mode and fm-public-followup suites, Bash syntax checks, diff checks, and bin/fm-lint.sh
* no-mistakes(ci): Fixed the guarded link-clear race by refusing an unlocked absence decision when a publisher already owns the metadata lock in a non-writable directory. Added a behavioral concurrency regression test. Verified with fm-x-mode, fm-public-followup, syntax/diff checks, and fm-lint
* no-mistakes(ci): Fixed the guarded-clear race by refusing all guarded clears when the metadata parent is non-writable, including apparent link absence. Added a behavioral regression with a publisher waiting to create the lock, updated remote-retirement expectations and verification docs. Passed fm-x-mode, fm-public-followup, fm-lint, documentation audience, Bash syntax, and diff checks
* fix(relay): bound the guarded remote link clear so it refuses instead of hanging
The guarded clear checks that the remote state directory is writable before
taking the metadata lock, but that check cannot close the window: the parent can
turn non-writable between the check and lock creation, and a lock held by a live
holder is indistinguishable from that at the acquire. `fm_lock_acquire_wait` is
an unbounded `while ! try; do sleep 0.1; done`, so either case retried forever
and `deliver` or `retire` wedged with nothing reported, instead of returning the
retained-for-reconciliation refusal the guard exists to produce. This path runs
unattended over the secondmate transport, where a wedge is worse than either
outcome the guard defines.
The guarded clear now acquires through `fm_lock_acquire_wait_bounded`
(FMX_LINK_CLEAR_LOCK_TIMEOUT, default 10 seconds) and refuses on timeout through
the existing failure path. Unguarded local callers keep the ordinary unbounded
wait, so local behavior is unchanged.
The bounded primitive's header no longer claims presentation-only scope, since
this is a second authorized caller; nothing else in the shared lock
infrastructure changed.
The regression holds the metadata lock with a genuinely live process while
leaving the state directory writable, so the refusal can only come from the
bound and never from the writability precondition. Against the unbounded wait it
does not terminate at all; with the bound it refuses, retains the registration,
writes no receipt, and leaves the remote link untouched.
* no-mistakes(review): Harden lock-timeout regression with independent deadline
* no-mistakes(review): Restore no-op guarded clears on read-only state
* no-mistakes(document): Clarify remote public-followup cleanup contract
* fix(bin): support process events under symlinked homes (#3484)
* fix(bin): resolve process-event state roots before validating them
The process-event module validated the caller's spelling of a home's state
root instead of the directory it operates on: it required the supplied path
to equal its own lexical normalization, which rejects any path reached
through a symlinked ancestor. On macOS both /tmp and $TMPDIR are symlinks,
so an operator home under either could never claim a source. Reconcile still
reported the runner started, while the detached runner died writing "cannot
claim source" to the discarded stderr, and the source silently never fired.
Resolve the state root to its physical directory once, then apply the
existing private-directory validation to that resolved directory and derive
every path, recorded claim identity, and later confinement check from it.
This keeps the confinement contract for the directory actually operated on
rather than only for callers that already spelled it physically, and removes
the window where an ancestor symlink could be repointed between check and
use. Homes already spelled physically behave identically.
This was the single cause of both deterministic macOS failures in
tests/fm-procevent.test.sh ("reconcile never claimed the registered source")
and tests/fm-procevent-when.test.sh ("the winning concurrent arm did not
produce an outcome"). The new case pins the behavior with an explicit
symlinked-ancestor home, so it fails without the fix on any platform rather
than only where the temp root happens to be a symlink.
* fix(bin): pin the external capture staging boundary to its physical path
The extension capture path pinned its registry staging boundary by comparing
`pwd -P` against the caller-spelled registry directory, so a home reached
through a symlinked ancestor still refused to start an extension-backed
source after the state root itself resolved correctly. That left such a home
half working: built-in sources ran while external ones failed.
The staging preparer now prints the physical registry directory it validated,
matching the inbox and reservation preparers beside it, and the start path
pins on that returned path. The new end-to-end case drives the shipped
file-signal package from a symlinked home spelling.
* no-mistakes(review): Propagate canonical process-event state roots
* no-mistakes(review): Propagate canonical state to process-event adapters
* no-mistakes(document): Document physical process-event state roots
* fix(pi): deliver captain outcomes as deterministic transcript entries (#3312)
* fix(pi): persist captain outcomes visibly
* no-mistakes(review): Recover captain outcomes after cold-start lock acquisition
* no-mistakes(document): Document cold-start captain-outcome recovery
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes(review): Prove immediate Pi captain-outcome transcript delivery
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* fix(pi): process captain outcomes through a sequence-keyed turn
PR #3312 made every captain-facing supervision outcome a durable, exact-once
visible transcript entry with the read cursor advancing only after that entry
exists. That is the display half of the delivery contract. Left alone it turns
a probabilistic silent loss into a deterministic one: the captain sees an
anchor line, and firstmate never acts, because nothing opens a turn and
nothing records whether main ever processed the outcome.
The 2026-08-31 timeline showed the two shapes this must survive on the
previous hidden-turn path: seven delivered decision outcomes each answered by
an empty assistant message (cursor advanced, no retry, unanswered for close
to three hours), and two answered by an unrelated prior reply. Both happened
because delivery advanced the cursor at enqueue and accepted whatever the
next assistant message was.
Add the processing half on top of the persistence half:
- bin/fm-branch-outcome.sh keeps a processed marker separate from the read
cursor (`unprocessed`, `mark-processed --through`, `processed-init`). It
only advances through an explicit sequence-bound acknowledgement, never
past the read cursor and never backwards; an absent marker reads as zero
and `processed-init` migrates delivered history once so an upgraded home
is not re-presented its past.
- After the visible entry for a captain outcome exists, the extension hands
every still-unprocessed captain row to main as one hidden, typed
`fm-branch-process` request listing each `[seq N] task: summary`, opening
exactly one main turn. Main closes it only by calling the new
`fm_branch_processed` tool with the highest sequence listed. An unrelated,
empty, or paraphrased answer leaves the sequence open, and the same request
is presented again at the end of the next main run and at session start.
The first two presentations of a sequence set open a turn of their own;
after that the request rides the captain's next prompt so an ignored
request cannot loop, and a session replacement resets that budget.
Routine outcomes stay turn-free.
- The regressions cover exactly those incident shapes against the real store
scripts: an empty answer and an unrelated prior answer neither advance the
marker nor stop re-presentation, the acknowledgement is refused beyond the
read cursor and outside lock ownership, a partial acknowledgement keeps the
newer sequence open, and #3312's own assertions now forbid an unkeyed turn
rather than any turn. The store suite pins the marker's bounds and the
migration; the real-SDK guard for appendEntry persistence and model
exclusion is unchanged.
Docs move the protocol from "no model turn" to "one sequence-keyed processing
turn closed only by its acknowledgement", and the verification record carries
the dated run against Pi 0.84.4.
* no-mistakes(review): Harden outcome listing and sequence-bound acknowledgements
* no-mistakes(review): Harden outcome state validation and request pacing
* no-mistakes(review): Reject unsafe sidecars and unterminated outcome stores
* no-mistakes(review): Validate canonical mark-read cursor state
* no-mistakes(review): Guard cursor advancement against corrupt processed state
* no-mistakes(review): Bind acknowledgements to active processing requests
* no-mistakes(review): Reset pacing when processing sequence membership changes
* no-mistakes(review): Enforce silent outcome invariants at storage boundary
* no-mistakes(document): Document hardened captain outcome processing contracts
---------
Co-authored-by: kunchenguid <kun@kunchenguid.com>
* feat: add bounded concurrent Bearings ledger collection (#3481)
* feat: bound Bearings remote ledger collection
* no-mistakes(review): Clarify default remote-ledger collection behavior
* no-mistakes(review): Detach reconcile delivery from watcher loop
* no-mistakes(review): Enforce bounded snapshot and request captures
* no-mistakes(review): Bound legacy summary capture before parsing
* no-mistakes(review): Bound primary remote ledger captures
* no-mistakes(document): Correct snapshot and reconcile documentation
* no-mistakes(lint): Fix ShellCheck quoting in bounded collector
* no-mistakes(ci): Fixed all three CI failures: updated the macOS Bearings assertion to 44 tests, made the home-summary test deterministic and aligned with default ledger consumption, and increased the asynchronous reconcile retirement wait for loaded CI. Verified both focused suites, all 44 Bearings tests, ShellCheck, actionlint, Bash parsing, and git diff checks
* test: await reconcile request retirement
* no-mistakes(review): Avoid empty reconcile queue process churn
* no-mistakes(review): Read ledger summaries from immutable snapshots
* no-mistakes(review): Reject multi-document home ledger streams
* no-mistakes(review): Coalesce durable reconcile requests per target
* no-mistakes(review): Unify reconcile keys and reject snapshot streams
* no-mistakes(review): Key reconcile requests by stable target ID
* no-mistakes(document): Document per-target reconcile request coalescing
* no-mistakes(lint): Remove unused snapshot summary file variable
* no-mistakes(ci): Adjusted the concurrent collector regression’s end-to-end timing ceiling to account for stock macOS process/jq overhead outside the three-second remote collection budget, while remaining below the 15-second serial-read floor. Verified with stock /bin/bash 3.2: all 44 Bearings tests pass; bash syntax and git diff checks pass
* no-mistakes(ci): Fixed legacy summary validation to require exactly one top-level JSON document and added behavioral regression coverage. Stabilized CI by conditionally waiting longer for durable reconcile delivery and synchronously stopping the fm-on worker tree before fixture cleanup. Removed a redundant flaky healthy-path timing assertion; the wedged-reader test still proves concurrent bounded collection. Verified fm-bearings-snapshot, fm-secondmate-reconcile, and fm-on tests, plus project ShellCheck, bash syntax, and git diff checks
* ci: rebalance portable serial test shards (#3489)
* fix(ci): rebalance the portable serial shards on measured durations
The "Behavior portable serial 3" shard ran 17-20 minutes against its
20-minute job cap and intermittently timed out seconds after a passing
test, on branches and on main alike.
Shards are packed longest-processing-time from per-script duration hints,
and those hints were last measured on 2026-08-21 at 116 scripts. The lane
has since grown to 139 scripts and from ~42 to ~63 minutes: 17 scripts had
no hint at all and fell back to the 20 s default, and several existing
hints were low by 2-5x (fm-watch-triage 142 s hinted vs 263 s measured,
fm-public-followup 36 s vs 197 s). The partition therefore looked
perfectly balanced in hint space, 734.6 s per shard, while really running
11.5, 13.6, 18.8 and 16.5 minutes. Script-count balance, which is what the
tests asserted, stayed normal throughout and hid it.
Refresh the hints from the timing artifacts of three green runs, taking
the slowest measurement of each script so the balance holds on a slow
runner, and split the lane across five shards instead of four. Replayed
against those runs' real per-script durations the worst shard is now
12.54 minutes, 63% of the unchanged 20-minute cap, and the serial lane's
wall clock drops from ~20 to ~12.5 minutes.
Bound the drift that caused this rather than relying on the hints being
refreshed by hand: the coverage guard now reports the unmeasured share as
serial_unhinted= and refuses past PORTABLE_SERIAL_MAX_UNHINTED_PERCENT,
which leaves room for newly added tests while making a stale table fail
the guard instead of silently pushing one shard into its cap.
No test changes what it asserts and no test stops running; only the
partition across shards changes.
* no-mistakes(document): Clarify conservative shard timing aggregate
* fix(pi): fall back on incomplete supervision branch prompts (#3491)
* fix(pi): fall back after settled branch errors
* no-mistakes(review): Detect provider errors across prompt compaction
* no-mistakes(review): Preserve in-flight branch state across selection changes
* fix(pi): re-probe supervision branch after cooldown (#3497)
* fix(pi): recover supervision branch after cooldown
* no-mistakes(review): Defer branch recovery until prompt settlement
* no-mistakes(document): Clarify supervision cooldown recovery contract
* fix(bin): remove legacy remote snapshot reads (#3501)
* refactor: remove legacy remote summary reads
* no-mistakes(document): Document ledger-only snapshot reads
* no-mistakes(ci): Fixed the snapshot test fixture so ledger refreshes use the same fake executable PATH as the snapshot consumer. This preserves observable endpoint freshness after removing legacy summary computation. Verified stock Bash parsing and all 44 Bearings tests pass under /bin/bash; git diff checks pass
* no-mistakes(ci): Fixed the CI-only snapshot fixture failure by ensuring the bounded-ledger refresh uses its fake tmux backend. This removes host tmux availability as a source of nondeterminism. Verified all 44 Bearings tests pass, Bash syntax passes, and git diff checks are clean
* no-mistakes(ci): Fixed CI nondeterminism in the Bearings fixture: all local ledger refreshes now use the fixture’s fake tmux backend when available, instead of depending on host tmux state. Verified stock /bin/bash syntax, git diff checks, and all 44 Bearings tests with a deliberately failing host tmux
* fix(pi): preserve watcher continuity across session replacement (#3498)
* fix(pi): rearm watcher after session replacement
* no-mistakes(review): Queue actionable closes across Pi session replacement
* no-mistakes(review): Stop replacement arm when handoff persistence fails
* no-mistakes(review): Preserve actionable wakes through branch and late child races
* no-mistakes(review): Surface late handoff failures without crashing Pi
* no-mistakes(review): Coordinate replacement delivery settlement and unique handoff tokens
* no-mistakes(review): Retry stale deliveries and release settled claims
* no-mistakes(review): Distinguish branch settlement and retry handoff cleanup
* no-mistakes(review): Deduplicate persistent handoff cleanup alerts
* no-mistakes(review): Acknowledge watcher follow-ups only when consumed
* no-mistakes(review): Persist idle follow-ups until agent consumption
* no-mistakes(review): Preserve pending outcomes when handoff persistence fails
* no-mistakes(review): Arm replacement before awaiting prior delivery settlement
* no-mistakes(review): Adopt pending handoffs after lock reclamation
* no-mistakes(review): Prevent stale generations from adopting replacement handoffs
* no-mistakes(review): Scope replacement handoffs by watcher state
* no-mistakes(document): Clarify replacement handoff documentation
* no-mistakes(ci): Fixed the failing branch-extension tests to model the new settlement-promise contract. Failure cases now assert that delivery ownership returns to the watcher instead of expecting direct extension fallback. Verified the updated branch suite, Pi watcher suite, shell syntax, and diff checks
* no-mistakes(review): Update branch settlement tests and preserve chunked outcomes
* no-mistakes(document): Document watcher-owned replacement handoffs
* no-mistakes(document): Verify replacement handoff documentation
* test(pi): cover watcher-owned branch fallback
* no-mistakes(document): Refresh watcher-owned fallback documentation
* fix(bin): resurface task statuses missed by wake handling (#3495)
* fix(bin): resurface terminal statuses lost after branch handling
* test(watch): canonicalize process-event fixture homes
* no-mistakes(review): Index branch outcomes by causal status position
* no-mistakes(review): Recover outcome indexes and deduplicate resurfaced statuses
* no-mistakes(review): Handle legacy ambiguity and oversized status diagnostics
* no-mistakes(review): Keep unclassifiable oversized statuses silent
* no-mistakes(document): Document lost-wake outcome backstop
* no-mistakes(document): Update outcome backstop documentation
* no-mistakes(ci): Fixed CI regressions in wake-drain: parseable reserved-key decisions can no longer bypass the durable decision-fold guard, and status output is prepared and receipt-committed before presentation to prevent repeated one-shot outcomes after later failures. Added a behavioral regression for receipt commit failure and retry. Targeted backstop, correlation-token, decision-cursor, open-decision, unread-status, syntax, and diff checks pass locally. Shard-4 failures appeared unrelated/flaky; the network-parallel test passed locally
* no-mistakes(ci): Fixed the Greptile P1 data-loss issue by committing presentation receipts only after prepared output reaches stdout. Added behavioral coverage proving output failure leaves the backstop retryable and receipt failure may duplicate but never lose a presentation. Relevant wake-drain suites and syntax/diff checks pass. The shard-4 Pi extension failure is unrelated to this PR and did not warrant changes
* no-mistakes(ci): Stabilized tests/fm-bootstrap-network-parallel.test.sh by replacing scheduler-sensitive equal-sleep timing with bounded synchronization between mocked fetch and remote probes. This preserves detection of real serialization while avoiding false failures under CI load. Verified with five consecutive test runs, bash syntax validation, ShellCheck, and git diff checks. The separate Pi stock-rendering failure reproduces locally but is unrelated environment/version drift
* no-mistakes(ci): Fixed Behavior portable serial 4 by adding fm-classify-lib.sh and fm-timeout-lib.sh to the broken-root Pi test fixture; fm-branch-outcome.sh now depends on them. Verified the full Pi branch-extension suite with real-Pi checks skipped, the wake-drain outcome-backstop suite, Bash syntax, and git diff checks. Greptile findings are already addressed at HEAD; the no-mistakes attestation failure is external head-SHA state
* fix(bin): collect follow-up results from remote work homes (#3503)
* fix(bin): deliver typed terminal results from remote work homes
A public commitment whose work is bound to a REMOTE secondmate home could
never receive its typed terminal result. `fm-public-followup.sh brief`
printed an emit command carrying this home's own absolute path and this
checkout's own script path, neither of which exists on the machine the
worker runs on, so the worker had nothing it could write to that the
owning home would ever read - and `consume` kept finding nothing while
the promise stayed open.
The brief is now route-aware: for a remote work home it prints that
route's own code root and home with `--stage-in`, so the typed event is
staged in the home where the work actually runs, and the closing
paragraph names the owning home as the one on the other machine instead
of pointing at the path above it. The owning home collects those staged
results over the same SSH route it reaches that secondmate on, because
the transport only runs outbound: `consume` pulls them into its own
inbox and reconciles them exactly as it reconciles a local report.
Collection is non-destructive until the result is durably held, so a
dropped connection cannot lose a terminal result, and a route that could
not be reached is named in `consume`'s output with the promise left open
rather than reported as an empty inbox.
A local work home is untouched: the brief still prints `--home` with this
home and this checkout's script, and the event still lands directly in
this home's typed terminal-result inbox.
This is the emit-side counterpart of the retire/clear fix in #3479 and
reuses the remote-route resolution that landed with it. Reconciling a
loop bound to a remote route now reaches that route, so the existing
remote cases drive `consume` through the same faked transport their
other steps already use.
* no-mistakes(review): Fail loudly on unresolved routes and invalid staging homes
* no-mistakes(review): Fail collection when remote outbox is unreadable
* no-mistakes(review): Surface reassigned remote routes during empty collection
* no-mistakes(review): Fail remote collection on invalid registrations
* no-mistakes(review): Reject unsafe registration entries during remote collection
* no-mistakes(review): Restore healthy empty remote collection behavior
* no-mistakes(review): Skip remote collection for delivered registrations
* no-mistakes(review): Skip delivered registrations before route validation
* no-mistakes(document): Document remote follow-up collection semantics
* fix(bin): exclude secondmates from home-summary validity (#3504)
* fix(bin): exclude secondmates from home-summary child inventory
kind=secondmate meta records never have backlog rows, so counting them in unowned_children or terminal_in_flight made a clean main home look invalid once earlier ledger checks passed.
* no-mistakes(review): Cover terminal secondmate in-flight exclusion
* no-mistakes(ci): Updated the stock macOS Bash CI snapshot expectation from 15 to 16 tests. Verified all 16 snapshot/fleet-view tests pass under Bash 3.2.57 and `git diff --check` succeeds
* fix(bin): self-heal outcome indexes on first drain (#3509)
* fix(bin): self-heal status-outcome indexes on every drain
Missing ready markers were skipping the lost-wake backstop on non-Pi homes because only the Pi branch ran processed-init. Drain now rebuilds those indexes under the outcome lock and fails closed only on a real store fault.
* no-mistakes(review): Guard held-lock initialization and fail marker writes
* no-mistakes(document): Document cross-harness outcome-index self-healing
* fix(bearings): keep active children underway during captain holds (#3505)
* fix(bearings): keep active children underway beside a captain hold
Project each readable home's active children into Underway independently of the home-level captain-decision classification so a hold no longer hides live work.
* no-mistakes(review): Preserve Underway repos and disclose child truncation
* no-mistakes(review): Fall back to task project for Underway repos
* no-mistakes(ci): Updated the stock macOS Bash CI assertion from 44 to 45 Bearings tests, matching the newly added behavioral regression. Verified all 45 tests pass under /bin/bash, Bash syntax checks pass, and git diff validation is clean
* fix(pi): settle watcher delivery on Pi accepting the follow-up (#3513)
* fix(pi): settle watcher delivery on Pi accepting the follow-up
A follow-up queued while main is streaming joins the running run without
ever raising before_agent_start, so waiting on that event before clearing
the successor pipeline (#3498) stalled every later actionable close: no
successor started, no wake was delivered or offered to the branch, and the
turn-end guard woke main to re-arm by hand after every close.
The pipeline now settles once Pi accepts the follow-up. Consumption is
observed at before_agent_start for an idle main and at the user
message_start for a streaming main, and decides only what a replacement
session (/new, /resume, /fork, reload) replays. An exhausted restoration
delivers its typed failure without launching an arm past the retry bound,
which the stall had hidden. The replacement-coordinator map is typed so the
strict no-emit typecheck passes again.
Tests: the doubles no longer raise before_agent_start for a streaming send,
a portable regression drives two actionable closes while main streams and
proves the successor chain plus consumption-scoped replay, and a
credential-free real-SDK probe pins Pi's event contract for both the
streaming and the idle follow-up.
Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a
* fix(pi): retry a verified successor that fails during wake delivery
A verified successor can exit while the wake it was started for is still
being delivered, most plausibly during a branch turn that holds the
settlement for minutes. Its failure close arrived while the pipeline's
single-flight guard was set, so the close handler skipped the retry, and
the pipeline's end no longer launched an arm, which left the live
generation with no watcher and no retry timer.
The close handler now records that failure when the child had reported
readiness and was not retired by the restoration itself, and the pipeline
runs the ordinary bounded, lock-checked retry for it once the delivery
settles. A restoration started for a later pending supersedes it, and an
exhausted restoration still hands repair to main without a further arm.
The regression holds a branch settlement open while the verified
successor exits with a failure and proves one retry watcher starts after
the settlement releases, none while it is held.
Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a
* fix(bin): bound repeat stale wakes for parked workers (#3532)
* fix(bin): bound repeat stale wakes for a parked but live worker
A worker parked on a declared wait - `paused:` for an external or pipeline
wait, or a verified `captain-held` transfer - kept waking firstmate far inside
FM_PAUSE_RESURFACE_SECS. Observed as five consecutive alarms on one
captain-held worker and dozens across a day on a pipeline wait, and reported
upstream as four wakes in 75 minutes against a 3600s window.
pause_state_class deliberately answers `none` for a still-live agent even under
a declared wait, so a worker genuinely waiting on a decision is never silenced.
That classification is correct and is left alone; it routes every parked but
live worker through surface_nonterminal_stale on first sight of each distinct
stale hash, and an idle parked pane still churns its hash on a clock or a token
counter without changing what is being waited on.
Two places let that churn re-alarm:
- surface_nonterminal_stale queued the wake BEFORE consulting whether a wait was
declared, then wrote `.paused-resurfaced-<key>` - the very throttle that should
have suppressed it. The throttle was never read on this path and was advanced
by the wake it should have prevented.
- The hash-change path cleared that throttle through clear_pause_tracking
whenever the classification came back `none`, so each tick also bought the same
declared wait a fresh window. Fixing only the first site changes nothing.
Read the throttle before anything is queued and advance it only on a wake that
really fires, and on the hash-change path reset only the per-hash bookkeeping
while the declaration still stands, via a clear_stale_hash_tracking split so
neither half of clear_pause_tracking is duplicated. The throttle is keyed to the
declaration, not to the pane.
First sight still wakes, so an inconclusive state is still inspected, and the
window's end still re-surfaces once, so a forgotten wait cannot rot invisibly -
noise traded for a bounded cadence, never for silence. The wake identity stays
the plain `stale: <win>` the away-mode handoff depends on.
Tests cover both observed forms and were confirmed to fail against three
deliberate breaks: each site reverted on its own, and a re-surface that never
fires again.
* fix(document): Clarify declared-wait wake cadence documentation
* fix(ci): Captain, fixed the stale-throttle inheritance: cadence markers now bind to the current wait declaration, so replacement paused and captain-held waits each emit their first plain `stale:` wake. Added behavioral coverage for both forms. Bite proof failed as expected when identity matching was removed, then passed after restoration. Full watcher triage suite, `bin/fm-lint.sh`, syntax checks, and diff checks pass. Changes remain uncommitted for the outer executor
* fix(ci): Captain, fixed the confirmed Greptile finding. `resurface_absorbed` now applies a throttle only when its stored declaration scope matches the current wait, so replacement `paused:` and `captain-held` waits surface immediately without changing classification. Added executable coverage for both absorbed forms. Bite proof failed before the fix at the intended assertion; afterward the full watcher triage suite, `bin/fm-lint.sh`, shell syntax checks, and `git diff --check` passed
* fix(bin): accept the away-mode daemon as the turn-end supervision owner (#3567)
* fix(turnend): accept the away-mode daemon as the supervision owner
While state/.afk exists the away-mode daemon owns supervision and runs
bin/fm-watch.sh one-shot: the watcher exits on every wake and the daemon
starts its replacement. The turn-end guard tested for a live watcher
process holding the watch lock at that instant, so a turn boundary that
landed in the hand-off blocked with "TURN WOULD END BLIND" while
supervision was completely healthy, costing a full handling turn each
time.
Reproduced with the real daemon wrapping the real watcher and the real
guard sampling the same home: 6 of 40 samples blocked, every one of them
with the daemon alive and the beacon 2-3 seconds old, and a new watcher
pid on each cycle. After the fix the same reproduction blocks 0 of 40,
and killing the daemon and its watcher (away mode still on, beacon still
fresh) blocks again.
The guard now accepts a live, identity-matched daemon holding this home
as proof of supervision while away mode is active. The identity match is
the same discipline the watcher lock uses, so a recycled pid or a lock
left by a killed daemon proves nothing. The fresh-beacon half of the
predicate is unchanged: a daemon that stops restarting its watcher still
blocks once the beacon passes grace, a home with no supervisor blocks
exactly as before, and with away mode off the strict watcher predicate is
untouched.
The predicate reads only durable state, so it behaves identically for
every primary harness and runtime backend.
* no-mistakes(document): clarify away-mode daemon supervision proof and test coverage
* no-mistakes(document): generalize stale turn-end predicate summary in architecture.md
* fix(backlog): omit --file from row probes for non-markdown backends (#3582)
* fix(backlog): omit markdown file for beads probes
* no-mistakes(document): Narrow backlog addressing doc to mutations for backend-aware probes
* no-mistakes(ci): Fixed the Greptile P2 review comment (the only failing check) on tests/fm-backlog-atomicity.test.sh. The comment correctly noted that an exported TASKS_AXI_BACKEND environment variable would inherit into the spawned scripts and, because fm_tasks_axi_backend gives it top precedence, override each test case's .tasks.toml backend fixture — making the backend-specific argv assertions fail for environmental reasons. Fix: unset TASKS_AXI_BACKEND in the test harness right after sourcing tests/lib.sh, with a comment explaining why, so every case deterministically exercises its declared backend (4 lines added; no production code touched). Verified: reproduced the leak before the fix (TASKS_AXI_BACKEND=beads made the markdown dispatch case fail with 'beads show failed', exactly the reported failure mode); after the fix the full suite passes (0 failures, exit 0) both with and without TASKS_AXI_BACKEND=beads exported. The added lines are shellcheck-clean (the only shellcheck note, SC1091 on the lib.sh source line, pre-exists this change)
* fix(bin): classify progress updates on requested work as routine (#3589)
The supervision branch's verdict rule escalated every outcome that
answered a captain request, so "the work started" and "still working"
notes reached the captain with nothing to look at. The rule now keeps a
finished result of requested work captain-facing, even when healthy, and
treats start or still-working updates that bring no new artifact,
finding, or decision as routine. The captain list for review-ready PRs,
ask-user findings, exhausted blockers, credentials, and destructive or
security-sensitive cases is unchanged, as are the unsolicited-routine,
silent-fleet-review, and doubt-chooses-captain rules.
The fm_branch_report tool description and the two docs that restated the
old unconditional rule now point at the prompt's "Verdict: routine or
captain" section as the one owner instead of carrying a second copy.
* fix(bin): preserve captain calls during teardown (#3595)
* fix(bin): never close a captain call during cleanup
A scout that held its own work item for the captain, which is what
captain-hold-lifecycle prefers ("hold the work item the question gates"),
was closed by bin/fm-teardown.sh's automatic backlog transition. The
completion gate passed, cleanup ran, and the captain's question moved to
Done with no recorded answer: the one thing the policy says must never
happen. `tasks-axi done` closes a held row silently, and nothing in
teardown asked whether the row was the captain's own call.
bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when
the task is still an open captain call, 1 when it is not, 2 when that
cannot be established. It reads the row through the transition library's
backend-aware probe, so it addresses the same backlog teardown does; the
script's other commands now address the configured data directory the
same way instead of FM_HOME, which also fixes captain holds in a home
with a relocated data directory.
Teardown asks `open` before any destructive step and refuses on 2. On 0
only the close changes: after cleanup and still under the task's own
lock, the row gets one "Deliverable of the finished work" line at the end
of its body and returns to Queued through `tasks-axi reopen`, keeping its
hold, so it lands in Captain's Call instead of reading as work under way.
--force does not lift this: it authorizes discarding unlanded work, never
the captain's question. The deliverable goes into the body because
`tasks-axi update --report` rewrites the title of a row that is not Done.
The crash window reuses the pending-close record teardown already stages:
a `mode=retain` line makes the existing replay record the deliverable and
reopen instead of closing, with the same validator, stale-generation
check, cleanup-incomplete marking, and non-blocking bootstrap lock as an
ordinary close. A retained row the captain answered first simply retires
the record. No parallel record type, recovery command, or second bootstrap
loop is introduced.
Regressions run the real executables: the captain-held scout survives
cleanup queued, held, with its deliverable and on the board, only
`answer` closes it, --force keeps it open, and an ordinary scout still
closes with its report; an interrupted cleanup leaves the row untouched
and the next session start retains it; a relocated backlog keeps the
retention in its one configured file; and a ship row whose hold cannot be
read refuses cleanup before anything destructive.
Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np
* no-mistakes(review): Serialize captain holds and fix backend-aware listing
* no-mistakes(document): Update captain-call retention documentation
* no-mistakes(document): Fix relocated captain-hold backlog diagnostics
* fix(bin): deliver secondmate outcomes to the parent channel (#3592)
* fix(bin): deliver every secondmate outcome on the parent channel from the recording scripts
A secondmate's captain-facing outcomes could miss: the mate model addressed
the captain in its own unread chat instead of appending to the parent
channel, and a PR-ready report, a finding, a decision, a blocker, and a
failure all depended on that one remembered append. Make delivery
structural, so the parent channel never depends on the model:
- bin/fm-parent-channel-lib.sh is the one owner of channel resolution and
exact-line append-once; the merge outcome path and the inactive-outcome
scan now publish through it instead of two private copies.
- bin/fm-inactive-reconcile.sh gains a ledger-first path that runs on every
watcher poll in a secondmate home: a direct child's whole terminal done or
failed line is delivered at once with its note, recorded PR, mode, merge
posture, and scout report pointer, keyed and receipted so it is delivered
once, and the inactive path yields to it. `report <task-id>` runs the same
delivery for a caller holding the child's meta lock.
- bin/fm-pr-check.sh publishes the PR-ready line with the canonical URL at
registration.
- bin/fm-captain-hold.sh publishes a hold and its answer, keyed by task id
and resolution-record count, with no new persisted state.
- bin/fm-teardown.sh delivers the child's final line before removing its
record and refuses, retaining every record, while the channel cannot be
written.
- The charter opens with the parent-channel rule and confines the mate's own
appends to judgement; AGENTS.md carries the carve-out at the persona
address rule and the escalation list.
docs/secondmate-parent-channel.md records the design and its coverage, and
docs/verification/secondmate-parent-channel.md records the live run with real
tmux panes and both real watchers delivering every line with no model.
Supersedes #3569.
* no-mistakes(review): Fix parent outcome retries and reconciliation locking
* no-mistakes(review): Prevent busy children from starving ledger delivery
* no-mistakes(review): Correct ledger metadata and hold occurrence handling
* no-mistakes(review): Disambiguate ledger outcomes and normalize hold reasons
* no-mistakes(review): Close ledger races and preserve teardown records
* no-mistakes(document): Correct parent-channel receipt and scanner documentation
* no-mistakes(lint): Quote done arguments for ShellCheck compliance
* no-mistakes(ci): Fixed both CI failures. Updated GOTMP teardown fixtures for the new final-outcome reporter and isolated them from host tmux state. Updated the PR security assertion to distinguish the accepted PR-ready line from duplicate merge outcomes. Verified with both failing test suites, bash syntax checks, and git diff checks
* no-mistakes(ci): Fixed Greptile’s duplicate-delivery race in bin/fm-inactive-reconcile.sh. Ledger events now claim matching already-delivered inactive receipts using the prior status fingerprint, preventing duplicate parent reports while preserving later same-state completions. Added behavioral regression coverage. Verified inactive-reconcile tests, project lint, documentation audience checks, syntax, and diff checks. Teardown tests passed relevant cases before the documented pre-existing herdr-preflight-missing-adapter failure
* fix(bin): sync remote second mates to primary commit (#3599)
* fix(bin): sync remote second-mate homes to the parent primary commit
Session start and remote launch pointed a remote second-mate home at whatever
Firstmate copy its own host kept, so a home that had already advanced past that
copy refused as a non-fast-forward and every other home stopped at the host's
older commit while the primary ran ahead.
The parent now resolves ITS primary default-branch commit with the existing
helper and hands that commit to the host on both paths. Because a remote home
is a standalone clone, the host imports that one commit before advancing -
already present, else from that host's Firstmate copy without moving it, else
from the home's own origin - and then runs the SAME ff_target guards a local
home gets, so dirty, diverged, feature-branch, and unresolvable targets skip
untouched and the ancestry rules keep one owner. An unimportable target now
names /updatefirstmate instead of failing opaquely, and a host still running an
older Firstmate copy is reported the same way rather than echoing a bare
refusal.
The host-local launch leg no longer re-runs its own secondmate sync, so the
spawn it drives cannot re-target that host's copy after the parent has already
converged the home.
/updatefirstmate is unchanged: it still refreshes the remote code root from that
host's origin and then syncs the home to that refreshed copy, which is what the
sync call with no target commit means.
* no-mistakes(document): Document primary-targeted remote secondmate synchronization
* fix(bin): separate captain intent from firstmate specs (#3597)
* fix(bin): split brief task into captain intent and firstmate spec
Keep no-mistakes --intent as the captain's ask plus later captain words, not the build spec or worker tradeoffs.
* fix(bin): stop task-subsection copies at the next heading
Promotion was swallowing the scout Setup contract into Firstmate spec, and pre-subsection briefs lost their # Task body.
* no-mistakes(review): Validate brief content and preserve nested specifications
* no-mistakes(review): Scope placeholder validation to scaffold-only subsection bodies
* no-mistakes(review): Ignore fenced subsection headings during brief validation
* no-mistakes(review): Preserve captain intent across scout promotion
* no-mistakes(review): Enforce safe intent boundaries for legacy promotions
* no-mistakes(review): Allow marked legacy intent and reject empty promotions
* no-mistakes(review): Scope task parsing and overlay legacy intent contracts
* no-mistakes(review): Overlay current intent contract for all no-mistakes spawns
* no-mistakes(review): Preserve later captain clarifications in intent overlays
* no-mistakes(document): Document brief intent enforcement and ownership
* no-mistakes(ci): Updated spawn-related test fixtures to use valid Captain intent and Firstmate spec subsections, corrected launch-path expectations to launch-brief.md, and resolved ShellCheck quoting findings. Verified with fm-lint.sh and 15 affected behavior tests, including real Herdr tests; all passed
* no-mistakes(ci): Updated stale spawn/promotion fixtures in the Muse, Orca, secondmate-harness, and public-followup suites to provide valid Captain's intent and Firstmate spec subsections. Verified full Orca and secondmate-harness suites, targeted public-followup promotion behavior, Bash syntax, diff checks, and fm-lint
* fix: start a fresh supervision branch for every main session (#3600)
* fix(pi): start a new supervision branch conversation per main session
The supervision branch reopened one recorded conversation forever, so
every main session start reloaded the current generated prompt and then
weeks of accumulated thread, where a superseded rule could still outweigh
today's.
The branch conversation is now scoped to one main session: the session
generation owns the recorded conversation, so a cold start, /new,
/resume, /fork, or a reload always builds a new one, while a rebuild
inside one session (a model or effort change) still continues that
session's own conversation.
The dialog mirror re-anchors with it. Its durable cursor records what the
previous branch conversation received, so a /resume or reload - which
keeps main's own session file - would otherwise leave the new branch
blind to dialog main itself still has. The reset is bounded by the
current main session, and the cursor keeps advancing incrementally within
it. The durable outcome store and its processed marker are untouched, so
unacknowledged captain-facing outcomes still re-present on the new main
session.
* no-mistakes(document): Document fresh Pi supervision conversations
* no-mistakes(ci): Fixed the flaky concurrent inbox failure. Lock acquisition now retries when a competing lock disappears between a failed claim and inspection. Added a behavioral regression covering that race. Verified the full inbox test four times, project lint, and git diff checks
* feat: restart second mates after instruction updates (#3614)
* feat(update): restart second mates whose instructions changed
/updatefirstmate pulled new bytes onto disk and then asked each advanced
second mate to re-read them. A running agent holds AGENTS.md and every
loaded skill frozen from launch and no verified harness offers a reload,
so that steer could not reach a loaded skill at all and left the mate
holding two contradictory copies of its own job description.
An eligible mate is now restarted instead, in the same home and endpoint,
through the existing transactional relaunch. The restart is gated on the
mate first writing down the open work it holds only in conversation - the
open-record half of /stow, never its memory sweeps - so an unregistered
captain call is flushed before the conversation is spent. Anything that
leaves the reload unprovable falls back to the old re-read message and is
reported as exactly that, never as a clean reload.
Remote mates take the same path: fm-remote-secondmate-control.sh gains a
relaunch verb whose host-local leg runs that same control plane, since the
mate is an ordinary local secondmate from its host's point of view. The
primary resolves the profile and passes it explicitly, because
config/secondmate-harness is not inherited and the file on that host
belongs to a different home.
fm-update.sh now splits its advanced live mates into a restart set and a
nudge residual, and both sets require a changed instruction surface, which
also closes the over-nudge against the session-start sweep. Restart is
stricter still: a bin/-only advance reloads itself on the next call, so it
never costs a conversation.
Colocated tests cover the gating, the persist-then-restart order, the
task-subset persist request, each unsafe fallback, the remote hop, and the
remote sync's new instruction-surface report.
* no-mistakes(review): Fix restart correlation, concurrent waits, and lifecycle reporting
* no-mistakes(review): Parallelize relaunches and classify replacement incarnations
* no-mistakes(review): Gate restart actions on live agent state
* no-mistakes(review): Handle failed restart workers without hanging
* no-mistakes(review): Nudge legacy remotes and preserve persist recovery
* no-mistakes(review): Document one-time secondmate restart rollout
* no-mistakes(review): Honor arrived replies and refresh remote profiles
* no-mistakes(review): Revert remote parent profile reconciliation
* no-mistakes(review): Reset remote profile defaults and honor published results
* no-mistakes(review): Preserve fallback nudges for unverifiable secondmates
* no-mistakes(document): Document second-mate restart update flow
* no-mistakes(lint): Fix ShellCheck warnings in restart scripts
* perf: accelerate local validation with bounded concurrency (#3644)
* perf(tests): route gate verification through the bounded concurrent runner
Local validation was the pipeline's dominant cost: across 67 recorded
no-mistakes agent sessions on this repo, 99.3% of command execution was
`bash tests/*.test.sh`, run strictly one script at a time, and 2% of those
calls were killed by an agent-guessed timeout and paid for twice.
Three changes, each measured:
- `.no-mistakes.yaml` pins `commands.test` to
`bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner
already owns changed-file selection, bounded concurrency, the refusal of
unproven scripts, and a generous automatic per-script bound, so the gate's
baseline is neither a serial chain nor a guessed timeout. It stays
intent-targeted - the Test step still runs its evidence agent on top - and
excludes the live-Herdr family the required Herdr lane owns.
- `bin/fm-test-run.sh` gives a plain list of script paths the same bounded
automatic scheduler and automatic bound that `--changed` gets. Naming several
subjects is how a verification round asks for exactly those scripts. The
curated selections are untouched: `--lane` still composes CI shards whose
serial lane must stay serial, `--family` is what the required Herdr lane runs,
and `--all` stays a deliberate complete regression.
- `pr-forge` is admitted to the concurrent-safe family registry on two
consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those,
and records `secondmate` and `session-bootstrap` as refused with the exact
script and reason each failed on, so the refusals are actionable rather than
silent.
Measured on this host, 0 failures on both sides:
verification round, 4 scripts 448s chained -> 231s through the runner (-48%)
pr-forge family 409.2s at 1 worker -> 237.9s at 4 (1.72x)
watcher-wake-lock family 1311.1s at 1 worker -> 539.3s at 4 (2.43x)
A fourth lever was implemented and then removed because the measurement
refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made
`fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s
unchanged, back to back. Those sleeps are not overhead added to the clock -
they are how a test waits for a subject moving on fm-watch.sh's own one-second
cadence - so sampling less often only delays detection. It also broke
`fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a
settled condition. CONTRIBUTING.md records that result so the experiment is not
repeated.
* no-mistakes(review): Separate concurrent runs by isolation proof family
* no-mistakes(review): Limit automatic timeouts to changed-file validation
* no-mistakes(document): Clarify validation concurrency documentation
* fix: copy PR URLs from durable records (#3648)
* fix: copy PR URLs from records or abstain, never assemble them
Supervision reported a plausible but dead PR link three times because its
prompt demanded a full https:// URL at a moment when only a PR number was
observable, so the model assembled an owner/repository from memory, and the PR
check then accepted that URL and wrote it into the task record, after which the
model kept defending its own tool-endorsed guess over the worker's real link.
Three changes close that chain without any live forge lookup, so private
forges are treated exactly like public ones:
- bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy
or abstain" section requires a URL to be copied verbatim from a dura…
Ye1806431561
pushed a commit
to Ye1806431561/firstmate
that referenced
this pull request
Sep 20, 2026
* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work MAIN answered a supervision-branch outcome for completed captain-requested work (implementation done, PR ready for review and merge approval) with "Captain, shipshape.", reading section 9's no-action reply as covering it and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no captain-facing response is owed". Section 9 now limits the shipshape reply to true no-ops (idle re-read, empty heartbeat, consequence-free acknowledgement) and requires a short outcome response naming what finished and what word is needed whenever requested work finishes or a result needs the captain's word, even when a transcript entry already shows the substance. The Pi protocol's re-emit rule now says it bounds repetition only, and carries a worked example of the ready-for-review outcome whose correct processing turn a shipshape reply fails. No executable contract evaluates the content of MAIN's captain-facing reply, so the regression is the protocol example in the owner doc rather than a text-match test. * no-mistakes(document): Clarify captain-facing outcomes versus no-ops * docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line The document step condensed the Pi protocol's re-emit rule and dropped the worked example of a finished, ready-for-review outcome whose correct processing turn a "Captain, shipshape." reply fails. That example is the contract's regression: no executable contract evaluates the content of MAIN's captain-facing reply, so the owner doc's example is the test case. Restore it directly under the re-emit rule, prefixed as a regression example that is kept verbatim and never condensed or summarized away. * no-mistakes(review): Clarify captain outcome and decision-word requirements * no-mistakes(document): Clarify captain-facing completion outcomes * docs(pi): require the PR URL in the visible captain-facing outcome reply Captain review on the regression example: drop the sample reply string and say only that the ready-for-review outcome requires relaying a captain-facing outcome response, not just "Captain, shipshape.". Fold in the visible-PR-handoff failure seen this session: after the branch outcome reporting this fix green, MAIN's visible reply was only "Awaiting your merge call." with no PR URL, leaning on the dim anchor. Section 9's URL rule now also covers a review or merge ask and names the visible reply as where the URL goes, sourced from the ready status, pr= metadata, or the supervision branch's summary and never left to a transcript entry. The Pi protocol adds the same-way failure and places the captain-facing text in the final visible assistant reply after the fm_branch_processed call, because Calm hides assistant text emitted in the same step as a tool call as a working note. Investigation verdict, evidence in the PR comment: no recent PR caused the handoff failure; Pi has hidden same-step pre-tool assistant text since kunchenguid#2339 (2026-08-13), kunchenguid#4655 changed only the Claude Code mod, and kunchenguid#4658 touched only remote report transfer. * no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs * no-mistakes(document): Clarify captain-facing supervision outcomes * docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule The consolidated section 9 URL rule narrowed its trigger to a review or merge ask, dropping the "whenever a PR is mentioned" catch-all from kunchenguid#3648 that keeps every PR URL copied from a durable record and never assembled from memory. Restore that trigger as a union with the review or merge ask so the one consolidated rule covers both.
peterOC26
added a commit
to peterOC26/firstmate
that referenced
this pull request
Sep 20, 2026
* fix: surface inbound Relay media to responding agents (#3442)
* fix: surface inbound Relay attachments to the responding agent
A Discord support thread's screenshots were never seen by the agent
handling the mention. The relay delivered them and the poll stashed
them: the reporter's images arrived on the `thread_starter` entry of
`in_reply_to_chain` while the mention's own media list was empty. The
gap was in the responder's playbook, which enumerated a fixed field
list (`request_id`, `text`, `in_reply_to`, `in_reply_to_chain`) and so
made every other field, attachments included, invisible.
Fix it where the gap is, in prose:
- Read the complete payload object rather than a fixed field list, so
media and later relay fields are never skipped again.
- Fetch and view attached media with the agent's own tools, on the
mention and on every chain entry, and call out the common shape where
only the thread starter carries the screenshots.
- Restrict those fetches to known-good platform media hosts over https
(Discord: cdn.discordapp.com, media.discordapp.net,
images-ext-1.discordapp.net, images-ext-2.discordapp.net; X:
pbs.twimg.com, video.twimg.com), report a blocked host instead of
working around it, and treat everything fetched as untrusted public
input on the same terms as the surrounding thread text.
The poll stays out of it and downloads nothing, so no third-party bytes
are pulled on the polling path.
The new test pins the contract the playbook depends on: a mention in the
incident's shape, with an empty top-level media list and screenshots on
the thread starter, must reach the inbox with the payload intact and its
media URLs unfetched.
* no-mistakes(review): Preserve media authority and enforce poll-only fetching
* no-mistakes(document): Clarify Relay attachment safety prose
* fix(bin): defer inactive reconciliation during startup (#3480)
* Defer inactive startup reconciliation
* no-mistakes(review): Queue deferred inactive reconciliation diagnostics durably
* no-mistakes(review): Require worker phases to cover startup requests
* no-mistakes(review): Make diagnostic wakes safely acknowledgeable
* no-mistakes(document): Document deferred startup phase coverage
* fix(bin): bound wake drain presentation lock waits (#3475)
* fix: bound status presentation lock waits
* no-mistakes(review): Distinguish malformed presentation locks from live contention
* no-mistakes(review): Bound no-ack drain queue lock acquisition
* no-mistakes(document): Document bounded presentation-lock drain behavior
* no-mistakes(lint): Annotate bounded lock output global
* no-mistakes(ci): Added deterministic regression coverage for successful bounded-lock acquisition after live contention, verifying helper-to-caller PID ownership handoff and caller release. Verified with bash syntax checks, git diff checks, and the full fm-wake-queue test suite
* fix(bin): retire public follow-ups in remote homes (#3479)
* fix(relay): close a public loop whose work lives in a remote secondmate home
A public-followup loop bound to a REMOTE secondmate could never be closed.
`clear_public_followup_link` (bin/fm-public-followup.sh:701) required an
absolute recorded `work_home_path` for a `secondmate:*` work home, but a remote
route has no local path on this machine, so registration records that field
empty (bin/fm-public-followup.sh:291). Every close ran that clear first, so
`retire` died with "could not clear the legacy X link ... retained for
reconciliation" forever, and `deliver` posted the public reply and then stranded
the loop at `posted`. `--force` never covered that step.
The clear now goes to the remote home over that route's SSH transport, running
`fm-x-followup.sh --clear <work-id>` through `bin/fm-on.sh`. The route is decided
from `data/secondmates.md` before any local path is consulted, so a same-named
local directory can never stand in for a remote home, and registrations already
on disk retire without needing a new field. `fm-on.sh` passes ssh's status
through, so 255 stays the established "delivered but completion unknown" result
this codebase already reconciles: the close is refused, the registration and the
remote link are left exactly as they were, and the message names the unknown
completion instead of claiming a definite failure.
Local secondmate and `main` work homes are untouched, and `--force` still
governs only the unresolved-obligation refusal.
Three regression cases drive a remote route end to end, faking only the ssh
binary at the FM_SSH_BIN seam and then running the real remote entrypoint
against a local checkout, so the clear that must reach the remote home actually
happens there.
* no-mistakes(review): Guard remote link clears by request identity
* no-mistakes(review): Fail guarded clears on unreadable remote state
* no-mistakes(review): Reject guarded clears on non-writable remote state
* no-mistakes(review): Allow no-link retirement in non-writable remote state
* no-mistakes(document): Correct public-followup verification guarantee count
* no-mistakes(ci): Fixed the guarded link-clear race by ensuring absence is decided under the metadata lock whenever publication is possible. Added a behavioral concurrency regression test. Verified with fm-x-mode and fm-public-followup suites, Bash syntax checks, diff checks, and bin/fm-lint.sh
* no-mistakes(ci): Fixed the guarded link-clear race by refusing an unlocked absence decision when a publisher already owns the metadata lock in a non-writable directory. Added a behavioral concurrency regression test. Verified with fm-x-mode, fm-public-followup, syntax/diff checks, and fm-lint
* no-mistakes(ci): Fixed the guarded-clear race by refusing all guarded clears when the metadata parent is non-writable, including apparent link absence. Added a behavioral regression with a publisher waiting to create the lock, updated remote-retirement expectations and verification docs. Passed fm-x-mode, fm-public-followup, fm-lint, documentation audience, Bash syntax, and diff checks
* fix(relay): bound the guarded remote link clear so it refuses instead of hanging
The guarded clear checks that the remote state directory is writable before
taking the metadata lock, but that check cannot close the window: the parent can
turn non-writable between the check and lock creation, and a lock held by a live
holder is indistinguishable from that at the acquire. `fm_lock_acquire_wait` is
an unbounded `while ! try; do sleep 0.1; done`, so either case retried forever
and `deliver` or `retire` wedged with nothing reported, instead of returning the
retained-for-reconciliation refusal the guard exists to produce. This path runs
unattended over the secondmate transport, where a wedge is worse than either
outcome the guard defines.
The guarded clear now acquires through `fm_lock_acquire_wait_bounded`
(FMX_LINK_CLEAR_LOCK_TIMEOUT, default 10 seconds) and refuses on timeout through
the existing failure path. Unguarded local callers keep the ordinary unbounded
wait, so local behavior is unchanged.
The bounded primitive's header no longer claims presentation-only scope, since
this is a second authorized caller; nothing else in the shared lock
infrastructure changed.
The regression holds the metadata lock with a genuinely live process while
leaving the state directory writable, so the refusal can only come from the
bound and never from the writability precondition. Against the unbounded wait it
does not terminate at all; with the bound it refuses, retains the registration,
writes no receipt, and leaves the remote link untouched.
* no-mistakes(review): Harden lock-timeout regression with independent deadline
* no-mistakes(review): Restore no-op guarded clears on read-only state
* no-mistakes(document): Clarify remote public-followup cleanup contract
* fix(bin): support process events under symlinked homes (#3484)
* fix(bin): resolve process-event state roots before validating them
The process-event module validated the caller's spelling of a home's state
root instead of the directory it operates on: it required the supplied path
to equal its own lexical normalization, which rejects any path reached
through a symlinked ancestor. On macOS both /tmp and $TMPDIR are symlinks,
so an operator home under either could never claim a source. Reconcile still
reported the runner started, while the detached runner died writing "cannot
claim source" to the discarded stderr, and the source silently never fired.
Resolve the state root to its physical directory once, then apply the
existing private-directory validation to that resolved directory and derive
every path, recorded claim identity, and later confinement check from it.
This keeps the confinement contract for the directory actually operated on
rather than only for callers that already spelled it physically, and removes
the window where an ancestor symlink could be repointed between check and
use. Homes already spelled physically behave identically.
This was the single cause of both deterministic macOS failures in
tests/fm-procevent.test.sh ("reconcile never claimed the registered source")
and tests/fm-procevent-when.test.sh ("the winning concurrent arm did not
produce an outcome"). The new case pins the behavior with an explicit
symlinked-ancestor home, so it fails without the fix on any platform rather
than only where the temp root happens to be a symlink.
* fix(bin): pin the external capture staging boundary to its physical path
The extension capture path pinned its registry staging boundary by comparing
`pwd -P` against the caller-spelled registry directory, so a home reached
through a symlinked ancestor still refused to start an extension-backed
source after the state root itself resolved correctly. That left such a home
half working: built-in sources ran while external ones failed.
The staging preparer now prints the physical registry directory it validated,
matching the inbox and reservation preparers beside it, and the start path
pins on that returned path. The new end-to-end case drives the shipped
file-signal package from a symlinked home spelling.
* no-mistakes(review): Propagate canonical process-event state roots
* no-mistakes(review): Propagate canonical state to process-event adapters
* no-mistakes(document): Document physical process-event state roots
* fix(pi): deliver captain outcomes as deterministic transcript entries (#3312)
* fix(pi): persist captain outcomes visibly
* no-mistakes(review): Recover captain outcomes after cold-start lock acquisition
* no-mistakes(document): Document cold-start captain-outcome recovery
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes(review): Prove immediate Pi captain-outcome transcript delivery
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* fix(pi): process captain outcomes through a sequence-keyed turn
PR #3312 made every captain-facing supervision outcome a durable, exact-once
visible transcript entry with the read cursor advancing only after that entry
exists. That is the display half of the delivery contract. Left alone it turns
a probabilistic silent loss into a deterministic one: the captain sees an
anchor line, and firstmate never acts, because nothing opens a turn and
nothing records whether main ever processed the outcome.
The 2026-08-31 timeline showed the two shapes this must survive on the
previous hidden-turn path: seven delivered decision outcomes each answered by
an empty assistant message (cursor advanced, no retry, unanswered for close
to three hours), and two answered by an unrelated prior reply. Both happened
because delivery advanced the cursor at enqueue and accepted whatever the
next assistant message was.
Add the processing half on top of the persistence half:
- bin/fm-branch-outcome.sh keeps a processed marker separate from the read
cursor (`unprocessed`, `mark-processed --through`, `processed-init`). It
only advances through an explicit sequence-bound acknowledgement, never
past the read cursor and never backwards; an absent marker reads as zero
and `processed-init` migrates delivered history once so an upgraded home
is not re-presented its past.
- After the visible entry for a captain outcome exists, the extension hands
every still-unprocessed captain row to main as one hidden, typed
`fm-branch-process` request listing each `[seq N] task: summary`, opening
exactly one main turn. Main closes it only by calling the new
`fm_branch_processed` tool with the highest sequence listed. An unrelated,
empty, or paraphrased answer leaves the sequence open, and the same request
is presented again at the end of the next main run and at session start.
The first two presentations of a sequence set open a turn of their own;
after that the request rides the captain's next prompt so an ignored
request cannot loop, and a session replacement resets that budget.
Routine outcomes stay turn-free.
- The regressions cover exactly those incident shapes against the real store
scripts: an empty answer and an unrelated prior answer neither advance the
marker nor stop re-presentation, the acknowledgement is refused beyond the
read cursor and outside lock ownership, a partial acknowledgement keeps the
newer sequence open, and #3312's own assertions now forbid an unkeyed turn
rather than any turn. The store suite pins the marker's bounds and the
migration; the real-SDK guard for appendEntry persistence and model
exclusion is unchanged.
Docs move the protocol from "no model turn" to "one sequence-keyed processing
turn closed only by its acknowledgement", and the verification record carries
the dated run against Pi 0.84.4.
* no-mistakes(review): Harden outcome listing and sequence-bound acknowledgements
* no-mistakes(review): Harden outcome state validation and request pacing
* no-mistakes(review): Reject unsafe sidecars and unterminated outcome stores
* no-mistakes(review): Validate canonical mark-read cursor state
* no-mistakes(review): Guard cursor advancement against corrupt processed state
* no-mistakes(review): Bind acknowledgements to active processing requests
* no-mistakes(review): Reset pacing when processing sequence membership changes
* no-mistakes(review): Enforce silent outcome invariants at storage boundary
* no-mistakes(document): Document hardened captain outcome processing contracts
---------
Co-authored-by: kunchenguid <kun@kunchenguid.com>
* feat: add bounded concurrent Bearings ledger collection (#3481)
* feat: bound Bearings remote ledger collection
* no-mistakes(review): Clarify default remote-ledger collection behavior
* no-mistakes(review): Detach reconcile delivery from watcher loop
* no-mistakes(review): Enforce bounded snapshot and request captures
* no-mistakes(review): Bound legacy summary capture before parsing
* no-mistakes(review): Bound primary remote ledger captures
* no-mistakes(document): Correct snapshot and reconcile documentation
* no-mistakes(lint): Fix ShellCheck quoting in bounded collector
* no-mistakes(ci): Fixed all three CI failures: updated the macOS Bearings assertion to 44 tests, made the home-summary test deterministic and aligned with default ledger consumption, and increased the asynchronous reconcile retirement wait for loaded CI. Verified both focused suites, all 44 Bearings tests, ShellCheck, actionlint, Bash parsing, and git diff checks
* test: await reconcile request retirement
* no-mistakes(review): Avoid empty reconcile queue process churn
* no-mistakes(review): Read ledger summaries from immutable snapshots
* no-mistakes(review): Reject multi-document home ledger streams
* no-mistakes(review): Coalesce durable reconcile requests per target
* no-mistakes(review): Unify reconcile keys and reject snapshot streams
* no-mistakes(review): Key reconcile requests by stable target ID
* no-mistakes(document): Document per-target reconcile request coalescing
* no-mistakes(lint): Remove unused snapshot summary file variable
* no-mistakes(ci): Adjusted the concurrent collector regression’s end-to-end timing ceiling to account for stock macOS process/jq overhead outside the three-second remote collection budget, while remaining below the 15-second serial-read floor. Verified with stock /bin/bash 3.2: all 44 Bearings tests pass; bash syntax and git diff checks pass
* no-mistakes(ci): Fixed legacy summary validation to require exactly one top-level JSON document and added behavioral regression coverage. Stabilized CI by conditionally waiting longer for durable reconcile delivery and synchronously stopping the fm-on worker tree before fixture cleanup. Removed a redundant flaky healthy-path timing assertion; the wedged-reader test still proves concurrent bounded collection. Verified fm-bearings-snapshot, fm-secondmate-reconcile, and fm-on tests, plus project ShellCheck, bash syntax, and git diff checks
* ci: rebalance portable serial test shards (#3489)
* fix(ci): rebalance the portable serial shards on measured durations
The "Behavior portable serial 3" shard ran 17-20 minutes against its
20-minute job cap and intermittently timed out seconds after a passing
test, on branches and on main alike.
Shards are packed longest-processing-time from per-script duration hints,
and those hints were last measured on 2026-08-21 at 116 scripts. The lane
has since grown to 139 scripts and from ~42 to ~63 minutes: 17 scripts had
no hint at all and fell back to the 20 s default, and several existing
hints were low by 2-5x (fm-watch-triage 142 s hinted vs 263 s measured,
fm-public-followup 36 s vs 197 s). The partition therefore looked
perfectly balanced in hint space, 734.6 s per shard, while really running
11.5, 13.6, 18.8 and 16.5 minutes. Script-count balance, which is what the
tests asserted, stayed normal throughout and hid it.
Refresh the hints from the timing artifacts of three green runs, taking
the slowest measurement of each script so the balance holds on a slow
runner, and split the lane across five shards instead of four. Replayed
against those runs' real per-script durations the worst shard is now
12.54 minutes, 63% of the unchanged 20-minute cap, and the serial lane's
wall clock drops from ~20 to ~12.5 minutes.
Bound the drift that caused this rather than relying on the hints being
refreshed by hand: the coverage guard now reports the unmeasured share as
serial_unhinted= and refuses past PORTABLE_SERIAL_MAX_UNHINTED_PERCENT,
which leaves room for newly added tests while making a stale table fail
the guard instead of silently pushing one shard into its cap.
No test changes what it asserts and no test stops running; only the
partition across shards changes.
* no-mistakes(document): Clarify conservative shard timing aggregate
* fix(pi): fall back on incomplete supervision branch prompts (#3491)
* fix(pi): fall back after settled branch errors
* no-mistakes(review): Detect provider errors across prompt compaction
* no-mistakes(review): Preserve in-flight branch state across selection changes
* fix(pi): re-probe supervision branch after cooldown (#3497)
* fix(pi): recover supervision branch after cooldown
* no-mistakes(review): Defer branch recovery until prompt settlement
* no-mistakes(document): Clarify supervision cooldown recovery contract
* fix(bin): remove legacy remote snapshot reads (#3501)
* refactor: remove legacy remote summary reads
* no-mistakes(document): Document ledger-only snapshot reads
* no-mistakes(ci): Fixed the snapshot test fixture so ledger refreshes use the same fake executable PATH as the snapshot consumer. This preserves observable endpoint freshness after removing legacy summary computation. Verified stock Bash parsing and all 44 Bearings tests pass under /bin/bash; git diff checks pass
* no-mistakes(ci): Fixed the CI-only snapshot fixture failure by ensuring the bounded-ledger refresh uses its fake tmux backend. This removes host tmux availability as a source of nondeterminism. Verified all 44 Bearings tests pass, Bash syntax passes, and git diff checks are clean
* no-mistakes(ci): Fixed CI nondeterminism in the Bearings fixture: all local ledger refreshes now use the fixture’s fake tmux backend when available, instead of depending on host tmux state. Verified stock /bin/bash syntax, git diff checks, and all 44 Bearings tests with a deliberately failing host tmux
* fix(pi): preserve watcher continuity across session replacement (#3498)
* fix(pi): rearm watcher after session replacement
* no-mistakes(review): Queue actionable closes across Pi session replacement
* no-mistakes(review): Stop replacement arm when handoff persistence fails
* no-mistakes(review): Preserve actionable wakes through branch and late child races
* no-mistakes(review): Surface late handoff failures without crashing Pi
* no-mistakes(review): Coordinate replacement delivery settlement and unique handoff tokens
* no-mistakes(review): Retry stale deliveries and release settled claims
* no-mistakes(review): Distinguish branch settlement and retry handoff cleanup
* no-mistakes(review): Deduplicate persistent handoff cleanup alerts
* no-mistakes(review): Acknowledge watcher follow-ups only when consumed
* no-mistakes(review): Persist idle follow-ups until agent consumption
* no-mistakes(review): Preserve pending outcomes when handoff persistence fails
* no-mistakes(review): Arm replacement before awaiting prior delivery settlement
* no-mistakes(review): Adopt pending handoffs after lock reclamation
* no-mistakes(review): Prevent stale generations from adopting replacement handoffs
* no-mistakes(review): Scope replacement handoffs by watcher state
* no-mistakes(document): Clarify replacement handoff documentation
* no-mistakes(ci): Fixed the failing branch-extension tests to model the new settlement-promise contract. Failure cases now assert that delivery ownership returns to the watcher instead of expecting direct extension fallback. Verified the updated branch suite, Pi watcher suite, shell syntax, and diff checks
* no-mistakes(review): Update branch settlement tests and preserve chunked outcomes
* no-mistakes(document): Document watcher-owned replacement handoffs
* no-mistakes(document): Verify replacement handoff documentation
* test(pi): cover watcher-owned branch fallback
* no-mistakes(document): Refresh watcher-owned fallback documentation
* fix(bin): resurface task statuses missed by wake handling (#3495)
* fix(bin): resurface terminal statuses lost after branch handling
* test(watch): canonicalize process-event fixture homes
* no-mistakes(review): Index branch outcomes by causal status position
* no-mistakes(review): Recover outcome indexes and deduplicate resurfaced statuses
* no-mistakes(review): Handle legacy ambiguity and oversized status diagnostics
* no-mistakes(review): Keep unclassifiable oversized statuses silent
* no-mistakes(document): Document lost-wake outcome backstop
* no-mistakes(document): Update outcome backstop documentation
* no-mistakes(ci): Fixed CI regressions in wake-drain: parseable reserved-key decisions can no longer bypass the durable decision-fold guard, and status output is prepared and receipt-committed before presentation to prevent repeated one-shot outcomes after later failures. Added a behavioral regression for receipt commit failure and retry. Targeted backstop, correlation-token, decision-cursor, open-decision, unread-status, syntax, and diff checks pass locally. Shard-4 failures appeared unrelated/flaky; the network-parallel test passed locally
* no-mistakes(ci): Fixed the Greptile P1 data-loss issue by committing presentation receipts only after prepared output reaches stdout. Added behavioral coverage proving output failure leaves the backstop retryable and receipt failure may duplicate but never lose a presentation. Relevant wake-drain suites and syntax/diff checks pass. The shard-4 Pi extension failure is unrelated to this PR and did not warrant changes
* no-mistakes(ci): Stabilized tests/fm-bootstrap-network-parallel.test.sh by replacing scheduler-sensitive equal-sleep timing with bounded synchronization between mocked fetch and remote probes. This preserves detection of real serialization while avoiding false failures under CI load. Verified with five consecutive test runs, bash syntax validation, ShellCheck, and git diff checks. The separate Pi stock-rendering failure reproduces locally but is unrelated environment/version drift
* no-mistakes(ci): Fixed Behavior portable serial 4 by adding fm-classify-lib.sh and fm-timeout-lib.sh to the broken-root Pi test fixture; fm-branch-outcome.sh now depends on them. Verified the full Pi branch-extension suite with real-Pi checks skipped, the wake-drain outcome-backstop suite, Bash syntax, and git diff checks. Greptile findings are already addressed at HEAD; the no-mistakes attestation failure is external head-SHA state
* fix(bin): collect follow-up results from remote work homes (#3503)
* fix(bin): deliver typed terminal results from remote work homes
A public commitment whose work is bound to a REMOTE secondmate home could
never receive its typed terminal result. `fm-public-followup.sh brief`
printed an emit command carrying this home's own absolute path and this
checkout's own script path, neither of which exists on the machine the
worker runs on, so the worker had nothing it could write to that the
owning home would ever read - and `consume` kept finding nothing while
the promise stayed open.
The brief is now route-aware: for a remote work home it prints that
route's own code root and home with `--stage-in`, so the typed event is
staged in the home where the work actually runs, and the closing
paragraph names the owning home as the one on the other machine instead
of pointing at the path above it. The owning home collects those staged
results over the same SSH route it reaches that secondmate on, because
the transport only runs outbound: `consume` pulls them into its own
inbox and reconciles them exactly as it reconciles a local report.
Collection is non-destructive until the result is durably held, so a
dropped connection cannot lose a terminal result, and a route that could
not be reached is named in `consume`'s output with the promise left open
rather than reported as an empty inbox.
A local work home is untouched: the brief still prints `--home` with this
home and this checkout's script, and the event still lands directly in
this home's typed terminal-result inbox.
This is the emit-side counterpart of the retire/clear fix in #3479 and
reuses the remote-route resolution that landed with it. Reconciling a
loop bound to a remote route now reaches that route, so the existing
remote cases drive `consume` through the same faked transport their
other steps already use.
* no-mistakes(review): Fail loudly on unresolved routes and invalid staging homes
* no-mistakes(review): Fail collection when remote outbox is unreadable
* no-mistakes(review): Surface reassigned remote routes during empty collection
* no-mistakes(review): Fail remote collection on invalid registrations
* no-mistakes(review): Reject unsafe registration entries during remote collection
* no-mistakes(review): Restore healthy empty remote collection behavior
* no-mistakes(review): Skip remote collection for delivered registrations
* no-mistakes(review): Skip delivered registrations before route validation
* no-mistakes(document): Document remote follow-up collection semantics
* fix(bin): exclude secondmates from home-summary validity (#3504)
* fix(bin): exclude secondmates from home-summary child inventory
kind=secondmate meta records never have backlog rows, so counting them in unowned_children or terminal_in_flight made a clean main home look invalid once earlier ledger checks passed.
* no-mistakes(review): Cover terminal secondmate in-flight exclusion
* no-mistakes(ci): Updated the stock macOS Bash CI snapshot expectation from 15 to 16 tests. Verified all 16 snapshot/fleet-view tests pass under Bash 3.2.57 and `git diff --check` succeeds
* fix(bin): self-heal outcome indexes on first drain (#3509)
* fix(bin): self-heal status-outcome indexes on every drain
Missing ready markers were skipping the lost-wake backstop on non-Pi homes because only the Pi branch ran processed-init. Drain now rebuilds those indexes under the outcome lock and fails closed only on a real store fault.
* no-mistakes(review): Guard held-lock initialization and fail marker writes
* no-mistakes(document): Document cross-harness outcome-index self-healing
* fix(bearings): keep active children underway during captain holds (#3505)
* fix(bearings): keep active children underway beside a captain hold
Project each readable home's active children into Underway independently of the home-level captain-decision classification so a hold no longer hides live work.
* no-mistakes(review): Preserve Underway repos and disclose child truncation
* no-mistakes(review): Fall back to task project for Underway repos
* no-mistakes(ci): Updated the stock macOS Bash CI assertion from 44 to 45 Bearings tests, matching the newly added behavioral regression. Verified all 45 tests pass under /bin/bash, Bash syntax checks pass, and git diff validation is clean
* fix(pi): settle watcher delivery on Pi accepting the follow-up (#3513)
* fix(pi): settle watcher delivery on Pi accepting the follow-up
A follow-up queued while main is streaming joins the running run without
ever raising before_agent_start, so waiting on that event before clearing
the successor pipeline (#3498) stalled every later actionable close: no
successor started, no wake was delivered or offered to the branch, and the
turn-end guard woke main to re-arm by hand after every close.
The pipeline now settles once Pi accepts the follow-up. Consumption is
observed at before_agent_start for an idle main and at the user
message_start for a streaming main, and decides only what a replacement
session (/new, /resume, /fork, reload) replays. An exhausted restoration
delivers its typed failure without launching an arm past the retry bound,
which the stall had hidden. The replacement-coordinator map is typed so the
strict no-emit typecheck passes again.
Tests: the doubles no longer raise before_agent_start for a streaming send,
a portable regression drives two actionable closes while main streams and
proves the successor chain plus consumption-scoped replay, and a
credential-free real-SDK probe pins Pi's event contract for both the
streaming and the idle follow-up.
Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a
* fix(pi): retry a verified successor that fails during wake delivery
A verified successor can exit while the wake it was started for is still
being delivered, most plausibly during a branch turn that holds the
settlement for minutes. Its failure close arrived while the pipeline's
single-flight guard was set, so the close handler skipped the retry, and
the pipeline's end no longer launched an arm, which left the live
generation with no watcher and no retry timer.
The close handler now records that failure when the child had reported
readiness and was not retired by the restoration itself, and the pipeline
runs the ordinary bounded, lock-checked retry for it once the delivery
settles. A restoration started for a later pending supersedes it, and an
exhausted restoration still hands repair to main without a further arm.
The regression holds a branch settlement open while the verified
successor exits with a failure and proves one retry watcher starts after
the settlement releases, none while it is held.
Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a
* fix(bin): bound repeat stale wakes for parked workers (#3532)
* fix(bin): bound repeat stale wakes for a parked but live worker
A worker parked on a declared wait - `paused:` for an external or pipeline
wait, or a verified `captain-held` transfer - kept waking firstmate far inside
FM_PAUSE_RESURFACE_SECS. Observed as five consecutive alarms on one
captain-held worker and dozens across a day on a pipeline wait, and reported
upstream as four wakes in 75 minutes against a 3600s window.
pause_state_class deliberately answers `none` for a still-live agent even under
a declared wait, so a worker genuinely waiting on a decision is never silenced.
That classification is correct and is left alone; it routes every parked but
live worker through surface_nonterminal_stale on first sight of each distinct
stale hash, and an idle parked pane still churns its hash on a clock or a token
counter without changing what is being waited on.
Two places let that churn re-alarm:
- surface_nonterminal_stale queued the wake BEFORE consulting whether a wait was
declared, then wrote `.paused-resurfaced-<key>` - the very throttle that should
have suppressed it. The throttle was never read on this path and was advanced
by the wake it should have prevented.
- The hash-change path cleared that throttle through clear_pause_tracking
whenever the classification came back `none`, so each tick also bought the same
declared wait a fresh window. Fixing only the first site changes nothing.
Read the throttle before anything is queued and advance it only on a wake that
really fires, and on the hash-change path reset only the per-hash bookkeeping
while the declaration still stands, via a clear_stale_hash_tracking split so
neither half of clear_pause_tracking is duplicated. The throttle is keyed to the
declaration, not to the pane.
First sight still wakes, so an inconclusive state is still inspected, and the
window's end still re-surfaces once, so a forgotten wait cannot rot invisibly -
noise traded for a bounded cadence, never for silence. The wake identity stays
the plain `stale: <win>` the away-mode handoff depends on.
Tests cover both observed forms and were confirmed to fail against three
deliberate breaks: each site reverted on its own, and a re-surface that never
fires again.
* fix(document): Clarify declared-wait wake cadence documentation
* fix(ci): Captain, fixed the stale-throttle inheritance: cadence markers now bind to the current wait declaration, so replacement paused and captain-held waits each emit their first plain `stale:` wake. Added behavioral coverage for both forms. Bite proof failed as expected when identity matching was removed, then passed after restoration. Full watcher triage suite, `bin/fm-lint.sh`, syntax checks, and diff checks pass. Changes remain uncommitted for the outer executor
* fix(ci): Captain, fixed the confirmed Greptile finding. `resurface_absorbed` now applies a throttle only when its stored declaration scope matches the current wait, so replacement `paused:` and `captain-held` waits surface immediately without changing classification. Added executable coverage for both absorbed forms. Bite proof failed before the fix at the intended assertion; afterward the full watcher triage suite, `bin/fm-lint.sh`, shell syntax checks, and `git diff --check` passed
* fix(bin): accept the away-mode daemon as the turn-end supervision owner (#3567)
* fix(turnend): accept the away-mode daemon as the supervision owner
While state/.afk exists the away-mode daemon owns supervision and runs
bin/fm-watch.sh one-shot: the watcher exits on every wake and the daemon
starts its replacement. The turn-end guard tested for a live watcher
process holding the watch lock at that instant, so a turn boundary that
landed in the hand-off blocked with "TURN WOULD END BLIND" while
supervision was completely healthy, costing a full handling turn each
time.
Reproduced with the real daemon wrapping the real watcher and the real
guard sampling the same home: 6 of 40 samples blocked, every one of them
with the daemon alive and the beacon 2-3 seconds old, and a new watcher
pid on each cycle. After the fix the same reproduction blocks 0 of 40,
and killing the daemon and its watcher (away mode still on, beacon still
fresh) blocks again.
The guard now accepts a live, identity-matched daemon holding this home
as proof of supervision while away mode is active. The identity match is
the same discipline the watcher lock uses, so a recycled pid or a lock
left by a killed daemon proves nothing. The fresh-beacon half of the
predicate is unchanged: a daemon that stops restarting its watcher still
blocks once the beacon passes grace, a home with no supervisor blocks
exactly as before, and with away mode off the strict watcher predicate is
untouched.
The predicate reads only durable state, so it behaves identically for
every primary harness and runtime backend.
* no-mistakes(document): clarify away-mode daemon supervision proof and test coverage
* no-mistakes(document): generalize stale turn-end predicate summary in architecture.md
* fix(backlog): omit --file from row probes for non-markdown backends (#3582)
* fix(backlog): omit markdown file for beads probes
* no-mistakes(document): Narrow backlog addressing doc to mutations for backend-aware probes
* no-mistakes(ci): Fixed the Greptile P2 review comment (the only failing check) on tests/fm-backlog-atomicity.test.sh. The comment correctly noted that an exported TASKS_AXI_BACKEND environment variable would inherit into the spawned scripts and, because fm_tasks_axi_backend gives it top precedence, override each test case's .tasks.toml backend fixture — making the backend-specific argv assertions fail for environmental reasons. Fix: unset TASKS_AXI_BACKEND in the test harness right after sourcing tests/lib.sh, with a comment explaining why, so every case deterministically exercises its declared backend (4 lines added; no production code touched). Verified: reproduced the leak before the fix (TASKS_AXI_BACKEND=beads made the markdown dispatch case fail with 'beads show failed', exactly the reported failure mode); after the fix the full suite passes (0 failures, exit 0) both with and without TASKS_AXI_BACKEND=beads exported. The added lines are shellcheck-clean (the only shellcheck note, SC1091 on the lib.sh source line, pre-exists this change)
* fix(bin): classify progress updates on requested work as routine (#3589)
The supervision branch's verdict rule escalated every outcome that
answered a captain request, so "the work started" and "still working"
notes reached the captain with nothing to look at. The rule now keeps a
finished result of requested work captain-facing, even when healthy, and
treats start or still-working updates that bring no new artifact,
finding, or decision as routine. The captain list for review-ready PRs,
ask-user findings, exhausted blockers, credentials, and destructive or
security-sensitive cases is unchanged, as are the unsolicited-routine,
silent-fleet-review, and doubt-chooses-captain rules.
The fm_branch_report tool description and the two docs that restated the
old unconditional rule now point at the prompt's "Verdict: routine or
captain" section as the one owner instead of carrying a second copy.
* fix(bin): preserve captain calls during teardown (#3595)
* fix(bin): never close a captain call during cleanup
A scout that held its own work item for the captain, which is what
captain-hold-lifecycle prefers ("hold the work item the question gates"),
was closed by bin/fm-teardown.sh's automatic backlog transition. The
completion gate passed, cleanup ran, and the captain's question moved to
Done with no recorded answer: the one thing the policy says must never
happen. `tasks-axi done` closes a held row silently, and nothing in
teardown asked whether the row was the captain's own call.
bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when
the task is still an open captain call, 1 when it is not, 2 when that
cannot be established. It reads the row through the transition library's
backend-aware probe, so it addresses the same backlog teardown does; the
script's other commands now address the configured data directory the
same way instead of FM_HOME, which also fixes captain holds in a home
with a relocated data directory.
Teardown asks `open` before any destructive step and refuses on 2. On 0
only the close changes: after cleanup and still under the task's own
lock, the row gets one "Deliverable of the finished work" line at the end
of its body and returns to Queued through `tasks-axi reopen`, keeping its
hold, so it lands in Captain's Call instead of reading as work under way.
--force does not lift this: it authorizes discarding unlanded work, never
the captain's question. The deliverable goes into the body because
`tasks-axi update --report` rewrites the title of a row that is not Done.
The crash window reuses the pending-close record teardown already stages:
a `mode=retain` line makes the existing replay record the deliverable and
reopen instead of closing, with the same validator, stale-generation
check, cleanup-incomplete marking, and non-blocking bootstrap lock as an
ordinary close. A retained row the captain answered first simply retires
the record. No parallel record type, recovery command, or second bootstrap
loop is introduced.
Regressions run the real executables: the captain-held scout survives
cleanup queued, held, with its deliverable and on the board, only
`answer` closes it, --force keeps it open, and an ordinary scout still
closes with its report; an interrupted cleanup leaves the row untouched
and the next session start retains it; a relocated backlog keeps the
retention in its one configured file; and a ship row whose hold cannot be
read refuses cleanup before anything destructive.
Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np
* no-mistakes(review): Serialize captain holds and fix backend-aware listing
* no-mistakes(document): Update captain-call retention documentation
* no-mistakes(document): Fix relocated captain-hold backlog diagnostics
* fix(bin): deliver secondmate outcomes to the parent channel (#3592)
* fix(bin): deliver every secondmate outcome on the parent channel from the recording scripts
A secondmate's captain-facing outcomes could miss: the mate model addressed
the captain in its own unread chat instead of appending to the parent
channel, and a PR-ready report, a finding, a decision, a blocker, and a
failure all depended on that one remembered append. Make delivery
structural, so the parent channel never depends on the model:
- bin/fm-parent-channel-lib.sh is the one owner of channel resolution and
exact-line append-once; the merge outcome path and the inactive-outcome
scan now publish through it instead of two private copies.
- bin/fm-inactive-reconcile.sh gains a ledger-first path that runs on every
watcher poll in a secondmate home: a direct child's whole terminal done or
failed line is delivered at once with its note, recorded PR, mode, merge
posture, and scout report pointer, keyed and receipted so it is delivered
once, and the inactive path yields to it. `report <task-id>` runs the same
delivery for a caller holding the child's meta lock.
- bin/fm-pr-check.sh publishes the PR-ready line with the canonical URL at
registration.
- bin/fm-captain-hold.sh publishes a hold and its answer, keyed by task id
and resolution-record count, with no new persisted state.
- bin/fm-teardown.sh delivers the child's final line before removing its
record and refuses, retaining every record, while the channel cannot be
written.
- The charter opens with the parent-channel rule and confines the mate's own
appends to judgement; AGENTS.md carries the carve-out at the persona
address rule and the escalation list.
docs/secondmate-parent-channel.md records the design and its coverage, and
docs/verification/secondmate-parent-channel.md records the live run with real
tmux panes and both real watchers delivering every line with no model.
Supersedes #3569.
* no-mistakes(review): Fix parent outcome retries and reconciliation locking
* no-mistakes(review): Prevent busy children from starving ledger delivery
* no-mistakes(review): Correct ledger metadata and hold occurrence handling
* no-mistakes(review): Disambiguate ledger outcomes and normalize hold reasons
* no-mistakes(review): Close ledger races and preserve teardown records
* no-mistakes(document): Correct parent-channel receipt and scanner documentation
* no-mistakes(lint): Quote done arguments for ShellCheck compliance
* no-mistakes(ci): Fixed both CI failures. Updated GOTMP teardown fixtures for the new final-outcome reporter and isolated them from host tmux state. Updated the PR security assertion to distinguish the accepted PR-ready line from duplicate merge outcomes. Verified with both failing test suites, bash syntax checks, and git diff checks
* no-mistakes(ci): Fixed Greptile’s duplicate-delivery race in bin/fm-inactive-reconcile.sh. Ledger events now claim matching already-delivered inactive receipts using the prior status fingerprint, preventing duplicate parent reports while preserving later same-state completions. Added behavioral regression coverage. Verified inactive-reconcile tests, project lint, documentation audience checks, syntax, and diff checks. Teardown tests passed relevant cases before the documented pre-existing herdr-preflight-missing-adapter failure
* fix(bin): sync remote second mates to primary commit (#3599)
* fix(bin): sync remote second-mate homes to the parent primary commit
Session start and remote launch pointed a remote second-mate home at whatever
Firstmate copy its own host kept, so a home that had already advanced past that
copy refused as a non-fast-forward and every other home stopped at the host's
older commit while the primary ran ahead.
The parent now resolves ITS primary default-branch commit with the existing
helper and hands that commit to the host on both paths. Because a remote home
is a standalone clone, the host imports that one commit before advancing -
already present, else from that host's Firstmate copy without moving it, else
from the home's own origin - and then runs the SAME ff_target guards a local
home gets, so dirty, diverged, feature-branch, and unresolvable targets skip
untouched and the ancestry rules keep one owner. An unimportable target now
names /updatefirstmate instead of failing opaquely, and a host still running an
older Firstmate copy is reported the same way rather than echoing a bare
refusal.
The host-local launch leg no longer re-runs its own secondmate sync, so the
spawn it drives cannot re-target that host's copy after the parent has already
converged the home.
/updatefirstmate is unchanged: it still refreshes the remote code root from that
host's origin and then syncs the home to that refreshed copy, which is what the
sync call with no target commit means.
* no-mistakes(document): Document primary-targeted remote secondmate synchronization
* fix(bin): separate captain intent from firstmate specs (#3597)
* fix(bin): split brief task into captain intent and firstmate spec
Keep no-mistakes --intent as the captain's ask plus later captain words, not the build spec or worker tradeoffs.
* fix(bin): stop task-subsection copies at the next heading
Promotion was swallowing the scout Setup contract into Firstmate spec, and pre-subsection briefs lost their # Task body.
* no-mistakes(review): Validate brief content and preserve nested specifications
* no-mistakes(review): Scope placeholder validation to scaffold-only subsection bodies
* no-mistakes(review): Ignore fenced subsection headings during brief validation
* no-mistakes(review): Preserve captain intent across scout promotion
* no-mistakes(review): Enforce safe intent boundaries for legacy promotions
* no-mistakes(review): Allow marked legacy intent and reject empty promotions
* no-mistakes(review): Scope task parsing and overlay legacy intent contracts
* no-mistakes(review): Overlay current intent contract for all no-mistakes spawns
* no-mistakes(review): Preserve later captain clarifications in intent overlays
* no-mistakes(document): Document brief intent enforcement and ownership
* no-mistakes(ci): Updated spawn-related test fixtures to use valid Captain intent and Firstmate spec subsections, corrected launch-path expectations to launch-brief.md, and resolved ShellCheck quoting findings. Verified with fm-lint.sh and 15 affected behavior tests, including real Herdr tests; all passed
* no-mistakes(ci): Updated stale spawn/promotion fixtures in the Muse, Orca, secondmate-harness, and public-followup suites to provide valid Captain's intent and Firstmate spec subsections. Verified full Orca and secondmate-harness suites, targeted public-followup promotion behavior, Bash syntax, diff checks, and fm-lint
* fix: start a fresh supervision branch for every main session (#3600)
* fix(pi): start a new supervision branch conversation per main session
The supervision branch reopened one recorded conversation forever, so
every main session start reloaded the current generated prompt and then
weeks of accumulated thread, where a superseded rule could still outweigh
today's.
The branch conversation is now scoped to one main session: the session
generation owns the recorded conversation, so a cold start, /new,
/resume, /fork, or a reload always builds a new one, while a rebuild
inside one session (a model or effort change) still continues that
session's own conversation.
The dialog mirror re-anchors with it. Its durable cursor records what the
previous branch conversation received, so a /resume or reload - which
keeps main's own session file - would otherwise leave the new branch
blind to dialog main itself still has. The reset is bounded by the
current main session, and the cursor keeps advancing incrementally within
it. The durable outcome store and its processed marker are untouched, so
unacknowledged captain-facing outcomes still re-present on the new main
session.
* no-mistakes(document): Document fresh Pi supervision conversations
* no-mistakes(ci): Fixed the flaky concurrent inbox failure. Lock acquisition now retries when a competing lock disappears between a failed claim and inspection. Added a behavioral regression covering that race. Verified the full inbox test four times, project lint, and git diff checks
* feat: restart second mates after instruction updates (#3614)
* feat(update): restart second mates whose instructions changed
/updatefirstmate pulled new bytes onto disk and then asked each advanced
second mate to re-read them. A running agent holds AGENTS.md and every
loaded skill frozen from launch and no verified harness offers a reload,
so that steer could not reach a loaded skill at all and left the mate
holding two contradictory copies of its own job description.
An eligible mate is now restarted instead, in the same home and endpoint,
through the existing transactional relaunch. The restart is gated on the
mate first writing down the open work it holds only in conversation - the
open-record half of /stow, never its memory sweeps - so an unregistered
captain call is flushed before the conversation is spent. Anything that
leaves the reload unprovable falls back to the old re-read message and is
reported as exactly that, never as a clean reload.
Remote mates take the same path: fm-remote-secondmate-control.sh gains a
relaunch verb whose host-local leg runs that same control plane, since the
mate is an ordinary local secondmate from its host's point of view. The
primary resolves the profile and passes it explicitly, because
config/secondmate-harness is not inherited and the file on that host
belongs to a different home.
fm-update.sh now splits its advanced live mates into a restart set and a
nudge residual, and both sets require a changed instruction surface, which
also closes the over-nudge against the session-start sweep. Restart is
stricter still: a bin/-only advance reloads itself on the next call, so it
never costs a conversation.
Colocated tests cover the gating, the persist-then-restart order, the
task-subset persist request, each unsafe fallback, the remote hop, and the
remote sync's new instruction-surface report.
* no-mistakes(review): Fix restart correlation, concurrent waits, and lifecycle reporting
* no-mistakes(review): Parallelize relaunches and classify replacement incarnations
* no-mistakes(review): Gate restart actions on live agent state
* no-mistakes(review): Handle failed restart workers without hanging
* no-mistakes(review): Nudge legacy remotes and preserve persist recovery
* no-mistakes(review): Document one-time secondmate restart rollout
* no-mistakes(review): Honor arrived replies and refresh remote profiles
* no-mistakes(review): Revert remote parent profile reconciliation
* no-mistakes(review): Reset remote profile defaults and honor published results
* no-mistakes(review): Preserve fallback nudges for unverifiable secondmates
* no-mistakes(document): Document second-mate restart update flow
* no-mistakes(lint): Fix ShellCheck warnings in restart scripts
* perf: accelerate local validation with bounded concurrency (#3644)
* perf(tests): route gate verification through the bounded concurrent runner
Local validation was the pipeline's dominant cost: across 67 recorded
no-mistakes agent sessions on this repo, 99.3% of command execution was
`bash tests/*.test.sh`, run strictly one script at a time, and 2% of those
calls were killed by an agent-guessed timeout and paid for twice.
Three changes, each measured:
- `.no-mistakes.yaml` pins `commands.test` to
`bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner
already owns changed-file selection, bounded concurrency, the refusal of
unproven scripts, and a generous automatic per-script bound, so the gate's
baseline is neither a serial chain nor a guessed timeout. It stays
intent-targeted - the Test step still runs its evidence agent on top - and
excludes the live-Herdr family the required Herdr lane owns.
- `bin/fm-test-run.sh` gives a plain list of script paths the same bounded
automatic scheduler and automatic bound that `--changed` gets. Naming several
subjects is how a verification round asks for exactly those scripts. The
curated selections are untouched: `--lane` still composes CI shards whose
serial lane must stay serial, `--family` is what the required Herdr lane runs,
and `--all` stays a deliberate complete regression.
- `pr-forge` is admitted to the concurrent-safe family registry on two
consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those,
and records `secondmate` and `session-bootstrap` as refused with the exact
script and reason each failed on, so the refusals are actionable rather than
silent.
Measured on this host, 0 failures on both sides:
verification round, 4 scripts 448s chained -> 231s through the runner (-48%)
pr-forge family 409.2s at 1 worker -> 237.9s at 4 (1.72x)
watcher-wake-lock family 1311.1s at 1 worker -> 539.3s at 4 (2.43x)
A fourth lever was implemented and then removed because the measurement
refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made
`fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s
unchanged, back to back. Those sleeps are not overhead added to the clock -
they are how a test waits for a subject moving on fm-watch.sh's own one-second
cadence - so sampling less often only delays detection. It also broke
`fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a
settled condition. CONTRIBUTING.md records that result so the experiment is not
repeated.
* no-mistakes(review): Separate concurrent runs by isolation proof family
* no-mistakes(review): Limit automatic timeouts to changed-file validation
* no-mistakes(document): Clarify validation concurrency documentation
* fix: copy PR URLs from durable records (#3648)
* fix: copy PR URLs from records or abstain, never assemble them
Supervision reported a plausible but dead PR link three times because its
prompt demanded a full https:// URL at a moment when only a PR number was
observable, so the model assembled an owner/repository from memory, and the PR
check then accepted that URL and wrote it into the task record, after which the
model kept defending its own tool-endorsed guess over the worker's real link.
Three changes close that chain without any live forge lookup, so private
forges are treated exactly like public ones:
- bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy
or abstain" section requires a URL to be copied verbatim from a durable
record (the done: PR <url> status line, pr= metadata, or the backlog note),
forbids assembling owner, repository, host, or number from memory, and has
the branch report only the identifier it actually holds when no record names
the URL yet, leaving the PR check unarmed until the worker's ready line
arrives. AGENTS.md section 7 and 9 carry the same copy-or-abstain rule for
main in place of the bare full-URL mandate.
- Worker briefs (bin/fm-brief.sh, ship and scout rules) require the full
https:// URL wherever a PR is mentioned - status line, terminal, or summary -
never a bare "PR 108", so the link is in view as early as the number is.
- bin/fm-pr-check.sh refuses, offline and before any side effect, a URL that
the task's own done lines contradict, printing both spellings; a log naming
no URL still records the argument as before. fm_pr_status_ready_urls in
bin/fm-pr-lib.sh owns reading those lines. The refusal also reaches
bin/fm-pr-merge.sh, so nothing merges under a contradicted URL.
Tests cover the offline refusal with zero side effects, the recorded spelling
being accepted, markdown-wrapped and punctuated URLs, working lines not
counting, the merge wrapper propagation, a self-hosted merge request with no
forge call, the prompt carrying the rule, and the brief carrying the worker
rule.
* no-mistakes(review): Remove stale PR URL enforcement
* no-mistakes(ci): Removed backlog notes as an accepted PR identity source. PR URLs may now be copied only from the task’s `done: PR <url>` status or canonical `pr=` metadata; otherwise supervision reports only the known identifier and leaves PR checking unarmed. Updated related guidance/docs and verified with branch-supervision tests, brief tests, ShellCheck, and `git diff --check`
* fix(bin): disable Claude feedback drafts for fleet launches (#3661)
* fix(bin): disable Claude's feedback-draft flow for fleet-launched agents
Scope --settings '{"feedbackDrafts":"off"}' to every Firstmate-launched
Claude crewmate and secondmate, so /bug and /feedback never queue or
submit a bug report on the captain's behalf. feedbackDrafts is the
documented settings key (Claude Code changelog 2.1.247); the
per-launch CLI flag never touches the captain's global settings.json.
Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3
* no-mistakes(review): Prevent managed settings from re-enabling Claude feedback drafts
* no-mistakes(document): Fix Claude feedback documentation formatting
* fix(bin): layer both feedback-draft controls for defense in depth
The prior --settings-only fix can be overridden by a managed Claude
settings policy (feedbackDrafts precedence). Keep CLAUDE_CODE_SEND_FEEDBACK=0
alongside --settings '{"feedbackDrafts":"off"}': either control alone
disables the SendFeedback tool, so a managed override of one still
leaves the other in force.
Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3
* no-mistakes(document): Document Claude feedback-draft suppression ownership
* feat(tests): run three more validation families concurrently (#3662)
* perf(tests): admit three more families to concurrent validation
The three families that `docs/fm-test-isolation-proof.md` recorded as refused
were not refused for concurrency. Each blocker was a test that decided a
property by wall clock, or a script filed where it cannot run. Fixing those
three things admits all three families and recovers 28.6 minutes of local
validation with no assertion removed or weakened.
- `tests/fm-backlog-handoff.test.sh` injected its pre-move crash by killing the
handoff, sleeping a fixed second, then delegating the move to the real
binary. Nothing ever killed the fake, so on a host slow enough for the case's
next assertions to take longer than a second, the orphan woke and completed
the very move the case requires left undone, and recovery then failed with
`Task "pre-move-crash" not found in this backlog`. Watching the two backlogs
during the injected crash showed exactly that, the item moving one second
after the crash. All four crash injections in the file now go through a new
`fm_fake_crash_injector` shim that signals the target and returns only once
it is observably gone, and the pre-move fake never delegates the move at all.
- `tests/fm-session-start.test.sh` proved the startup digest does not block on
a slow current-state read by timing the whole digest against a fixed
eight-second sleep, which a loaded host exceeds without the property being
violated. It now holds that read open until the case releases it and asserts,
the moment the digest returns, that the read has not finished. A digest that
waited would wait indefinitely rather than for an interval a slow host can
out-run, so the assertion is stronger than the bound it replaces. Its scan
budget moves to the maximum, because the old value left two seconds of margin
over the fixed sleep and measured the host rather than the deadline that
`tests/fm-inactive-reconcile.test.sh` owns.
- `fm-backend-herdr-focus-flash-e2e` was filed in the family map's catch-all,
which put it in the portable serial lane, where Linux CI gate-skips it: that
real-Herdr regression was running nowhere. It moves to `real-herdr-gated` and
the required Herdr lane. `fm-claude-stop-autoarm-live-e2e` gate-skips on its
opt-in variable and moves to `live-harness-optin`.
The 28 remaining ungrouped scripts become an enumerated `standalone` family
instead of admitting `unclassified` itself. `unclassified` is the family map's
`*)` arm, so admitting it would silently grant concurrency to every test added
afterwards, which is exactly the population with no proof. A new test still
lands in `unclassified` and stays serial, and `tests/fm-test-run.test.sh`
covers that split behaviorally.
Each family passes two consecutive four-worker proofs with zero failures. On
the production runner, `secondmate` goes 1233.1s to 453.4s, `session-bootstrap`
756.4s to 286.4s, and `standalone` 724.6s to 261.1s: 2.71x overall and 1713.2s
recovered. The whole suite runs 177 scripts in 52.6 minutes of wall clock
against 121 minutes of summed script time.
* no-mistakes(document): Refresh concurrent validation and shard documentation
* no-mistakes(ci): Fixed the real-Herdr focus-flash E2E race exposed by reclassification. Part C now starts its persistent child atomically via `pane run` and verifies stable child identity through Herdr’s public `process-info` interface, avoiding the racy send-text/send-keys sequence and platform-specific `ps` matching. Verified with bash syntax checking, ShellCheck, git diff checks, and the complete E2E test on Herdr 0.8.2
* feat: structure no-mistakes ask-user escalations (#3670)
* feat(brief): structure no-mistakes ask-user escalation as event + snapshot file
Crewmates escalating a no-mistakes ask-user gate now report one status
event naming every finding id plus a snapshot file holding the gate's
axi finding records verbatim (id, severity, file, line, description,
authority), using the same shape even for a single finding. The status
line never paraphrases. The format is defined once in fm-dod-lib.sh and
rendered into both the scout and ship rule 6 in fm-brief.sh, so a
promoted scout - whose rule 6 fm-promote.sh preserves unchanged - gets
the identical contract as a freshly-spawned no-mistakes ship worker.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PpiWaDerbYavTLPPtEjQei
* no-mistakes(review): Preserve ask-user escalation output contract
* no-mistakes(review): Align escalation format test expectation
* no-mistakes(review): Scope ask-user escalation instructions correctly
* no-mistakes(review): Remove ask-user from generic decision rules
---------
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* fix(bin): require self-sufficient no-mistakes intent (#3671)
* fix(bin): require a self-sufficient no-mistakes intent
A no-mistakes worker's --intent is only as useful as the string it
passes. PR #3604 shipped with an intent that was only "do 1, 2, 3, 7
from the report": the real contract lived in a private scout report and
never reached --intent, so nobody holding that string plus the codebase
could have derived the specification.
This is pure instruction at the contract's one owner; no spawn-side or
promotion-side check is added.
- bin/fm-dod-lib.sh: the generated no-mistakes Definition of done now
states that the --intent string must be self-sufficient (the string
plus the codebase reconstructs roughly the same specification) and
tells the worker to write the substance of any report, decision,…
sanis
added a commit
to sanis/firstmate
that referenced
this pull request
Sep 21, 2026
…dpoint reclaim (#28) * fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (#4424) * fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (#42) * fix(pr-merge): read a plan-gated 403 on branch rules as no merge queue github_read_queue_method left status=unreadable for every failed rules read, including a 403 whose body is GitHub's own "Upgrade to GitHub Pro or make this repository public" message. A repository whose plan cannot expose branch rules cannot have a merge_queue rule either, so that specific 403 now resolves to status=none instead of unreadable - unblocking the away-merge grant on private repos without GitHub Pro. Any other failure (auth, rate limit, network, 404, unrelated 403) still reads as unreadable. * no-mistakes(document): Update stale away-merge queue-grant comment for plan-gated 403 --------- Co-authored-by: NewAiCoder <claude@theinbtw.com> * no-mistakes(review): Fix misleading away-queue-grant comment in fm-pr-merge and its test * no-mistakes(document): Update architecture.md for plan-gated-403 merge queue exception --------- Co-authored-by: NewAiCoder <claude@theinbtw.com> * fix(bin): select suites that read a changed top-level test fixture (#4246) * fix(tests): select readers of a changed top-level test fixture bin/fm-test-run.sh --changed recognised shared test helpers by an explicit list, tests/lib.sh|tests/*-helpers.sh|tests/fixtures.sh. A top-level tests/*-fixture.sh matched none of those, fell through to the tests/* catch-all, and was marked unmapped, so selection aborted with "no changed-test mapping for source path" and the run selected nothing at all. tests/herdr-client-pair-fixture.sh and tests/remote-herdr-fixture.sh are real shared fixtures with real consumers, so any branch touching one of them left a validation pipeline driving --changed with a hard abort rather than a narrowed selection. Extend the helper arm to tests/*-fixture.sh rather than routing it through the tests/fixtures/*/* arm. Both arms resolve consumers with the same reference scan, and that scan is what selects the right suites here: it finds exactly the tests that read the fixture. The fixtures/ arm adds only a directory-keying step, which has nothing to key on for a top-level file, so the helper arm is the same behaviour with no extra machinery. A tests/ path nothing reads still reaches the catch-all and still refuses loudly. Refs https://github.com/kunchenguid/firstmate/issues/4100 * no-mistakes(test): order nested fixtures arm before top-level fixture glob * no-mistakes(document): document tests/ shared-file mapping contract and arm order * no-mistakes(review): drop vacuous test phase, correct header claim, restore comment * fix(bin): treat Claude Code's default external-imports flags as never asked, not declined (#4387) * fix(bin): read Claude Code's default external-imports flags as never asked, not declined (#4378) fm-claude-trust.sh refused the whole trust registration whenever the project-root entry carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code writes that value only on an explicit "No, disable". Claude Code's default project entry carries Approved and WarningShown both false before the dialog is ever shown, so every such project refused every spawn. Only Approved === false with WarningShown === true — the pair the dialog writes on a decline — now counts as a decline. false/false behaves like an absent flag: trust is registered and no import consent is manufactured. New case test_project_root_entry_default_import_flags_are_not_a_decline fails on b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31, bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * no-mistakes(review): Correct harness doc's external-imports decline predicate --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(bin): keep operator-address labels out of no-mistakes intent (#4445) * fix(brief): keep operator address out of composed intent Teach raw-word authoring for intent sections and mid-task relays, with a neutral [captain] provenance marker for legacy mixed tasks. Keep headings and contract prose outside the serialized intent body. The legacy selector already excluded the old speaker labels from its output; preserve that read compatibility. The reproduced leak comes from adding labels inside a modern intent body, not from the legacy selector. Do not scrub actual request content. Add exact serialized-input and generated-contract regressions, retaining refusal of unmarked legacy tasks and coverage of scout promotion. Fixes https://github.com/kunchenguid/firstmate/issues/3882 * no-mistakes(review): Refuse operator-address lines in Captain's intent body * no-mistakes(document): Document operator-address refusal in intent contract comments * fix: classify OpenCode ellipsis hint as idle (#4451) * fix(composer): recognize Grok 1.0.5's oversized titled bottom border as a proven empty composer (#4455) * fix(composer): accept Grok title overhang * no-mistakes(review): summary: named Grok overhang constant, doc caveat, restored tmux typed-title coverage * fix(bin): translate Stop hook timeout signals into durable auto-arm failure (#4474) * fix(bin): recover Claude auto-arm after timeout * no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list * fix(spawn): establish Claude task channel authority (#4464) * fix(spawn): establish Claude task channel authority * no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference * fix(bin): refuse fm-control.sh exit when the composer holds unproven or pending text (#4458) * fix: guard relaunch exit against pending input * no-mistakes(review): Verifying test run in progress * no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard * no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now * fix(spawn): establish crewmate identity first (#4481) * fix(bin): reconcile redundant secondmate divergence during updates (#4460) * fix: reconcile diverged secondmate updates * no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md * no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md * feat: enable gpt-5.6-luna max reasoning for crew dispatch (#4497) * fix(dispatch): support Codex Luna max effort * no-mistakes(review): use portable CODEX_HOME path in codex effort reference * feat(calm): render smooth Unicode swell with asymmetric two-color sail (#4498) * feat(calm): render smooth Unicode swell * feat(calm): make sails asymmetric * feat(calm): use quarter sail glyph * no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer * no-mistakes(document): docs: sync calm wave phase doc comment * no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。 * fix(bin): supersede stale scout delivery text in brief.md on promotion (#4491) * fix: supersede scout delivery brief on promotion * fix: preserve ship safety contract after promotion * no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch * fix(bin): make captain holds work on hosts with an older JSON::PP, and stop cleanup dropping accents from a held body (#4471) * fix(bin): let captain holds work on hosts with an older JSON::PP Holding a task for the captain, and the cleanup that keeps a captain-held row open, both fail outright on any host whose JSON::PP defaults allow_nonref off - 2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`, but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older library rejects that whole value with "must be object or array". The consequence is fleet-wide on such a host, not one broken command: a worker there cannot formally record a decision for the captain at all. It can only mention the decision in passing in a status line, where it can be missed - which is how a real decision goes unrecorded. The hold reports that the task lost its hold-set stamp; the cleanup cannot return the row to Queued. Both call sites now ask for allow_nonref explicitly rather than inheriting whatever the installed library defaults to. The second one is worth naming: its `/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string case that fails, so the guard selects for the failing input rather than protecting against it. The regression case forces the older default back off for every perl the commands spawn, then drives both paths - holding a task that carries a body, and tearing down a captain-held row whose deliverable must still be appended. It also probes that the simulation genuinely rejects a bare scalar, so the case cannot pass vacuously on a lenient host. Each half was verified failing on its own unfixed call site with that site's real error message. Suites: fm-captain-hold-lifecycle 51 cases, fm-backlog-atomicity 99 cases, 0 failures. Verification limit: the mechanism is reproduced and tested, but neither fix is verified against a real JSON::PP 2.27202 host, because none is in the loop. This laptop runs 4.06, where the bug does not manifest. `bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a brace-delimited object before decoding, so allow_nonref never applies. * fix(bin): stop cleanup silently dropping accented characters from a held body Cleanup rewrites a captain-held row's body to append the finished work's deliverable, and the decoder it reads that body with printed decoded characters to a stream with no `:raw` layer. A character at or below U+00FF then came out as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the accent. `fm_backlog_retain` writes that body straight back through `--body-file`, and nothing reported an error - the character was simply gone from a row still waiting on the captain. The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus `utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already used. Review of the parent commit found this on one of the lines that commit already changed. It predates that change. The test asserts bytes rather than decoded strings, because comparing strings cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any character above U+00FF makes perl print the whole string as UTF-8, so one body carrying both an accent and an em dash passes even unfixed and proves nothing. Verified failing before the fix on the accented row, passing after. Suites: fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures. * no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc * no-mistakes(review): drop whole-file UTF-8 check from retained-body test * no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc * no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc * fix(bin): read codex 0.154's idle braille starfield rows as composer furniture (#4532) * fix(composer): read codex 0.154's idle starfield and status footer as furniture codex-cli 0.154.0 animates a braille "starfield" around its idle composer: on the row above the bold `›` prompt row, on the `›` row behind the SGR-2 dim `Ask Codex to do anything` placeholder, and on the row below it, then draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`). The cells are truecolor greys on both sides of the ghost luminance ceiling, so the brighter ones survive ghost stripping, and the rows below the glyph carry no structural edge. The shared classifier selected the bare `›` shape, extended its wrap region over the two rows beneath the glyph, read the survivors and the footer as wrapped typed input, and answered `pending`; the steering doorbell defers on exactly that verdict, so no doorbell ever reached an idle codex 0.154 pane. bin/fm-composer-lib.sh now recognises that furniture by shape, declared once next to the idle placeholders and reached from the two wrap-region boundary points: - a row whose non-whitespace content is entirely braille cells (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it never counts as wrapped typed content and bounds a bare composer's wrap region; braille behind the glyph row's content is stripped before the emptiness decision when nothing else follows the glyph; a row mixing braille with other text stays typed content; - the codex status footer bounds the wrap region exactly as omp's status row does, anchored on the effort token, a spaced middle dot, and a `~` or `/` path cell, so a typed `fix · tests` stays composer input; - `^Ask Codex to do anything$` joins the verified idle-placeholder set; the ghost strip remains what proves that row empty, and the bare-row rule that bright placeholder text is real input is unchanged. Unchanged: the strict blank-row rule, the styled=0 degradation (a plain cmux/orca capture of this screen still reads `unknown`, never `pending`), FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape. tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte with the divergence (letters in place of the starfield read `pending`) and the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh is the default-on live guard (token-free, skips explicitly without codex or tmux) that launches the installed codex idle and asserts `empty` through both the tmux and the cursorless styled reads, naming codex --version on failure. docs/verification/runtime-backends.md records the dated Herdr evidence: `pending` before, `empty` after, on the captured screen. * no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry --------- Co-authored-by: Todd Billings <todd@usdvcapital.com> * fix(bin): refuse empty text steers in fm-send (#4259) * fix(bin): refuse empty text steers in fm-send A marked secondmate request sent with an empty message delivered only marker and correlation bytes and minted a pending-reply expectation the parent could never see resolved, stalling the fleet with no loud error (#4255). Fail closed on an empty or whitespace-only message on the text path, mirroring the existing --resolve-key refusal. * chore: retain ambient Pi-lens autoformat as its own commit Formatting-only edits produced by ambient Pi-lens autoformat during the msg-loss investigation, kept separate from the behavioural change in c23acba6 so the fix stays reviewable on its own. AGENTS.md is deliberately excluded: its only autoformat edit stripped the trailing space from the documented FM_OPERATIONAL_PREFIX value, which bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11 records as permanent compatibility. Documenting that constant without its trailing space makes the doc wrong about the contract, so that one line was restored rather than retained. * fix(calm): paint the working ship one yellow over all-blue water (#4554) On rose-pine-moon the two-color water (cyan crests over blue troughs) read as a pink stripe over aqua, the yellow left sail and mast clashed with the red right sail, and the hull carried a blue interior run. Every water cell is now blue so the swell reads through glyph height alone, and both sail halves, the mast, and the whole hull are one yellow run. Geometry, cadence, animation, direction flip, resize clamping, and the narrow fallback are unchanged. Update the unit and real-TUI color assertions to the new palette and the Calm docs that described the old one. * fix(bin): stop aging a second mate's active turn from its launch (#4270) * fix(watch): stop aging a second mate's active turn from its launch The parent watcher's second-mate wake-loop stall check exempts a mate that is demonstrably inside an active turn, but secondmate_in_active_turn asked busy_turn_over_age first and returned "not in a turn" whenever that said the bound was crossed. busy_turn_over_age ages from state/<task>.turn-ended, falling back to state/<task>.meta. A second mate's turns end in its own home, so the parent never gets a turn-ended mark for it and the fallback ages the mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago was therefore permanently "over age", the busy pane was never consulted, and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false wake-loop stall. The gate now bounds the busy exemption by <idle> - how long the queue's drain position has not moved - which is evidence this home actually holds. A busy mate stays exempt while the queue has been frozen for less than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so the bound that stops a busy pane from proving liveness forever is kept rather than removed. busy_turn_over_age is untouched; its remaining callers are the ordinary crew busy-pane bound. The regression pins the case that actually broke: a mate whose launch record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn must not escalate, while the same mate with its queue frozen past the bound still publishes exactly one notification. The existing coverage only exercised a freshly launched mate, which passes either way. Reaching that alert now costs a pane capture inside the gate, so the three checkpoints in this suite that assert an alert move from a 1s to a 4s bound - the value the neighbouring active-turn cases already use. The bound is a ceiling, not a wait: the checkpoint returns on the first actionable wake. On a loaded machine a 1s bound missed the alert repeatedly; at 4s it did not miss in 20 runs under the same load. * no-mistakes(review): scope the second-mate active-turn regression test's coverage claim * no-mistakes(document): fix stale second-mate active-turn comments in fm-watch * feat(bin): add read-only PR blocker and reviewer discovery commands (#4278) * feat(bin): add read-only PR blocker and reviewer-discovery commands Two focused, opt-in commands that read GitHub and never write to it. fm-pr-state.sh reports what still blocks one pull request from the author's side: a closed or merged state, draft state, unknown or conflicting mergeability, absent or failing required checks, and a blocking CHANGES_REQUESTED decision explained by each reviewer's latest verdict, marked STALE when it was left at a superseded head. A pull request that only awaits an approval is not reported as blocked, and advisory checks are omitted. Every reading is taken against one exact head; a push that lands mid-read invalidates the whole result rather than mixing two snapshots. fm-pr-reviewers.sh suggests reviewers from the most recent commits to the pull request's exact changed paths, counting each commit once, resolving handles through GitHub's own commit author.login mapping, and excluding the author and Bot accounts. Both stay read-only: no review request, no approval, no merge. Unresolved review-thread state is left unreported because the REST API does not expose it and unattended commands may not use GraphQL. Closes #3731 * no-mistakes(review): accept only PR URLs and stop at terminal state * no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate * no-mistakes(review): stop attributing readings to unverified heads * no-mistakes(review): narrow readiness contract to checks that have reported * no-mistakes(review): read the pull request once, drop the head guard * no-mistakes(document): scope pr-forge isolation proof to its measured members * no-mistakes(document): record uncovered pr-forge members and their pending proof * docs(isolation-proof): re-prove pr-forge at its full membership tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the pr-forge family in this branch, and script_allows_concurrency grants four workers by family membership alone, so both ran concurrently on a proof measured before they existed. Re-proved the family at all eight members: two consecutive runs, 0 failures, each begun with the one-minute load average below 6.0 so the result measures isolation rather than contention. A third run taken between them is disclosed rather than recorded, because it started while the previous run's workers were still decaying. The new durations are not comparable with the six-member measurement above them, so they are not presented as evidence about the two new members, and that record's 1.72x four-worker figure is left as a statement about its own run rather than restated as current. * no-mistakes(review): disclose gh error-text coupling at its matching site and tests * fix(bin): teach validation-round pauses in generated briefs (#2752) * fix(bin): teach validation-round pauses in briefs * no-mistakes(document): Point classifier comments to authoritative pause examples * docs(readme): add star history chart (#4558) * fix(bin): refuse teardown when a task's endpoint close fails (#4510) * fix(teardown): refuse a cleanup whose endpoint close failed bin/fm-teardown.sh discarded both the exit status and the stderr of every fm_backend_kill call, so a close that genuinely failed was indistinguishable from one that succeeded. Teardown continued past it, deleted the task's durable records, returned its worktree, and reported the cleanup as completed. The deleted metadata is the only record of which endpoint belongs to the task, so such a close did not merely leave a stray session behind, it stranded one: nothing was left on disk naming it. The adapters could not carry that signal either. Driven against the real code, every backend arm returned 0 for a genuine failure exactly as it did for an already-exited endpoint, so there was nothing for the four call sites to propagate even once they stopped swallowing it. The tmux arm now resolves a close that did not succeed against the window's exact recorded identity, since kill-window fails the same way for a window that is gone and one that is still there. The Orca arm reports a close its missing CLI never attempted. Both stay silent for an endpoint that is already legitimately gone, and the remaining arms are unchanged: their close-command timing cannot be established without the real Zellij, Orca, and cmux binaries, and a gate that refused ordinary cleanup of an already-exited session would be worse than the defect. docs/verification/runtime-backends.md records what each backend can prove. A reported close failure now reaches teardown's existing retain-and-stop refusal before the records naming the endpoint are removed, matching where the Herdr confirmed-gone gates already sit for the same hazard, and the retained records let a rerun finish once the close works. * no-mistakes(review): refuse unreadable tmux close re-read; honor --force override * no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close * no-mistakes(document): document endpoint-close refusal in its backend and retirement owners * no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree * feat(calm): add flag-gated Claude Code Calm mode (#4565) * feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that brings Calm to Claude Code: the sailboat replaces the stock working row through a Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration, and canonically classified operational user rows draw at zero height. /calm is registered by the hooks module itself and toggles the same per-home config/calm preference the Pi extension uses, so one choice applies on either harness; rows redraw retroactively on toggle and stay hidden across claude --continue. The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS flag is on. Nothing sets that flag in any settings file, and the plugin carries no command file, skill, agent, or classic hook, so it is a complete no-op while the flag is off. The trusted project auto-loads it through an .agents/skills symlink, the only path Claude Code scans for project plugins. Extract the working-ship geometry, bounce track, cadences, and freeze/resume state into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module imports from outside the plugin folder) and have the Pi widget paint that core's frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify operational rows through a port of bin/fm-operational-input.sh's classify command guarded by a corpus parity test against the shell owner. Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering, Raster packing, policy, classifier parity), the mod's own claude plugin test suites behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op, the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272. Docs: record the version-scoped Claude Code evidence and the three bounded gaps in docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md, and make the shared preference, layout, and contributor notes harness-neutral. * no-mistakes(review): Preserve colliding final replies and strengthen parser parity * no-mistakes(review): Preserve final replies and strengthen canonical parity checks * no-mistakes(review): Require exact function-hooks opt-in before Calm activation * no-mistakes(review): Clarify Calm module loading and activation boundaries * no-mistakes(review): Reset Calm presentation state across session starts * no-mistakes(document): Refresh Calm session lifecycle documentation * feat(calm): paint the Claude Code working ship in Claude's own theme colors The captain picked the "Claude native" palette for the Claude Code mod's Raster: every water cell takes the spinner blue of the active theme family (#93a5ff dark, #5769f7 light) and the whole boat takes the Claude orange of the stock spinner (#d77757), one water color and one boat color. The family follows the `theme` setting's prefix, read at load through $.config.list and re-read on a config.set of that row, with `auto` and custom themes falling back to the dark set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte. Rename the shared sprite's color classes from hue names to `water` and `boat`, since each harness now maps them to its own colors; geometry, motion, cadence, and the activation gate are untouched. Tests cover both palettes' packing and the family rule under Node, and the plugin kit drives every theme value, a theme change mid-session, the Calm-off pass-through, and inertness of the menu read while the flag is off. The docs describe the Claude Code colors and record the guard passing on 2.1.273. * no-mistakes(review): Use light palette for unresolved Claude themes * no-mistakes(document): Refresh Claude Calm verification evidence * fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586) * fix(watch): honour a declared wait before wedge-escalating a quiet pane wedge_timer_check escalated on elapsed idle time alone. Nothing asked whether the worker had already said why its pane was quiet, so a lane that declared a bounded external wait climbed the escalation ladder for as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every repeat carried demand-deep-inspection - which by its own wording forbids re-absorbing on the run-step or pane state, so the supervisor could not use the evidence that was there either. The generated brief promises that declaring `paused:` buys the long recheck cadence instead of a wedge, but the timer was still reachable while that declaration stood: a crew that declares a wait and then has an active run or busy pane attributed to it is handed to the timer as provably-working. The declaration is what the worker said about its own silence, so it now outranks a liveness verdict that only says something is running. The consult runs in the at-threshold branch that was about to escalate, beside the worktree walk already there, and costs one status-line read. Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS recheck the declared-wait absorber already uses, so the wait is still rechecked and cannot rot invisibly. Which verb declared it decides the wording, because the two block on different people: a `paused:` wait is owed by an external dependency and asks the reader to confirm it still holds, while a `captain-held:` transfer is owed by the captain reading the recheck and asks them to answer or release the hold. A hold is not rechecked at all while the away-posture record exists, as on every other captain-held path, and that absorb arms no throttle so the recheck is owed in full on return. A declared clearing time that has already passed stops counting, and a lane that never declared one keeps the identical escalation schedule, reason, count and demand-deep-inspection wording, so detection and its worst-case time are unchanged. The deferral restarts the idle timer rather than cancelling it, so a lane that stops waiting escalates again within one threshold. A lane quiet because its own validation run is parked at a gate awaiting a human decision is deliberately out of scope: reading that state needs a signal carrying who the wait is on and what clears it, rather than one inferred from a parked verdict that also covers gates awaiting the crewmate itself. Tests pin both directions for each case and were each confirmed to fail with the consult removed. * no-mistakes(document): docs: honour declared waits in stale-escalation docs * fix(bin): report verified PR state for passed runs (#4624) * fix(bin): derive passed PR state from PR record A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists. For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim. Fixes #4607 * no-mistakes(review): Add bounded GitLab merge-request state reads * no-mistakes(review): Preserve network-free inactive crew-state scans * no-mistakes(document): Document PR record readers in shared library * fix: restore published contribution follow-up (Fixes #4469) (#4627) * fix: restore published contribution follow-up (Fixes #4469) * fix(review): Fix contribution freshness and merge actor routing * fix(review): Restore issue triage and scope contribution follow-up * fix(test): test: assert one wake per contribution signal * fix(document): Document contribution follow-up * fix: restore truthful terminal delivery evidence * fix(review): Disclose unsupported contributions and deduplicate watcher wakes * fix(review): Preserve unmeasured unsupported contributions across Bearings * fix(review): Deduplicate shared contribution wakes and isolate diagnostics * fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower * fix(bin): make remote report transfers explicit and fail-open (#4658) * fix(bin): make a remote-reply document gap self-clearing and re-attemptable A remote mate's undelivered document raised a keyed `blocked` decision that nothing could ever resolve, and any `data/*.md` substring in any mirrored line was an unconditional fetch instruction. A mate announcing a report it had not written yet therefore manufactured a permanent, factually false blocker, and its own explanation of the false alarm manufactured more. The reader has no permanence vocabulary: a report still being written refuses exactly like a path that will never exist. So an undelivered document is now a durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`, re-attempted on the next delta and on the channel's own quiet poll, and retired with a matching `resolved` line naming the local copy once it arrives. The cursor still advances and no delta stalls on one bad pointer. Only a structured `report=data/....md` pointer now offers a document, so a path merely mentioned in prose - including one under another home's mirror tree, which is provably not that mate's to serve - is never fetched. Offers are deduplicated across the whole delta, the escalation names each missing document once and carries the reader's own reason instead of discarding it, and a strictly increasing notice ordinal keeps a later escalation from being swallowed as duplicate bytes. A mirrored line still lands once whichever pointer form it was first written under. * no-mistakes(review): Require structured pointer token boundaries * no-mistakes(review): Unify boundary-safe pointer extraction and rewriting * fix(bin): identify a mirrored line independently of its delivery state Two defects in the boundary-safe pointer work. The at-most-once check compared only the all-remote and all-local renderings of a line, so it could not recognize a mixed one. A line offering two documents where only the first was deliverable mirrored as local-plus-remote; once the second arrived, a cursor-loss whole-log recapture rendered the same line all-local, matched neither alternate, and mirrored a second time. A line's identity is now the canonical form every boundary-valid pointer would take once delivered, derived by the same parser that does extraction and rewriting, so it no longer depends on which documents happened to be deliverable at the time. The pointer map was passed to awk through the process environment. A delta may carry up to the configured 1 MiB bound, and an expanded map of delivered pointers can exceed the platform's exec argument limit, so awk would fail to start; because no caller checked, the empty result would have been appended as blank lines while the cursor advanced past dropped status content. The map now travels in a file, and every call site checks the exit status and stops the ingest rather than committing a delta it could not render. Both passes now run once per stream instead of twice per line. * no-mistakes(review): Abort ingest when document pointer extraction fails * no-mistakes(review): Exclude structured cross-home pointers from document transfer * fix(bin): fail open on an undeliverable remote document instead of tracking it Narrow the remote-reply document fix to the scope the diagnosis actually requires, as decided after measuring a simpler alternative. A document the reader cannot deliver now fails open. The mate's line is mirrored with its own pointer, the cursor advances, and one unkeyed note carries the reader's reason. A note never enters the open-decision fold, so it cannot stand open the way the original keyed block did - which removes the never-clearing false blocker by construction rather than by resolving it. That makes the durable self-clearing obligation unnecessary, so it goes: the per-mate pending-documents record, its notice ordinal and resolved announcements, and the poll-side retry. Canonical line identity goes too, and with it a way to silently drop a genuine status line; mirroring is back to at-most-once on exact bytes. The cross-home exclusion goes as well: under fail-open a cross-home report= either fails harmlessly or is a nested remote report this mate genuinely holds, which is now relayed again. Kept: fetching only on a structured report= pointer, the boundary-correct parser, the file-based rewrite map, and checked extraction and rewrite exit status. The parser now scans behind a sentinel byte so a rejected candidate can no longer give the text right after it a false leading boundary. The reported incident is covered end to end: a report path announced in prose before it exists raises no decision, and the report still arrives through the ledger publisher's structured offer once written. * no-mistakes(review): Preserve source-line identity across remote reply replays * no-mistakes(document): Document remote reply transfer and replay semantics * no-mistakes(lint): Fix staging truncation lint checks * fix(calm): preserve substantive mid-turn responses (#4655) * Preserve substantive Calm mid-turn text * no-mistakes(review): Distinguish newline-preserved replies from short narration * no-mistakes(document): Document Calm mid-turn preservation boundaries * no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check` * fix(bin): preserve PR merge polls across volume remounts (#4656) * fix(bin): re-record PR poll identity after a volume device renumber (Fixes #4260) A volume remount can renumber the state filesystem's st_dev while every inode and byte stays the same; APFS does this across a reboot. A poll registration records its sidecar and check as device:inode, so every poll armed before the remount failed strict validation and the watcher refused all of them as unauthenticated state checks until each was re-armed by hand. There are two device comparisons. fm_pr_private_file_valid compares a live file's device with the state directory's device read in the same invocation: it refuses a file that is not on the state directory's own filesystem and already survives a renumber, so it is unchanged. The registration's recorded identity versus the live identity (from #556, reused by the #932 retirement receipt) binds the registration to the exact files published in its own transaction; its device part is what breaks. When strict capture fails, the watcher now proves the device is the only difference: every other artifact check passes (template bytes, both hashes, private mode, single link, live device, metadata), both recorded identities name one device, and each recorded inode equals its live inode. Only then, under the task's control lock, does it rewrite the two identity lines, repeating the whole proof and comparing the registration's file identity and bytes just before the rename, and then capture strictly again. A swapped, altered, re-moded, relinked, split-device, or foreign-device artifact still fails a proof and is still refused, and a pending retirement receipt blocks the rewrite. Reproduction: on macOS a poll armed on an APFS disk image that was detached and re-attached behind another image moved st_dev 16777239 -> 16777243 with inodes, bytes, mode, and link count unchanged; the real watcher refused it on main and reports its merge with this change. The portable regression test rewrites a real registration's recorded device and drives the watcher. Not changed here: the status presentation cursor keys rows by its own device:inode identity in bin/fm-classify-lib.sh, a different helper that needs its own fix; a retirement receipt left by a reboot between its publication and removal still names the old device and stays refused; custom check trust binds only a content hash and is unaffected. * fix(review): Serialize PR poll publication writers * fix(review): Bound PR poll publication lock scope * fix(bin): keep contribution records when the poll budget runs out (follow-up to #4627) (#4661) A budget that expires partway through an observation no longer records an error or prints the unavailable wake; the URL keeps its prior record and is observed first next poll. forge() flags budget exhaustion at the point it refuses, or when a read is killed at the budget's own deadline, so a genuine forge failure still records the error and wakes. Each distinct URL is now observed once per poll and applied to every owning task. * fix(bin): clear parent pending-replies on local secondmate retirement (#4680) * fix(bin): clear parent pending-replies on local secondmate retirement Local secondmate teardown left resolved parent pending-reply records behind after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced retirement while any reply for that id is still unresolved, and delete every matching record plus its delivery confirmation after a successful local or remote retirement, matching the remote cleanup path. * no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup * no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung * no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern * no-mistakes(review): Pending-replies Basename und corr_id abgleichen * no-mistakes(document): Clarify forced retirement pending-reply cleanup --------- Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc> * fix(bin): accept Orca's composite worktree id when tearing down a task (#4677) * fix(bin): accept Orca's composite worktree id at teardown Teardown refused every Orca-backed task because the endpoint validator checked orca_worktree_id with the simple-atom rule meant for tmux-style window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca returns that id as `<orca id>::<absolute worktree path>`, so the colon and slashes in every real value made validation fail and finished Orca tasks could never be cleaned up. Validate the field as the composite it is: both halves of the first `::` split present, the path half absolute, and no embedded newline, carriage return, or tab. The terminal field keeps the atom check, which is correct for it, and no other backend's validation changes. The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca never returns, which is why the suite passed a check the real value fails. They now carry the composite form, so the tests exercise the real value. * no-mistakes(document): name Orca's repo id in the composite worktree id * no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points * feat(bin): add opt-in typed dispatch resolution (#4692) * feat(bin): add opt-in typed dispatch resolution through typesafe.ai Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or scout profile from a written brief with typesafe.ai's System One model: one Choice question over the rules' `when` texts, then the confidence floor, the rule's `approval` and `floor`, each profile's `provider` and `floor`, one quota-axi snapshot, and the spendPriority argmax all in code. It is off unless TYPESAFE_API_KEY is in the environment or the home's gitignored .env; off means one stderr line, exit 0, and no network call, so firstmate dispatches exactly as before. The key reaches curl on a file descriptor, never argv. Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates the four new optional dispatch fields. Document the schema, the operator contract, the AGENTS.md intake step, and the live and benchmark evidence. * no-mistakes(review): Harden typed dispatch resolution and quota bounds * no-mistakes(review): Validate dispatch floors and ranking evidence * no-mistakes(review): Tighten dispatch response and floor evidence * no-mistakes(review): Neutralize none matching and resolve defaults locally * no-mistakes(review): Preserve providerless profiles outside typed resolution * no-mistakes(review): Validate response usage and reject duplicate profiles * no-mistakes(review): Escalate unverifiable floors and validate probabilities * no-mistakes(review): Validate probability mass and unknown profile floors * no-mistakes(review): Simplify resolver interface and preserve fallback routing * no-mistakes(review): Fix constants and rank partial quota evidence * no-mistakes(review): Add authoritative provider mapping and enforce explicit providers * no-mistakes(review): Declare provider for documented Pi profile * no-mistakes(review): Validate provider identifiers and support Gemini dispatch * no-mistakes(review): Strictly anchor provider identifiers * no-mistakes(review): Validate selectors and preserve fallback candidate evidence * no-mistakes(review): Gate typed validation and harden resolver evidence * no-mistakes(review): Preserve opt-in routing and harden candidate evidence * no-mistakes(review): Prioritize known exhaustion over quota uncertainty * no-mistakes(review): Isolate API secrets and preserve no-key diagnostics * no-mistakes(review): Fallback safely when dispatch rules are absent * no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets * no-mistakes(document): Document typed dispatch safety and fallback behavior * fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753) * test: reproduce buried status declarations in shared readers * fix: share status event reads and preserve open blockers * fix: retain terminal scout and ship status declarations * no-mistakes(review): Fix status chronology, legacy completions, and reader performance * no-mistakes(review): Share terminal decision reconciliation across fleet snapshots * no-mistakes(review): Unify terminal supersession across cached folds and consumers * no-mistakes(review): Filter per-key status history while preserving terminal chronology * no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells * no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses * no-mistakes(document): Document latest-event status read and kind-scoped fold cursor * no-mistakes(lint): Quote literal done in test for-lists for SC1010 * ci: expect 19 snapshot/fleet-view tests This branch adds a fleet-snapshot regression, so the stock macOS Bash lane's hardcoded guard of 18 'ok - ' lines fails on the new count. Bump the guard and its message to 19. * no-mistakes(review): Restore multiline child outcome reporting * no-mistakes(review): Select ledger terminal events through bounded shared reader * no-mistakes(review): Report newest open decision instead of preferring blocked * no-mistakes(review): Require colon before ship/scout terminal supersession in fold * no-mistakes(review): Gate socket-down override on latest event; drop lock matrix * no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions * no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold * no-mistakes(test): Update fleet-view expectations to newest-open-decision rule * no-mistakes(document): Align status-read docs with fold-resolved crew state * no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers * no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree * test: fold terminal-cleanup snapshot coverage into the completed-scout case Keep the ship/scout/secondmate supersession assertions without adding a nineteenth top-level fleet-view test, so CI can stay at the upstream suite count. * no-mistakes(document): Clarify socket-down override expiry in architecture doc * ci: retrigger flaky contribution check * fix(bin): launch codex crewmates with codex's hook layer disabled (#4689) * fix(spawn): launch codex crewmates with codex's hook layer disabled A freshly launched Codex worker never reached its instructions. Codex stopped it on an interactive "Hooks need review" modal whose selection sits on "Review hooks", which is neither trusting nor declining. Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow navigation, so the selection cannot be moved, and pre-accepting the prompt by writing Codex's own trust store would record an operator consent that was never given. The hooks are the machine's own ~/.codex/hooks.json plus any project's .codex/hooks.json. A crewmate needs neither: its turn-end signal is the -c notify= program on the same launch, and Firstmate's project hooks are primary-session infrastructure that stands down in a child worktree. Crewmate and scout launches now pass --disable hooks. That is the opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted hooks; disabling the feature runs none of them and leaves the operator's ~/.codex untouched. An unknown feature name is a hard Codex error, so a release that drops the flag fails the launch loudly instead of silently restoring the modal. A secondmate is a primary in its own home and keeps the project hooks its turn-end guard and session-start digest ride on. Verified on codex-cli 0.151.0: the modal is gone and the turn-end notification still lands. This unblocks the second review that every finished pull request is supposed to get. Fixes kunchenguid/firstmate#4673 * no-mistakes(review): Fix contradictory hook count in Codex verification record * fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710) * fix(bin): settle terminal contributions and wake once per read-failure episode A contribution whose last good observation is merged or closed is final: poll no longer re-reads it, projection keeps it fresh, and a stale error recorded beside it is cleared once. A genuine forge-read failure on an open contribution still records its error on every cycle but prints the unavailable wake only when it starts a failure episode; a successful read ends the episode. Open PRs linked from done tasks keep being observed. The false unavailable beside a complete observation was budget exhaustion mid-observation, already fixed by #4661. * fix(review): Settle terminal contribution owners * fix(review): Deduplicate shared contribution failure episodes * fix(test): Preserve settled terminal contribution records * fix: select authoritative no-mistakes runs (#4476) * fix(crew-state): select authoritative validation runs by identity Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row. Refs: https://github.com/kunchenguid/firstmate/issues/3215 * fix(review): Resolve same-branch run identities beyond capped history * fix(review): Fix run-selection compatibility, races, and worker-state fallbacks * fix(review): Limit run validation to the requested branch * fix(test): Anchor AXI fixtures and document remaining live evidence gaps * fix(document): Clarify run selection documentation and capture ownership * fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation * fix: distinguish captain outcomes from no-op updates (#4738) * fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work MAIN answered a supervision-branch outcome for completed captain-requested work (implementation done, PR ready for review and merge approval) with "Captain, shipshape.", reading section 9's no-action reply as covering it and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no captain-facing response is owed". Section 9 now limits the shipshape reply to true no-ops (idle re-read, empty heartbeat, consequence-free acknowledgement) and requires a short outcome response naming what finished and what word is needed whenever requested work finishes or a result needs the captain's word, even when a transcript entry already shows the substance. The Pi protocol's re-emit rule now says it bounds repetition only, and carries a worked example of the ready-for-review outcome whose correct processing turn a shipshape reply fails. No executable contract evaluates the content of MAIN's captain-facing reply, so the regression is the protocol example in the owner doc rather than a text-match test. * no-mistakes(document): Clarify captain-facing outcomes versus no-ops * docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line The document step condensed the Pi protocol's re-emit rule and dropped the worked example of a finished, ready-for-review outcome whose correct processing turn a "Captain, shipshape." reply fails. That example is the contract's regression: no executable contract evaluates the content of MAIN's captain-facing reply, so the owner doc's example is the test case. Restore it directly under the re-emit rule, prefixed as a regression example that is kept verbatim and never condensed or summarized away. * no-mistakes(review): Clarify captain outcome and decision-word requirements * no-mistakes(document): Clarify captain-facing completion outcomes * docs(pi): require the PR URL in the visible captain-facing outcome reply Captain review on the regression example: drop the sample reply string and say only that the ready-for-review outcome requires relaying a captain-facing outcome response, not just "Captain, shipshape.". Fold in the visible-PR-handoff failure seen this session: after the branch outcome reporting this fix green, MAIN's visible reply was only "Awaiting your merge call." with no PR URL, leaning on the dim anchor. Section 9's URL rule now also covers a review or merge ask and names the visible reply as where the URL goes, sourced from the ready status, pr= metadata, or the supervision branch's summary and never left to a transcript entry. The Pi protocol adds the same-way failure and places the captain-facing text in the final visible assistant reply after the fm_branch_processed call, because Calm hides assistant text emitted in the same step as a tool call as a working note. Investigation verdict, evidence in the PR comment: no recent PR caused the handoff failure; Pi has hidden same-step pre-tool assistant text since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and #4658 touched only remote report transfer. * no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs * no-mistakes(document): Clarify captain-facing supervision outcomes * docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule The consolidated section 9 URL rule narrowed its trigger to a review or merge ask, dropping the "whenever a PR is mentioned" catch-all from #3648 that keeps every PR URL copied from a durable record and never assembled from memory. Restore that trigger as a union with the review or merge ask so the one consolidated rule covers both. * fix(bin): let non-owner Claude Stops exit safely (#4777) * Fix foreign-owner turn-end supervision loop * no-mistakes(review): Scope foreign-owner safe exit to Claude guard * no-mistakes(document): Document Claude foreign-owner safe exit * fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778) Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty indexed array as an unbound variable and aborts the shell. In signal_turnend_panes_churned() the missing_keys loop was reachable with an empty array whenever every churned key already held a fresh .churn-since-* marker (a second churning turn-end inside an open deferral window), so each watcher cycle died about half a minute in and supervision restarted endlessly. The created_keys rollback loops had the same latent crash on their error paths. Audit of bin/ for the same pattern found one more confirmed-reachable case: remote_handoff's noncanonical-body scan iterates to_move, which is empty when a retried remote handoff finds every key already staged in the outbox. All other "${arr[@]}" sites are either count-guarded, guaranteed non-empty by construction, or unreachable while empty. Guard the three reachable expansions with the repo's existing "${arr[@]+...}" idiom. New regression test drives a real watcher through the all-marked churn path; the macos-stock-bash CI lane runs it under real /bin/bash 3.2 via FM_TEST_ONLY. * Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783) The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: require complete captain-facing final responses (#4779) * docs: require complete final responses across harnesses * no-mistakes(document): Document complete final replies for Grok Bot * docs: point Grok replies to the shared contract owner * no-mistakes(review): Clarify final recap without batching decision asks * fix: preserve substantive mid-turn text in Pi Calm (#4788) * fix(calm): preserve substantive Pi mid-turn text * no-mistakes(review): Preserve substantive Pi Calm text per block * no-mistakes(test): Cover shared Calm preservation boundaries behaviorally * no-mistakes(document): Consolidate Calm preservation documentation * fix: harden mail checks and rebalance full-coverage CI (#4800) * Improve CI reliability and rebalance full-coverage validation * no-mistakes(document): Clarify lint partition documentation * fix(bin): answer Kimi 2.0.0 folder-trust dialog during spawn (#4799) * Handle Kimi workspace trust dialog * no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers * no-mistakes(review): Gate Kimi ready on any trust marker and clean captures * no-mistakes(review): Read visible pane for Kimi trust and ready gates * no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate * no-mistakes(review): Harden Kimi viewport capture and trust dialog detection * no-mistakes(document): Document Kimi spawn refusal on cmux and Orca * fix(bin): report a dead-agent record once instead of escalating forever (#4775) * fix(bin): report a record whose agent is gone once instead of escalating forever The wedge escalation path never asked whether there was still an agent to be wedged. A wedge is something stuck that might recover, so re-alarming it earns its cost; an agent that is gone never moves again, its pane never churns, the idle timer never resets, and the escalate path clears its own timer and re-arms with nothing bounding the count. Observed on a live fleet: two finished lanes reached 226 and 203 consecutive escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400 notifications a day from two lanes with no agent running at all. On one, fm-control.sh exit answered already-stopped and fm-crew-state.sh read "failed - run failed". Closing the Herdr pane did not stop it either: with the pane genuinely gone and herdr pane read returning pane_not_found, the count kept climbing, because the poll is driven by the record's window= line rather than by the pane. The cost is not the repetition but that it drowns the alarms that matter. fm_backend_agent_state already separates a thinking agent from a gone one at process level. In t…
sctru
added a commit
to sctru/firstmate
that referenced
this pull request
Sep 21, 2026
* fix(bin): safely unregister custom checks (#3369)
* fix(bin): add a safe owner for custom-check retirement
Agents were improvising rm of check files with unset STATE/ID, which wedges
headless panes. Unregister validates the id and state directory first.
Co-authored-by: Cursor <cursoragent@cursor.com>
* no-mistakes(review): Refuse explicitly empty custom-check state overrides
* no-mistakes(document): Document custom-check retirement safety contract
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* refactor(quota): extract mid-task polling and candidate selection into dedicated scripts (#3221)
* Add quota exhaustion detection and safe fallback helpers
- bin/fm-procevent-quota.sh: generic procevent adapter that arms a
recurring quota-axi --json poll and wakes firstmate when a tracked
provider's effectivePercentRemaining drops below a threshold or its
runway.status becomes exhausted_now.
- bin/fm-quota-choose.sh: worker-side helper that picks the first ranked
harness:model candidate with positive effectivePercentRemaining.
- AGENTS.md and .agents/skills/quota-array-dispatch/SKILL.md: document
the new helpers and the mid-task quota-exhaustion wake path.
- tests/fm-quota-choose.test.sh: unit tests with a mocked quota-axi JSON
source.
* no-mistakes(review): Fix quota polling and scope bounds
* no-mistakes(review): Enforce safe default quota selection
* no-mistakes(review): Handle decimal quota values safely
* no-mistakes(review): Fail closed on invalid quota inputs
* no-mistakes(review): Reject empty quota candidate segments
* no-mistakes(review): Harden quota parsing and timeout ownership
* no-mistakes(review): Reuse captured quota snapshots consistently
* no-mistakes(review): Match quota using explicit candidate providers
* no-mistakes(review): Centralize fail-closed quota schema validation
* no-mistakes(review): Reject out-of-range quota percentages
* no-mistakes(review): Validate quota runway status enum
* no-mistakes(review): Tighten quota scope and status contracts
* no-mistakes(review): Preserve unknown quota and exact product bounds
* no-mistakes(review): Preserve provider-level unknown quota
* no-mistakes(review): Reuse canonical verified harness validation
* no-mistakes(document): Document mid-task quota handling
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* fix(docs): restore default routing contract, keep quota helper optional
Restore the AGENTS.md section 4 always-loaded routing paragraph the PR
had deleted, so the standing TOON-first intake, spendPriority ranker,
every-candidate accounting, and load-trigger contract stay exactly as
before this PR. The mid-task quota wake is optional and must not alter
default routing.
Restore the quota-array-dispatch skill ownership line to section 4 as
the always-loaded intake boundary owner; keep the worker-side helper
section as an addition only, without rewiring ownership or load
triggers to section 13.
* fix(bin): use harness-keyed quota matching in optional helper
Revert fm-quota-choose.sh from harness:provider:model tuples back to
harness:model candidates with harness-keyed provider matching, per the
resolved ask-user finding. The helper is optional; authoritative
multi-provider routing (provider discovery from the harness catalog and
quota matching by that explicit provider) stays owned by AGENTS.md
section 4 and the quota-array-dispatch skill intake procedure, not the
helper.
Document the multi-provider limitation in the helper header and the
quota-array-dispatch skill: the helper maps each harness to one primary
provider family only, so a candidate whose established provider differs
from that primary family is checked against the wrong quota row. Use it
only when the brief fixed the candidate order and every candidate's
provider is the harness's primary family.
The helper still consumes one already-captured default-TOON or JSON
snapshot via stdin or --snapshot and never calls quota-axi itself, so
it selects from the same quota state as the intake.
* no-mistakes(review): Fix Muse quota mapping and helper contract docs
* no-mistakes(review): Reject known-empty quotas and map quota tests explicitly
* no-mistakes(review): Preserve unmeasured candidates and enforce snapshot reuse
* no-mistakes(review): Fix quota retirement and dependent regression coverage
* no-mistakes(review): Accept zero-row quota TOON snapshots
* no-mistakes(review): Enforce quota semantics status consistency
* no-mistakes(review): Veto dispatch on any exhausted applicable scope
* no-mistakes(review): Record exhausted quota scope in wake details
* no-mistakes(review): Fix quota help and control dependency coverage
* no-mistakes(review): Decode quoted TOON fields and document quota wakes
* no-mistakes(review): Validate zero-row TOON and map timeout coverage
* no-mistakes(review): Reject multi-value JSON and malformed TOON envelopes
* no-mistakes(review): Validate complete nonzero TOON envelopes
* no-mistakes(review): Accept producer-shaped quota TOON envelopes
* no-mistakes(review): Support empty quota arrays and validate counted rows
* no-mistakes(review): Harden TOON completion, scopes, and quoted fields
* no-mistakes(review): Preserve unknown-headroom exhaustion and reject trailing fields
* no-mistakes(review): Allow unknown headroom under known semantics
* no-mistakes(review): Reject noncanonical quota identities
* no-mistakes(review): Preserve empty quota polling and validate attention identities
* no-mistakes(review): Reject noncanonical provider watches
* no-mistakes(review): Validate all candidates before quota selection
* no-mistakes(document): Correct quota helper safety documentation
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* fix: surface comments on Lavish annotations (#3371)
* fix(bin): keep typed Lavish comments when an element is also annotated
read preferred element text over prompt, so an annotate-and-comment
item dropped the captain's words. Surface prompt as its own field.
Co-authored-by: Cursor <cursoragent@cursor.com>
* no-mistakes(review): Filter non-comment prompts from Lavish reader output
* no-mistakes(document): Clarify Lavish comment presentation contract
* no-mistakes(ci): Fixed Lavish reader comment provenance: non-choice prompts are now emitted even when identical to element text. Added observable regression coverage for identical selector+comment input while retaining pure annotation/message coverage. Reader cases, bash syntax, and diff checks pass. Full fm-procevent suite stops earlier at unrelated “reconcile never claimed” setup failure
* no-mistakes(ci): Fixed duplicate pure-annotation prompts by emitting `prompt:` only when it differs from captured element text. Updated behavioral coverage for selector+comment, pure annotation, and pure message cases. Focused reader regressions, syntax checks, and diff checks pass. Full suite remains blocked by the pre-existing “reconcile never claimed the registered source” failure
* fix(bin): always emit Lavish comments and use real annotation fixtures
Stop inferring comment provenance from prompt==text. Real pure
annotations have no prompt, so always-emit does not duplicate.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: support first public-followup registration on Bash 3.2 (#3420)
* Fix public-followup register crashing on empty lock arrays under bash 3.2.
bash 3.2 with set -u treats "${arr[@]}" on an empty array as unbound, so the first register in a fresh home aborted before taking the registry lock.
The empty-lock regression also runs under the existing stock macOS Bash CI lane so pre-fix code would fail there.
* no-mistakes(document): Document stock Bash registration coverage
* no-mistakes(ci): Pinned the stock macOS Bash CI lane to tasks-axi@0.2.5, eliminating dependency drift. Verified workflow YAML parsing, git diff checks, and the focused regression under /bin/bash 3.2.57 with tasks-axi 0.2.5
* no-mistakes(ci): Fixed the flaky portable CI test: it treated exited zombie processes as live because `kill -0` succeeds for zombies. The watcher and descendant assertions now check process state and regard zombies as exited. Verified `tests/fm-pr-check-security.test.sh`, ShellCheck, `git diff --check`, and the focused Bash public-followup regression
* fix(bin): isolate new Herdr server environments (#2792)
* fix(herdr): isolate server launch environment
* no-mistakes(review): Clear inherited supervision model from Herdr launches
* no-mistakes(document): Document Herdr server launch environment isolation
* fix: surface inbound Relay media to responding agents (#3442)
* fix: surface inbound Relay attachments to the responding agent
A Discord support thread's screenshots were never seen by the agent
handling the mention. The relay delivered them and the poll stashed
them: the reporter's images arrived on the `thread_starter` entry of
`in_reply_to_chain` while the mention's own media list was empty. The
gap was in the responder's playbook, which enumerated a fixed field
list (`request_id`, `text`, `in_reply_to`, `in_reply_to_chain`) and so
made every other field, attachments included, invisible.
Fix it where the gap is, in prose:
- Read the complete payload object rather than a fixed field list, so
media and later relay fields are never skipped again.
- Fetch and view attached media with the agent's own tools, on the
mention and on every chain entry, and call out the common shape where
only the thread starter carries the screenshots.
- Restrict those fetches to known-good platform media hosts over https
(Discord: cdn.discordapp.com, media.discordapp.net,
images-ext-1.discordapp.net, images-ext-2.discordapp.net; X:
pbs.twimg.com, video.twimg.com), report a blocked host instead of
working around it, and treat everything fetched as untrusted public
input on the same terms as the surrounding thread text.
The poll stays out of it and downloads nothing, so no third-party bytes
are pulled on the polling path.
The new test pins the contract the playbook depends on: a mention in the
incident's shape, with an empty top-level media list and screenshots on
the thread starter, must reach the inbox with the payload intact and its
media URLs unfetched.
* no-mistakes(review): Preserve media authority and enforce poll-only fetching
* no-mistakes(document): Clarify Relay attachment safety prose
* fix(bin): defer inactive reconciliation during startup (#3480)
* Defer inactive startup reconciliation
* no-mistakes(review): Queue deferred inactive reconciliation diagnostics durably
* no-mistakes(review): Require worker phases to cover startup requests
* no-mistakes(review): Make diagnostic wakes safely acknowledgeable
* no-mistakes(document): Document deferred startup phase coverage
* fix(bin): bound wake drain presentation lock waits (#3475)
* fix: bound status presentation lock waits
* no-mistakes(review): Distinguish malformed presentation locks from live contention
* no-mistakes(review): Bound no-ack drain queue lock acquisition
* no-mistakes(document): Document bounded presentation-lock drain behavior
* no-mistakes(lint): Annotate bounded lock output global
* no-mistakes(ci): Added deterministic regression coverage for successful bounded-lock acquisition after live contention, verifying helper-to-caller PID ownership handoff and caller release. Verified with bash syntax checks, git diff checks, and the full fm-wake-queue test suite
* fix(bin): retire public follow-ups in remote homes (#3479)
* fix(relay): close a public loop whose work lives in a remote secondmate home
A public-followup loop bound to a REMOTE secondmate could never be closed.
`clear_public_followup_link` (bin/fm-public-followup.sh:701) required an
absolute recorded `work_home_path` for a `secondmate:*` work home, but a remote
route has no local path on this machine, so registration records that field
empty (bin/fm-public-followup.sh:291). Every close ran that clear first, so
`retire` died with "could not clear the legacy X link ... retained for
reconciliation" forever, and `deliver` posted the public reply and then stranded
the loop at `posted`. `--force` never covered that step.
The clear now goes to the remote home over that route's SSH transport, running
`fm-x-followup.sh --clear <work-id>` through `bin/fm-on.sh`. The route is decided
from `data/secondmates.md` before any local path is consulted, so a same-named
local directory can never stand in for a remote home, and registrations already
on disk retire without needing a new field. `fm-on.sh` passes ssh's status
through, so 255 stays the established "delivered but completion unknown" result
this codebase already reconciles: the close is refused, the registration and the
remote link are left exactly as they were, and the message names the unknown
completion instead of claiming a definite failure.
Local secondmate and `main` work homes are untouched, and `--force` still
governs only the unresolved-obligation refusal.
Three regression cases drive a remote route end to end, faking only the ssh
binary at the FM_SSH_BIN seam and then running the real remote entrypoint
against a local checkout, so the clear that must reach the remote home actually
happens there.
* no-mistakes(review): Guard remote link clears by request identity
* no-mistakes(review): Fail guarded clears on unreadable remote state
* no-mistakes(review): Reject guarded clears on non-writable remote state
* no-mistakes(review): Allow no-link retirement in non-writable remote state
* no-mistakes(document): Correct public-followup verification guarantee count
* no-mistakes(ci): Fixed the guarded link-clear race by ensuring absence is decided under the metadata lock whenever publication is possible. Added a behavioral concurrency regression test. Verified with fm-x-mode and fm-public-followup suites, Bash syntax checks, diff checks, and bin/fm-lint.sh
* no-mistakes(ci): Fixed the guarded link-clear race by refusing an unlocked absence decision when a publisher already owns the metadata lock in a non-writable directory. Added a behavioral concurrency regression test. Verified with fm-x-mode, fm-public-followup, syntax/diff checks, and fm-lint
* no-mistakes(ci): Fixed the guarded-clear race by refusing all guarded clears when the metadata parent is non-writable, including apparent link absence. Added a behavioral regression with a publisher waiting to create the lock, updated remote-retirement expectations and verification docs. Passed fm-x-mode, fm-public-followup, fm-lint, documentation audience, Bash syntax, and diff checks
* fix(relay): bound the guarded remote link clear so it refuses instead of hanging
The guarded clear checks that the remote state directory is writable before
taking the metadata lock, but that check cannot close the window: the parent can
turn non-writable between the check and lock creation, and a lock held by a live
holder is indistinguishable from that at the acquire. `fm_lock_acquire_wait` is
an unbounded `while ! try; do sleep 0.1; done`, so either case retried forever
and `deliver` or `retire` wedged with nothing reported, instead of returning the
retained-for-reconciliation refusal the guard exists to produce. This path runs
unattended over the secondmate transport, where a wedge is worse than either
outcome the guard defines.
The guarded clear now acquires through `fm_lock_acquire_wait_bounded`
(FMX_LINK_CLEAR_LOCK_TIMEOUT, default 10 seconds) and refuses on timeout through
the existing failure path. Unguarded local callers keep the ordinary unbounded
wait, so local behavior is unchanged.
The bounded primitive's header no longer claims presentation-only scope, since
this is a second authorized caller; nothing else in the shared lock
infrastructure changed.
The regression holds the metadata lock with a genuinely live process while
leaving the state directory writable, so the refusal can only come from the
bound and never from the writability precondition. Against the unbounded wait it
does not terminate at all; with the bound it refuses, retains the registration,
writes no receipt, and leaves the remote link untouched.
* no-mistakes(review): Harden lock-timeout regression with independent deadline
* no-mistakes(review): Restore no-op guarded clears on read-only state
* no-mistakes(document): Clarify remote public-followup cleanup contract
* fix(bin): support process events under symlinked homes (#3484)
* fix(bin): resolve process-event state roots before validating them
The process-event module validated the caller's spelling of a home's state
root instead of the directory it operates on: it required the supplied path
to equal its own lexical normalization, which rejects any path reached
through a symlinked ancestor. On macOS both /tmp and $TMPDIR are symlinks,
so an operator home under either could never claim a source. Reconcile still
reported the runner started, while the detached runner died writing "cannot
claim source" to the discarded stderr, and the source silently never fired.
Resolve the state root to its physical directory once, then apply the
existing private-directory validation to that resolved directory and derive
every path, recorded claim identity, and later confinement check from it.
This keeps the confinement contract for the directory actually operated on
rather than only for callers that already spelled it physically, and removes
the window where an ancestor symlink could be repointed between check and
use. Homes already spelled physically behave identically.
This was the single cause of both deterministic macOS failures in
tests/fm-procevent.test.sh ("reconcile never claimed the registered source")
and tests/fm-procevent-when.test.sh ("the winning concurrent arm did not
produce an outcome"). The new case pins the behavior with an explicit
symlinked-ancestor home, so it fails without the fix on any platform rather
than only where the temp root happens to be a symlink.
* fix(bin): pin the external capture staging boundary to its physical path
The extension capture path pinned its registry staging boundary by comparing
`pwd -P` against the caller-spelled registry directory, so a home reached
through a symlinked ancestor still refused to start an extension-backed
source after the state root itself resolved correctly. That left such a home
half working: built-in sources ran while external ones failed.
The staging preparer now prints the physical registry directory it validated,
matching the inbox and reservation preparers beside it, and the start path
pins on that returned path. The new end-to-end case drives the shipped
file-signal package from a symlinked home spelling.
* no-mistakes(review): Propagate canonical process-event state roots
* no-mistakes(review): Propagate canonical state to process-event adapters
* no-mistakes(document): Document physical process-event state roots
* fix(pi): deliver captain outcomes as deterministic transcript entries (#3312)
* fix(pi): persist captain outcomes visibly
* no-mistakes(review): Recover captain outcomes after cold-start lock acquisition
* no-mistakes(document): Document cold-start captain-outcome recovery
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes(review): Prove immediate Pi captain-outcome transcript delivery
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* fix(pi): process captain outcomes through a sequence-keyed turn
PR #3312 made every captain-facing supervision outcome a durable, exact-once
visible transcript entry with the read cursor advancing only after that entry
exists. That is the display half of the delivery contract. Left alone it turns
a probabilistic silent loss into a deterministic one: the captain sees an
anchor line, and firstmate never acts, because nothing opens a turn and
nothing records whether main ever processed the outcome.
The 2026-08-31 timeline showed the two shapes this must survive on the
previous hidden-turn path: seven delivered decision outcomes each answered by
an empty assistant message (cursor advanced, no retry, unanswered for close
to three hours), and two answered by an unrelated prior reply. Both happened
because delivery advanced the cursor at enqueue and accepted whatever the
next assistant message was.
Add the processing half on top of the persistence half:
- bin/fm-branch-outcome.sh keeps a processed marker separate from the read
cursor (`unprocessed`, `mark-processed --through`, `processed-init`). It
only advances through an explicit sequence-bound acknowledgement, never
past the read cursor and never backwards; an absent marker reads as zero
and `processed-init` migrates delivered history once so an upgraded home
is not re-presented its past.
- After the visible entry for a captain outcome exists, the extension hands
every still-unprocessed captain row to main as one hidden, typed
`fm-branch-process` request listing each `[seq N] task: summary`, opening
exactly one main turn. Main closes it only by calling the new
`fm_branch_processed` tool with the highest sequence listed. An unrelated,
empty, or paraphrased answer leaves the sequence open, and the same request
is presented again at the end of the next main run and at session start.
The first two presentations of a sequence set open a turn of their own;
after that the request rides the captain's next prompt so an ignored
request cannot loop, and a session replacement resets that budget.
Routine outcomes stay turn-free.
- The regressions cover exactly those incident shapes against the real store
scripts: an empty answer and an unrelated prior answer neither advance the
marker nor stop re-presentation, the acknowledgement is refused beyond the
read cursor and outside lock ownership, a partial acknowledgement keeps the
newer sequence open, and #3312's own assertions now forbid an unkeyed turn
rather than any turn. The store suite pins the marker's bounds and the
migration; the real-SDK guard for appendEntry persistence and model
exclusion is unchanged.
Docs move the protocol from "no model turn" to "one sequence-keyed processing
turn closed only by its acknowledgement", and the verification record carries
the dated run against Pi 0.84.4.
* no-mistakes(review): Harden outcome listing and sequence-bound acknowledgements
* no-mistakes(review): Harden outcome state validation and request pacing
* no-mistakes(review): Reject unsafe sidecars and unterminated outcome stores
* no-mistakes(review): Validate canonical mark-read cursor state
* no-mistakes(review): Guard cursor advancement against corrupt processed state
* no-mistakes(review): Bind acknowledgements to active processing requests
* no-mistakes(review): Reset pacing when processing sequence membership changes
* no-mistakes(review): Enforce silent outcome invariants at storage boundary
* no-mistakes(document): Document hardened captain outcome processing contracts
---------
Co-authored-by: kunchenguid <kun@kunchenguid.com>
* feat: add bounded concurrent Bearings ledger collection (#3481)
* feat: bound Bearings remote ledger collection
* no-mistakes(review): Clarify default remote-ledger collection behavior
* no-mistakes(review): Detach reconcile delivery from watcher loop
* no-mistakes(review): Enforce bounded snapshot and request captures
* no-mistakes(review): Bound legacy summary capture before parsing
* no-mistakes(review): Bound primary remote ledger captures
* no-mistakes(document): Correct snapshot and reconcile documentation
* no-mistakes(lint): Fix ShellCheck quoting in bounded collector
* no-mistakes(ci): Fixed all three CI failures: updated the macOS Bearings assertion to 44 tests, made the home-summary test deterministic and aligned with default ledger consumption, and increased the asynchronous reconcile retirement wait for loaded CI. Verified both focused suites, all 44 Bearings tests, ShellCheck, actionlint, Bash parsing, and git diff checks
* test: await reconcile request retirement
* no-mistakes(review): Avoid empty reconcile queue process churn
* no-mistakes(review): Read ledger summaries from immutable snapshots
* no-mistakes(review): Reject multi-document home ledger streams
* no-mistakes(review): Coalesce durable reconcile requests per target
* no-mistakes(review): Unify reconcile keys and reject snapshot streams
* no-mistakes(review): Key reconcile requests by stable target ID
* no-mistakes(document): Document per-target reconcile request coalescing
* no-mistakes(lint): Remove unused snapshot summary file variable
* no-mistakes(ci): Adjusted the concurrent collector regression’s end-to-end timing ceiling to account for stock macOS process/jq overhead outside the three-second remote collection budget, while remaining below the 15-second serial-read floor. Verified with stock /bin/bash 3.2: all 44 Bearings tests pass; bash syntax and git diff checks pass
* no-mistakes(ci): Fixed legacy summary validation to require exactly one top-level JSON document and added behavioral regression coverage. Stabilized CI by conditionally waiting longer for durable reconcile delivery and synchronously stopping the fm-on worker tree before fixture cleanup. Removed a redundant flaky healthy-path timing assertion; the wedged-reader test still proves concurrent bounded collection. Verified fm-bearings-snapshot, fm-secondmate-reconcile, and fm-on tests, plus project ShellCheck, bash syntax, and git diff checks
* ci: rebalance portable serial test shards (#3489)
* fix(ci): rebalance the portable serial shards on measured durations
The "Behavior portable serial 3" shard ran 17-20 minutes against its
20-minute job cap and intermittently timed out seconds after a passing
test, on branches and on main alike.
Shards are packed longest-processing-time from per-script duration hints,
and those hints were last measured on 2026-08-21 at 116 scripts. The lane
has since grown to 139 scripts and from ~42 to ~63 minutes: 17 scripts had
no hint at all and fell back to the 20 s default, and several existing
hints were low by 2-5x (fm-watch-triage 142 s hinted vs 263 s measured,
fm-public-followup 36 s vs 197 s). The partition therefore looked
perfectly balanced in hint space, 734.6 s per shard, while really running
11.5, 13.6, 18.8 and 16.5 minutes. Script-count balance, which is what the
tests asserted, stayed normal throughout and hid it.
Refresh the hints from the timing artifacts of three green runs, taking
the slowest measurement of each script so the balance holds on a slow
runner, and split the lane across five shards instead of four. Replayed
against those runs' real per-script durations the worst shard is now
12.54 minutes, 63% of the unchanged 20-minute cap, and the serial lane's
wall clock drops from ~20 to ~12.5 minutes.
Bound the drift that caused this rather than relying on the hints being
refreshed by hand: the coverage guard now reports the unmeasured share as
serial_unhinted= and refuses past PORTABLE_SERIAL_MAX_UNHINTED_PERCENT,
which leaves room for newly added tests while making a stale table fail
the guard instead of silently pushing one shard into its cap.
No test changes what it asserts and no test stops running; only the
partition across shards changes.
* no-mistakes(document): Clarify conservative shard timing aggregate
* fix(pi): fall back on incomplete supervision branch prompts (#3491)
* fix(pi): fall back after settled branch errors
* no-mistakes(review): Detect provider errors across prompt compaction
* no-mistakes(review): Preserve in-flight branch state across selection changes
* fix(pi): re-probe supervision branch after cooldown (#3497)
* fix(pi): recover supervision branch after cooldown
* no-mistakes(review): Defer branch recovery until prompt settlement
* no-mistakes(document): Clarify supervision cooldown recovery contract
* fix(bin): remove legacy remote snapshot reads (#3501)
* refactor: remove legacy remote summary reads
* no-mistakes(document): Document ledger-only snapshot reads
* no-mistakes(ci): Fixed the snapshot test fixture so ledger refreshes use the same fake executable PATH as the snapshot consumer. This preserves observable endpoint freshness after removing legacy summary computation. Verified stock Bash parsing and all 44 Bearings tests pass under /bin/bash; git diff checks pass
* no-mistakes(ci): Fixed the CI-only snapshot fixture failure by ensuring the bounded-ledger refresh uses its fake tmux backend. This removes host tmux availability as a source of nondeterminism. Verified all 44 Bearings tests pass, Bash syntax passes, and git diff checks are clean
* no-mistakes(ci): Fixed CI nondeterminism in the Bearings fixture: all local ledger refreshes now use the fixture’s fake tmux backend when available, instead of depending on host tmux state. Verified stock /bin/bash syntax, git diff checks, and all 44 Bearings tests with a deliberately failing host tmux
* fix(pi): preserve watcher continuity across session replacement (#3498)
* fix(pi): rearm watcher after session replacement
* no-mistakes(review): Queue actionable closes across Pi session replacement
* no-mistakes(review): Stop replacement arm when handoff persistence fails
* no-mistakes(review): Preserve actionable wakes through branch and late child races
* no-mistakes(review): Surface late handoff failures without crashing Pi
* no-mistakes(review): Coordinate replacement delivery settlement and unique handoff tokens
* no-mistakes(review): Retry stale deliveries and release settled claims
* no-mistakes(review): Distinguish branch settlement and retry handoff cleanup
* no-mistakes(review): Deduplicate persistent handoff cleanup alerts
* no-mistakes(review): Acknowledge watcher follow-ups only when consumed
* no-mistakes(review): Persist idle follow-ups until agent consumption
* no-mistakes(review): Preserve pending outcomes when handoff persistence fails
* no-mistakes(review): Arm replacement before awaiting prior delivery settlement
* no-mistakes(review): Adopt pending handoffs after lock reclamation
* no-mistakes(review): Prevent stale generations from adopting replacement handoffs
* no-mistakes(review): Scope replacement handoffs by watcher state
* no-mistakes(document): Clarify replacement handoff documentation
* no-mistakes(ci): Fixed the failing branch-extension tests to model the new settlement-promise contract. Failure cases now assert that delivery ownership returns to the watcher instead of expecting direct extension fallback. Verified the updated branch suite, Pi watcher suite, shell syntax, and diff checks
* no-mistakes(review): Update branch settlement tests and preserve chunked outcomes
* no-mistakes(document): Document watcher-owned replacement handoffs
* no-mistakes(document): Verify replacement handoff documentation
* test(pi): cover watcher-owned branch fallback
* no-mistakes(document): Refresh watcher-owned fallback documentation
* fix(bin): resurface task statuses missed by wake handling (#3495)
* fix(bin): resurface terminal statuses lost after branch handling
* test(watch): canonicalize process-event fixture homes
* no-mistakes(review): Index branch outcomes by causal status position
* no-mistakes(review): Recover outcome indexes and deduplicate resurfaced statuses
* no-mistakes(review): Handle legacy ambiguity and oversized status diagnostics
* no-mistakes(review): Keep unclassifiable oversized statuses silent
* no-mistakes(document): Document lost-wake outcome backstop
* no-mistakes(document): Update outcome backstop documentation
* no-mistakes(ci): Fixed CI regressions in wake-drain: parseable reserved-key decisions can no longer bypass the durable decision-fold guard, and status output is prepared and receipt-committed before presentation to prevent repeated one-shot outcomes after later failures. Added a behavioral regression for receipt commit failure and retry. Targeted backstop, correlation-token, decision-cursor, open-decision, unread-status, syntax, and diff checks pass locally. Shard-4 failures appeared unrelated/flaky; the network-parallel test passed locally
* no-mistakes(ci): Fixed the Greptile P1 data-loss issue by committing presentation receipts only after prepared output reaches stdout. Added behavioral coverage proving output failure leaves the backstop retryable and receipt failure may duplicate but never lose a presentation. Relevant wake-drain suites and syntax/diff checks pass. The shard-4 Pi extension failure is unrelated to this PR and did not warrant changes
* no-mistakes(ci): Stabilized tests/fm-bootstrap-network-parallel.test.sh by replacing scheduler-sensitive equal-sleep timing with bounded synchronization between mocked fetch and remote probes. This preserves detection of real serialization while avoiding false failures under CI load. Verified with five consecutive test runs, bash syntax validation, ShellCheck, and git diff checks. The separate Pi stock-rendering failure reproduces locally but is unrelated environment/version drift
* no-mistakes(ci): Fixed Behavior portable serial 4 by adding fm-classify-lib.sh and fm-timeout-lib.sh to the broken-root Pi test fixture; fm-branch-outcome.sh now depends on them. Verified the full Pi branch-extension suite with real-Pi checks skipped, the wake-drain outcome-backstop suite, Bash syntax, and git diff checks. Greptile findings are already addressed at HEAD; the no-mistakes attestation failure is external head-SHA state
* fix(bin): collect follow-up results from remote work homes (#3503)
* fix(bin): deliver typed terminal results from remote work homes
A public commitment whose work is bound to a REMOTE secondmate home could
never receive its typed terminal result. `fm-public-followup.sh brief`
printed an emit command carrying this home's own absolute path and this
checkout's own script path, neither of which exists on the machine the
worker runs on, so the worker had nothing it could write to that the
owning home would ever read - and `consume` kept finding nothing while
the promise stayed open.
The brief is now route-aware: for a remote work home it prints that
route's own code root and home with `--stage-in`, so the typed event is
staged in the home where the work actually runs, and the closing
paragraph names the owning home as the one on the other machine instead
of pointing at the path above it. The owning home collects those staged
results over the same SSH route it reaches that secondmate on, because
the transport only runs outbound: `consume` pulls them into its own
inbox and reconciles them exactly as it reconciles a local report.
Collection is non-destructive until the result is durably held, so a
dropped connection cannot lose a terminal result, and a route that could
not be reached is named in `consume`'s output with the promise left open
rather than reported as an empty inbox.
A local work home is untouched: the brief still prints `--home` with this
home and this checkout's script, and the event still lands directly in
this home's typed terminal-result inbox.
This is the emit-side counterpart of the retire/clear fix in #3479 and
reuses the remote-route resolution that landed with it. Reconciling a
loop bound to a remote route now reaches that route, so the existing
remote cases drive `consume` through the same faked transport their
other steps already use.
* no-mistakes(review): Fail loudly on unresolved routes and invalid staging homes
* no-mistakes(review): Fail collection when remote outbox is unreadable
* no-mistakes(review): Surface reassigned remote routes during empty collection
* no-mistakes(review): Fail remote collection on invalid registrations
* no-mistakes(review): Reject unsafe registration entries during remote collection
* no-mistakes(review): Restore healthy empty remote collection behavior
* no-mistakes(review): Skip remote collection for delivered registrations
* no-mistakes(review): Skip delivered registrations before route validation
* no-mistakes(document): Document remote follow-up collection semantics
* fix(bin): exclude secondmates from home-summary validity (#3504)
* fix(bin): exclude secondmates from home-summary child inventory
kind=secondmate meta records never have backlog rows, so counting them in unowned_children or terminal_in_flight made a clean main home look invalid once earlier ledger checks passed.
* no-mistakes(review): Cover terminal secondmate in-flight exclusion
* no-mistakes(ci): Updated the stock macOS Bash CI snapshot expectation from 15 to 16 tests. Verified all 16 snapshot/fleet-view tests pass under Bash 3.2.57 and `git diff --check` succeeds
* fix(bin): self-heal outcome indexes on first drain (#3509)
* fix(bin): self-heal status-outcome indexes on every drain
Missing ready markers were skipping the lost-wake backstop on non-Pi homes because only the Pi branch ran processed-init. Drain now rebuilds those indexes under the outcome lock and fails closed only on a real store fault.
* no-mistakes(review): Guard held-lock initialization and fail marker writes
* no-mistakes(document): Document cross-harness outcome-index self-healing
* fix(bearings): keep active children underway during captain holds (#3505)
* fix(bearings): keep active children underway beside a captain hold
Project each readable home's active children into Underway independently of the home-level captain-decision classification so a hold no longer hides live work.
* no-mistakes(review): Preserve Underway repos and disclose child truncation
* no-mistakes(review): Fall back to task project for Underway repos
* no-mistakes(ci): Updated the stock macOS Bash CI assertion from 44 to 45 Bearings tests, matching the newly added behavioral regression. Verified all 45 tests pass under /bin/bash, Bash syntax checks pass, and git diff validation is clean
* fix(pi): settle watcher delivery on Pi accepting the follow-up (#3513)
* fix(pi): settle watcher delivery on Pi accepting the follow-up
A follow-up queued while main is streaming joins the running run without
ever raising before_agent_start, so waiting on that event before clearing
the successor pipeline (#3498) stalled every later actionable close: no
successor started, no wake was delivered or offered to the branch, and the
turn-end guard woke main to re-arm by hand after every close.
The pipeline now settles once Pi accepts the follow-up. Consumption is
observed at before_agent_start for an idle main and at the user
message_start for a streaming main, and decides only what a replacement
session (/new, /resume, /fork, reload) replays. An exhausted restoration
delivers its typed failure without launching an arm past the retry bound,
which the stall had hidden. The replacement-coordinator map is typed so the
strict no-emit typecheck passes again.
Tests: the doubles no longer raise before_agent_start for a streaming send,
a portable regression drives two actionable closes while main streams and
proves the successor chain plus consumption-scoped replay, and a
credential-free real-SDK probe pins Pi's event contract for both the
streaming and the idle follow-up.
Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a
* fix(pi): retry a verified successor that fails during wake delivery
A verified successor can exit while the wake it was started for is still
being delivered, most plausibly during a branch turn that holds the
settlement for minutes. Its failure close arrived while the pipeline's
single-flight guard was set, so the close handler skipped the retry, and
the pipeline's end no longer launched an arm, which left the live
generation with no watcher and no retry timer.
The close handler now records that failure when the child had reported
readiness and was not retired by the restoration itself, and the pipeline
runs the ordinary bounded, lock-checked retry for it once the delivery
settles. A restoration started for a later pending supersedes it, and an
exhausted restoration still hands repair to main without a further arm.
The regression holds a branch settlement open while the verified
successor exits with a failure and proves one retry watcher starts after
the settlement releases, none while it is held.
Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a
* fix(bin): bound repeat stale wakes for parked workers (#3532)
* fix(bin): bound repeat stale wakes for a parked but live worker
A worker parked on a declared wait - `paused:` for an external or pipeline
wait, or a verified `captain-held` transfer - kept waking firstmate far inside
FM_PAUSE_RESURFACE_SECS. Observed as five consecutive alarms on one
captain-held worker and dozens across a day on a pipeline wait, and reported
upstream as four wakes in 75 minutes against a 3600s window.
pause_state_class deliberately answers `none` for a still-live agent even under
a declared wait, so a worker genuinely waiting on a decision is never silenced.
That classification is correct and is left alone; it routes every parked but
live worker through surface_nonterminal_stale on first sight of each distinct
stale hash, and an idle parked pane still churns its hash on a clock or a token
counter without changing what is being waited on.
Two places let that churn re-alarm:
- surface_nonterminal_stale queued the wake BEFORE consulting whether a wait was
declared, then wrote `.paused-resurfaced-<key>` - the very throttle that should
have suppressed it. The throttle was never read on this path and was advanced
by the wake it should have prevented.
- The hash-change path cleared that throttle through clear_pause_tracking
whenever the classification came back `none`, so each tick also bought the same
declared wait a fresh window. Fixing only the first site changes nothing.
Read the throttle before anything is queued and advance it only on a wake that
really fires, and on the hash-change path reset only the per-hash bookkeeping
while the declaration still stands, via a clear_stale_hash_tracking split so
neither half of clear_pause_tracking is duplicated. The throttle is keyed to the
declaration, not to the pane.
First sight still wakes, so an inconclusive state is still inspected, and the
window's end still re-surfaces once, so a forgotten wait cannot rot invisibly -
noise traded for a bounded cadence, never for silence. The wake identity stays
the plain `stale: <win>` the away-mode handoff depends on.
Tests cover both observed forms and were confirmed to fail against three
deliberate breaks: each site reverted on its own, and a re-surface that never
fires again.
* fix(document): Clarify declared-wait wake cadence documentation
* fix(ci): Captain, fixed the stale-throttle inheritance: cadence markers now bind to the current wait declaration, so replacement paused and captain-held waits each emit their first plain `stale:` wake. Added behavioral coverage for both forms. Bite proof failed as expected when identity matching was removed, then passed after restoration. Full watcher triage suite, `bin/fm-lint.sh`, syntax checks, and diff checks pass. Changes remain uncommitted for the outer executor
* fix(ci): Captain, fixed the confirmed Greptile finding. `resurface_absorbed` now applies a throttle only when its stored declaration scope matches the current wait, so replacement `paused:` and `captain-held` waits surface immediately without changing classification. Added executable coverage for both absorbed forms. Bite proof failed before the fix at the intended assertion; afterward the full watcher triage suite, `bin/fm-lint.sh`, shell syntax checks, and `git diff --check` passed
* fix(bin): accept the away-mode daemon as the turn-end supervision owner (#3567)
* fix(turnend): accept the away-mode daemon as the supervision owner
While state/.afk exists the away-mode daemon owns supervision and runs
bin/fm-watch.sh one-shot: the watcher exits on every wake and the daemon
starts its replacement. The turn-end guard tested for a live watcher
process holding the watch lock at that instant, so a turn boundary that
landed in the hand-off blocked with "TURN WOULD END BLIND" while
supervision was completely healthy, costing a full handling turn each
time.
Reproduced with the real daemon wrapping the real watcher and the real
guard sampling the same home: 6 of 40 samples blocked, every one of them
with the daemon alive and the beacon 2-3 seconds old, and a new watcher
pid on each cycle. After the fix the same reproduction blocks 0 of 40,
and killing the daemon and its watcher (away mode still on, beacon still
fresh) blocks again.
The guard now accepts a live, identity-matched daemon holding this home
as proof of supervision while away mode is active. The identity match is
the same discipline the watcher lock uses, so a recycled pid or a lock
left by a killed daemon proves nothing. The fresh-beacon half of the
predicate is unchanged: a daemon that stops restarting its watcher still
blocks once the beacon passes grace, a home with no supervisor blocks
exactly as before, and with away mode off the strict watcher predicate is
untouched.
The predicate reads only durable state, so it behaves identically for
every primary harness and runtime backend.
* no-mistakes(document): clarify away-mode daemon supervision proof and test coverage
* no-mistakes(document): generalize stale turn-end predicate summary in architecture.md
* fix(backlog): omit --file from row probes for non-markdown backends (#3582)
* fix(backlog): omit markdown file for beads probes
* no-mistakes(document): Narrow backlog addressing doc to mutations for backend-aware probes
* no-mistakes(ci): Fixed the Greptile P2 review comment (the only failing check) on tests/fm-backlog-atomicity.test.sh. The comment correctly noted that an exported TASKS_AXI_BACKEND environment variable would inherit into the spawned scripts and, because fm_tasks_axi_backend gives it top precedence, override each test case's .tasks.toml backend fixture — making the backend-specific argv assertions fail for environmental reasons. Fix: unset TASKS_AXI_BACKEND in the test harness right after sourcing tests/lib.sh, with a comment explaining why, so every case deterministically exercises its declared backend (4 lines added; no production code touched). Verified: reproduced the leak before the fix (TASKS_AXI_BACKEND=beads made the markdown dispatch case fail with 'beads show failed', exactly the reported failure mode); after the fix the full suite passes (0 failures, exit 0) both with and without TASKS_AXI_BACKEND=beads exported. The added lines are shellcheck-clean (the only shellcheck note, SC1091 on the lib.sh source line, pre-exists this change)
* fix(bin): classify progress updates on requested work as routine (#3589)
The supervision branch's verdict rule escalated every outcome that
answered a captain request, so "the work started" and "still working"
notes reached the captain with nothing to look at. The rule now keeps a
finished result of requested work captain-facing, even when healthy, and
treats start or still-working updates that bring no new artifact,
finding, or decision as routine. The captain list for review-ready PRs,
ask-user findings, exhausted blockers, credentials, and destructive or
security-sensitive cases is unchanged, as are the unsolicited-routine,
silent-fleet-review, and doubt-chooses-captain rules.
The fm_branch_report tool description and the two docs that restated the
old unconditional rule now point at the prompt's "Verdict: routine or
captain" section as the one owner instead of carrying a second copy.
* fix(bin): preserve captain calls during teardown (#3595)
* fix(bin): never close a captain call during cleanup
A scout that held its own work item for the captain, which is what
captain-hold-lifecycle prefers ("hold the work item the question gates"),
was closed by bin/fm-teardown.sh's automatic backlog transition. The
completion gate passed, cleanup ran, and the captain's question moved to
Done with no recorded answer: the one thing the policy says must never
happen. `tasks-axi done` closes a held row silently, and nothing in
teardown asked whether the row was the captain's own call.
bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when
the task is still an open captain call, 1 when it is not, 2 when that
cannot be established. It reads the row through the transition library's
backend-aware probe, so it addresses the same backlog teardown does; the
script's other commands now address the configured data directory the
same way instead of FM_HOME, which also fixes captain holds in a home
with a relocated data directory.
Teardown asks `open` before any destructive step and refuses on 2. On 0
only the close changes: after cleanup and still under the task's own
lock, the row gets one "Deliverable of the finished work" line at the end
of its body and returns to Queued through `tasks-axi reopen`, keeping its
hold, so it lands in Captain's Call instead of reading as work under way.
--force does not lift this: it authorizes discarding unlanded work, never
the captain's question. The deliverable goes into the body because
`tasks-axi update --report` rewrites the title of a row that is not Done.
The crash window reuses the pending-close record teardown already stages:
a `mode=retain` line makes the existing replay record the deliverable and
reopen instead of closing, with the same validator, stale-generation
check, cleanup-incomplete marking, and non-blocking bootstrap lock as an
ordinary close. A retained row the captain answered first simply retires
the record. No parallel record type, recovery command, or second bootstrap
loop is introduced.
Regressions run the real executables: the captain-held scout survives
cleanup queued, held, with its deliverable and on the board, only
`answer` closes it, --force keeps it open, and an ordinary scout still
closes with its report; an interrupted cleanup leaves the row untouched
and the next session start retains it; a relocated backlog keeps the
retention in its one configured file; and a ship row whose hold cannot be
read refuses cleanup before anything destructive.
Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np
* no-mistakes(review): Serialize captain holds and fix backend-aware listing
* no-mistakes(document): Update captain-call retention documentation
* no-mistakes(document): Fix relocated captain-hold backlog diagnostics
* fix(bin): deliver secondmate outcomes to the parent channel (#3592)
* fix(bin): deliver every secondmate outcome on the parent channel from the recording scripts
A secondmate's captain-facing outcomes could miss: the mate model addressed
the captain in its own unread chat instead of appending to the parent
channel, and a PR-ready report, a finding, a decision, a blocker, and a
failure all depended on that one remembered append. Make delivery
structural, so the parent channel never depends on the model:
- bin/fm-parent-channel-lib.sh is the one owner of channel resolution and
exact-line append-once; the merge outcome path and the inactive-outcome
scan now publish through it instead of two private copies.
- bin/fm-inactive-reconcile.sh gains a ledger-first path that runs on every
watcher poll in a secondmate home: a direct child's whole terminal done or
failed line is delivered at once with its note, recorded PR, mode, merge
posture, and scout report pointer, keyed and receipted so it is delivered
once, and the inactive path yields to it. `report <task-id>` runs the same
delivery for a caller holding the child's meta lock.
- bin/fm-pr-check.sh publishes the PR-ready line with the canonical URL at
registration.
- bin/fm-captain-hold.sh publishes a hold and its answer, keyed by task id
and resolution-record count, with no new persisted state.
- bin/fm-teardown.sh delivers the child's final line before removing its
record and refuses, retaining every record, while the channel cannot be
written.
- The charter opens with the parent-channel rule and confines the mate's own
appends to judgement; AGENTS.md carries the carve-out at the persona
address rule and the escalation list.
docs/secondmate-parent-channel.md records the design and its coverage, and
docs/verification/secondmate-parent-channel.md records the live run with real
tmux panes and both real watchers delivering every line with no model.
Supersedes #3569.
* no-mistakes(review): Fix parent outcome retries and reconciliation locking
* no-mistakes(review): Prevent busy children from starving ledger delivery
* no-mistakes(review): Correct ledger metadata and hold occurrence handling
* no-mistakes(review): Disambiguate ledger outcomes and normalize hold reasons
* no-mistakes(review): Close ledger races and preserve teardown records
* no-mistakes(document): Correct parent-channel receipt and scanner documentation
* no-mistakes(lint): Quote done arguments for ShellCheck compliance
* no-mistakes(ci): Fixed both CI failures. Updated GOTMP teardown fixtures for the new final-outcome reporter and isolated them from host tmux state. Updated the PR security assertion to distinguish the accepted PR-ready line from duplicate merge outcomes. Verified with both failing test suites, bash syntax checks, and git diff checks
* no-mistakes(ci): Fixed Greptile’s duplicate-delivery race in bin/fm-inactive-reconcile.sh. Ledger events now claim matching already-delivered inactive receipts using the prior status fingerprint, preventing duplicate parent reports while preserving later same-state completions. Added behavioral regression coverage. Verified inactive-reconcile tests, project lint, documentation audience checks, syntax, and diff checks. Teardown tests passed relevant cases before the documented pre-existing herdr-preflight-missing-adapter failure
* fix(bin): sync remote second mates to primary commit (#3599)
* fix(bin): sync remote second-mate homes to the parent primary commit
Session start and remote launch pointed a remote second-mate home at whatever
Firstmate copy its own host kept, so a home that had already advanced past that
copy refused as a non-fast-forward and every other home stopped at the host's
older commit while the primary ran ahead.
The parent now resolves ITS primary default-branch commit with the existing
helper and hands that commit to the host on both paths. Because a remote home
is a standalone clone, the host imports that one commit before advancing -
already present, else from that host's Firstmate copy without moving it, else
from the home's own origin - and then runs the SAME ff_target guards a local
home gets, so dirty, diverged, feature-branch, and unresolvable targets skip
untouched and the ancestry rules keep one owner. An unimportable target now
names /updatefirstmate instead of failing opaquely, and a host still running an
older Firstmate copy is reported the same way rather than echoing a bare
refusal.
The host-local launch leg no longer re-runs its own secondmate sync, so the
spawn it drives cannot re-target that host's copy after the parent has already
converged the home.
/updatefirstmate is unchanged: it still refreshes the remote code root from that
host's origin and then syncs the home to that refreshed copy, which is what the
sync call with no target commit means.
* no-mistakes(document): Document primary-targeted remote secondmate synchronization
* fix(bin): separate captain intent from firstmate specs (#3597)
* fix(bin): split brief task into captain intent and firstmate spec
Keep no-mistakes --intent as the captain's ask plus later captain words, not the build spec or worker tradeoffs.
* fix(bin): stop task-subsection copies at the next heading
Promotion was swallowing the scout Setup contract into Firstmate spec, and pre-subsection briefs lost their # Task body.
* no-mistakes(review): Validate brief content and preserve nested specifications
* no-mistakes(review): Scope placeholder validation to scaffold-only subsection bodies
* no-mistakes(review): Ignore fenced subsection headings during brief validation
* no-mistakes(review): Preserve captain intent across scout promotion
* no-mistakes(review): Enforce safe intent boundaries for legacy promotions
* no-mistakes(review): Allow marked legacy intent and reject empty promotions
* no-mistakes(review): Scope task parsing and overlay legacy intent contracts
* no-mistakes(review): Overlay current intent contract for all no-mistakes spawns
* no-mistakes(review): Preserve later captain clarifications in intent overlays
* no-mistakes(document): Document brief intent enforcement and ownership
* no-mistakes(ci): Updated spawn-related test fixtures to use valid Captain intent and Firstmate spec subsections, corrected launch-path expectations to launch-brief.md, and resolved ShellCheck quoting findings. Verified with fm-lint.sh and 15 affected behavior tests, including real Herdr tests; all passed
* no-mistakes(ci): Updated stale spawn/promotion fixtures in the Muse, Orca, secondmate-harness, and public-followup suites to provide valid Captain's intent and Firstmate spec subsections. Verified full Orca and secondmate-harness suites, targeted public-followup promotion behavior, Bash syntax, diff checks, and fm-lint
* fix: start a fresh supervision branch for every main session (#3600)
* fix(pi): start a new supervision branch conversation per main session
The supervision branch reopened one recorded conversation forever, so
every main session start reloaded the current generated prompt and then
weeks of accumulated thread, where a superseded rule could still outweigh
today's.
The branch conversation is now scoped to one main session: the session
generation owns the recorded conversation, so a cold start, /new,
/resume, /fork, or a reload always builds a new one, while a rebuild
inside one session (a model or effort change) still continues that
session's own conversation.
The dialog mirror re-anchors with it. Its durable cursor records what the
previous branch conversation received, so a /resume or reload - which
keeps main's own session file - would otherwise leave the new branch
blind to dialog main itself still has. The reset is bounded by the
current main session, and the cursor keeps advancing incrementally within
it. The durable outcome store and its processed marker are untouched, so
unacknowledged captain-facing outcomes still re-present on the new main
session.
* no-mistakes(document): Document fresh Pi supervision conversations
* no-mistakes(ci): Fixed the flaky concurrent inbox failure. Lock acquisition now retries when a competing lock disappears between a failed claim and inspection. Added a behavioral regression covering that race. Verified the full inbox test four times, project lint, and git diff checks
* feat: restart second mates after instruction updates (#3614)
* feat(update): restart second mates whose instructions changed
/updatefirstmate pulled new bytes onto disk and then asked each advanced
second mate to re-read them. A running agent holds AGENTS.md and every
loaded skill frozen from launch and no verified harness offers a reload,
so that steer could not reach a loaded skill at all and left the mate
holding two contradictory copies of its own job description.
An eligible mate is now restarted instead, in the same home and endpoint,
through the existing transactional relaunch. The restart is gated on the
mate first writing down the open work it holds only in conversation - the
open-record half of /stow, never its memory sweeps - so an unregistered
captain call is flushed before the conversation is spent. Anything that
leaves the reload unprovable falls back to the old re-read message and is
reported as exactly that, never as a clean reload.
Remote mates take the same path: fm-remote-secondmate-control.sh gains a
relaunch verb whose host-local leg runs that same control plane, since the
mate is an ordinary local secondmate from its host's point of view. The
primary resolves the profile and passes it explicitly, because
config/secondmate-harness is not inherited and the file on that host
belongs to a different home.
fm-update.sh now splits its advanced live mates into a restart set and a
nudge residual, and both sets require a changed instruction surface, which
also closes the over-nudge against the session-start sweep. Restart is
stricter still: a bin/-only advance reloads itself on the next call, so it
never costs a conversation.
Colocated tests cover the gating, the persist-then-restart order, the
task-subset persist request, each unsafe fallback, the remote hop, and the
remote sync's new instruction-surface report.
* no-mistakes(review): Fix restart correlation, concurrent waits, and lifecycle reporting
* no-mistakes(review): Parallelize relaunches and classify replacement incarnations
* no-mistakes(review): Gate restart actions on live agent state
* no-mistakes(review): Handle failed restart workers without hanging
* no-mistakes(review): Nudge legacy remotes and preserve persist recovery
* no-mistakes(review): Document one-time secondmate restart rollout
* no-mistakes(review): Honor arrived replies and refresh remote profiles
* no-mistakes(review): Revert remote parent profile reconciliation
* no-mistakes(review): Reset remote profile defaults and honor published results
* no-mistakes(review): Preserve fallback nudges for unverifiable secondmates
* no-mistakes(document): Document second-mate restart update flow
* no-mistakes(lint): Fix ShellCheck warnings in restart scripts
* perf: accelerate local validation with bounded concurrency (#3644)
* perf(tests): route gate verification through the bounded concurrent runner
Local validation was the pipeline's dominant cost: across 67 recorded
no-mistakes agent sessions on this repo, 99.3% of command execution was
`bash tests/*.test.sh`, run strictly one script at a time, and 2% of those
calls were killed by an agent-guessed timeout and paid for twice.
Three changes, each measured:
- `.no-mistakes.yaml` pins `commands.test` to
`bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner
already owns changed-file selection, bounded concurrency, the refusal of
unproven scripts, and a generous automatic per-script bound, so the gate's
baseline is neither a serial chain nor a guessed timeout. It stays
intent-targeted - the Test step still runs its evidence agent on top - and
excludes the live-Herdr family the required Herdr lane owns.
- `bin/fm-test-run.sh` gives a plain list of script paths the same bounded
automatic scheduler and automatic bound that `--changed` gets. Naming several
subjects is how a verification round asks for exactly those scripts. The
curated selections are untouched: `--lane` still composes CI shards whose
serial lane must stay serial, `--family` is what the required Herdr lane runs,
and `--all` stays a deliberate complete regression.
- `pr-forge` is admitted to the concurrent-safe family registry on two
consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those,
and records `secondmate` and `session-bootstrap` as refused with the exact
script and reason each failed on, so the refusals are actionable rather than
silent.
Measured on this host, 0 failures on both sides:
verification round, 4 scripts 448s chained -> 231s through the runner (-48%)
pr-forge family 409.2s at 1 worker -> 237.9s at 4 (1.72x)
watcher-wake-lock family 1311.1s at 1 worker -> 539.3s at 4 (2.43x)
A fourth lever was implemented and then removed because the measurement
refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made
`fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s
unchanged, back to back. Those sleeps are not overhead added to the clock -
they are how a test waits for a subject moving on fm-watch.sh's own one-second
cadence - so sampling less often only delays detection. It also broke
`fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a
settled condition. CONTRIBUTING.md records that result so the experiment is not
repeated.
* no-mistakes(review): Separate concurrent runs by isolation proof family
* no-mistakes(review): Limit automatic timeouts to changed-file validation
* no-mistakes(document): Clarify validation concurrency documentation
* fix: copy PR URLs from durable records (#3648)
* fix: copy PR URLs from records or abstain, never assemble them
Supervision reported a plausible but dead PR link three times because its
prompt demanded a full https:// URL at a moment when only a PR number was
observable, so the model assembled an owner/repository from memory, and the PR
check then accepted that URL and wrote it into the task record, after which the
model kept defending its own tool-endorsed guess over the worker's real link.
Three changes close that chain without any live forge lookup, so private
forges are treated exactly like public ones:
- bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy
or abstain" section requires a URL to be copied verbatim from a …
doitdigital0495
added a commit
to doitdigital0495/firstmate
that referenced
this pull request
Sep 21, 2026
…ent checks (#20) * fix(bin): treat Claude Code's default external-imports flags as never asked, not declined (#4387) * fix(bin): read Claude Code's default external-imports flags as never asked, not declined (#4378) fm-claude-trust.sh refused the whole trust registration whenever the project-root entry carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code writes that value only on an explicit "No, disable". Claude Code's default project entry carries Approved and WarningShown both false before the dialog is ever shown, so every such project refused every spawn. Only Approved === false with WarningShown === true — the pair the dialog writes on a decline — now counts as a decline. false/false behaves like an absent flag: trust is registered and no import consent is manufactured. New case test_project_root_entry_default_import_flags_are_not_a_decline fails on b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31, bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * no-mistakes(review): Correct harness doc's external-imports decline predicate --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(bin): keep operator-address labels out of no-mistakes intent (#4445) * fix(brief): keep operator address out of composed intent Teach raw-word authoring for intent sections and mid-task relays, with a neutral [captain] provenance marker for legacy mixed tasks. Keep headings and contract prose outside the serialized intent body. The legacy selector already excluded the old speaker labels from its output; preserve that read compatibility. The reproduced leak comes from adding labels inside a modern intent body, not from the legacy selector. Do not scrub actual request content. Add exact serialized-input and generated-contract regressions, retaining refusal of unmarked legacy tasks and coverage of scout promotion. Fixes https://github.com/kunchenguid/firstmate/issues/3882 * no-mistakes(review): Refuse operator-address lines in Captain's intent body * no-mistakes(document): Document operator-address refusal in intent contract comments * fix: classify OpenCode ellipsis hint as idle (#4451) * fix(composer): recognize Grok 1.0.5's oversized titled bottom border as a proven empty composer (#4455) * fix(composer): accept Grok title overhang * no-mistakes(review): summary: named Grok overhang constant, doc caveat, restored tmux typed-title coverage * fix(bin): translate Stop hook timeout signals into durable auto-arm failure (#4474) * fix(bin): recover Claude auto-arm after timeout * no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list * fix(spawn): establish Claude task channel authority (#4464) * fix(spawn): establish Claude task channel authority * no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference * fix(bin): refuse fm-control.sh exit when the composer holds unproven or pending text (#4458) * fix: guard relaunch exit against pending input * no-mistakes(review): Verifying test run in progress * no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard * no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now * fix(spawn): establish crewmate identity first (#4481) * fix(bin): reconcile redundant secondmate divergence during updates (#4460) * fix: reconcile diverged secondmate updates * no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md * no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md * feat: enable gpt-5.6-luna max reasoning for crew dispatch (#4497) * fix(dispatch): support Codex Luna max effort * no-mistakes(review): use portable CODEX_HOME path in codex effort reference * feat(calm): render smooth Unicode swell with asymmetric two-color sail (#4498) * feat(calm): render smooth Unicode swell * feat(calm): make sails asymmetric * feat(calm): use quarter sail glyph * no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer * no-mistakes(document): docs: sync calm wave phase doc comment * no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。 * fix(bin): supersede stale scout delivery text in brief.md on promotion (#4491) * fix: supersede scout delivery brief on promotion * fix: preserve ship safety contract after promotion * no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch * fix(bin): make captain holds work on hosts with an older JSON::PP, and stop cleanup dropping accents from a held body (#4471) * fix(bin): let captain holds work on hosts with an older JSON::PP Holding a task for the captain, and the cleanup that keeps a captain-held row open, both fail outright on any host whose JSON::PP defaults allow_nonref off - 2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`, but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older library rejects that whole value with "must be object or array". The consequence is fleet-wide on such a host, not one broken command: a worker there cannot formally record a decision for the captain at all. It can only mention the decision in passing in a status line, where it can be missed - which is how a real decision goes unrecorded. The hold reports that the task lost its hold-set stamp; the cleanup cannot return the row to Queued. Both call sites now ask for allow_nonref explicitly rather than inheriting whatever the installed library defaults to. The second one is worth naming: its `/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string case that fails, so the guard selects for the failing input rather than protecting against it. The regression case forces the older default back off for every perl the commands spawn, then drives both paths - holding a task that carries a body, and tearing down a captain-held row whose deliverable must still be appended. It also probes that the simulation genuinely rejects a bare scalar, so the case cannot pass vacuously on a lenient host. Each half was verified failing on its own unfixed call site with that site's real error message. Suites: fm-captain-hold-lifecycle 51 cases, fm-backlog-atomicity 99 cases, 0 failures. Verification limit: the mechanism is reproduced and tested, but neither fix is verified against a real JSON::PP 2.27202 host, because none is in the loop. This laptop runs 4.06, where the bug does not manifest. `bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a brace-delimited object before decoding, so allow_nonref never applies. * fix(bin): stop cleanup silently dropping accented characters from a held body Cleanup rewrites a captain-held row's body to append the finished work's deliverable, and the decoder it reads that body with printed decoded characters to a stream with no `:raw` layer. A character at or below U+00FF then came out as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the accent. `fm_backlog_retain` writes that body straight back through `--body-file`, and nothing reported an error - the character was simply gone from a row still waiting on the captain. The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus `utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already used. Review of the parent commit found this on one of the lines that commit already changed. It predates that change. The test asserts bytes rather than decoded strings, because comparing strings cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any character above U+00FF makes perl print the whole string as UTF-8, so one body carrying both an accent and an em dash passes even unfixed and proves nothing. Verified failing before the fix on the accented row, passing after. Suites: fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures. * no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc * no-mistakes(review): drop whole-file UTF-8 check from retained-body test * no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc * no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc * fix(bin): read codex 0.154's idle braille starfield rows as composer furniture (#4532) * fix(composer): read codex 0.154's idle starfield and status footer as furniture codex-cli 0.154.0 animates a braille "starfield" around its idle composer: on the row above the bold `›` prompt row, on the `›` row behind the SGR-2 dim `Ask Codex to do anything` placeholder, and on the row below it, then draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`). The cells are truecolor greys on both sides of the ghost luminance ceiling, so the brighter ones survive ghost stripping, and the rows below the glyph carry no structural edge. The shared classifier selected the bare `›` shape, extended its wrap region over the two rows beneath the glyph, read the survivors and the footer as wrapped typed input, and answered `pending`; the steering doorbell defers on exactly that verdict, so no doorbell ever reached an idle codex 0.154 pane. bin/fm-composer-lib.sh now recognises that furniture by shape, declared once next to the idle placeholders and reached from the two wrap-region boundary points: - a row whose non-whitespace content is entirely braille cells (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it never counts as wrapped typed content and bounds a bare composer's wrap region; braille behind the glyph row's content is stripped before the emptiness decision when nothing else follows the glyph; a row mixing braille with other text stays typed content; - the codex status footer bounds the wrap region exactly as omp's status row does, anchored on the effort token, a spaced middle dot, and a `~` or `/` path cell, so a typed `fix · tests` stays composer input; - `^Ask Codex to do anything$` joins the verified idle-placeholder set; the ghost strip remains what proves that row empty, and the bare-row rule that bright placeholder text is real input is unchanged. Unchanged: the strict blank-row rule, the styled=0 degradation (a plain cmux/orca capture of this screen still reads `unknown`, never `pending`), FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape. tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte with the divergence (letters in place of the starfield read `pending`) and the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh is the default-on live guard (token-free, skips explicitly without codex or tmux) that launches the installed codex idle and asserts `empty` through both the tmux and the cursorless styled reads, naming codex --version on failure. docs/verification/runtime-backends.md records the dated Herdr evidence: `pending` before, `empty` after, on the captured screen. * no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry --------- Co-authored-by: Todd Billings <todd@usdvcapital.com> * fix(bin): refuse empty text steers in fm-send (#4259) * fix(bin): refuse empty text steers in fm-send A marked secondmate request sent with an empty message delivered only marker and correlation bytes and minted a pending-reply expectation the parent could never see resolved, stalling the fleet with no loud error (#4255). Fail closed on an empty or whitespace-only message on the text path, mirroring the existing --resolve-key refusal. * chore: retain ambient Pi-lens autoformat as its own commit Formatting-only edits produced by ambient Pi-lens autoformat during the msg-loss investigation, kept separate from the behavioural change in c23acba6 so the fix stays reviewable on its own. AGENTS.md is deliberately excluded: its only autoformat edit stripped the trailing space from the documented FM_OPERATIONAL_PREFIX value, which bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11 records as permanent compatibility. Documenting that constant without its trailing space makes the doc wrong about the contract, so that one line was restored rather than retained. * fix(calm): paint the working ship one yellow over all-blue water (#4554) On rose-pine-moon the two-color water (cyan crests over blue troughs) read as a pink stripe over aqua, the yellow left sail and mast clashed with the red right sail, and the hull carried a blue interior run. Every water cell is now blue so the swell reads through glyph height alone, and both sail halves, the mast, and the whole hull are one yellow run. Geometry, cadence, animation, direction flip, resize clamping, and the narrow fallback are unchanged. Update the unit and real-TUI color assertions to the new palette and the Calm docs that described the old one. * fix(bin): stop aging a second mate's active turn from its launch (#4270) * fix(watch): stop aging a second mate's active turn from its launch The parent watcher's second-mate wake-loop stall check exempts a mate that is demonstrably inside an active turn, but secondmate_in_active_turn asked busy_turn_over_age first and returned "not in a turn" whenever that said the bound was crossed. busy_turn_over_age ages from state/<task>.turn-ended, falling back to state/<task>.meta. A second mate's turns end in its own home, so the parent never gets a turn-ended mark for it and the fallback ages the mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago was therefore permanently "over age", the busy pane was never consulted, and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false wake-loop stall. The gate now bounds the busy exemption by <idle> - how long the queue's drain position has not moved - which is evidence this home actually holds. A busy mate stays exempt while the queue has been frozen for less than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so the bound that stops a busy pane from proving liveness forever is kept rather than removed. busy_turn_over_age is untouched; its remaining callers are the ordinary crew busy-pane bound. The regression pins the case that actually broke: a mate whose launch record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn must not escalate, while the same mate with its queue frozen past the bound still publishes exactly one notification. The existing coverage only exercised a freshly launched mate, which passes either way. Reaching that alert now costs a pane capture inside the gate, so the three checkpoints in this suite that assert an alert move from a 1s to a 4s bound - the value the neighbouring active-turn cases already use. The bound is a ceiling, not a wait: the checkpoint returns on the first actionable wake. On a loaded machine a 1s bound missed the alert repeatedly; at 4s it did not miss in 20 runs under the same load. * no-mistakes(review): scope the second-mate active-turn regression test's coverage claim * no-mistakes(document): fix stale second-mate active-turn comments in fm-watch * feat(bin): add read-only PR blocker and reviewer discovery commands (#4278) * feat(bin): add read-only PR blocker and reviewer-discovery commands Two focused, opt-in commands that read GitHub and never write to it. fm-pr-state.sh reports what still blocks one pull request from the author's side: a closed or merged state, draft state, unknown or conflicting mergeability, absent or failing required checks, and a blocking CHANGES_REQUESTED decision explained by each reviewer's latest verdict, marked STALE when it was left at a superseded head. A pull request that only awaits an approval is not reported as blocked, and advisory checks are omitted. Every reading is taken against one exact head; a push that lands mid-read invalidates the whole result rather than mixing two snapshots. fm-pr-reviewers.sh suggests reviewers from the most recent commits to the pull request's exact changed paths, counting each commit once, resolving handles through GitHub's own commit author.login mapping, and excluding the author and Bot accounts. Both stay read-only: no review request, no approval, no merge. Unresolved review-thread state is left unreported because the REST API does not expose it and unattended commands may not use GraphQL. Closes #3731 * no-mistakes(review): accept only PR URLs and stop at terminal state * no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate * no-mistakes(review): stop attributing readings to unverified heads * no-mistakes(review): narrow readiness contract to checks that have reported * no-mistakes(review): read the pull request once, drop the head guard * no-mistakes(document): scope pr-forge isolation proof to its measured members * no-mistakes(document): record uncovered pr-forge members and their pending proof * docs(isolation-proof): re-prove pr-forge at its full membership tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the pr-forge family in this branch, and script_allows_concurrency grants four workers by family membership alone, so both ran concurrently on a proof measured before they existed. Re-proved the family at all eight members: two consecutive runs, 0 failures, each begun with the one-minute load average below 6.0 so the result measures isolation rather than contention. A third run taken between them is disclosed rather than recorded, because it started while the previous run's workers were still decaying. The new durations are not comparable with the six-member measurement above them, so they are not presented as evidence about the two new members, and that record's 1.72x four-worker figure is left as a statement about its own run rather than restated as current. * no-mistakes(review): disclose gh error-text coupling at its matching site and tests * fix(bin): teach validation-round pauses in generated briefs (#2752) * fix(bin): teach validation-round pauses in briefs * no-mistakes(document): Point classifier comments to authoritative pause examples * docs(readme): add star history chart (#4558) * fix(bin): refuse teardown when a task's endpoint close fails (#4510) * fix(teardown): refuse a cleanup whose endpoint close failed bin/fm-teardown.sh discarded both the exit status and the stderr of every fm_backend_kill call, so a close that genuinely failed was indistinguishable from one that succeeded. Teardown continued past it, deleted the task's durable records, returned its worktree, and reported the cleanup as completed. The deleted metadata is the only record of which endpoint belongs to the task, so such a close did not merely leave a stray session behind, it stranded one: nothing was left on disk naming it. The adapters could not carry that signal either. Driven against the real code, every backend arm returned 0 for a genuine failure exactly as it did for an already-exited endpoint, so there was nothing for the four call sites to propagate even once they stopped swallowing it. The tmux arm now resolves a close that did not succeed against the window's exact recorded identity, since kill-window fails the same way for a window that is gone and one that is still there. The Orca arm reports a close its missing CLI never attempted. Both stay silent for an endpoint that is already legitimately gone, and the remaining arms are unchanged: their close-command timing cannot be established without the real Zellij, Orca, and cmux binaries, and a gate that refused ordinary cleanup of an already-exited session would be worse than the defect. docs/verification/runtime-backends.md records what each backend can prove. A reported close failure now reaches teardown's existing retain-and-stop refusal before the records naming the endpoint are removed, matching where the Herdr confirmed-gone gates already sit for the same hazard, and the retained records let a rerun finish once the close works. * no-mistakes(review): refuse unreadable tmux close re-read; honor --force override * no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close * no-mistakes(document): document endpoint-close refusal in its backend and retirement owners * no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree * feat(calm): add flag-gated Claude Code Calm mode (#4565) * feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that brings Calm to Claude Code: the sailboat replaces the stock working row through a Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration, and canonically classified operational user rows draw at zero height. /calm is registered by the hooks module itself and toggles the same per-home config/calm preference the Pi extension uses, so one choice applies on either harness; rows redraw retroactively on toggle and stay hidden across claude --continue. The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS flag is on. Nothing sets that flag in any settings file, and the plugin carries no command file, skill, agent, or classic hook, so it is a complete no-op while the flag is off. The trusted project auto-loads it through an .agents/skills symlink, the only path Claude Code scans for project plugins. Extract the working-ship geometry, bounce track, cadences, and freeze/resume state into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module imports from outside the plugin folder) and have the Pi widget paint that core's frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify operational rows through a port of bin/fm-operational-input.sh's classify command guarded by a corpus parity test against the shell owner. Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering, Raster packing, policy, classifier parity), the mod's own claude plugin test suites behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op, the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272. Docs: record the version-scoped Claude Code evidence and the three bounded gaps in docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md, and make the shared preference, layout, and contributor notes harness-neutral. * no-mistakes(review): Preserve colliding final replies and strengthen parser parity * no-mistakes(review): Preserve final replies and strengthen canonical parity checks * no-mistakes(review): Require exact function-hooks opt-in before Calm activation * no-mistakes(review): Clarify Calm module loading and activation boundaries * no-mistakes(review): Reset Calm presentation state across session starts * no-mistakes(document): Refresh Calm session lifecycle documentation * feat(calm): paint the Claude Code working ship in Claude's own theme colors The captain picked the "Claude native" palette for the Claude Code mod's Raster: every water cell takes the spinner blue of the active theme family (#93a5ff dark, #5769f7 light) and the whole boat takes the Claude orange of the stock spinner (#d77757), one water color and one boat color. The family follows the `theme` setting's prefix, read at load through $.config.list and re-read on a config.set of that row, with `auto` and custom themes falling back to the dark set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte. Rename the shared sprite's color classes from hue names to `water` and `boat`, since each harness now maps them to its own colors; geometry, motion, cadence, and the activation gate are untouched. Tests cover both palettes' packing and the family rule under Node, and the plugin kit drives every theme value, a theme change mid-session, the Calm-off pass-through, and inertness of the menu read while the flag is off. The docs describe the Claude Code colors and record the guard passing on 2.1.273. * no-mistakes(review): Use light palette for unresolved Claude themes * no-mistakes(document): Refresh Claude Calm verification evidence * fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586) * fix(watch): honour a declared wait before wedge-escalating a quiet pane wedge_timer_check escalated on elapsed idle time alone. Nothing asked whether the worker had already said why its pane was quiet, so a lane that declared a bounded external wait climbed the escalation ladder for as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every repeat carried demand-deep-inspection - which by its own wording forbids re-absorbing on the run-step or pane state, so the supervisor could not use the evidence that was there either. The generated brief promises that declaring `paused:` buys the long recheck cadence instead of a wedge, but the timer was still reachable while that declaration stood: a crew that declares a wait and then has an active run or busy pane attributed to it is handed to the timer as provably-working. The declaration is what the worker said about its own silence, so it now outranks a liveness verdict that only says something is running. The consult runs in the at-threshold branch that was about to escalate, beside the worktree walk already there, and costs one status-line read. Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS recheck the declared-wait absorber already uses, so the wait is still rechecked and cannot rot invisibly. Which verb declared it decides the wording, because the two block on different people: a `paused:` wait is owed by an external dependency and asks the reader to confirm it still holds, while a `captain-held:` transfer is owed by the captain reading the recheck and asks them to answer or release the hold. A hold is not rechecked at all while the away-posture record exists, as on every other captain-held path, and that absorb arms no throttle so the recheck is owed in full on return. A declared clearing time that has already passed stops counting, and a lane that never declared one keeps the identical escalation schedule, reason, count and demand-deep-inspection wording, so detection and its worst-case time are unchanged. The deferral restarts the idle timer rather than cancelling it, so a lane that stops waiting escalates again within one threshold. A lane quiet because its own validation run is parked at a gate awaiting a human decision is deliberately out of scope: reading that state needs a signal carrying who the wait is on and what clears it, rather than one inferred from a parked verdict that also covers gates awaiting the crewmate itself. Tests pin both directions for each case and were each confirmed to fail with the consult removed. * no-mistakes(document): docs: honour declared waits in stale-escalation docs * fix(bin): report verified PR state for passed runs (#4624) * fix(bin): derive passed PR state from PR record A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists. For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim. Fixes #4607 * no-mistakes(review): Add bounded GitLab merge-request state reads * no-mistakes(review): Preserve network-free inactive crew-state scans * no-mistakes(document): Document PR record readers in shared library * fix: restore published contribution follow-up (Fixes #4469) (#4627) * fix: restore published contribution follow-up (Fixes #4469) * fix(review): Fix contribution freshness and merge actor routing * fix(review): Restore issue triage and scope contribution follow-up * fix(test): test: assert one wake per contribution signal * fix(document): Document contribution follow-up * fix: restore truthful terminal delivery evidence * fix(review): Disclose unsupported contributions and deduplicate watcher wakes * fix(review): Preserve unmeasured unsupported contributions across Bearings * fix(review): Deduplicate shared contribution wakes and isolate diagnostics * fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower * fix(bin): make remote report transfers explicit and fail-open (#4658) * fix(bin): make a remote-reply document gap self-clearing and re-attemptable A remote mate's undelivered document raised a keyed `blocked` decision that nothing could ever resolve, and any `data/*.md` substring in any mirrored line was an unconditional fetch instruction. A mate announcing a report it had not written yet therefore manufactured a permanent, factually false blocker, and its own explanation of the false alarm manufactured more. The reader has no permanence vocabulary: a report still being written refuses exactly like a path that will never exist. So an undelivered document is now a durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`, re-attempted on the next delta and on the channel's own quiet poll, and retired with a matching `resolved` line naming the local copy once it arrives. The cursor still advances and no delta stalls on one bad pointer. Only a structured `report=data/....md` pointer now offers a document, so a path merely mentioned in prose - including one under another home's mirror tree, which is provably not that mate's to serve - is never fetched. Offers are deduplicated across the whole delta, the escalation names each missing document once and carries the reader's own reason instead of discarding it, and a strictly increasing notice ordinal keeps a later escalation from being swallowed as duplicate bytes. A mirrored line still lands once whichever pointer form it was first written under. * no-mistakes(review): Require structured pointer token boundaries * no-mistakes(review): Unify boundary-safe pointer extraction and rewriting * fix(bin): identify a mirrored line independently of its delivery state Two defects in the boundary-safe pointer work. The at-most-once check compared only the all-remote and all-local renderings of a line, so it could not recognize a mixed one. A line offering two documents where only the first was deliverable mirrored as local-plus-remote; once the second arrived, a cursor-loss whole-log recapture rendered the same line all-local, matched neither alternate, and mirrored a second time. A line's identity is now the canonical form every boundary-valid pointer would take once delivered, derived by the same parser that does extraction and rewriting, so it no longer depends on which documents happened to be deliverable at the time. The pointer map was passed to awk through the process environment. A delta may carry up to the configured 1 MiB bound, and an expanded map of delivered pointers can exceed the platform's exec argument limit, so awk would fail to start; because no caller checked, the empty result would have been appended as blank lines while the cursor advanced past dropped status content. The map now travels in a file, and every call site checks the exit status and stops the ingest rather than committing a delta it could not render. Both passes now run once per stream instead of twice per line. * no-mistakes(review): Abort ingest when document pointer extraction fails * no-mistakes(review): Exclude structured cross-home pointers from document transfer * fix(bin): fail open on an undeliverable remote document instead of tracking it Narrow the remote-reply document fix to the scope the diagnosis actually requires, as decided after measuring a simpler alternative. A document the reader cannot deliver now fails open. The mate's line is mirrored with its own pointer, the cursor advances, and one unkeyed note carries the reader's reason. A note never enters the open-decision fold, so it cannot stand open the way the original keyed block did - which removes the never-clearing false blocker by construction rather than by resolving it. That makes the durable self-clearing obligation unnecessary, so it goes: the per-mate pending-documents record, its notice ordinal and resolved announcements, and the poll-side retry. Canonical line identity goes too, and with it a way to silently drop a genuine status line; mirroring is back to at-most-once on exact bytes. The cross-home exclusion goes as well: under fail-open a cross-home report= either fails harmlessly or is a nested remote report this mate genuinely holds, which is now relayed again. Kept: fetching only on a structured report= pointer, the boundary-correct parser, the file-based rewrite map, and checked extraction and rewrite exit status. The parser now scans behind a sentinel byte so a rejected candidate can no longer give the text right after it a false leading boundary. The reported incident is covered end to end: a report path announced in prose before it exists raises no decision, and the report still arrives through the ledger publisher's structured offer once written. * no-mistakes(review): Preserve source-line identity across remote reply replays * no-mistakes(document): Document remote reply transfer and replay semantics * no-mistakes(lint): Fix staging truncation lint checks * fix(calm): preserve substantive mid-turn responses (#4655) * Preserve substantive Calm mid-turn text * no-mistakes(review): Distinguish newline-preserved replies from short narration * no-mistakes(document): Document Calm mid-turn preservation boundaries * no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check` * fix(bin): preserve PR merge polls across volume remounts (#4656) * fix(bin): re-record PR poll identity after a volume device renumber (Fixes #4260) A volume remount can renumber the state filesystem's st_dev while every inode and byte stays the same; APFS does this across a reboot. A poll registration records its sidecar and check as device:inode, so every poll armed before the remount failed strict validation and the watcher refused all of them as unauthenticated state checks until each was re-armed by hand. There are two device comparisons. fm_pr_private_file_valid compares a live file's device with the state directory's device read in the same invocation: it refuses a file that is not on the state directory's own filesystem and already survives a renumber, so it is unchanged. The registration's recorded identity versus the live identity (from #556, reused by the #932 retirement receipt) binds the registration to the exact files published in its own transaction; its device part is what breaks. When strict capture fails, the watcher now proves the device is the only difference: every other artifact check passes (template bytes, both hashes, private mode, single link, live device, metadata), both recorded identities name one device, and each recorded inode equals its live inode. Only then, under the task's control lock, does it rewrite the two identity lines, repeating the whole proof and comparing the registration's file identity and bytes just before the rename, and then capture strictly again. A swapped, altered, re-moded, relinked, split-device, or foreign-device artifact still fails a proof and is still refused, and a pending retirement receipt blocks the rewrite. Reproduction: on macOS a poll armed on an APFS disk image that was detached and re-attached behind another image moved st_dev 16777239 -> 16777243 with inodes, bytes, mode, and link count unchanged; the real watcher refused it on main and reports its merge with this change. The portable regression test rewrites a real registration's recorded device and drives the watcher. Not changed here: the status presentation cursor keys rows by its own device:inode identity in bin/fm-classify-lib.sh, a different helper that needs its own fix; a retirement receipt left by a reboot between its publication and removal still names the old device and stays refused; custom check trust binds only a content hash and is unaffected. * fix(review): Serialize PR poll publication writers * fix(review): Bound PR poll publication lock scope * fix(bin): keep contribution records when the poll budget runs out (follow-up to #4627) (#4661) A budget that expires partway through an observation no longer records an error or prints the unavailable wake; the URL keeps its prior record and is observed first next poll. forge() flags budget exhaustion at the point it refuses, or when a read is killed at the budget's own deadline, so a genuine forge failure still records the error and wakes. Each distinct URL is now observed once per poll and applied to every owning task. * fix(bin): clear parent pending-replies on local secondmate retirement (#4680) * fix(bin): clear parent pending-replies on local secondmate retirement Local secondmate teardown left resolved parent pending-reply records behind after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced retirement while any reply for that id is still unresolved, and delete every matching record plus its delivery confirmation after a successful local or remote retirement, matching the remote cleanup path. * no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup * no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung * no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern * no-mistakes(review): Pending-replies Basename und corr_id abgleichen * no-mistakes(document): Clarify forced retirement pending-reply cleanup --------- Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc> * fix(bin): accept Orca's composite worktree id when tearing down a task (#4677) * fix(bin): accept Orca's composite worktree id at teardown Teardown refused every Orca-backed task because the endpoint validator checked orca_worktree_id with the simple-atom rule meant for tmux-style window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca returns that id as `<orca id>::<absolute worktree path>`, so the colon and slashes in every real value made validation fail and finished Orca tasks could never be cleaned up. Validate the field as the composite it is: both halves of the first `::` split present, the path half absolute, and no embedded newline, carriage return, or tab. The terminal field keeps the atom check, which is correct for it, and no other backend's validation changes. The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca never returns, which is why the suite passed a check the real value fails. They now carry the composite form, so the tests exercise the real value. * no-mistakes(document): name Orca's repo id in the composite worktree id * no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points * feat(bin): add opt-in typed dispatch resolution (#4692) * feat(bin): add opt-in typed dispatch resolution through typesafe.ai Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or scout profile from a written brief with typesafe.ai's System One model: one Choice question over the rules' `when` texts, then the confidence floor, the rule's `approval` and `floor`, each profile's `provider` and `floor`, one quota-axi snapshot, and the spendPriority argmax all in code. It is off unless TYPESAFE_API_KEY is in the environment or the home's gitignored .env; off means one stderr line, exit 0, and no network call, so firstmate dispatches exactly as before. The key reaches curl on a file descriptor, never argv. Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates the four new optional dispatch fields. Document the schema, the operator contract, the AGENTS.md intake step, and the live and benchmark evidence. * no-mistakes(review): Harden typed dispatch resolution and quota bounds * no-mistakes(review): Validate dispatch floors and ranking evidence * no-mistakes(review): Tighten dispatch response and floor evidence * no-mistakes(review): Neutralize none matching and resolve defaults locally * no-mistakes(review): Preserve providerless profiles outside typed resolution * no-mistakes(review): Validate response usage and reject duplicate profiles * no-mistakes(review): Escalate unverifiable floors and validate probabilities * no-mistakes(review): Validate probability mass and unknown profile floors * no-mistakes(review): Simplify resolver interface and preserve fallback routing * no-mistakes(review): Fix constants and rank partial quota evidence * no-mistakes(review): Add authoritative provider mapping and enforce explicit providers * no-mistakes(review): Declare provider for documented Pi profile * no-mistakes(review): Validate provider identifiers and support Gemini dispatch * no-mistakes(review): Strictly anchor provider identifiers * no-mistakes(review): Validate selectors and preserve fallback candidate evidence * no-mistakes(review): Gate typed validation and harden resolver evidence * no-mistakes(review): Preserve opt-in routing and harden candidate evidence * no-mistakes(review): Prioritize known exhaustion over quota uncertainty * no-mistakes(review): Isolate API secrets and preserve no-key diagnostics * no-mistakes(review): Fallback safely when dispatch rules are absent * no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets * no-mistakes(document): Document typed dispatch safety and fallback behavior * fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753) * test: reproduce buried status declarations in shared readers * fix: share status event reads and preserve open blockers * fix: retain terminal scout and ship status declarations * no-mistakes(review): Fix status chronology, legacy completions, and reader performance * no-mistakes(review): Share terminal decision reconciliation across fleet snapshots * no-mistakes(review): Unify terminal supersession across cached folds and consumers * no-mistakes(review): Filter per-key status history while preserving terminal chronology * no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells * no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses * no-mistakes(document): Document latest-event status read and kind-scoped fold cursor * no-mistakes(lint): Quote literal done in test for-lists for SC1010 * ci: expect 19 snapshot/fleet-view tests This branch adds a fleet-snapshot regression, so the stock macOS Bash lane's hardcoded guard of 18 'ok - ' lines fails on the new count. Bump the guard and its message to 19. * no-mistakes(review): Restore multiline child outcome reporting * no-mistakes(review): Select ledger terminal events through bounded shared reader * no-mistakes(review): Report newest open decision instead of preferring blocked * no-mistakes(review): Require colon before ship/scout terminal supersession in fold * no-mistakes(review): Gate socket-down override on latest event; drop lock matrix * no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions * no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold * no-mistakes(test): Update fleet-view expectations to newest-open-decision rule * no-mistakes(document): Align status-read docs with fold-resolved crew state * no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers * no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree * test: fold terminal-cleanup snapshot coverage into the completed-scout case Keep the ship/scout/secondmate supersession assertions without adding a nineteenth top-level fleet-view test, so CI can stay at the upstream suite count. * no-mistakes(document): Clarify socket-down override expiry in architecture doc * ci: retrigger flaky contribution check * fix(bin): launch codex crewmates with codex's hook layer disabled (#4689) * fix(spawn): launch codex crewmates with codex's hook layer disabled A freshly launched Codex worker never reached its instructions. Codex stopped it on an interactive "Hooks need review" modal whose selection sits on "Review hooks", which is neither trusting nor declining. Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow navigation, so the selection cannot be moved, and pre-accepting the prompt by writing Codex's own trust store would record an operator consent that was never given. The hooks are the machine's own ~/.codex/hooks.json plus any project's .codex/hooks.json. A crewmate needs neither: its turn-end signal is the -c notify= program on the same launch, and Firstmate's project hooks are primary-session infrastructure that stands down in a child worktree. Crewmate and scout launches now pass --disable hooks. That is the opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted hooks; disabling the feature runs none of them and leaves the operator's ~/.codex untouched. An unknown feature name is a hard Codex error, so a release that drops the flag fails the launch loudly instead of silently restoring the modal. A secondmate is a primary in its own home and keeps the project hooks its turn-end guard and session-start digest ride on. Verified on codex-cli 0.151.0: the modal is gone and the turn-end notification still lands. This unblocks the second review that every finished pull request is supposed to get. Fixes kunchenguid/firstmate#4673 * no-mistakes(review): Fix contradictory hook count in Codex verification record * fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710) * fix(bin): settle terminal contributions and wake once per read-failure episode A contribution whose last good observation is merged or closed is final: poll no longer re-reads it, projection keeps it fresh, and a stale error recorded beside it is cleared once. A genuine forge-read failure on an open contribution still records its error on every cycle but prints the unavailable wake only when it starts a failure episode; a successful read ends the episode. Open PRs linked from done tasks keep being observed. The false unavailable beside a complete observation was budget exhaustion mid-observation, already fixed by #4661. * fix(review): Settle terminal contribution owners * fix(review): Deduplicate shared contribution failure episodes * fix(test): Preserve settled terminal contribution records * fix: select authoritative no-mistakes runs (#4476) * fix(crew-state): select authoritative validation runs by identity Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row. Refs: https://github.com/kunchenguid/firstmate/issues/3215 * fix(review): Resolve same-branch run identities beyond capped history * fix(review): Fix run-selection compatibility, races, and worker-state fallbacks * fix(review): Limit run validation to the requested branch * fix(test): Anchor AXI fixtures and document remaining live evidence gaps * fix(document): Clarify run selection documentation and capture ownership * fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation * fix: distinguish captain outcomes from no-op updates (#4738) * fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work MAIN answered a supervision-branch outcome for completed captain-requested work (implementation done, PR ready for review and merge approval) with "Captain, shipshape.", reading section 9's no-action reply as covering it and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no captain-facing response is owed". Section 9 now limits the shipshape reply to true no-ops (idle re-read, empty heartbeat, consequence-free acknowledgement) and requires a short outcome response naming what finished and what word is needed whenever requested work finishes or a result needs the captain's word, even when a transcript entry already shows the substance. The Pi protocol's re-emit rule now says it bounds repetition only, and carries a worked example of the ready-for-review outcome whose correct processing turn a shipshape reply fails. No executable contract evaluates the content of MAIN's captain-facing reply, so the regression is the protocol example in the owner doc rather than a text-match test. * no-mistakes(document): Clarify captain-facing outcomes versus no-ops * docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line The document step condensed the Pi protocol's re-emit rule and dropped the worked example of a finished, ready-for-review outcome whose correct processing turn a "Captain, shipshape." reply fails. That example is the contract's regression: no executable contract evaluates the content of MAIN's captain-facing reply, so the owner doc's example is the test case. Restore it directly under the re-emit rule, prefixed as a regression example that is kept verbatim and never condensed or summarized away. * no-mistakes(review): Clarify captain outcome and decision-word requirements * no-mistakes(document): Clarify captain-facing completion outcomes * docs(pi): require the PR URL in the visible captain-facing outcome reply Captain review on the regression example: drop the sample reply string and say only that the ready-for-review outcome requires relaying a captain-facing outcome response, not just "Captain, shipshape.". Fold in the visible-PR-handoff failure seen this session: after the branch outcome reporting this fix green, MAIN's visible reply was only "Awaiting your merge call." with no PR URL, leaning on the dim anchor. Section 9's URL rule now also covers a review or merge ask and names the visible reply as where the URL goes, sourced from the ready status, pr= metadata, or the supervision branch's summary and never left to a transcript entry. The Pi protocol adds the same-way failure and places the captain-facing text in the final visible assistant reply after the fm_branch_processed call, because Calm hides assistant text emitted in the same step as a tool call as a working note. Investigation verdict, evidence in the PR comment: no recent PR caused the handoff failure; Pi has hidden same-step pre-tool assistant text since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and #4658 touched only remote report transfer. * no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs * no-mistakes(document): Clarify captain-facing supervision outcomes * docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule The consolidated section 9 URL rule narrowed its trigger to a review or merge ask, dropping the "whenever a PR is mentioned" catch-all from #3648 that keeps every PR URL copied from a durable record and never assembled from memory. Restore that trigger as a union with the review or merge ask so the one consolidated rule covers both. * fix(bin): let non-owner Claude Stops exit safely (#4777) * Fix foreign-owner turn-end supervision loop * no-mistakes(review): Scope foreign-owner safe exit to Claude guard * no-mistakes(document): Document Claude foreign-owner safe exit * fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778) Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty indexed array as an unbound variable and aborts the shell. In signal_turnend_panes_churned() the missing_keys loop was reachable with an empty array whenever every churned key already held a fresh .churn-since-* marker (a second churning turn-end inside an open deferral window), so each watcher cycle died about half a minute in and supervision restarted endlessly. The created_keys rollback loops had the same latent crash on their error paths. Audit of bin/ for the same pattern found one more confirmed-reachable case: remote_handoff's noncanonical-body scan iterates to_move, which is empty when a retried remote handoff finds every key already staged in the outbox. All other "${arr[@]}" sites are either count-guarded, guaranteed non-empty by construction, or unreachable while empty. Guard the three reachable expansions with the repo's existing "${arr[@]+...}" idiom. New regression test drives a real watcher through the all-marked churn path; the macos-stock-bash CI lane runs it under real /bin/bash 3.2 via FM_TEST_ONLY. * Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783) The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: require complete captain-facing final responses (#4779) * docs: require complete final responses across harnesses * no-mistakes(document): Document complete final replies for Grok Bot * docs: point Grok replies to the shared contract owner * no-mistakes(review): Clarify final recap without batching decision asks * fix: preserve substantive mid-turn text in Pi Calm (#4788) * fix(calm): preserve substantive Pi mid-turn text * no-mistakes(review): Preserve substantive Pi Calm text per block * no-mistakes(test): Cover shared Calm preservation boundaries behaviorally * no-mistakes(document): Consolidate Calm preservation documentation * fix: harden mail checks and rebalance full-coverage CI (#4800) * Improve CI reliability and rebalance full-coverage validation * no-mistakes(document): Clarify lint partition documentation * fix(bin): answer Kimi 2.0.0 folder-trust dialog during spawn (#4799) * Handle Kimi workspace trust dialog * no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers * no-mistakes(review): Gate Kimi ready on any trust marker and clean captures * no-mistakes(review): Read visible pane for Kimi trust and ready gates * no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate * no-mistakes(review): Harden Kimi viewport capture and trust dialog detection * no-mistakes(document): Document Kimi spawn refusal on cmux and Orca * fix(bin): report a dead-agent record once instead of escalating forever (#4775) * fix(bin): report a record whose agent is gone once instead of escalating forever The wedge escalation path never asked whether there was still an agent to be wedged. A wedge is something stuck that might recover, so re-alarming it earns its cost; an agent that is gone never moves again, its pane never churns, the idle timer never resets, and the escalate path clears its own timer and re-arms with nothing bounding the count. Observed on a live fleet: two finished lanes reached 226 and 203 consecutive escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400 notifications a day from two lanes with no agent running at all. On one, fm-control.sh exit answered already-stopped and fm-crew-state.sh read "failed - run failed". Closing the Herdr pane did not stop it either: with the pane genuinely gone and herdr pane read returning pane_not_found, the count kept climbing, because the poll is driven by the record's window= line rather than by the pane. The cost is not the repetition but that it drowns the alarms that matter. fm_backend_agent_state already separates a thinking agent from a gone one at process level. In the branch that was about to escalate, read it once and treat only its two recovery-grade verdicts - dead (endpoint present, no agent in it) and missing (endpoint authoritatively absent) - as proof, reporting that record once and not re-escalating it while it stays that way. Every other verdict, including alive, ambiguous, unreadable, unverified, and a read that failed outright, keeps the identical schedule, reason, and escalation count, so a genuinely wedged live agent is unaffected. The probe costs at most one backend read per window per threshold, the same budget the declared-wait consult and the worktree write probe already take. The report decides nothing about the record's fate: both lanes still held unlanded work and teardown refusing them was correct, so retiring, relaunching, or cleaning up stays with the supervisor. The once-only marker is owned entirely by that function and is dropped by the same read the moment the endpoint stops reading gone, so a replacement launched into the same window escalates normally and its own later death is reported again. Related, and not closed by this: #4412, #4482, #4316. Tests drive the real watcher against a record whose endpoint does not exist and pin both directions: dead and missing report once and never advance the count across later thresholds, while alive, ambiguous, and unreadable endpoints keep escalating with the identical reason and a climbing count. * fix(bin): bind the once-only dead report to the pane it reported Review of the parent commit found a reachable sequence where a later death in the same window lost its promised report. The marker was keyed on the verdict string alone and dropped only when a threshold probe read a non-gone verdict, but probes run only at thresholds: a replacement launched into the same window that dies without ever being probed alive - it crashes at startup, or works and then crashes - was absorbed by the previous death's marker. The pane's first sight yielded only the generic stale wake and every later threshold matched the stale marker, so the second death never got the detailed once-report that both the function's own comment and docs/architecture.md promise. Record the verdict together with the pane hash it was reported for, and absorb a repeat only while both still match. A replacement churns the pane, which resets the stale suppressor, wedge timer, and escalation count while no reset site touches this marker, so the pane half is what tells the second death apart from the first. The live-probe drop stays as it was. Clearing the marker at those reset sites instead would re-open unbounded re-alarming for a dead pane whose display ever ticks, which is the exact defect th…
friesentius
pushed a commit
to friesentius/firstmate
that referenced
this pull request
Sep 21, 2026
* fix: copy PR URLs from records or abstain, never assemble them Supervision reported a plausible but dead PR link three times because its prompt demanded a full https:// URL at a moment when only a PR number was observable, so the model assembled an owner/repository from memory, and the PR check then accepted that URL and wrote it into the task record, after which the model kept defending its own tool-endorsed guess over the worker's real link. Three changes close that chain without any live forge lookup, so private forges are treated exactly like public ones: - bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy or abstain" section requires a URL to be copied verbatim from a durable record (the done: PR <url> status line, pr= metadata, or the backlog note), forbids assembling owner, repository, host, or number from memory, and has the branch report only the identifier it actually holds when no record names the URL yet, leaving the PR check unarmed until the worker's ready line arrives. AGENTS.md section 7 and 9 carry the same copy-or-abstain rule for main in place of the bare full-URL mandate. - Worker briefs (bin/fm-brief.sh, ship and scout rules) require the full https:// URL wherever a PR is mentioned - status line, terminal, or summary - never a bare "PR 108", so the link is in view as early as the number is. - bin/fm-pr-check.sh refuses, offline and before any side effect, a URL that the task's own done lines contradict, printing both spellings; a log naming no URL still records the argument as before. fm_pr_status_ready_urls in bin/fm-pr-lib.sh owns reading those lines. The refusal also reaches bin/fm-pr-merge.sh, so nothing merges under a contradicted URL. Tests cover the offline refusal with zero side effects, the recorded spelling being accepted, markdown-wrapped and punctuated URLs, working lines not counting, the merge wrapper propagation, a self-hosted merge request with no forge call, the prompt carrying the rule, and the brief carrying the worker rule. * no-mistakes(review): Remove stale PR URL enforcement * no-mistakes(ci): Removed backlog notes as an accepted PR identity source. PR URLs may now be copied only from the task’s `done: PR <url>` status or canonical `pr=` metadata; otherwise supervision reports only the known identifier and leaves PR checking unarmed. Updated related guidance/docs and verified with branch-supervision tests, brief tests, ShellCheck, and `git diff --check`
friesentius
pushed a commit
to friesentius/firstmate
that referenced
this pull request
Sep 21, 2026
* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work MAIN answered a supervision-branch outcome for completed captain-requested work (implementation done, PR ready for review and merge approval) with "Captain, shipshape.", reading section 9's no-action reply as covering it and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no captain-facing response is owed". Section 9 now limits the shipshape reply to true no-ops (idle re-read, empty heartbeat, consequence-free acknowledgement) and requires a short outcome response naming what finished and what word is needed whenever requested work finishes or a result needs the captain's word, even when a transcript entry already shows the substance. The Pi protocol's re-emit rule now says it bounds repetition only, and carries a worked example of the ready-for-review outcome whose correct processing turn a shipshape reply fails. No executable contract evaluates the content of MAIN's captain-facing reply, so the regression is the protocol example in the owner doc rather than a text-match test. * no-mistakes(document): Clarify captain-facing outcomes versus no-ops * docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line The document step condensed the Pi protocol's re-emit rule and dropped the worked example of a finished, ready-for-review outcome whose correct processing turn a "Captain, shipshape." reply fails. That example is the contract's regression: no executable contract evaluates the content of MAIN's captain-facing reply, so the owner doc's example is the test case. Restore it directly under the re-emit rule, prefixed as a regression example that is kept verbatim and never condensed or summarized away. * no-mistakes(review): Clarify captain outcome and decision-word requirements * no-mistakes(document): Clarify captain-facing completion outcomes * docs(pi): require the PR URL in the visible captain-facing outcome reply Captain review on the regression example: drop the sample reply string and say only that the ready-for-review outcome requires relaying a captain-facing outcome response, not just "Captain, shipshape.". Fold in the visible-PR-handoff failure seen this session: after the branch outcome reporting this fix green, MAIN's visible reply was only "Awaiting your merge call." with no PR URL, leaning on the dim anchor. Section 9's URL rule now also covers a review or merge ask and names the visible reply as where the URL goes, sourced from the ready status, pr= metadata, or the supervision branch's summary and never left to a transcript entry. The Pi protocol adds the same-way failure and places the captain-facing text in the final visible assistant reply after the fm_branch_processed call, because Calm hides assistant text emitted in the same step as a tool call as a working note. Investigation verdict, evidence in the PR comment: no recent PR caused the handoff failure; Pi has hidden same-step pre-tool assistant text since kunchenguid#2339 (2026-08-13), kunchenguid#4655 changed only the Claude Code mod, and kunchenguid#4658 touched only remote report transfer. * no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs * no-mistakes(document): Clarify captain-facing supervision outcomes * docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule The consolidated section 9 URL rule narrowed its trigger to a review or merge ask, dropping the "whenever a PR is mentioned" catch-all from kunchenguid#3648 that keeps every PR URL copied from a durable record and never assembled from memory. Restore that trigger as a union with the review or merge ask so the one consolidated rule covers both.
friesentius
pushed a commit
to friesentius/firstmate
that referenced
this pull request
Sep 21, 2026
* fix: copy PR URLs from records or abstain, never assemble them Supervision reported a plausible but dead PR link three times because its prompt demanded a full https:// URL at a moment when only a PR number was observable, so the model assembled an owner/repository from memory, and the PR check then accepted that URL and wrote it into the task record, after which the model kept defending its own tool-endorsed guess over the worker's real link. Three changes close that chain without any live forge lookup, so private forges are treated exactly like public ones: - bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy or abstain" section requires a URL to be copied verbatim from a durable record (the done: PR <url> status line, pr= metadata, or the backlog note), forbids assembling owner, repository, host, or number from memory, and has the branch report only the identifier it actually holds when no record names the URL yet, leaving the PR check unarmed until the worker's ready line arrives. AGENTS.md section 7 and 9 carry the same copy-or-abstain rule for main in place of the bare full-URL mandate. - Worker briefs (bin/fm-brief.sh, ship and scout rules) require the full https:// URL wherever a PR is mentioned - status line, terminal, or summary - never a bare "PR 108", so the link is in view as early as the number is. - bin/fm-pr-check.sh refuses, offline and before any side effect, a URL that the task's own done lines contradict, printing both spellings; a log naming no URL still records the argument as before. fm_pr_status_ready_urls in bin/fm-pr-lib.sh owns reading those lines. The refusal also reaches bin/fm-pr-merge.sh, so nothing merges under a contradicted URL. Tests cover the offline refusal with zero side effects, the recorded spelling being accepted, markdown-wrapped and punctuated URLs, working lines not counting, the merge wrapper propagation, a self-hosted merge request with no forge call, the prompt carrying the rule, and the brief carrying the worker rule. * no-mistakes(review): Remove stale PR URL enforcement * no-mistakes(ci): Removed backlog notes as an accepted PR identity source. PR URLs may now be copied only from the task’s `done: PR <url>` status or canonical `pr=` metadata; otherwise supervision reports only the known identifier and leaves PR checking unarmed. Updated related guidance/docs and verified with branch-supervision tests, brief tests, ShellCheck, and `git diff --check`
friesentius
pushed a commit
to friesentius/firstmate
that referenced
this pull request
Sep 21, 2026
* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work MAIN answered a supervision-branch outcome for completed captain-requested work (implementation done, PR ready for review and merge approval) with "Captain, shipshape.", reading section 9's no-action reply as covering it and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no captain-facing response is owed". Section 9 now limits the shipshape reply to true no-ops (idle re-read, empty heartbeat, consequence-free acknowledgement) and requires a short outcome response naming what finished and what word is needed whenever requested work finishes or a result needs the captain's word, even when a transcript entry already shows the substance. The Pi protocol's re-emit rule now says it bounds repetition only, and carries a worked example of the ready-for-review outcome whose correct processing turn a shipshape reply fails. No executable contract evaluates the content of MAIN's captain-facing reply, so the regression is the protocol example in the owner doc rather than a text-match test. * no-mistakes(document): Clarify captain-facing outcomes versus no-ops * docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line The document step condensed the Pi protocol's re-emit rule and dropped the worked example of a finished, ready-for-review outcome whose correct processing turn a "Captain, shipshape." reply fails. That example is the contract's regression: no executable contract evaluates the content of MAIN's captain-facing reply, so the owner doc's example is the test case. Restore it directly under the re-emit rule, prefixed as a regression example that is kept verbatim and never condensed or summarized away. * no-mistakes(review): Clarify captain outcome and decision-word requirements * no-mistakes(document): Clarify captain-facing completion outcomes * docs(pi): require the PR URL in the visible captain-facing outcome reply Captain review on the regression example: drop the sample reply string and say only that the ready-for-review outcome requires relaying a captain-facing outcome response, not just "Captain, shipshape.". Fold in the visible-PR-handoff failure seen this session: after the branch outcome reporting this fix green, MAIN's visible reply was only "Awaiting your merge call." with no PR URL, leaning on the dim anchor. Section 9's URL rule now also covers a review or merge ask and names the visible reply as where the URL goes, sourced from the ready status, pr= metadata, or the supervision branch's summary and never left to a transcript entry. The Pi protocol adds the same-way failure and places the captain-facing text in the final visible assistant reply after the fm_branch_processed call, because Calm hides assistant text emitted in the same step as a tool call as a working note. Investigation verdict, evidence in the PR comment: no recent PR caused the handoff failure; Pi has hidden same-step pre-tool assistant text since kunchenguid#2339 (2026-08-13), kunchenguid#4655 changed only the Claude Code mod, and kunchenguid#4658 touched only remote report transfer. * no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs * no-mistakes(document): Clarify captain-facing supervision outcomes * docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule The consolidated section 9 URL rule narrowed its trigger to a review or merge ask, dropping the "whenever a PR is mentioned" catch-all from kunchenguid#3648 that keeps every PR URL copied from a durable record and never assembled from memory. Restore that trigger as a union with the review or merge ask so the one consolidated rule covers both.
friesentius
added a commit
to friesentius/firstmate
that referenced
this pull request
Sep 21, 2026
…g handling (#6) * feat: restart second mates after instruction updates (#3614) * feat(update): restart second mates whose instructions changed /updatefirstmate pulled new bytes onto disk and then asked each advanced second mate to re-read them. A running agent holds AGENTS.md and every loaded skill frozen from launch and no verified harness offers a reload, so that steer could not reach a loaded skill at all and left the mate holding two contradictory copies of its own job description. An eligible mate is now restarted instead, in the same home and endpoint, through the existing transactional relaunch. The restart is gated on the mate first writing down the open work it holds only in conversation - the open-record half of /stow, never its memory sweeps - so an unregistered captain call is flushed before the conversation is spent. Anything that leaves the reload unprovable falls back to the old re-read message and is reported as exactly that, never as a clean reload. Remote mates take the same path: fm-remote-secondmate-control.sh gains a relaunch verb whose host-local leg runs that same control plane, since the mate is an ordinary local secondmate from its host's point of view. The primary resolves the profile and passes it explicitly, because config/secondmate-harness is not inherited and the file on that host belongs to a different home. fm-update.sh now splits its advanced live mates into a restart set and a nudge residual, and both sets require a changed instruction surface, which also closes the over-nudge against the session-start sweep. Restart is stricter still: a bin/-only advance reloads itself on the next call, so it never costs a conversation. Colocated tests cover the gating, the persist-then-restart order, the task-subset persist request, each unsafe fallback, the remote hop, and the remote sync's new instruction-surface report. * no-mistakes(review): Fix restart correlation, concurrent waits, and lifecycle reporting * no-mistakes(review): Parallelize relaunches and classify replacement incarnations * no-mistakes(review): Gate restart actions on live agent state * no-mistakes(review): Handle failed restart workers without hanging * no-mistakes(review): Nudge legacy remotes and preserve persist recovery * no-mistakes(review): Document one-time secondmate restart rollout * no-mistakes(review): Honor arrived replies and refresh remote profiles * no-mistakes(review): Revert remote parent profile reconciliation * no-mistakes(review): Reset remote profile defaults and honor published results * no-mistakes(review): Preserve fallback nudges for unverifiable secondmates * no-mistakes(document): Document second-mate restart update flow * no-mistakes(lint): Fix ShellCheck warnings in restart scripts * perf: accelerate local validation with bounded concurrency (#3644) * perf(tests): route gate verification through the bounded concurrent runner Local validation was the pipeline's dominant cost: across 67 recorded no-mistakes agent sessions on this repo, 99.3% of command execution was `bash tests/*.test.sh`, run strictly one script at a time, and 2% of those calls were killed by an agent-guessed timeout and paid for twice. Three changes, each measured: - `.no-mistakes.yaml` pins `commands.test` to `bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner already owns changed-file selection, bounded concurrency, the refusal of unproven scripts, and a generous automatic per-script bound, so the gate's baseline is neither a serial chain nor a guessed timeout. It stays intent-targeted - the Test step still runs its evidence agent on top - and excludes the live-Herdr family the required Herdr lane owns. - `bin/fm-test-run.sh` gives a plain list of script paths the same bounded automatic scheduler and automatic bound that `--changed` gets. Naming several subjects is how a verification round asks for exactly those scripts. The curated selections are untouched: `--lane` still composes CI shards whose serial lane must stay serial, `--family` is what the required Herdr lane runs, and `--all` stays a deliberate complete regression. - `pr-forge` is admitted to the concurrent-safe family registry on two consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those, and records `secondmate` and `session-bootstrap` as refused with the exact script and reason each failed on, so the refusals are actionable rather than silent. Measured on this host, 0 failures on both sides: verification round, 4 scripts 448s chained -> 231s through the runner (-48%) pr-forge family 409.2s at 1 worker -> 237.9s at 4 (1.72x) watcher-wake-lock family 1311.1s at 1 worker -> 539.3s at 4 (2.43x) A fourth lever was implemented and then removed because the measurement refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made `fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s unchanged, back to back. Those sleeps are not overhead added to the clock - they are how a test waits for a subject moving on fm-watch.sh's own one-second cadence - so sampling less often only delays detection. It also broke `fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a settled condition. CONTRIBUTING.md records that result so the experiment is not repeated. * no-mistakes(review): Separate concurrent runs by isolation proof family * no-mistakes(review): Limit automatic timeouts to changed-file validation * no-mistakes(document): Clarify validation concurrency documentation * fix: copy PR URLs from durable records (#3648) * fix: copy PR URLs from records or abstain, never assemble them Supervision reported a plausible but dead PR link three times because its prompt demanded a full https:// URL at a moment when only a PR number was observable, so the model assembled an owner/repository from memory, and the PR check then accepted that URL and wrote it into the task record, after which the model kept defending its own tool-endorsed guess over the worker's real link. Three changes close that chain without any live forge lookup, so private forges are treated exactly like public ones: - bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy or abstain" section requires a URL to be copied verbatim from a durable record (the done: PR <url> status line, pr= metadata, or the backlog note), forbids assembling owner, repository, host, or number from memory, and has the branch report only the identifier it actually holds when no record names the URL yet, leaving the PR check unarmed until the worker's ready line arrives. AGENTS.md section 7 and 9 carry the same copy-or-abstain rule for main in place of the bare full-URL mandate. - Worker briefs (bin/fm-brief.sh, ship and scout rules) require the full https:// URL wherever a PR is mentioned - status line, terminal, or summary - never a bare "PR 108", so the link is in view as early as the number is. - bin/fm-pr-check.sh refuses, offline and before any side effect, a URL that the task's own done lines contradict, printing both spellings; a log naming no URL still records the argument as before. fm_pr_status_ready_urls in bin/fm-pr-lib.sh owns reading those lines. The refusal also reaches bin/fm-pr-merge.sh, so nothing merges under a contradicted URL. Tests cover the offline refusal with zero side effects, the recorded spelling being accepted, markdown-wrapped and punctuated URLs, working lines not counting, the merge wrapper propagation, a self-hosted merge request with no forge call, the prompt carrying the rule, and the brief carrying the worker rule. * no-mistakes(review): Remove stale PR URL enforcement * no-mistakes(ci): Removed backlog notes as an accepted PR identity source. PR URLs may now be copied only from the task’s `done: PR <url>` status or canonical `pr=` metadata; otherwise supervision reports only the known identifier and leaves PR checking unarmed. Updated related guidance/docs and verified with branch-supervision tests, brief tests, ShellCheck, and `git diff --check` * fix(bin): disable Claude feedback drafts for fleet launches (#3661) * fix(bin): disable Claude's feedback-draft flow for fleet-launched agents Scope --settings '{"feedbackDrafts":"off"}' to every Firstmate-launched Claude crewmate and secondmate, so /bug and /feedback never queue or submit a bug report on the captain's behalf. feedbackDrafts is the documented settings key (Claude Code changelog 2.1.247); the per-launch CLI flag never touches the captain's global settings.json. Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3 * no-mistakes(review): Prevent managed settings from re-enabling Claude feedback drafts * no-mistakes(document): Fix Claude feedback documentation formatting * fix(bin): layer both feedback-draft controls for defense in depth The prior --settings-only fix can be overridden by a managed Claude settings policy (feedbackDrafts precedence). Keep CLAUDE_CODE_SEND_FEEDBACK=0 alongside --settings '{"feedbackDrafts":"off"}': either control alone disables the SendFeedback tool, so a managed override of one still leaves the other in force. Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3 * no-mistakes(document): Document Claude feedback-draft suppression ownership * feat(tests): run three more validation families concurrently (#3662) * perf(tests): admit three more families to concurrent validation The three families that `docs/fm-test-isolation-proof.md` recorded as refused were not refused for concurrency. Each blocker was a test that decided a property by wall clock, or a script filed where it cannot run. Fixing those three things admits all three families and recovers 28.6 minutes of local validation with no assertion removed or weakened. - `tests/fm-backlog-handoff.test.sh` injected its pre-move crash by killing the handoff, sleeping a fixed second, then delegating the move to the real binary. Nothing ever killed the fake, so on a host slow enough for the case's next assertions to take longer than a second, the orphan woke and completed the very move the case requires left undone, and recovery then failed with `Task "pre-move-crash" not found in this backlog`. Watching the two backlogs during the injected crash showed exactly that, the item moving one second after the crash. All four crash injections in the file now go through a new `fm_fake_crash_injector` shim that signals the target and returns only once it is observably gone, and the pre-move fake never delegates the move at all. - `tests/fm-session-start.test.sh` proved the startup digest does not block on a slow current-state read by timing the whole digest against a fixed eight-second sleep, which a loaded host exceeds without the property being violated. It now holds that read open until the case releases it and asserts, the moment the digest returns, that the read has not finished. A digest that waited would wait indefinitely rather than for an interval a slow host can out-run, so the assertion is stronger than the bound it replaces. Its scan budget moves to the maximum, because the old value left two seconds of margin over the fixed sleep and measured the host rather than the deadline that `tests/fm-inactive-reconcile.test.sh` owns. - `fm-backend-herdr-focus-flash-e2e` was filed in the family map's catch-all, which put it in the portable serial lane, where Linux CI gate-skips it: that real-Herdr regression was running nowhere. It moves to `real-herdr-gated` and the required Herdr lane. `fm-claude-stop-autoarm-live-e2e` gate-skips on its opt-in variable and moves to `live-harness-optin`. The 28 remaining ungrouped scripts become an enumerated `standalone` family instead of admitting `unclassified` itself. `unclassified` is the family map's `*)` arm, so admitting it would silently grant concurrency to every test added afterwards, which is exactly the population with no proof. A new test still lands in `unclassified` and stays serial, and `tests/fm-test-run.test.sh` covers that split behaviorally. Each family passes two consecutive four-worker proofs with zero failures. On the production runner, `secondmate` goes 1233.1s to 453.4s, `session-bootstrap` 756.4s to 286.4s, and `standalone` 724.6s to 261.1s: 2.71x overall and 1713.2s recovered. The whole suite runs 177 scripts in 52.6 minutes of wall clock against 121 minutes of summed script time. * no-mistakes(document): Refresh concurrent validation and shard documentation * no-mistakes(ci): Fixed the real-Herdr focus-flash E2E race exposed by reclassification. Part C now starts its persistent child atomically via `pane run` and verifies stable child identity through Herdr’s public `process-info` interface, avoiding the racy send-text/send-keys sequence and platform-specific `ps` matching. Verified with bash syntax checking, ShellCheck, git diff checks, and the complete E2E test on Herdr 0.8.2 * feat: structure no-mistakes ask-user escalations (#3670) * feat(brief): structure no-mistakes ask-user escalation as event + snapshot file Crewmates escalating a no-mistakes ask-user gate now report one status event naming every finding id plus a snapshot file holding the gate's axi finding records verbatim (id, severity, file, line, description, authority), using the same shape even for a single finding. The status line never paraphrases. The format is defined once in fm-dod-lib.sh and rendered into both the scout and ship rule 6 in fm-brief.sh, so a promoted scout - whose rule 6 fm-promote.sh preserves unchanged - gets the identical contract as a freshly-spawned no-mistakes ship worker. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PpiWaDerbYavTLPPtEjQei * no-mistakes(review): Preserve ask-user escalation output contract * no-mistakes(review): Align escalation format test expectation * no-mistakes(review): Scope ask-user escalation instructions correctly * no-mistakes(review): Remove ask-user from generic decision rules --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> * fix(bin): require self-sufficient no-mistakes intent (#3671) * fix(bin): require a self-sufficient no-mistakes intent A no-mistakes worker's --intent is only as useful as the string it passes. PR #3604 shipped with an intent that was only "do 1, 2, 3, 7 from the report": the real contract lived in a private scout report and never reached --intent, so nobody holding that string plus the codebase could have derived the specification. This is pure instruction at the contract's one owner; no spawn-side or promotion-side check is added. - bin/fm-dod-lib.sh: the generated no-mistakes Definition of done now states that the --intent string must be self-sufficient (the string plus the codebase reconstructs roughly the same specification) and tells the worker to write the substance of any report, decision, or PR the captain's intent refers to into --intent rather than the pointer, while Firstmate build instructions and the worker's own decisions still stay out. The spawn-time overlay points back at that rule so its "supersedes" wording cannot cancel it, and the header's owner statement carries the rule. - AGENTS.md section 11 and bin/fm-brief.sh's header ask Firstmate to include the substance of referenced material when filling ## Captain's intent, and section 11 points at the owner of the rule. - tests/fm-brief.test.sh and tests/fm-task-delivery.test.sh assert the rendered brief and launch contract carry the rule. Claude-Session: https://claude.ai/code/session_01YMhEe42q7BAAoN6RxNuzim * no-mistakes(document): Replace incident-specific intent test commentary * fix: accelerate local Bearings snapshot composition (#3499) * Speed local fleet snapshot composition * no-mistakes(review): Stabilize task inventory during concurrent snapshot composition * no-mistakes(document): Document local snapshot observation concurrency * no-mistakes(ci): Fixed CI failures by making empty task manifests compatible with stock macOS Bash 3.2, snapshotting task metadata before concurrent observations to prevent generation drift, strengthening the behavioral race regression, and updating the stock-Bash Bearings test count to 45. Verified fleet snapshot tests (15), Bearings tests (45), workflow lint tests, project lint, Bash 3.2 parsing, and diff checks * no-mistakes(ci): Fixed the Linux CI failure caused by passing large backlog/task JSON through jq command-line arguments, which exceeded the per-argument size limit. Both inventory projections now stream large JSON inputs through stdin. Verified with fm-bearings-snapshot.test.sh (45 tests), fm-fleet-snapshot-view.test.sh (15 tests), Bash syntax, and git diff checks * no-mistakes(ci): Fixed concurrent task teardown during metadata capture: vanished metadata is now omitted while genuine copy failures remain fatal. Added a deterministic public Bearings regression test and updated CI’s expected test count. Verified with the full Bearings suite, workflow-lint suite, Bash syntax checks, and git diff checks * no-mistakes(ci): Fixed PR-caused CI and review issues: streamed large fleet JSON through jq stdin to avoid Linux argument limits, kept crew-state reads bound to captured metadata generations, and strengthened the behavioral race test. Bearings (46 tests), fleet snapshot (15 tests), crew-state, backend, lint, Bash syntax, and diff checks pass locally. Serial shard 5’s unrelated task-inbox segmentation fault appears infrastructural/flaky * no-mistakes(ci): Fixed endpoint-state generation crossing by validating captured spawn_gen before and after local endpoint probes, falling back to exact metadata identity for legacy tasks. Stale probe results now become unknown instead of false unhealthy state. Added a behavioral relaunch-race regression test. Verified the full Bearings snapshot suite, shellcheck, bash syntax, and git diff checks * fix(snapshot): keep live observations generation-coherent * no-mistakes(review): Keep secondmate observations generation-bound without copying reports * no-mistakes(document): Document generation-coherent snapshot observations * test(bearings): measure local read overlap instead of wall-clock budget The large-local-snapshot regression asserted that a whole snapshot composed in under five seconds. That bound measures how loaded the host is, not whether the per-task reads actually overlap, so it failed intermittently on a contended machine: one run in six on a box at load 16-20, landing exactly on the five second boundary. Time a serialized run and a concurrent run of the same workload instead and require the concurrent one to save at least two seconds. Both runs pay the same composition overhead, so the difference isolates the overlap this change delivers. Five one-second reads serialize into five seconds and overlap into about one, and re-serializing the reads collapses the saving to roughly zero, so the assertion still fails loudly if the concurrency regresses. Also bump the pinned Bearings test count to 48, since rebasing onto the current default branch picked up its captain-hold test. * no-mistakes(review): Restore JSON-derived decision flags * no-mistakes(review): Unify status-derived snapshot observations * no-mistakes(ci): Updated the stock macOS Bash CI check’s Bearings test count from 48 to 49. Verified the full Bearings suite passes and emits exactly 49 TAP successes; git diff checks pass * fix: prevent stale supervision wake loops (#3672) * fix(bin): stop the supervision branch's stale-ack and ghost-report loops Clean-slate implementation of the four authorized recommendations from the supervision-ghost-retrigger analysis (items 1, 2, 3, and 7), in their minimal form, superseding PR #3604: - fm_branch_report refuses a task the wake being handled never named. The extension fixes the reportable task set from the eligible rows before each prompt (signal and stale rows resolve to their tasks, a heartbeat allows any task with a live record, fleet is always allowed), so a report typed from memory about a task whose records teardown already removed is never stored or delivered. - An acknowledgement that consumes nothing says "nothing was acknowledged through N" and prints the exact --ack-through / --recovery-generation command for the current presented wake, instead of "re-run the drain", which re-fed the same stale acknowledgement in a loop. - bin/fm-guard.sh no longer tells the branch actor to drain queued wakes while it is handling them; it names the granted rows instead. - Teardown removes state/.<task>.branch-outcome-index for ordinary tasks and descendants; the index rebuild and the append-side index write both skip a task with neither a live record nor a status log, so the branch's report of a teardown it just performed is stored without recreating the index. No new locking, no spawn-generation binding, and no retired-task refusal: the branch can still report the outcome of a task it just tore down, and the teardown test now proves that path end to end. * fix(bin): narrow the branch report scope and guard silence to the minimal form Apply the four review decisions on the clean-slate branch: - A signal or stale prompt may report only the tasks its own rows resolve to; fleet is refused there too. A heartbeat review is not scoped by task at all, so the extension no longer tracks live task records and refuses nothing by task id during a fleet review. - The outcome-index rebuild no longer skips retired tasks; the append-side skip alone keeps a torn-down task's index from being recreated. - bin/fm-guard.sh keeps the queued-wakes warning silent for the branch actor instead of printing a replacement note. * no-mistakes(document): Align supervision docs with scoped wake handling * fix(bin): avoid fleet snapshot argument limits (#3677) * Fix fleet snapshot large JSON transport * no-mistakes(review): Captain: file-back fleet snapshot transport safely * no-mistakes(review): Captain: file-back parent summary aggregation * no-mistakes(ci): Rebased the PR's three commits onto f4d7875824ecc5e274b4bb896f10c1e1f207b7e4 and resolved the fleet snapshot conflict while preserving the base's task-observation lifecycle. Fixed Greptile's valid finding by recursively removing the private mktemp transport directory, so future transport files cannot cause cleanup to fail. Verified with tests/fm-home-summary-refresh.test.sh, bin/fm-lint.sh, git diff --check, and ancestry checks. All passed; the fix remains as an uncommitted worktree change for the outer executor * fix(bin): attribute active runs with unfetched pipeline heads (#3681) * fix(bin): recognize active pipeline fix rounds with unfetched run heads A no-mistakes fix round advances the run head beyond the submitted head, and the pipeline commits in its own checkout, so the task copy never receives the new commit object. fm-crew-state's strict head rule rejected the active row, the coarse runs-list scan skipped it and matched the older failed row at the submitted head, and an active validation read as failed (observed on model-routing-benchmark-hardening: active head ac61c64b vs task copy at fb47636d). fm_nm_runs_status_for_worktree in bin/fm-nm-run-lib.sh now owns runs-ledger attribution: the branch's newest row alone decides, and a newest row whose head cannot resolve locally is recognized only as a provable pipeline-owned continuation - active (running) and anchored by the immediately older row for the same branch having ended at exactly this worktree's HEAD. The reader keeps the axi TOON as full detail for that proven same-branch run. Unanchored, ancestor-anchored, and terminal unresolvable rows stay unattributed, so branch-name coincidence and other tasks' runs never match, and fm_nm_head_matches_worktree keeps its exact prior semantics for teardown (verified by the full teardown suite). Tests: reproduction regression for the unfetched active fix head (reads working via full run-step detail), coarse-path continuation when axi answers another branch, and negative controls for the unanchored active row and the unresolvable terminal row with the historical fallback preserved. Ported onto upstream/main f4d78758, where #3194 independently added the branch_sync custody exemption on the full axi-status path: both mechanisms now coexist, each owning one surface (TOON custody on the full path, the runs ledger on the coarse path). The port deletes the superseded coarse scan-and-skip (nm_runs_status_for_branch) and its now caller-less helpers (fm_nm_head_resolvable, nm_coarse_head_matches_worktree), renames the exemption comment's "the one exemption" phrasing now that a second complementary exemption exists, and points the stale FM_CREW_STATE_RUNS_LIMIT comment at fm_nm_runs_status_for_worktree (judge follow-up #1). The parent coarse-guard test's fixture is the ledger-anchored continuation shape, so its expectation flips to the fixed behavior (working via run-step, never the older failed row); a new mismatched-anchor coarse negative control preserves that guard's original no-anchor protection (pane answers, never the older row). * no-mistakes(document): Clarify pipeline attribution documentation * fix(bin): pre-register claude workspace trust at spawn time (#3663) * fix(bin): pre-register claude workspace trust for task worktrees A claude crewmate launched into a fresh task worktree met Claude Code's interactive workspace-trust dialog before it ever read its brief, and firstmate could not answer it: the key plane carries only Enter, Escape, and C-c with no arrow navigation, and the dialog's selection starts on "No, exit", so the documented Enter recipe ended the session instead of accepting it. Two workers wedged this way and were unblocked only by hand-seeding the trust store per path. --dangerously-skip-permissions does not cover that gate. `claude --help` records the dialog as skipped only in non-interactive mode, through -p or a non-TTY stdout, and a crewmate pane is interactive, so there is no launch flag to reach for. fm-spawn now pre-registers the worktree through bin/fm-claude-trust.sh in the existing claude branch, before the project settings that the same gate would otherwise block, and refuses the spawn when that write fails rather than launching a worker that would wedge. The scope test is the safety property and is structural rather than a path policy: the path must be a linked git worktree, sharing the spawning project's common dir, whose top level is exactly the resolved argument. Git is the ground truth, so the argument is never trusted on its own word, and a primary checkout, an unrelated repo, a worktree subdirectory, a plain directory, and a home directory are each refused rather than warned about or skipped. A treehouse or orca path prefix was deliberately avoided because treehouse's root is configurable, which would make a prefix both wrong and a new policy surface. One structural test covers both worktree providers. tests/fm-claude-trust.test.sh pins both halves, including a case where HOME is itself a valid linked worktree so the home guard is proven load-bearing rather than passing vacuously, plus the spawn-level proof that a claude spawn trusts its worktree and launches with the brief pointed at the same store. The adapter reference no longer tells a firstmate to press Enter on that dialog, and the shared trust reference now names every harness surface: which harnesses gate, which suppress at launch, which dodge the gate, which now pre-registers, and that a claude secondmate is excluded by design. The spawn fixture runs each spawn against a throwaway HOME so the suite cannot write the developer's real store, isolating through HOME rather than CLAUDE_CONFIG_DIR because the spawn forwards a set CLAUDE_CONFIG_DIR onto the launch command that launch-shape assertions read. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HNEN2GLnew27HFyfi4ms4v * fix(bin): create the staged trust store exclusively The staged store was written to a predictable pid-based path with a plain write, which follows a symlink. Where the Claude config directory is writable by another local account, that account could pre-create the path as a symlink and redirect the write into another file the launching user owns. The staged name now carries random bytes and is created with an exclusive "wx" open, so an existing path is refused outright instead of followed. The happy-path test also asserts no staged store survives the rename. The durability comment now states the residual window plainly: the readback proves the entry landed, not that it survives, because a vendor session that rewrites the whole store afterwards can still drop it and no lock closes that window when the writer is Claude itself. The worker then meets the dialog and stalls, which reaches firstmate as the ordinary stale wake rather than as silent success. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HNEN2GLnew27HFyfi4ms4v * no-mistakes(review): neutralise CDPATH in claude trust scope guard * no-mistakes(review): sandbox HOME in spawn tests, drop out-of-scope artifacts * no-mistakes(review): refuse unresolvable git dir, compact store, fix secondmate doc * no-mistakes(review): clear git env overrides, resolve symlinked store target * no-mistakes(review): degrade without node, fix Pi gate claim, record trust proof * no-mistakes(review): refuse without node, pin CLAUDE_CONFIG_DIR in spawn tests * no-mistakes(review): refuse relative config dir and concurrent store modification * no-mistakes(review): correct orca worktree claim, clean staged store on failure * no-mistakes(review): restore pretty-printed store, correct trust dialog docs * no-mistakes(review): arm trust gate before busy state to avoid orphans * no-mistakes(document): record claude trust pre-registration in its owner docs * no-mistakes(document): note orca limit for claude trust pre-registration * no-mistakes(ci): Fixed the Greptile P1 on bin/fm-spawn.sh by moving the Claude trust gate earlier rather than adding cleanup machinery. Diagnosis: Greptile reported that when Claude trust registration fails on tmux/Zellij/cmux/non-projected Herdr, the exit runs after the backend endpoint and /tmp/fm-<id> were created, and the abort trap cleans neither. The endpoint half is pre-existing, deliberate architecture — the two refusals immediately above the gate (the 60s `treehouse get` timeout at fm-spawn.sh:2550 and `validate_spawn_worktree` at :2487) also exit with the endpoint live and direct the operator with "inspect window $T"; spawn_abort_cleanup only reclaims orca endpoints (already covered via ORCA_ABORT_CLEANUP) and herdr projections. The temp-root half was genuinely introduced by this PR: the gate was placed beside the busy-state arm, ~30 lines after `mkdir -p "$TASK_TMP/gotmp"`, and fm-teardown can only find that root through `tasktmp=` in a meta record a refused spawn never publishes. Root-cause fix (smallest correct change, no new subsystem): - bin/fm-spawn.sh — moved the `claude*` trust gate from inside the busy-arm block up to the first point $WT is known, immediately after the `freshen_spawn_worktree_base` block and before TASK_TMP creation, the STATE setup, and the relaunch `clear_relaunch_harness_wiring` retirement. A refusal now leaves no temp root, no retired relaunch wiring, and no busy record; only the endpoint remains, in the same class as the two refusals just above it. - bin/fm-spawn.sh — the refusal message now ends with "inspect window $T", matching the existing convention so control/teardown can identify the endpoint. $T is set for every backend on the non-secondmate path. - bin/fm-spawn.sh:196 — header note corrected from "before any state is armed" to "before any per-task state exists". - tests/fm-claude-trust.test.sh — the existing refused-spawn test's own comment claimed "before any task state exists" but only asserted busy state. Renamed to test_refused_spawn_leaves_no_task_state and added an assertion that /tmp/fm-<id> is absent, with the task id suffixed by the test process pid so the assertion reads only this run's path (a stale /tmp/fm-refusedspawn from the fixed-id version was in fact present on this box). No assertions on implementation source bytes. Verification run locally: - The new assertion fails against the pre-fix bin/fm-spawn.sh ("not ok - a refused spawn stranded a temp root no teardown can find") and passes after — a real before/after regression proof. - tests/fm-claude-trust.test.sh: 20/20 ok. - tests/fm-backend.test.sh, fm-backend-orca, fm-control-relaunch, fm-spawn-dispatch-profile, fm-trace-context-spawn, fm-gotmp: all pass. - tests/fm-backlog-atomicity.test.sh: rc=0, 79 assertions ok. - bin/fm-lint.sh (repo's single lint owner, pinned ShellCheck 0.11.0 + actionlint 1.7.12): clean. - No /tmp/fm-refusedspawn* leftovers after the runs. Scope respected: no trust subsystem, no policy layer, no config surface, no endpoint-cleanup mechanism added; the change is an ordering move plus one error-message clause and the test that pins it. Adapter references and docs made no ordering claim, so none needed updating. Changes are left uncommitted in the worktree for the outer executor --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * fix: restart every live second mate after updates (#3690) * feat(update): restart every live second mate after a successful update /updatefirstmate only restarted a second mate when that pass advanced its AGENTS.md or .agents/skills. An already-current home was skipped entirely, a bin/-only advance was steered instead, and a remote host that could not report its instruction diff was downgraded to a re-read. A running agent also freezes its launch-time wiring - turn-end hooks, harness flags, per-harness feature switches - and none of that is derivable from a file diff, so an unchanged tracked surface is not evidence the agent is already on the current behavior. Restart is now unconditional on a successful update of that home. Every live second mate the pass leaves on the target commit is restarted, whether it advanced or was already there. The safety contract is unchanged: open records are persisted before the agent is replaced, nothing is forced, stashed, or discarded, a home the pass had to skip is not restarted at all, and a mate whose runtime cannot prove a restart keeps the honest re-read path and is never reported as reloaded. bin/fm-ff-lib.sh gains a settled-state hook that fires for a home left at the base whether it advanced or was already there, and never for a skipped one; the instruction-gated hook the session-start convergence sweep uses is untouched. Regressions: fm-update pins the already-current mate into the restart set and the unprovable one into the nudge set, and fm-secondmate-restart drives both real commands end to end - an already-current home is named, persisted, and genuinely replaced with its checkout untouched, while the unprovable one keeps its running agent. * no-mistakes(document): Document unconditional secondmate restarts * fix(bin): close pending-reply decisions via resolve-key (#3696) * fix(bin): close reserved pending-reply keys via fm-send --resolve-key fm-send wrote answered: notes that the reserved-key fold ignores, so operator closes exited 0 while OPEN DECISIONS kept the decision open. Speak the owning library's close vocabulary on that path, and refuse when a reserved close cannot take effect. * no-mistakes(review): Safely quote manual decision-close recovery commands * no-mistakes(review): Reject unclosable overlong decision keys before sending * no-mistakes(review): Remove contract suffix from open decisions hint * no-mistakes(document): Document resolve-key line-cap refusal * fix(bin): prevent false missed-reply escalations (#3697) * fix(bin): stop false missed-reply escalations for same-basename self-home answers A healthy secondmate that wrote corr= to its own state/<id>.status never matched the parent channel, so recovery confirmed and the record escalated as pending-reply-missed. Make the report helper resolve the parent channel itself, skip parent-replies.status as wrong-home, put a readable sighting path on the missed line, and restatement-copy only that same-basename self-home file onto the parent channel. * no-mistakes(review): Resolve late replies before recovery escalation * no-mistakes(review): Tighten reply routing and regression coverage * no-mistakes(review): Preserve reply paths and require explicit home * no-mistakes(review): Encode wrong-home paths before persistence * no-mistakes(document): Document corrected secondmate reply routing * no-mistakes(lint): Fix pending-reply ShellCheck warnings * feat: add verified Gemini crewmate runtime (#3695) * feat(harness): verify gemini as a crewmate runtime adapter Adds Gemini CLI as a fourth dispatch target alongside claude, codex, and grok, scoped to crewmate and scout work only. Every axis was proven against gemini-cli 0.58.0 rather than inferred; docs/verification/runtime-backends.md carries the dated evidence and names what stayed unverified. Busy state is semantic, not rendered: BeforeAgent opens a turn and AfterAgent and SessionEnd close it. AfterAgent also fires on a manual interrupt, so a cancelled turn closes its own record. Three findings shaped the wiring rather than a config line: - --skip-trust and GEMINI_CLI_TRUST_WORKSPACE=true are presented by the CLI as equivalents and are not. A controlled A/B showed --skip-trust leaves project configuration unloaded, so workspace skills never load. - The worktree's .gemini/settings.json is the PROJECT's committed settings file, unlike claude's settings.local.json. Firstmate's hooks therefore go to a firstmate-owned state/<id>.gemini-settings.json reached through GEMINI_CLI_SYSTEM_SETTINGS_PATH, which also works untrusted and merges with a project's own hooks instead of replacing them. - The shipped CLI is a node bundle whose live process reports comm as MainThread, so ancestry cannot see it. GEMINI_CLI=1 is load-bearing and is tested before an inherited CLAUDECODE, and pane liveness identifies gemini from the script argument through the new bin/fm-gemini-lib.sh. Gemini is refused for secondmates: it has no primary supervision protocol. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L * test: clear gemini's marker in launch and detection expectations Every non-gemini launch now clears GEMINI_CLI the way it already clears cursor's markers, so the two tests that pin the exact launch prefix are updated to match. The harness-detection tests that scrub foreign markers before probing ancestry scrub GEMINI_CLI too, so running the suite from inside a gemini session cannot produce a false verdict. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L * docs: classify the gemini harness reference The documentation inventory is the single classification owner for maintained prose surfaces, and every surface must appear in it exactly once. The new harness reference is agent-runtime, matching its siblings. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L * no-mistakes(review): Narrow Gemini ancestry detection * no-mistakes(review): Restrict Gemini hooks to canonical launches * no-mistakes(document): Document Gemini adapter support boundaries * no-mistakes(ci): Fixed Gemini process identity when interpreter or script paths contain whitespace. Tmux liveness now uses NUL-delimited /proc argv on Linux, with the existing flattened ps fallback elsewhere. Added a real-process regression test. Verified with the Gemini harness test suite, full fm-lint, ShellCheck, and git diff --check. The CI and Require no-mistakes runs were action_required/attestation outcomes rather than code failures --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(teardown): conclude parked runs advanced past task copy (#3704) * conclude parked runs the pipeline advanced past the task copy A no-mistakes fix round commits in the daemon's own gate-repo clone, so a run parked at a gate can carry a head whose object the task copy never received. Teardown's strict object-local identity rule then declined to conclude the run, and cleanup left it parked forever holding a fleet slot (observed 2026-09-03; the same masking condition PR 3681 fixed on the read path, now closing the teardown half its scope boundary deferred). task_status_is_own_parked_run now falls back - only when the reported head resolves to no local object - to the one shared runs-ledger attribution rule fm_nm_runs_status_for_worktree (bin/fm-nm-run-lib.sh), whose anchored continuation proof binds the branch's newest active row to this worktree's exact submitted head. Foreign branches, stale history, terminal rows, ancestor-only anchors, diverged newer rows, and ambiguous multi-row shapes all still refuse, and runs that are actively running, fixing, or in CI remain untouched: only the parked-at-a-gate determination ever reaches the abort. No sqlite access, no fetches into another task copy, no custody changes, no duplicated matching logic. * tighten the parked-run ledger fallback and pin both judge corrections The teardown ledger fallback now authorizes concluding this task's parked run only when the shared runs-ledger rule's proved answer is the explicitly active word (running): a terminal newest row - even anchored at exactly the worktree's head - is finished history and never an abort authorization. The read path may classify the same owner's answer; teardown's abort must never fire for a run that already ended. Two bounded pre-validation corrections from the implementation review: - a fetched-object counterfactual pins the strict-rule path: a pipeline fix head fetched into the task copy aborts through object-local identity alone, with an empty ledger and a proof the runs query never fired; - a negative fixture pins the tightened boundary: an unresolvable reported head with a terminal newest same-branch row anchored at the worktree head engages the ledger fallback and still refuses, so the refusal is the terminal-word boundary and not an earlier guard. * no-mistakes(review): Bind teardown ledger fallback to validated run heads * no-mistakes(review): Restore validated advanced-head ledger continuation * no-mistakes(review): Reject invalid ledger dates and terminal statuses * no-mistakes(document): Document teardown ledger scan limit * fix: classify captain holds from structured state (#3508) * Keep parked and aged undated captain holds off live Captain's Call. Bearings was treating undated parked-style holds as live calls; mark those phrasings deferred and project holds older than a configurable 14-day since date as Charted Next gates instead. * no-mistakes(review): Bound parked marker matching to lexical tokens * no-mistakes(review): Age undated holds from durable hold-set dates * no-mistakes(review): Reset re-held timestamps and scan full bodies * no-mistakes(review): Preserve timestamp precision and prioritize parked suppression * no-mistakes(document): Document undated captain-hold aging * no-mistakes(ci): Fixed stock Bash CI test-count expectations (16 snapshot, 45 Bearings). Prevented fresh holds on old tasks from aging via stale `since` dates by aging only stamped holds. Added behavioral regressions and verified both suites plus Bash 3.2 parsing * no-mistakes(ci): account for rebased snapshot regression * no-mistakes(review): Restore legacy hold aging and mandate wrapper * no-mistakes(review): Restrict hold stamps to canonical leading lines * no-mistakes(review): Exclude historical answers and deduplicate revealed holds * no-mistakes(document): Correct captain-hold projection documentation * no-mistakes(ci): Rebased onto 8988af2 and resolved Bearings conflicts. Fixed the hold timestamp race by persisting and verifying the timestamp before publishing the captain hold; failures now leave the task unheld. Added behavioral coverage for ordering and failure handling. Preserved the required parked-phrase projection behavior. Relevant snapshot, Bearings, lifecycle, syntax, and ShellCheck validations pass * no-mistakes(review): Bound current prose before historical resolutions * no-mistakes(review): Preserve hold age across interrupted answers * no-mistakes(review): Preserve leading hold stamps until answer closure * no-mistakes(review): Normalize answer bodies on matching retries * no-mistakes(review): Document concurrent re-hold age-basis limitation * no-mistakes(document): Refresh captain hold lifecycle documentation * no-mistakes(ci): Fixed both CI failures. Updated the macOS Bash snapshot expectation from 45 to 46 Bearings tests. Narrowed parked-style deferral matching to explicit hold-reason prefixes while preserving legacy explicit markers and preventing contextual prose from hiding active decisions. Added behavioral regression coverage. Verified with stock Bash 3.2: 17 fleet snapshot tests and 46 Bearings tests pass; full lint and workflow validation also pass * no-mistakes(ci): Fixed Greptile’s P1 finding by restricting parked-style deferral phrases to complete hold-reason markers. Contextual reasons beginning with “not urgent,” “queued opportunity,” or “captain-gated” now remain visible decisions. Added behavioral coverage through the real fleet and Bearings snapshot paths and updated documentation. Verified both snapshot suites under Bash 3.2 (17 fleet tests and 46 Bearings tests), syntax checks, and git diff checks. The no-mistakes attestation failure is external/stale and requires the outer pipeline to refresh it for the new head * no-mistakes(ci): Fixed parked-style undated captain holds disappearing from the default Bearings board. They now project to Charted Next with omitted[] disclosure, while --all-decisions reveals them and removes the safety gate. Added behavioral coverage for the reported “not urgent” case and aligned documentation. Verified fm-bearings-snapshot, fleet snapshot view, and captain-hold lifecycle tests; shellcheck, bash syntax, and git diff checks pass * no-mistakes(test): Stabilize concurrency budget and provision timeout tests * no-mistakes(document): Correct captain hold documentation details * no-mistakes(ci): Fixed hold-reason parsing so commas in contextual reasons are preserved and do not incorrectly defer live Captain's Call decisions. Added end-to-end fleet/Bearings regression coverage. Reworked the flaky Herdr timeout test to assert observable late-launch behavior rather than process-ID liveness. Verified both snapshot suites, Herdr test 5 consecutive times, shell syntax, shellcheck, and git diff checks * Restore the Herdr lab timeout test to its main version. The stabilization rounds reworked tests/fm-herdr-lab.test.sh while chasing a load-induced flake, replacing the fake server's wall-clock delay with a SIGSTOP'd process and asserting that the blocked process is gone after a timed-out provision. A stopped process does not die from SIGTERM, so that assertion fails on Linux and the portable parallel shard stayed red. That test is unrelated to the undated captain-hold projection this branch delivers and was identical to main before these rounds, so restore main's version exactly. It still proves that a timed-out provision cancels its late launch before teardown. * no-mistakes(review): Preserve metadata-like prose in captain hold reasons * no-mistakes(review): Resurface due dated captain holds * no-mistakes(review): Distinguish parked holds from explicit deferrals * no-mistakes(review): Invalidate legacy secondmate summary caches * no-mistakes(review): Keep blocked deferred holds in Charted Next * no-mistakes(review): Count blocked deferred holds in omission disclosure * no-mistakes(document): Correct captain-hold projection documentation * no-mistakes(ci): Fixed both CI failures. Updated the macOS Bearings test count to 51. Preserved the v1 summary schema for compatibility while rejecting hold-bearing summaries missing the new aging fields, preventing stale caches from restoring noisy calls. Verified fleet snapshot, Bearings snapshot (51 tests), home-summary refresh, secondmate reconciliation, Bash 3.2 parsing, and diff checks * no-mistakes(ci): Fixed Greptile’s valid finding: `--all-decisions` now reveals deferred/aged captain holds even when blocked, for both main and secondmate homes, and removes their duplicate Charted Next gates. Added behavioral regression coverage and updated documentation. The prose-classifier finding was not applied because exact complete-phrase matching is explicitly required by the author intent; contextual wording remains live. Verified with Bearings and fleet snapshot tests, `bin/fm-lint.sh`, Bash syntax checking, and `git diff --check` * no-mistakes(ci): Fixed the actionable-state bug in Bearings: an arrived parked-style hold is live only when it is not explicitly non-actionable, so blocked due holds remain gated by default and are revealed by --all-decisions. Added behavioral regression coverage for that case. Preserved complete-reason parked-style classification as required by the author intent. Verified with tests/fm-bearings-snapshot.test.sh, bin/fm-lint.sh, and git diff --check * Show why a revealed captain hold is deferred. Under --all-decisions a deferred hold is revealed and its Charted Next gate is removed, but the revealed row carried only the bare hold reason. A date-deferred or blocked hold therefore read exactly like a genuine live decision, because the until date, the age, and the blocking work only ever appeared on the gate row that the reveal replaces. Annotate a row that is revealed because it is deferred with the same vocabulary the gate uses - until <date>, held <n>d, and the blocking work - so the expanded view reads as deferred-but-shown. A genuinely live call is left unannotated, and the default board is unchanged. * Classify captain holds from structured fields alone. Bucket membership was decided by several independent expressions, and two of them matched hold reason or body prose. That produced a recurring class of defects: holds that fell through every bucket and vanished from the board, and live decisions silently suppressed because their wording happened to contain a marker word - a reason of "non-deferred release choice" matched DEFERRED and disappeared. Replace all of it with one total classifier over structured fields only: hold_kind, state, hold_until, unresolved_blocker_ids, and the machine-written hold-set timestamp. Every captain hold gets exactly one hold_bucket - blocked, dated, aged, or live - so no hold can fall through and none can match two. captain_actionable is exactly the live bucket, and the --all-decisions reveal is a property of the bucket rather than a second filter. No hold reason or body prose is matched anywhere in the projection, so wording can no longer hide, reveal, or reclassify a decision. A hold that is superseded or no longer required is closed through the hold lifecycle instead of lingering as an open hold flagged by a keyword. * no-mistakes(review): Preserve working captain holds across bucket surfaces * no-mistakes(review): Reject pre-classifier secondmate summary caches * no-mistakes(review): Preserve complete live hold summaries * no-mistakes(review): Clarify working hold decision bucket semantics * no-mistakes(review): Reveal bounded remote holds and preserve blocker notes * no-mistakes(review): Make blocker overflow explicit in hold summaries * no-mistakes(document): Correct captain-hold projection documentation * no-mistakes(ci): Updated the stock macOS Bash CI snapshot expectation from 17 to 18 tests. Verified the suite under Bash 3.2.57: all 18 tests pass. `git diff --check` also passes * no-mistakes(ci): Updated the stock macOS Bash CI expectation from 51 to 53 Bearings tests. Verified all 53 pass under Bash 3.2.57; git diff --check passes * fix(pi): keep supervision outcome delivery responsive (#3767) * fix(pi): deliver supervision outcomes off Pi's render thread The supervision branch runs inside the captain's own Pi process, and Pi runs extensions, their tools, and their event handlers on the single JavaScript thread that also draws the TUI and reads the keyboard. Every delivered outcome ran roughly five bash script invocations plus several `ps` calls through spawnSync on that thread, so the TUI could not repaint or echo a keystroke for the whole chain - the subsecond freeze the captain saw every time a routine or captain-facing outcome arrived. Convert the delivery path's subprocess calls to an awaited spawn behind a serializing queue. lib/fm-async-exec.ts is the single owner of the awaited-spawn replacement and returns the same capture shape and failure verdicts spawnSync returned. Awaiting yields the thread, so what the single thread used to guarantee for free is now an explicit queue: every delivery, acknowledgement, and turn-boundary reconciliation runs as one unit of it, preserving the durable append before anything visible, one delivery at a time in sequence order, the read cursor advanced before the next reader sees a row, and one ownership activation per generation. Cancellation is preserved by the generation and lock-ownership rechecks the awaits are placed around. Two reads stay synchronous because Pi's own API is synchronous there, not as an optimization: its bash spawn hook is typed as a plain function, and the watcher reads offer.accepted the moment its dispatch event returns, so a session that does not own the fleet lock must still refuse a wake without waiting. Both walk the lock's process ancestry in full every time, never cached, because reparenting and pid reuse can invalidate a remembered chain and that answer decides ownership rather than hinting at it. The store scripts and their durability contracts are unchanged. Measured through the real fm_branch_report tool and real bin/ scripts with a 1 ms interval timer, the largest block of the JS thread falls from 273 to 2.0 ms for a routine outcome, 286 to 2.0 ms for a captain outcome, and 134 to 1.9 ms for main's acknowledgement, against a 1.3-2.2 ms idle floor. In a real Pi 0.82.0 TUI the worst keystroke echo while two outcomes arrive falls from 676.9 ms to 36.8 ms, against a 22.6 ms extension-free floor. Regressions: a delivery must leave the event loop running (zero timer ticks before this change, in 250 ms), interleaved reports stay ordered and exactly once, a session replaced mid-delivery neither loses nor duplicates an outcome, and a failing store script surfaces without losing or doubling one. The real-TUI half is an opt-in live guard that types into an isolated Pi pane while outcomes are delivered and fails if echo leaves the class of the same machine's own floor. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013bzoWyr2EcJGBKuoUjVRSp * no-mistakes(review): Revalidate ownership and bound asynchronous subprocess output * no-mistakes(ci): Fixed CI defects: routine outcomes now persist a sequence-keyed delivery receipt before awaiting cursor advancement, preventing duplicate delivery after mark-read failure. Corrected the session-replacement test to exercise an actual asynchronous ps ancestry lookup. Targeted behavioral tests, strict Pi typecheck, ShellCheck, and diff checks pass. The full extension test remains locally blocked by an unrelated stock-render assertion under the installed Pi runtime. The no-mistakes attestation failure is external pipeline state (test was previously skipped), not a source defect * fix(pi): keep the declined routine receipt out and skip the renderer case below its Pi floor Four follow-ups on the same branch, plus one revert. Revert the routine-delivery receipt a CI auto-fix round added. It introduced a new persisted `fm-branch-routine-delivery` entry, written into the captain's transcript for every routine note, to deduplicate a note whose cursor write failed. That is a change to the delivery contract, which this task is not authorized to make: the approved work is the asynchronous conversion with the existing durability contract preserved. The ownership re-read and output bounding from the review round are kept - both are genuine asynchronous correctness, not contract changes - as is that round's use of a real parent pid so the replacement regression traverses an actual ps subprocess. Record the routine gap instead of closing it. A routine note is a plain message with no sequence-keyed record, so a mark-read failure after delivery makes the next reconciliation send it once more; a captain row cannot duplicate that way because its visible entry is found by store sequence. That asymmetry predates moving delivery off the render thread. It is now stated at the call site and in the delivery-contract docs, tracked as fm-pi-routine-delivery-idempotency-followup-r1, and pinned by a regression that proves the routine note is re-delivered exactly once more and never again, the captain entry stays single, and the store keeps both rows. Give the stock-renderer case a Pi version floor. It compares the extension's renderers against Pi's stock rendering, so its verdict only means anything against the contract those renderers target: since 0.84.4 the stock renderer no longer supplies an implicit reset at multiline boundaries and the extension emits that reset itself, so an older installed Pi differs legitimately. It now names the installed version and the floor and skips, while a package whose version cannot be read at all still fails. Make the responsiveness regression's second signal a fraction rather than a millisecond budget. A loaded machine that deschedules the process inflates an absolute stall budget into a false failure, but it inflates the delivery's own wall time too, so requiring the worst stall to be a minority of that wall time holds under load. Synchronous delivery sits near 1.0 there whatever the load, and the tick-count signal still reads zero on it. Replace the test-family mapping for the Pi extension libraries with per-script targeting. Routing them to whole families - or leaving them unmapped, which widens through the reference scan to each referencing suite's entire family - selected dozens of suites with nothing to do with Pi and pulled an unrelated flake into the run. The changed-file selection drops from 112 scripts to 61. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013bzoWyr2EcJGBKuoUjVRSp * no-mistakes(document): Clarify asynchronous execution documentation --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * fix: avoid duplicate AGENTS.md governance for marked projects (#3763) * fix(memory): honor explicit project maintenance guidance * no-mistakes(test): Blocked by pre-existing Bash and Muse fixture failures * no-mistakes(test): Remove accidentally tracked test attribution report * no-mistakes(ci): Restricted the marker to the exact first line, preventing fenced examples from suppressing governance, and corrected the documentation. Regression failed before the fix; all 18 helper tests, focused ShellCheck, documentation validation, and diff checks pass. CI and Require no-mistakes report action_required with zero jobs executed; those external checks remain unresolved * fix: protect primary checkout when spawning from linked homes (#3783) * fix(spawn): refuse repository primary from linked spawning homes Compare the resolved task git directory with the spawning repository's common git directory before refreshing a fresh copy or relaunching a task. This protects the primary even when the spawning project is a linked home. Keep pooled copies accepted and preserve recorded work on relaunch. Fixes #3741. Verification for the pipeline PR body: - Red on origin/main 1820316b66ac2c68e244dd04a02512859ee8c1f4 with the new regression and unchanged production code: bin/fm-test-run.sh tests/fm-spawn-pool-base-freshen.test.sh exited 1 with "linked spawning home accepted primary as a disposable copy". - Green after the guard: the complete pool-base-freshen and control-relaunch suites passed through bin/fm-test-run.sh, covering primary and symlink refusal before fetch/reset, spawning-directory refusal, scout acceptance, and committed plus unfinished work preserved during linked-home relaunch. - The worktree-settle suite passed on pristine main and the final branch. An earlier loaded-host run exceeded its five-second assertion (6s); the final retry passed without changing code or the assertion. - Test fixture commits ran with GIT_CONFIG_COUNT=1, GIT_CONFIG_KEY_0=commit.gpgsign, GIT_CONFIG_VALUE_0=false. - bin/fm-lint.sh and /bin/bash -n for all three changed scripts passed. The upstream cwd-selection cause remains outside this change. * no-mistakes(document): Clarify spawn isolation ownership and relaunch preservation * fix(bin): stop reading an unanswered backend probe as a dead endpoint (#3785) * fix: read a failed herdr CLI as unreachable, not a gone backend target The no-run fallback in bin/fm-crew-state.sh collapsed every failed pane capture into 'backend target gone', which downstream consumers treat as positive death evidence - so a herdr CLI that errors or stalls under load briefly scored dozens of live claims dead on a busy box. Only a successful herdr answer proving the pane absent (fm_backend_agent_state's 'missing', backed by pane get answering pane_not_found) may now read as gone; every other verdict reports 'backend unreachable' with the endpoint state, which is never positive death evidence. Adds a behavior test: an always-failing fake herdr reads unknown/unreachable, never gone. * test: pin the herdr suite's ambient home to a marker-free fixture FM_HOME defaults to the suite's own root when unset, and any secondmate- marked checkout (every treehouse crew home carries .fm-secondmate-home) flips the default workspace label to 2ndmate-*, so the ambiguous-label placement test found zero firstmate matches and fell into the create path instead of refusing (expected exit 3, got 1) - deterministically green in CI, deterministically red from a crew home. Export a marker-free ambient FM_HOME fixture; per-test FM_HOME prefixes still override it. * fix: classify herdr endpoint answers instead of every non-missing verdict Review decision (firstmate, 2026-09-05): a failed pane capture is not itself evidence of death, but neither is every non-missing classifier verdict a failed answer. missing (pane get answered pane_not_found) and dead (pane present, agent_not_found husk) keep gone-class text so a stale-claim sweep may still reclaim them; an alive answer falls through to the normal busy/state flow instead of being discarded when only the heavy 200-line scrollback read failed; only when the cheap pane get / agent get calls themselves fail to answer does the line read 'backend unreachable'. Adds the two missing cases: alive with a failed scrollback read stays live, and a husk pane still reads gone. * no-mistakes(review): route tmux through agent-state classifier; drop test stall * no-mistakes(review): narrow inaccurate tmux socket and alive-arm fallback comments * no-mistakes(document): document classifier-backed endpoint verdicts in crew-state contract * fix(bin): preserve subshell lock ownership on Bash 3.2 (#3789) * fix: distinguish subshell wake-lock owners on stock Bash Restore distinct process ownership for issue #3743 using the existing PID helper, consistently across lock publication, reclaim, release, role checks, and bounded handoff. The existing wake-queue regression fails on pristine upstream Bash 3.2 with rc=13. The complete suite now passes on Bash 3.2.57 and Bash 5.3.15, with added coverage for ownership when BASHPID is unset. Canonical lint and stock-Bash syntax checks pass. * no-mistakes(document): Correct lock grace-period documentation * no-mistakes(ci): Captain, fixed all 14 SC2031 false positives with nine ShellCheck source-boundary annotations across three tests. Full CI-mode lint and the complete wake-queue suite on stock Bash 3.2 passed. Runtime behavior is unchanged * fix(bin): resolve captain holds and legacy teardowns on non-markdown backends (#3782) * fix(bin): close legacy records on the Beads backend honestly Two pre-Beads reads blocked honest closure of leftover records: 1. fm-captain-hold.sh complete/verify resolved attested legacy hold ids only against the live backend and the pre-collapse derived identity, so a home whose holds fm-hold-migration rehomed under fm- ids failed with an empty-name absence message (the resolve failure was swallowed by the command substitution feeding verify_hold_durable). Resolution now falls …
ionnich
added a commit
to ionnich/firstmate
that referenced
this pull request
Sep 22, 2026
* feat(calm): render smooth Unicode swell with asymmetric two-color sail (#4498)
* feat(calm): render smooth Unicode swell
* feat(calm): make sails asymmetric
* feat(calm): use quarter sail glyph
* no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer
* no-mistakes(document): docs: sync calm wave phase doc comment
* no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。
* fix(bin): supersede stale scout delivery text in brief.md on promotion (#4491)
* fix: supersede scout delivery brief on promotion
* fix: preserve ship safety contract after promotion
* no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch
* fix(bin): make captain holds work on hosts with an older JSON::PP, and stop cleanup dropping accents from a held body (#4471)
* fix(bin): let captain holds work on hosts with an older JSON::PP
Holding a task for the captain, and the cleanup that keeps a captain-held row
open, both fail outright on any host whose JSON::PP defaults allow_nonref off -
2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`,
but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older
library rejects that whole value with "must be object or array".
The consequence is fleet-wide on such a host, not one broken command: a worker
there cannot formally record a decision for the captain at all. It can only
mention the decision in passing in a status line, where it can be missed - which
is how a real decision goes unrecorded. The hold reports that the task lost its
hold-set stamp; the cleanup cannot return the row to Queued.
Both call sites now ask for allow_nonref explicitly rather than inheriting
whatever the installed library defaults to. The second one is worth naming: its
`/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string
case that fails, so the guard selects for the failing input rather than
protecting against it.
The regression case forces the older default back off for every perl the commands
spawn, then drives both paths - holding a task that carries a body, and tearing
down a captain-held row whose deliverable must still be appended. It also probes
that the simulation genuinely rejects a bare scalar, so the case cannot pass
vacuously on a lenient host. Each half was verified failing on its own unfixed
call site with that site's real error message. Suites: fm-captain-hold-lifecycle
51 cases, fm-backlog-atomicity 99 cases, 0 failures.
Verification limit: the mechanism is reproduced and tested, but neither fix is
verified against a real JSON::PP 2.27202 host, because none is in the loop. This
laptop runs 4.06, where the bug does not manifest.
`bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a
brace-delimited object before decoding, so allow_nonref never applies.
* fix(bin): stop cleanup silently dropping accented characters from a held body
Cleanup rewrites a captain-held row's body to append the finished work's
deliverable, and the decoder it reads that body with printed decoded characters
to a stream with no `:raw` layer. A character at or below U+00FF then came out
as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the
accent. `fm_backlog_retain` writes that body straight back through
`--body-file`, and nothing reported an error - the character was simply gone
from a row still waiting on the captain.
The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus
`utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already
used.
Review of the parent commit found this on one of the lines that commit already
changed. It predates that change.
The test asserts bytes rather than decoded strings, because comparing strings
cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any
character above U+00FF makes perl print the whole string as UTF-8, so one body
carrying both an accent and an em dash passes even unfixed and proves nothing.
Verified failing before the fix on the accented row, passing after. Suites:
fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures.
* no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc
* no-mistakes(review): drop whole-file UTF-8 check from retained-body test
* no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc
* no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc
* fix(bin): read codex 0.154's idle braille starfield rows as composer furniture (#4532)
* fix(composer): read codex 0.154's idle starfield and status footer as furniture
codex-cli 0.154.0 animates a braille "starfield" around its idle composer:
on the row above the bold `›` prompt row, on the `›` row behind the SGR-2
dim `Ask Codex to do anything` placeholder, and on the row below it, then
draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`).
The cells are truecolor greys on both sides of the ghost luminance ceiling,
so the brighter ones survive ghost stripping, and the rows below the glyph
carry no structural edge. The shared classifier selected the bare `›` shape,
extended its wrap region over the two rows beneath the glyph, read the
survivors and the footer as wrapped typed input, and answered `pending`;
the steering doorbell defers on exactly that verdict, so no doorbell ever
reached an idle codex 0.154 pane.
bin/fm-composer-lib.sh now recognises that furniture by shape, declared
once next to the idle placeholders and reached from the two wrap-region
boundary points:
- a row whose non-whitespace content is entirely braille cells
(U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it
never counts as wrapped typed content and bounds a bare composer's wrap
region; braille behind the glyph row's content is stripped before the
emptiness decision when nothing else follows the glyph; a row mixing
braille with other text stays typed content;
- the codex status footer bounds the wrap region exactly as omp's status
row does, anchored on the effort token, a spaced middle dot, and a `~` or
`/` path cell, so a typed `fix · tests` stays composer input;
- `^Ask Codex to do anything$` joins the verified idle-placeholder set; the
ghost strip remains what proves that row empty, and the bare-row rule that
bright placeholder text is real input is unchanged.
Unchanged: the strict blank-row rule, the styled=0 degradation (a plain
cmux/orca capture of this screen still reads `unknown`, never `pending`),
FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape.
tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte
with the divergence (letters in place of the starfield read `pending`) and
the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh
is the default-on live guard (token-free, skips explicitly without codex or
tmux) that launches the installed codex idle and asserts `empty` through
both the tmux and the cursorless styled reads, naming codex --version on
failure. docs/verification/runtime-backends.md records the dated Herdr
evidence: `pending` before, `empty` after, on the captured screen.
* no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry
---------
Co-authored-by: Todd Billings <todd@usdvcapital.com>
* fix(bin): refuse empty text steers in fm-send (#4259)
* fix(bin): refuse empty text steers in fm-send
A marked secondmate request sent with an empty message delivered only
marker and correlation bytes and minted a pending-reply expectation the
parent could never see resolved, stalling the fleet with no loud error
(#4255). Fail closed on an empty or whitespace-only message on the text
path, mirroring the existing --resolve-key refusal.
* chore: retain ambient Pi-lens autoformat as its own commit
Formatting-only edits produced by ambient Pi-lens autoformat during the
msg-loss investigation, kept separate from the behavioural change in
c23acba6 so the fix stays reviewable on its own.
AGENTS.md is deliberately excluded: its only autoformat edit stripped the
trailing space from the documented FM_OPERATIONAL_PREFIX value, which
bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11
records as permanent compatibility. Documenting that constant without its
trailing space makes the doc wrong about the contract, so that one line was
restored rather than retained.
* fix(calm): paint the working ship one yellow over all-blue water (#4554)
On rose-pine-moon the two-color water (cyan crests over blue troughs) read as
a pink stripe over aqua, the yellow left sail and mast clashed with the red
right sail, and the hull carried a blue interior run. Every water cell is now
blue so the swell reads through glyph height alone, and both sail halves, the
mast, and the whole hull are one yellow run. Geometry, cadence, animation,
direction flip, resize clamping, and the narrow fallback are unchanged.
Update the unit and real-TUI color assertions to the new palette and the Calm
docs that described the old one.
* fix(bin): stop aging a second mate's active turn from its launch (#4270)
* fix(watch): stop aging a second mate's active turn from its launch
The parent watcher's second-mate wake-loop stall check exempts a mate that
is demonstrably inside an active turn, but secondmate_in_active_turn asked
busy_turn_over_age first and returned "not in a turn" whenever that said
the bound was crossed.
busy_turn_over_age ages from state/<task>.turn-ended, falling back to
state/<task>.meta. A second mate's turns end in its own home, so the
parent never gets a turn-ended mark for it and the fallback ages the
mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago
was therefore permanently "over age", the busy pane was never consulted,
and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false
wake-loop stall.
The gate now bounds the busy exemption by <idle> - how long the queue's
drain position has not moved - which is evidence this home actually
holds. A busy mate stays exempt while the queue has been frozen for less
than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so
the bound that stops a busy pane from proving liveness forever is kept
rather than removed. busy_turn_over_age is untouched; its remaining
callers are the ordinary crew busy-pane bound.
The regression pins the case that actually broke: a mate whose launch
record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn
must not escalate, while the same mate with its queue frozen past the
bound still publishes exactly one notification. The existing coverage
only exercised a freshly launched mate, which passes either way.
Reaching that alert now costs a pane capture inside the gate, so the
three checkpoints in this suite that assert an alert move from a 1s to a
4s bound - the value the neighbouring active-turn cases already use. The
bound is a ceiling, not a wait: the checkpoint returns on the first
actionable wake. On a loaded machine a 1s bound missed the alert
repeatedly; at 4s it did not miss in 20 runs under the same load.
* no-mistakes(review): scope the second-mate active-turn regression test's coverage claim
* no-mistakes(document): fix stale second-mate active-turn comments in fm-watch
* feat(bin): add read-only PR blocker and reviewer discovery commands (#4278)
* feat(bin): add read-only PR blocker and reviewer-discovery commands
Two focused, opt-in commands that read GitHub and never write to it.
fm-pr-state.sh reports what still blocks one pull request from the
author's side: a closed or merged state, draft state, unknown or
conflicting mergeability, absent or failing required checks, and a
blocking CHANGES_REQUESTED decision explained by each reviewer's latest
verdict, marked STALE when it was left at a superseded head. A pull
request that only awaits an approval is not reported as blocked, and
advisory checks are omitted. Every reading is taken against one exact
head; a push that lands mid-read invalidates the whole result rather
than mixing two snapshots.
fm-pr-reviewers.sh suggests reviewers from the most recent commits to
the pull request's exact changed paths, counting each commit once,
resolving handles through GitHub's own commit author.login mapping, and
excluding the author and Bot accounts.
Both stay read-only: no review request, no approval, no merge.
Unresolved review-thread state is left unreported because the REST API
does not expose it and unattended commands may not use GraphQL.
Closes #3731
* no-mistakes(review): accept only PR URLs and stop at terminal state
* no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate
* no-mistakes(review): stop attributing readings to unverified heads
* no-mistakes(review): narrow readiness contract to checks that have reported
* no-mistakes(review): read the pull request once, drop the head guard
* no-mistakes(document): scope pr-forge isolation proof to its measured members
* no-mistakes(document): record uncovered pr-forge members and their pending proof
* docs(isolation-proof): re-prove pr-forge at its full membership
tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the
pr-forge family in this branch, and script_allows_concurrency grants
four workers by family membership alone, so both ran concurrently on a
proof measured before they existed.
Re-proved the family at all eight members: two consecutive runs, 0
failures, each begun with the one-minute load average below 6.0 so the
result measures isolation rather than contention. A third run taken
between them is disclosed rather than recorded, because it started
while the previous run's workers were still decaying.
The new durations are not comparable with the six-member measurement
above them, so they are not presented as evidence about the two new
members, and that record's 1.72x four-worker figure is left as a
statement about its own run rather than restated as current.
* no-mistakes(review): disclose gh error-text coupling at its matching site and tests
* fix(bin): teach validation-round pauses in generated briefs (#2752)
* fix(bin): teach validation-round pauses in briefs
* no-mistakes(document): Point classifier comments to authoritative pause examples
* docs(readme): add star history chart (#4558)
* fix(bin): refuse teardown when a task's endpoint close fails (#4510)
* fix(teardown): refuse a cleanup whose endpoint close failed
bin/fm-teardown.sh discarded both the exit status and the stderr of every
fm_backend_kill call, so a close that genuinely failed was indistinguishable
from one that succeeded. Teardown continued past it, deleted the task's durable
records, returned its worktree, and reported the cleanup as completed. The
deleted metadata is the only record of which endpoint belongs to the task, so
such a close did not merely leave a stray session behind, it stranded one:
nothing was left on disk naming it.
The adapters could not carry that signal either. Driven against the real code,
every backend arm returned 0 for a genuine failure exactly as it did for an
already-exited endpoint, so there was nothing for the four call sites to
propagate even once they stopped swallowing it.
The tmux arm now resolves a close that did not succeed against the window's
exact recorded identity, since kill-window fails the same way for a window that
is gone and one that is still there. The Orca arm reports a close its missing
CLI never attempted. Both stay silent for an endpoint that is already
legitimately gone, and the remaining arms are unchanged: their close-command
timing cannot be established without the real Zellij, Orca, and cmux binaries,
and a gate that refused ordinary cleanup of an already-exited session would be
worse than the defect. docs/verification/runtime-backends.md records what each
backend can prove.
A reported close failure now reaches teardown's existing retain-and-stop
refusal before the records naming the endpoint are removed, matching where the
Herdr confirmed-gone gates already sit for the same hazard, and the retained
records let a rerun finish once the close works.
* no-mistakes(review): refuse unreadable tmux close re-read; honor --force override
* no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close
* no-mistakes(document): document endpoint-close refusal in its backend and retirement owners
* no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree
* feat(calm): add flag-gated Claude Code Calm mode (#4565)
* feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag
Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that
brings Calm to Claude Code: the sailboat replaces the stock working row through a
Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration,
and canonically classified operational user rows draw at zero height. /calm is
registered by the hooks module itself and toggles the same per-home config/calm
preference the Pi extension uses, so one choice applies on either harness; rows
redraw retroactively on toggle and stay hidden across claude --continue.
The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS
flag is on. Nothing sets that flag in any settings file, and the plugin carries no
command file, skill, agent, or classic hook, so it is a complete no-op while the
flag is off. The trusted project auto-loads it through an .agents/skills symlink,
the only path Claude Code scans for project plugins.
Extract the working-ship geometry, bounce track, cadences, and freeze/resume state
into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module
imports from outside the plugin folder) and have the Pi widget paint that core's
frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify
operational rows through a port of bin/fm-operational-input.sh's classify command
guarded by a corpus parity test against the shell owner.
Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering,
Raster packing, policy, classifier parity), the mod's own claude plugin test suites
behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op,
the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272.
Docs: record the version-scoped Claude Code evidence and the three bounded gaps in
docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md,
and make the shared preference, layout, and contributor notes harness-neutral.
* no-mistakes(review): Preserve colliding final replies and strengthen parser parity
* no-mistakes(review): Preserve final replies and strengthen canonical parity checks
* no-mistakes(review): Require exact function-hooks opt-in before Calm activation
* no-mistakes(review): Clarify Calm module loading and activation boundaries
* no-mistakes(review): Reset Calm presentation state across session starts
* no-mistakes(document): Refresh Calm session lifecycle documentation
* feat(calm): paint the Claude Code working ship in Claude's own theme colors
The captain picked the "Claude native" palette for the Claude Code mod's Raster:
every water cell takes the spinner blue of the active theme family (#93a5ff dark,
#5769f7 light) and the whole boat takes the Claude orange of the stock spinner
(#d77757), one water color and one boat color. The family follows the `theme`
setting's prefix, read at load through $.config.list and re-read on a
config.set of that row, with `auto` and custom themes falling back to the dark
set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte.
Rename the shared sprite's color classes from hue names to `water` and `boat`,
since each harness now maps them to its own colors; geometry, motion, cadence,
and the activation gate are untouched.
Tests cover both palettes' packing and the family rule under Node, and the
plugin kit drives every theme value, a theme change mid-session, the Calm-off
pass-through, and inertness of the menu read while the flag is off. The docs
describe the Claude Code colors and record the guard passing on 2.1.273.
* no-mistakes(review): Use light palette for unresolved Claude themes
* no-mistakes(document): Refresh Claude Calm verification evidence
* fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586)
* fix(watch): honour a declared wait before wedge-escalating a quiet pane
wedge_timer_check escalated on elapsed idle time alone. Nothing asked
whether the worker had already said why its pane was quiet, so a lane
that declared a bounded external wait climbed the escalation ladder for
as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every
repeat carried demand-deep-inspection - which by its own wording forbids
re-absorbing on the run-step or pane state, so the supervisor could not
use the evidence that was there either.
The generated brief promises that declaring `paused:` buys the long
recheck cadence instead of a wedge, but the timer was still reachable
while that declaration stood: a crew that declares a wait and then has an
active run or busy pane attributed to it is handed to the timer as
provably-working. The declaration is what the worker said about its own
silence, so it now outranks a liveness verdict that only says something
is running.
The consult runs in the at-threshold branch that was about to escalate,
beside the worktree walk already there, and costs one status-line read.
Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS
recheck the declared-wait absorber already uses, so the wait is still
rechecked and cannot rot invisibly. Which verb declared it decides the
wording, because the two block on different people: a `paused:` wait is
owed by an external dependency and asks the reader to confirm it still
holds, while a `captain-held:` transfer is owed by the captain reading
the recheck and asks them to answer or release the hold. A hold is not
rechecked at all while the away-posture record exists, as on every other
captain-held path, and that absorb arms no throttle so the recheck is
owed in full on return.
A declared clearing time that has already passed stops counting, and a
lane that never declared one keeps the identical escalation schedule,
reason, count and demand-deep-inspection wording, so detection and its
worst-case time are unchanged. The deferral restarts the idle timer
rather than cancelling it, so a lane that stops waiting escalates again
within one threshold.
A lane quiet because its own validation run is parked at a gate awaiting
a human decision is deliberately out of scope: reading that state needs a
signal carrying who the wait is on and what clears it, rather than one
inferred from a parked verdict that also covers gates awaiting the
crewmate itself.
Tests pin both directions for each case and were each confirmed to fail
with the consult removed.
* no-mistakes(document): docs: honour declared waits in stale-escalation docs
* fix(bin): report verified PR state for passed runs (#4624)
* fix(bin): derive passed PR state from PR record
A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists.
For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim.
Fixes #4607
* no-mistakes(review): Add bounded GitLab merge-request state reads
* no-mistakes(review): Preserve network-free inactive crew-state scans
* no-mistakes(document): Document PR record readers in shared library
* fix: restore published contribution follow-up (Fixes #4469) (#4627)
* fix: restore published contribution follow-up (Fixes #4469)
* fix(review): Fix contribution freshness and merge actor routing
* fix(review): Restore issue triage and scope contribution follow-up
* fix(test): test: assert one wake per contribution signal
* fix(document): Document contribution follow-up
* fix: restore truthful terminal delivery evidence
* fix(review): Disclose unsupported contributions and deduplicate watcher wakes
* fix(review): Preserve unmeasured unsupported contributions across Bearings
* fix(review): Deduplicate shared contribution wakes and isolate diagnostics
* fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower
* fix(bin): make remote report transfers explicit and fail-open (#4658)
* fix(bin): make a remote-reply document gap self-clearing and re-attemptable
A remote mate's undelivered document raised a keyed `blocked` decision that
nothing could ever resolve, and any `data/*.md` substring in any mirrored line
was an unconditional fetch instruction. A mate announcing a report it had not
written yet therefore manufactured a permanent, factually false blocker, and
its own explanation of the false alarm manufactured more.
The reader has no permanence vocabulary: a report still being written refuses
exactly like a path that will never exist. So an undelivered document is now a
durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`,
re-attempted on the next delta and on the channel's own quiet poll, and retired
with a matching `resolved` line naming the local copy once it arrives. The
cursor still advances and no delta stalls on one bad pointer.
Only a structured `report=data/....md` pointer now offers a document, so a path
merely mentioned in prose - including one under another home's mirror tree,
which is provably not that mate's to serve - is never fetched. Offers are
deduplicated across the whole delta, the escalation names each missing document
once and carries the reader's own reason instead of discarding it, and a
strictly increasing notice ordinal keeps a later escalation from being
swallowed as duplicate bytes. A mirrored line still lands once whichever
pointer form it was first written under.
* no-mistakes(review): Require structured pointer token boundaries
* no-mistakes(review): Unify boundary-safe pointer extraction and rewriting
* fix(bin): identify a mirrored line independently of its delivery state
Two defects in the boundary-safe pointer work.
The at-most-once check compared only the all-remote and all-local renderings
of a line, so it could not recognize a mixed one. A line offering two documents
where only the first was deliverable mirrored as local-plus-remote; once the
second arrived, a cursor-loss whole-log recapture rendered the same line
all-local, matched neither alternate, and mirrored a second time. A line's
identity is now the canonical form every boundary-valid pointer would take once
delivered, derived by the same parser that does extraction and rewriting, so it
no longer depends on which documents happened to be deliverable at the time.
The pointer map was passed to awk through the process environment. A delta may
carry up to the configured 1 MiB bound, and an expanded map of delivered
pointers can exceed the platform's exec argument limit, so awk would fail to
start; because no caller checked, the empty result would have been appended as
blank lines while the cursor advanced past dropped status content. The map now
travels in a file, and every call site checks the exit status and stops the
ingest rather than committing a delta it could not render.
Both passes now run once per stream instead of twice per line.
* no-mistakes(review): Abort ingest when document pointer extraction fails
* no-mistakes(review): Exclude structured cross-home pointers from document transfer
* fix(bin): fail open on an undeliverable remote document instead of tracking it
Narrow the remote-reply document fix to the scope the diagnosis actually
requires, as decided after measuring a simpler alternative.
A document the reader cannot deliver now fails open. The mate's line is
mirrored with its own pointer, the cursor advances, and one unkeyed note
carries the reader's reason. A note never enters the open-decision fold, so it
cannot stand open the way the original keyed block did - which removes the
never-clearing false blocker by construction rather than by resolving it.
That makes the durable self-clearing obligation unnecessary, so it goes: the
per-mate pending-documents record, its notice ordinal and resolved
announcements, and the poll-side retry. Canonical line identity goes too, and
with it a way to silently drop a genuine status line; mirroring is back to
at-most-once on exact bytes. The cross-home exclusion goes as well: under
fail-open a cross-home report= either fails harmlessly or is a nested remote
report this mate genuinely holds, which is now relayed again.
Kept: fetching only on a structured report= pointer, the boundary-correct
parser, the file-based rewrite map, and checked extraction and rewrite exit
status. The parser now scans behind a sentinel byte so a rejected candidate can
no longer give the text right after it a false leading boundary.
The reported incident is covered end to end: a report path announced in prose
before it exists raises no decision, and the report still arrives through the
ledger publisher's structured offer once written.
* no-mistakes(review): Preserve source-line identity across remote reply replays
* no-mistakes(document): Document remote reply transfer and replay semantics
* no-mistakes(lint): Fix staging truncation lint checks
* fix(calm): preserve substantive mid-turn responses (#4655)
* Preserve substantive Calm mid-turn text
* no-mistakes(review): Distinguish newline-preserved replies from short narration
* no-mistakes(document): Document Calm mid-turn preservation boundaries
* no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`
* fix(bin): preserve PR merge polls across volume remounts (#4656)
* fix(bin): re-record PR poll identity after a volume device renumber (Fixes #4260)
A volume remount can renumber the state filesystem's st_dev while every
inode and byte stays the same; APFS does this across a reboot. A poll
registration records its sidecar and check as device:inode, so every poll
armed before the remount failed strict validation and the watcher refused
all of them as unauthenticated state checks until each was re-armed by hand.
There are two device comparisons. fm_pr_private_file_valid compares a live
file's device with the state directory's device read in the same invocation:
it refuses a file that is not on the state directory's own filesystem and
already survives a renumber, so it is unchanged. The registration's recorded
identity versus the live identity (from #556, reused by the #932 retirement
receipt) binds the registration to the exact files published in its own
transaction; its device part is what breaks.
When strict capture fails, the watcher now proves the device is the only
difference: every other artifact check passes (template bytes, both hashes,
private mode, single link, live device, metadata), both recorded identities
name one device, and each recorded inode equals its live inode. Only then,
under the task's control lock, does it rewrite the two identity lines,
repeating the whole proof and comparing the registration's file identity and
bytes just before the rename, and then capture strictly again. A swapped,
altered, re-moded, relinked, split-device, or foreign-device artifact still
fails a proof and is still refused, and a pending retirement receipt blocks
the rewrite.
Reproduction: on macOS a poll armed on an APFS disk image that was detached
and re-attached behind another image moved st_dev 16777239 -> 16777243 with
inodes, bytes, mode, and link count unchanged; the real watcher refused it on
main and reports its merge with this change. The portable regression test
rewrites a real registration's recorded device and drives the watcher.
Not changed here: the status presentation cursor keys rows by its own
device:inode identity in bin/fm-classify-lib.sh, a different helper that
needs its own fix; a retirement receipt left by a reboot between its
publication and removal still names the old device and stays refused; custom
check trust binds only a content hash and is unaffected.
* fix(review): Serialize PR poll publication writers
* fix(review): Bound PR poll publication lock scope
* fix(bin): keep contribution records when the poll budget runs out (follow-up to #4627) (#4661)
A budget that expires partway through an observation no longer records an
error or prints the unavailable wake; the URL keeps its prior record and is
observed first next poll. forge() flags budget exhaustion at the point it
refuses, or when a read is killed at the budget's own deadline, so a genuine
forge failure still records the error and wakes. Each distinct URL is now
observed once per poll and applied to every owning task.
* fix(bin): clear parent pending-replies on local secondmate retirement (#4680)
* fix(bin): clear parent pending-replies on local secondmate retirement
Local secondmate teardown left resolved parent pending-reply records behind
after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced
retirement while any reply for that id is still unresolved, and delete every
matching record plus its delivery confirmation after a successful local or
remote retirement, matching the remote cleanup path.
* no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup
* no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung
* no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern
* no-mistakes(review): Pending-replies Basename und corr_id abgleichen
* no-mistakes(document): Clarify forced retirement pending-reply cleanup
---------
Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>
* fix(bin): accept Orca's composite worktree id when tearing down a task (#4677)
* fix(bin): accept Orca's composite worktree id at teardown
Teardown refused every Orca-backed task because the endpoint validator
checked orca_worktree_id with the simple-atom rule meant for tmux-style
window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca
returns that id as `<orca id>::<absolute worktree path>`, so the colon and
slashes in every real value made validation fail and finished Orca tasks
could never be cleaned up.
Validate the field as the composite it is: both halves of the first `::`
split present, the path half absolute, and no embedded newline, carriage
return, or tab. The terminal field keeps the atom check, which is correct
for it, and no other backend's validation changes.
The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca
never returns, which is why the suite passed a check the real value fails.
They now carry the composite form, so the tests exercise the real value.
* no-mistakes(document): name Orca's repo id in the composite worktree id
* no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points
* feat(bin): add opt-in typed dispatch resolution (#4692)
* feat(bin): add opt-in typed dispatch resolution through typesafe.ai
Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or
scout profile from a written brief with typesafe.ai's System One model:
one Choice question over the rules' `when` texts, then the confidence
floor, the rule's `approval` and `floor`, each profile's `provider` and
`floor`, one quota-axi snapshot, and the spendPriority argmax all in code.
It is off unless TYPESAFE_API_KEY is in the environment or the home's
gitignored .env; off means one stderr line, exit 0, and no network call,
so firstmate dispatches exactly as before. The key reaches curl on a file
descriptor, never argv.
Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and
the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new
tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates
the four new optional dispatch fields. Document the schema, the operator
contract, the AGENTS.md intake step, and the live and benchmark evidence.
* no-mistakes(review): Harden typed dispatch resolution and quota bounds
* no-mistakes(review): Validate dispatch floors and ranking evidence
* no-mistakes(review): Tighten dispatch response and floor evidence
* no-mistakes(review): Neutralize none matching and resolve defaults locally
* no-mistakes(review): Preserve providerless profiles outside typed resolution
* no-mistakes(review): Validate response usage and reject duplicate profiles
* no-mistakes(review): Escalate unverifiable floors and validate probabilities
* no-mistakes(review): Validate probability mass and unknown profile floors
* no-mistakes(review): Simplify resolver interface and preserve fallback routing
* no-mistakes(review): Fix constants and rank partial quota evidence
* no-mistakes(review): Add authoritative provider mapping and enforce explicit providers
* no-mistakes(review): Declare provider for documented Pi profile
* no-mistakes(review): Validate provider identifiers and support Gemini dispatch
* no-mistakes(review): Strictly anchor provider identifiers
* no-mistakes(review): Validate selectors and preserve fallback candidate evidence
* no-mistakes(review): Gate typed validation and harden resolver evidence
* no-mistakes(review): Preserve opt-in routing and harden candidate evidence
* no-mistakes(review): Prioritize known exhaustion over quota uncertainty
* no-mistakes(review): Isolate API secrets and preserve no-key diagnostics
* no-mistakes(review): Fallback safely when dispatch rules are absent
* no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets
* no-mistakes(document): Document typed dispatch safety and fallback behavior
* fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753)
* test: reproduce buried status declarations in shared readers
* fix: share status event reads and preserve open blockers
* fix: retain terminal scout and ship status declarations
* no-mistakes(review): Fix status chronology, legacy completions, and reader performance
* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots
* no-mistakes(review): Unify terminal supersession across cached folds and consumers
* no-mistakes(review): Filter per-key status history while preserving terminal chronology
* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells
* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses
* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor
* no-mistakes(lint): Quote literal done in test for-lists for SC1010
* ci: expect 19 snapshot/fleet-view tests
This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.
* no-mistakes(review): Restore multiline child outcome reporting
* no-mistakes(review): Select ledger terminal events through bounded shared reader
* no-mistakes(review): Report newest open decision instead of preferring blocked
* no-mistakes(review): Require colon before ship/scout terminal supersession in fold
* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix
* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions
* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold
* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule
* no-mistakes(document): Align status-read docs with fold-resolved crew state
* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers
* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree
* test: fold terminal-cleanup snapshot coverage into the completed-scout case
Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.
* no-mistakes(document): Clarify socket-down override expiry in architecture doc
* ci: retrigger flaky contribution check
* fix(bin): launch codex crewmates with codex's hook layer disabled (#4689)
* fix(spawn): launch codex crewmates with codex's hook layer disabled
A freshly launched Codex worker never reached its instructions. Codex
stopped it on an interactive "Hooks need review" modal whose selection
sits on "Review hooks", which is neither trusting nor declining.
Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow
navigation, so the selection cannot be moved, and pre-accepting the
prompt by writing Codex's own trust store would record an operator
consent that was never given.
The hooks are the machine's own ~/.codex/hooks.json plus any project's
.codex/hooks.json. A crewmate needs neither: its turn-end signal is the
-c notify= program on the same launch, and Firstmate's project hooks are
primary-session infrastructure that stands down in a child worktree.
Crewmate and scout launches now pass --disable hooks. That is the
opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted
hooks; disabling the feature runs none of them and leaves the operator's
~/.codex untouched. An unknown feature name is a hard Codex error, so a
release that drops the flag fails the launch loudly instead of silently
restoring the modal. A secondmate is a primary in its own home and keeps
the project hooks its turn-end guard and session-start digest ride on.
Verified on codex-cli 0.151.0: the modal is gone and the turn-end
notification still lands.
This unblocks the second review that every finished pull request is supposed to get.
Fixes kunchenguid/firstmate#4673
* no-mistakes(review): Fix contradictory hook count in Codex verification record
* fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710)
* fix(bin): settle terminal contributions and wake once per read-failure episode
A contribution whose last good observation is merged or closed is final:
poll no longer re-reads it, projection keeps it fresh, and a stale error
recorded beside it is cleared once. A genuine forge-read failure on an open
contribution still records its error on every cycle but prints the
unavailable wake only when it starts a failure episode; a successful read
ends the episode. Open PRs linked from done tasks keep being observed.
The false unavailable beside a complete observation was budget exhaustion
mid-observation, already fixed by #4661.
* fix(review): Settle terminal contribution owners
* fix(review): Deduplicate shared contribution failure episodes
* fix(test): Preserve settled terminal contribution records
* fix: select authoritative no-mistakes runs (#4476)
* fix(crew-state): select authoritative validation runs by identity
Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row.
Refs: https://github.com/kunchenguid/firstmate/issues/3215
* fix(review): Resolve same-branch run identities beyond capped history
* fix(review): Fix run-selection compatibility, races, and worker-state fallbacks
* fix(review): Limit run validation to the requested branch
* fix(test): Anchor AXI fixtures and document remaining live evidence gaps
* fix(document): Clarify run selection documentation and capture ownership
* fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation
* fix: distinguish captain outcomes from no-op updates (#4738)
* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work
MAIN answered a supervision-branch outcome for completed captain-requested
work (implementation done, PR ready for review and merge approval) with
"Captain, shipshape.", reading section 9's no-action reply as covering it
and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no
captain-facing response is owed".
Section 9 now limits the shipshape reply to true no-ops (idle re-read,
empty heartbeat, consequence-free acknowledgement) and requires a short
outcome response naming what finished and what word is needed whenever
requested work finishes or a result needs the captain's word, even when a
transcript entry already shows the substance. The Pi protocol's re-emit
rule now says it bounds repetition only, and carries a worked example of
the ready-for-review outcome whose correct processing turn a shipshape
reply fails.
No executable contract evaluates the content of MAIN's captain-facing
reply, so the regression is the protocol example in the owner doc rather
than a text-match test.
* no-mistakes(document): Clarify captain-facing outcomes versus no-ops
* docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line
The document step condensed the Pi protocol's re-emit rule and dropped the
worked example of a finished, ready-for-review outcome whose correct
processing turn a "Captain, shipshape." reply fails. That example is the
contract's regression: no executable contract evaluates the content of
MAIN's captain-facing reply, so the owner doc's example is the test case.
Restore it directly under the re-emit rule, prefixed as a regression
example that is kept verbatim and never condensed or summarized away.
* no-mistakes(review): Clarify captain outcome and decision-word requirements
* no-mistakes(document): Clarify captain-facing completion outcomes
* docs(pi): require the PR URL in the visible captain-facing outcome reply
Captain review on the regression example: drop the sample reply string
and say only that the ready-for-review outcome requires relaying a
captain-facing outcome response, not just "Captain, shipshape.".
Fold in the visible-PR-handoff failure seen this session: after the
branch outcome reporting this fix green, MAIN's visible reply was only
"Awaiting your merge call." with no PR URL, leaning on the dim anchor.
Section 9's URL rule now also covers a review or merge ask and names the
visible reply as where the URL goes, sourced from the ready status, pr=
metadata, or the supervision branch's summary and never left to a
transcript entry. The Pi protocol adds the same-way failure and places
the captain-facing text in the final visible assistant reply after the
fm_branch_processed call, because Calm hides assistant text emitted in
the same step as a tool call as a working note.
Investigation verdict, evidence in the PR comment: no recent PR caused
the handoff failure; Pi has hidden same-step pre-tool assistant text
since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and
#4658 touched only remote report transfer.
* no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs
* no-mistakes(document): Clarify captain-facing supervision outcomes
* docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule
The consolidated section 9 URL rule narrowed its trigger to a review or
merge ask, dropping the "whenever a PR is mentioned" catch-all from
#3648 that keeps every PR URL copied from a durable record and never
assembled from memory. Restore that trigger as a union with the review
or merge ask so the one consolidated rule covers both.
* fix(bin): let non-owner Claude Stops exit safely (#4777)
* Fix foreign-owner turn-end supervision loop
* no-mistakes(review): Scope foreign-owner safe exit to Claude guard
* no-mistakes(document): Document Claude foreign-owner safe exit
* fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778)
Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty
indexed array as an unbound variable and aborts the shell. In
signal_turnend_panes_churned() the missing_keys loop was reachable with
an empty array whenever every churned key already held a fresh
.churn-since-* marker (a second churning turn-end inside an open
deferral window), so each watcher cycle died about half a minute in and
supervision restarted endlessly. The created_keys rollback loops had the
same latent crash on their error paths.
Audit of bin/ for the same pattern found one more confirmed-reachable
case: remote_handoff's noncanonical-body scan iterates to_move, which is
empty when a retried remote handoff finds every key already staged in
the outbox. All other "${arr[@]}" sites are either count-guarded,
guaranteed non-empty by construction, or unreachable while empty.
Guard the three reachable expansions with the repo's existing
"${arr[@]+...}" idiom. New regression test drives a real watcher
through the all-marked churn path; the macos-stock-bash CI lane runs it
under real /bin/bash 3.2 via FM_TEST_ONLY.
* Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783)
The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: require complete captain-facing final responses (#4779)
* docs: require complete final responses across harnesses
* no-mistakes(document): Document complete final replies for Grok Bot
* docs: point Grok replies to the shared contract owner
* no-mistakes(review): Clarify final recap without batching decision asks
* fix: preserve substantive mid-turn text in Pi Calm (#4788)
* fix(calm): preserve substantive Pi mid-turn text
* no-mistakes(review): Preserve substantive Pi Calm text per block
* no-mistakes(test): Cover shared Calm preservation boundaries behaviorally
* no-mistakes(document): Consolidate Calm preservation documentation
* fix: harden mail checks and rebalance full-coverage CI (#4800)
* Improve CI reliability and rebalance full-coverage validation
* no-mistakes(document): Clarify lint partition documentation
* fix(bin): answer Kimi 2.0.0 folder-trust dialog during spawn (#4799)
* Handle Kimi workspace trust dialog
* no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers
* no-mistakes(review): Gate Kimi ready on any trust marker and clean captures
* no-mistakes(review): Read visible pane for Kimi trust and ready gates
* no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate
* no-mistakes(review): Harden Kimi viewport capture and trust dialog detection
* no-mistakes(document): Document Kimi spawn refusal on cmux and Orca
* fix(bin): report a dead-agent record once instead of escalating forever (#4775)
* fix(bin): report a record whose agent is gone once instead of escalating forever
The wedge escalation path never asked whether there was still an agent to be
wedged. A wedge is something stuck that might recover, so re-alarming it earns
its cost; an agent that is gone never moves again, its pane never churns, the
idle timer never resets, and the escalate path clears its own timer and re-arms
with nothing bounding the count.
Observed on a live fleet: two finished lanes reached 226 and 203 consecutive
escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400
notifications a day from two lanes with no agent running at all. On one,
fm-control.sh exit answered already-stopped and fm-crew-state.sh read
"failed - run failed". Closing the Herdr pane did not stop it either: with the
pane genuinely gone and herdr pane read returning pane_not_found, the count kept
climbing, because the poll is driven by the record's window= line rather than by
the pane. The cost is not the repetition but that it drowns the alarms that
matter.
fm_backend_agent_state already separates a thinking agent from a gone one at
process level. In the branch that was about to escalate, read it once and treat
only its two recovery-grade verdicts - dead (endpoint present, no agent in it)
and missing (endpoint authoritatively absent) - as proof, reporting that record
once and not re-escalating it while it stays that way. Every other verdict,
including alive, ambiguous, unreadable, unverified, and a read that failed
outright, keeps the identical schedule, reason, and escalation count, so a
genuinely wedged live agent is unaffected. The probe costs at most one backend
read per window per threshold, the same budget the declared-wait consult and the
worktree write probe already take.
The report decides nothing about the record's fate: both lanes still held
unlanded work and teardown refusing them was correct, so retiring, relaunching,
or cleaning up stays with the supervisor. The once-only marker is owned entirely
by that function and is dropped by the same read the moment the endpoint stops
reading gone, so a replacement launched into the same window escalates normally
and its own later death is reported again.
Related, and not closed by this: #4412, #4482, #4316.
Tests drive the real watcher against a record whose endpoint does not exist and
pin both directions: dead and missing report once and never advance the count
across later thresholds, while alive, ambiguous, and unreadable endpoints keep
escalating with the identical reason and a climbing count.
* fix(bin): bind the once-only dead report to the pane it reported
Review of the parent commit found a reachable sequence where a later death in
the same window lost its promised report. The marker was keyed on the verdict
string alone and dropped only when a threshold probe read a non-gone verdict,
but probes run only at thresholds: a replacement launched into the same window
that dies without ever being probed alive - it crashes at startup, or works and
then crashes - was absorbed by the previous death's marker. The pane's first
sight yielded only the generic stale wake and every later threshold matched the
stale marker, so the second death never got the detailed once-report that both
the function's own comment and docs/architecture.md promise.
Record the verdict together with the pane hash it was reported for, and absorb a
repeat only while both still match. A replacement churns the pane, which resets
the stale suppressor, wedge timer, and escalation count while no reset site
touches this marker, so the pane half is what tells the second death apart from
the first. The live-probe drop stays as it was.
Clearing the marker at those reset sites instead would re-open unbounded
re-alarming for a dead pane whose display ever ticks, which is the exact defect
the parent commit exists to close.
The noise bound is unchanged: an unchanged dead pane still absorbs on every
later threshold and never advances the escalation count, and every verdict short
of proof still escalates exactly as before.
* no-mistakes(review): Key the dead-record once-marker on the busy incarnation token
* no-mistakes(document): Document dead-record escalation cap in stale-pane config entry
* no-mistakes(document): Add busy-state inventory line to AGENTS.md
* no-mistakes(document): Document dead-record probe on busy-turn-bound wedge path
* fix(bin): create captain-hold rows when Beads requires due (#4854)
Captain holds have no due semantics and are a hold kind, not a Beads issue
type. The create path now waives due.required and maps to native type task.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: disable compact adviser for spawned agents (#4877)
* feat(bin): launch every spawned agent with the compact adviser disabled
Every crewmate, scout, and secondmate Firstmate launches now starts with
COMPACT_ADVISER_DISABLE=1, on a fresh spawn and on a relaunch alike, so an
unattended session never activates the compact adviser.
The value is unconditional: no configuration file gates it and there is no
override, unlike the trace carrier beside it.
Three carriers deliver it, because no single one covers every launch shape.
The pane shell receives an export beside GOTMPDIR, so the agent's own children
inherit it too.
The launch command carries an explicit assignment, prepended outermost so it
wins over any ambient value the pane already held.
The cleared launch environment sets it again at the `env -i` boundary and keeps
COMPACT_ADVISER_DISABLE in the fixed operational floor, which is what preserves
the switch when config/launch-env-allowlist empties the environment, and what
delivers it on a remote host that never had the value.
bin/fm-control.sh relaunch, the bootstrap secondmate relaunch, and the remote
secondmate transport all rebuild their launch through bin/fm-spawn.sh, so they
inherit the same floor.
The captain's own primary session is untouched.
The two new suites drive the real spawn and then execute the launch command the
pane actually received, with the harness replaced by a probe that prints its own
environment, rather than matching script text.
They cover ship and secondmate launches with the allowlist absent and enabled,
the pane export and its ordering, fm-control.sh relaunch, and the full parent to
remote-host chain.
* no-mistakes(review): Export compact-adviser disable across compound launches
* no-mistakes(document): Document spawned-agent compact-adviser environment guarantee
* fix(bin): preserve Claude lock ownership after helper recycling (#4894)
* fix(bin): let a background Claude session keep owning its session lock
Session-lock ownership was decided by process ancestry alone. Under an
unattended Claude session the model loop runs in a transient bg-spare
bridged to the front-end by a shared daemon; when that bridge is
recycled the contiguous claude-named ancestry from a hook to the
recorded owner breaks while the owner pid stays alive, so the Stop
auto-arm stood down as a foreign live owner, the turn-end guard ended
every turn with its read-only diagnostic, and fm-lock.sh refused - a
self-sustaining outage until restart.
Ownership is now ancestry membership OR a trusted same-session id,
never id-first:
- fm-session-lock-lib.sh accepts CLAUDE_CODE_SESSION_ID only when
CLAUDE_PID is a Claude-shaped member of the current contiguous run,
compares it against the id recorded in state/.lock-session, and
requires the recorded pid to still be a live harness. No id, no
sidecar, an untrusted id, a different id, or a dead recorded pid
leaves the ancestry verdict unchanged. Ids are never read from ps
argv.
- fm-lock.sh accepts a same-session holder at both refusal sites,
writes, refreshes, and clears the sidecar only under its claim lock
(including the early already-mine exit, skipped only while the
deferred startup sweep leases that lock), keeps it byte-identical
across a same-session confirmation, records CLAUDE_PID on lock line 1
for a session with a trusted id so a shared daemon or front-end that
outlives the session never keeps a dead session's lock alive, never
rewrites a live line 1 on a same-session confirmation, and names t…
bcriswell
added a commit
to bcriswell/firstmate
that referenced
this pull request
Sep 23, 2026
* feat(bin): add opt-in typed dispatch resolution (#4692)
* feat(bin): add opt-in typed dispatch resolution through typesafe.ai
Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or
scout profile from a written brief with typesafe.ai's System One model:
one Choice question over the rules' `when` texts, then the confidence
floor, the rule's `approval` and `floor`, each profile's `provider` and
`floor`, one quota-axi snapshot, and the spendPriority argmax all in code.
It is off unless TYPESAFE_API_KEY is in the environment or the home's
gitignored .env; off means one stderr line, exit 0, and no network call,
so firstmate dispatches exactly as before. The key reaches curl on a file
descriptor, never argv.
Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and
the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new
tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates
the four new optional dispatch fields. Document the schema, the operator
contract, the AGENTS.md intake step, and the live and benchmark evidence.
* no-mistakes(review): Harden typed dispatch resolution and quota bounds
* no-mistakes(review): Validate dispatch floors and ranking evidence
* no-mistakes(review): Tighten dispatch response and floor evidence
* no-mistakes(review): Neutralize none matching and resolve defaults locally
* no-mistakes(review): Preserve providerless profiles outside typed resolution
* no-mistakes(review): Validate response usage and reject duplicate profiles
* no-mistakes(review): Escalate unverifiable floors and validate probabilities
* no-mistakes(review): Validate probability mass and unknown profile floors
* no-mistakes(review): Simplify resolver interface and preserve fallback routing
* no-mistakes(review): Fix constants and rank partial quota evidence
* no-mistakes(review): Add authoritative provider mapping and enforce explicit providers
* no-mistakes(review): Declare provider for documented Pi profile
* no-mistakes(review): Validate provider identifiers and support Gemini dispatch
* no-mistakes(review): Strictly anchor provider identifiers
* no-mistakes(review): Validate selectors and preserve fallback candidate evidence
* no-mistakes(review): Gate typed validation and harden resolver evidence
* no-mistakes(review): Preserve opt-in routing and harden candidate evidence
* no-mistakes(review): Prioritize known exhaustion over quota uncertainty
* no-mistakes(review): Isolate API secrets and preserve no-key diagnostics
* no-mistakes(review): Fallback safely when dispatch rules are absent
* no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets
* no-mistakes(document): Document typed dispatch safety and fallback behavior
* fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753)
* test: reproduce buried status declarations in shared readers
* fix: share status event reads and preserve open blockers
* fix: retain terminal scout and ship status declarations
* no-mistakes(review): Fix status chronology, legacy completions, and reader performance
* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots
* no-mistakes(review): Unify terminal supersession across cached folds and consumers
* no-mistakes(review): Filter per-key status history while preserving terminal chronology
* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells
* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses
* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor
* no-mistakes(lint): Quote literal done in test for-lists for SC1010
* ci: expect 19 snapshot/fleet-view tests
This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.
* no-mistakes(review): Restore multiline child outcome reporting
* no-mistakes(review): Select ledger terminal events through bounded shared reader
* no-mistakes(review): Report newest open decision instead of preferring blocked
* no-mistakes(review): Require colon before ship/scout terminal supersession in fold
* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix
* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions
* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold
* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule
* no-mistakes(document): Align status-read docs with fold-resolved crew state
* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers
* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree
* test: fold terminal-cleanup snapshot coverage into the completed-scout case
Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.
* no-mistakes(document): Clarify socket-down override expiry in architecture doc
* ci: retrigger flaky contribution check
* fix(bin): launch codex crewmates with codex's hook layer disabled (#4689)
* fix(spawn): launch codex crewmates with codex's hook layer disabled
A freshly launched Codex worker never reached its instructions. Codex
stopped it on an interactive "Hooks need review" modal whose selection
sits on "Review hooks", which is neither trusting nor declining.
Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow
navigation, so the selection cannot be moved, and pre-accepting the
prompt by writing Codex's own trust store would record an operator
consent that was never given.
The hooks are the machine's own ~/.codex/hooks.json plus any project's
.codex/hooks.json. A crewmate needs neither: its turn-end signal is the
-c notify= program on the same launch, and Firstmate's project hooks are
primary-session infrastructure that stands down in a child worktree.
Crewmate and scout launches now pass --disable hooks. That is the
opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted
hooks; disabling the feature runs none of them and leaves the operator's
~/.codex untouched. An unknown feature name is a hard Codex error, so a
release that drops the flag fails the launch loudly instead of silently
restoring the modal. A secondmate is a primary in its own home and keeps
the project hooks its turn-end guard and session-start digest ride on.
Verified on codex-cli 0.151.0: the modal is gone and the turn-end
notification still lands.
This unblocks the second review that every finished pull request is supposed to get.
Fixes kunchenguid/firstmate#4673
* no-mistakes(review): Fix contradictory hook count in Codex verification record
* fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710)
* fix(bin): settle terminal contributions and wake once per read-failure episode
A contribution whose last good observation is merged or closed is final:
poll no longer re-reads it, projection keeps it fresh, and a stale error
recorded beside it is cleared once. A genuine forge-read failure on an open
contribution still records its error on every cycle but prints the
unavailable wake only when it starts a failure episode; a successful read
ends the episode. Open PRs linked from done tasks keep being observed.
The false unavailable beside a complete observation was budget exhaustion
mid-observation, already fixed by #4661.
* fix(review): Settle terminal contribution owners
* fix(review): Deduplicate shared contribution failure episodes
* fix(test): Preserve settled terminal contribution records
* fix: select authoritative no-mistakes runs (#4476)
* fix(crew-state): select authoritative validation runs by identity
Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row.
Refs: https://github.com/kunchenguid/firstmate/issues/3215
* fix(review): Resolve same-branch run identities beyond capped history
* fix(review): Fix run-selection compatibility, races, and worker-state fallbacks
* fix(review): Limit run validation to the requested branch
* fix(test): Anchor AXI fixtures and document remaining live evidence gaps
* fix(document): Clarify run selection documentation and capture ownership
* fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation
* fix: distinguish captain outcomes from no-op updates (#4738)
* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work
MAIN answered a supervision-branch outcome for completed captain-requested
work (implementation done, PR ready for review and merge approval) with
"Captain, shipshape.", reading section 9's no-action reply as covering it
and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no
captain-facing response is owed".
Section 9 now limits the shipshape reply to true no-ops (idle re-read,
empty heartbeat, consequence-free acknowledgement) and requires a short
outcome response naming what finished and what word is needed whenever
requested work finishes or a result needs the captain's word, even when a
transcript entry already shows the substance. The Pi protocol's re-emit
rule now says it bounds repetition only, and carries a worked example of
the ready-for-review outcome whose correct processing turn a shipshape
reply fails.
No executable contract evaluates the content of MAIN's captain-facing
reply, so the regression is the protocol example in the owner doc rather
than a text-match test.
* no-mistakes(document): Clarify captain-facing outcomes versus no-ops
* docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line
The document step condensed the Pi protocol's re-emit rule and dropped the
worked example of a finished, ready-for-review outcome whose correct
processing turn a "Captain, shipshape." reply fails. That example is the
contract's regression: no executable contract evaluates the content of
MAIN's captain-facing reply, so the owner doc's example is the test case.
Restore it directly under the re-emit rule, prefixed as a regression
example that is kept verbatim and never condensed or summarized away.
* no-mistakes(review): Clarify captain outcome and decision-word requirements
* no-mistakes(document): Clarify captain-facing completion outcomes
* docs(pi): require the PR URL in the visible captain-facing outcome reply
Captain review on the regression example: drop the sample reply string
and say only that the ready-for-review outcome requires relaying a
captain-facing outcome response, not just "Captain, shipshape.".
Fold in the visible-PR-handoff failure seen this session: after the
branch outcome reporting this fix green, MAIN's visible reply was only
"Awaiting your merge call." with no PR URL, leaning on the dim anchor.
Section 9's URL rule now also covers a review or merge ask and names the
visible reply as where the URL goes, sourced from the ready status, pr=
metadata, or the supervision branch's summary and never left to a
transcript entry. The Pi protocol adds the same-way failure and places
the captain-facing text in the final visible assistant reply after the
fm_branch_processed call, because Calm hides assistant text emitted in
the same step as a tool call as a working note.
Investigation verdict, evidence in the PR comment: no recent PR caused
the handoff failure; Pi has hidden same-step pre-tool assistant text
since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and
#4658 touched only remote report transfer.
* no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs
* no-mistakes(document): Clarify captain-facing supervision outcomes
* docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule
The consolidated section 9 URL rule narrowed its trigger to a review or
merge ask, dropping the "whenever a PR is mentioned" catch-all from
#3648 that keeps every PR URL copied from a durable record and never
assembled from memory. Restore that trigger as a union with the review
or merge ask so the one consolidated rule covers both.
* fix(bin): let non-owner Claude Stops exit safely (#4777)
* Fix foreign-owner turn-end supervision loop
* no-mistakes(review): Scope foreign-owner safe exit to Claude guard
* no-mistakes(document): Document Claude foreign-owner safe exit
* fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778)
Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty
indexed array as an unbound variable and aborts the shell. In
signal_turnend_panes_churned() the missing_keys loop was reachable with
an empty array whenever every churned key already held a fresh
.churn-since-* marker (a second churning turn-end inside an open
deferral window), so each watcher cycle died about half a minute in and
supervision restarted endlessly. The created_keys rollback loops had the
same latent crash on their error paths.
Audit of bin/ for the same pattern found one more confirmed-reachable
case: remote_handoff's noncanonical-body scan iterates to_move, which is
empty when a retried remote handoff finds every key already staged in
the outbox. All other "${arr[@]}" sites are either count-guarded,
guaranteed non-empty by construction, or unreachable while empty.
Guard the three reachable expansions with the repo's existing
"${arr[@]+...}" idiom. New regression test drives a real watcher
through the all-marked churn path; the macos-stock-bash CI lane runs it
under real /bin/bash 3.2 via FM_TEST_ONLY.
* Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783)
The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: require complete captain-facing final responses (#4779)
* docs: require complete final responses across harnesses
* no-mistakes(document): Document complete final replies for Grok Bot
* docs: point Grok replies to the shared contract owner
* no-mistakes(review): Clarify final recap without batching decision asks
* fix: preserve substantive mid-turn text in Pi Calm (#4788)
* fix(calm): preserve substantive Pi mid-turn text
* no-mistakes(review): Preserve substantive Pi Calm text per block
* no-mistakes(test): Cover shared Calm preservation boundaries behaviorally
* no-mistakes(document): Consolidate Calm preservation documentation
* fix: harden mail checks and rebalance full-coverage CI (#4800)
* Improve CI reliability and rebalance full-coverage validation
* no-mistakes(document): Clarify lint partition documentation
* fix(bin): answer Kimi 2.0.0 folder-trust dialog during spawn (#4799)
* Handle Kimi workspace trust dialog
* no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers
* no-mistakes(review): Gate Kimi ready on any trust marker and clean captures
* no-mistakes(review): Read visible pane for Kimi trust and ready gates
* no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate
* no-mistakes(review): Harden Kimi viewport capture and trust dialog detection
* no-mistakes(document): Document Kimi spawn refusal on cmux and Orca
* fix(bin): report a dead-agent record once instead of escalating forever (#4775)
* fix(bin): report a record whose agent is gone once instead of escalating forever
The wedge escalation path never asked whether there was still an agent to be
wedged. A wedge is something stuck that might recover, so re-alarming it earns
its cost; an agent that is gone never moves again, its pane never churns, the
idle timer never resets, and the escalate path clears its own timer and re-arms
with nothing bounding the count.
Observed on a live fleet: two finished lanes reached 226 and 203 consecutive
escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400
notifications a day from two lanes with no agent running at all. On one,
fm-control.sh exit answered already-stopped and fm-crew-state.sh read
"failed - run failed". Closing the Herdr pane did not stop it either: with the
pane genuinely gone and herdr pane read returning pane_not_found, the count kept
climbing, because the poll is driven by the record's window= line rather than by
the pane. The cost is not the repetition but that it drowns the alarms that
matter.
fm_backend_agent_state already separates a thinking agent from a gone one at
process level. In the branch that was about to escalate, read it once and treat
only its two recovery-grade verdicts - dead (endpoint present, no agent in it)
and missing (endpoint authoritatively absent) - as proof, reporting that record
once and not re-escalating it while it stays that way. Every other verdict,
including alive, ambiguous, unreadable, unverified, and a read that failed
outright, keeps the identical schedule, reason, and escalation count, so a
genuinely wedged live agent is unaffected. The probe costs at most one backend
read per window per threshold, the same budget the declared-wait consult and the
worktree write probe already take.
The report decides nothing about the record's fate: both lanes still held
unlanded work and teardown refusing them was correct, so retiring, relaunching,
or cleaning up stays with the supervisor. The once-only marker is owned entirely
by that function and is dropped by the same read the moment the endpoint stops
reading gone, so a replacement launched into the same window escalates normally
and its own later death is reported again.
Related, and not closed by this: #4412, #4482, #4316.
Tests drive the real watcher against a record whose endpoint does not exist and
pin both directions: dead and missing report once and never advance the count
across later thresholds, while alive, ambiguous, and unreadable endpoints keep
escalating with the identical reason and a climbing count.
* fix(bin): bind the once-only dead report to the pane it reported
Review of the parent commit found a reachable sequence where a later death in
the same window lost its promised report. The marker was keyed on the verdict
string alone and dropped only when a threshold probe read a non-gone verdict,
but probes run only at thresholds: a replacement launched into the same window
that dies without ever being probed alive - it crashes at startup, or works and
then crashes - was absorbed by the previous death's marker. The pane's first
sight yielded only the generic stale wake and every later threshold matched the
stale marker, so the second death never got the detailed once-report that both
the function's own comment and docs/architecture.md promise.
Record the verdict together with the pane hash it was reported for, and absorb a
repeat only while both still match. A replacement churns the pane, which resets
the stale suppressor, wedge timer, and escalation count while no reset site
touches this marker, so the pane half is what tells the second death apart from
the first. The live-probe drop stays as it was.
Clearing the marker at those reset sites instead would re-open unbounded
re-alarming for a dead pane whose display ever ticks, which is the exact defect
the parent commit exists to close.
The noise bound is unchanged: an unchanged dead pane still absorbs on every
later threshold and never advances the escalation count, and every verdict short
of proof still escalates exactly as before.
* no-mistakes(review): Key the dead-record once-marker on the busy incarnation token
* no-mistakes(document): Document dead-record escalation cap in stale-pane config entry
* no-mistakes(document): Add busy-state inventory line to AGENTS.md
* no-mistakes(document): Document dead-record probe on busy-turn-bound wedge path
* fix(bin): create captain-hold rows when Beads requires due (#4854)
Captain holds have no due semantics and are a hold kind, not a Beads issue
type. The create path now waives due.required and maps to native type task.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: disable compact adviser for spawned agents (#4877)
* feat(bin): launch every spawned agent with the compact adviser disabled
Every crewmate, scout, and secondmate Firstmate launches now starts with
COMPACT_ADVISER_DISABLE=1, on a fresh spawn and on a relaunch alike, so an
unattended session never activates the compact adviser.
The value is unconditional: no configuration file gates it and there is no
override, unlike the trace carrier beside it.
Three carriers deliver it, because no single one covers every launch shape.
The pane shell receives an export beside GOTMPDIR, so the agent's own children
inherit it too.
The launch command carries an explicit assignment, prepended outermost so it
wins over any ambient value the pane already held.
The cleared launch environment sets it again at the `env -i` boundary and keeps
COMPACT_ADVISER_DISABLE in the fixed operational floor, which is what preserves
the switch when config/launch-env-allowlist empties the environment, and what
delivers it on a remote host that never had the value.
bin/fm-control.sh relaunch, the bootstrap secondmate relaunch, and the remote
secondmate transport all rebuild their launch through bin/fm-spawn.sh, so they
inherit the same floor.
The captain's own primary session is untouched.
The two new suites drive the real spawn and then execute the launch command the
pane actually received, with the harness replaced by a probe that prints its own
environment, rather than matching script text.
They cover ship and secondmate launches with the allowlist absent and enabled,
the pane export and its ordering, fm-control.sh relaunch, and the full parent to
remote-host chain.
* no-mistakes(review): Export compact-adviser disable across compound launches
* no-mistakes(document): Document spawned-agent compact-adviser environment guarantee
* fix(bin): preserve Claude lock ownership after helper recycling (#4894)
* fix(bin): let a background Claude session keep owning its session lock
Session-lock ownership was decided by process ancestry alone. Under an
unattended Claude session the model loop runs in a transient bg-spare
bridged to the front-end by a shared daemon; when that bridge is
recycled the contiguous claude-named ancestry from a hook to the
recorded owner breaks while the owner pid stays alive, so the Stop
auto-arm stood down as a foreign live owner, the turn-end guard ended
every turn with its read-only diagnostic, and fm-lock.sh refused - a
self-sustaining outage until restart.
Ownership is now ancestry membership OR a trusted same-session id,
never id-first:
- fm-session-lock-lib.sh accepts CLAUDE_CODE_SESSION_ID only when
CLAUDE_PID is a Claude-shaped member of the current contiguous run,
compares it against the id recorded in state/.lock-session, and
requires the recorded pid to still be a live harness. No id, no
sidecar, an untrusted id, a different id, or a dead recorded pid
leaves the ancestry verdict unchanged. Ids are never read from ps
argv.
- fm-lock.sh accepts a same-session holder at both refusal sites,
writes, refreshes, and clears the sidecar only under its claim lock
(including the early already-mine exit, skipped only while the
deferred startup sweep leases that lock), keeps it byte-identical
across a same-session confirmation, records CLAUDE_PID on lock line 1
for a session with a trusted id so a shared daemon or front-end that
outlives the session never keeps a dead session's lock alive, never
rewrites a live line 1 on a same-session confirmation, and names the
recorded id in the live-owner refusal.
- The .lock line-1 format is unchanged, so every reader that takes the
whole first line as the pid keeps working; the guard's foreign-owner
exit is unchanged and inherits the fix through the shared predicate.
Tests: the ancestry suite drives the ancestry and id signals apart in a
deterministic process table (asserting the divergence) and runs a real
orphaned front-end/daemon/pty-host/spare tree through six phases with
the real lock, auto-arm, and guard scripts; the foreign-owner repro
keeps its negative control and adds a same-id positive control.
Disclosure: no live unattended Claude background session ran on the
verifying machine. The topology is documented by the real process
listings in #3902, #2314, #3398, and #4066; coverage is the structural
predicate plus the executable fixtures, not a live pass.
Residual: bin/fm-sessionstart-nudge.sh keeps its own private ancestry
walk (it only decides whether to print a nudge) and may nudge on a
resume in the recycled case.
Out of scope, deliberately: no structured lock format, no guard budget
changes, no daemon-identity rejection, no fork lineage.
* no-mistakes(review): Wait for claim lock; revert failed sidecars
* no-mistakes(review): Revalidate ownership after wait; restore sidecars
* no-mistakes(review): Roll back sidecar by publication phase
* no-mistakes(review): Restore sidecar only if lock line is unchanged
* no-mistakes(review): Trust session ids without a spelling allowlist
* no-mistakes(review): Disarm sidecar rollback before backup cleanup
* no-mistakes(document): Updated session-lock ownership documentation
* feat: park main under the away posture on Pi (#4889)
* feat: park main under the away posture on Pi
While the away-posture record exists on a Pi primary, the supervision branch
takes every actionable wake, no processing turn opens on main, captain rows
accumulate for the return brief, and main's standing authority relocates to
the branch through the existing guarded scripts.
- lib/fm-branch-dispatch.ts: read the record at every routing decision; while
it exists claim check, decision-owned, and heartbeat rows too, keeping the
two broken-queue vetoes; expose checkSeqs so a claimed check row lifts task
scoping.
- fm-primary-pi-watch.ts: offer every actionable row under the record; a
declined wake and every watcher-failure alarm still reach main.
- fm-branch-supervision.ts: drop the legacy .afk decline; append a fixed
POSTURE: AWAY tail carrying the record's read-back verbatim per wake; open no
processing request while the record exists, re-checked immediately before a
request would open and at every run boundary; present the accumulated rows
at the first run boundary after archive.
- fm-lease-lib.sh: fm_lease_forbid_branch passes the branch for opted-in
actions only while fm-afk-contract.sh validate succeeds on a confirmed live
record; PR merge, fresh spawn, and decision answer opt in, local landing
never does.
- fm-send.sh: a --resolve-key naming an open needs-decision or captain-held
task is a decision answer and meets the partition; blocked: keys stay
steering.
- fm-spawn.sh: enforce the record's spend cap for a fresh ordinary spawn by
either actor; relaunches and secondmates exempt.
- fm-branch-prompt.sh: fixed Postures section and the verbatim
ask-user-authority policy; the prefix stays byte-stable.
- fm-afk-return.sh: count what the away session handled from the store.
- docs, afk skill, AGENTS.md stub: main parked on Pi, green merge gate
absolute while away.
- tests: watcher and branch extension suites, fleet-record, merge, and
decision-answer suites cover the relocation, the vetoes, the tail, the
parked processing turn, the cancellation, the re-presentation, and the
spend cap; dated live-guard evidence recorded.
* no-mistakes(review): Refuse branch merge after preflight archive race
* no-mistakes(review): Fix away wake, spawn, and processing races
* no-mistakes(review): Suppress parked processing; narrow away-only rejection
* no-mistakes(review): Abort dedicated processing; gate branch spawn once
* no-mistakes(review): Stamp away-only on the dispatch offer
* no-mistakes(review): Treat invalid away records as spend-cap absence
* no-mistakes(review): Drop spawn test hook; abort processing-opened runs
* no-mistakes(review): Bind abort to opening prompt; cap-read absence
* no-mistakes(review): Limit away branch spawn to queued work only
* no-mistakes(document): Correct AFK posture documentation
* ci: standardize workflow timeouts into three tiers (#4910)
* ci: simplify CI job timeouts to a three-tier policy
Replace the scattered per-job timeout values (10m parallel, 25m lint, 30m
serial, 10m macOS) with three readable tiers, each a hang tripwire with
headroom rather than a packing estimate:
- fast (5m): coverage guard, repo invariants, timing aggregate
- normal (30m, one shared budget): lint partitions, portable parallel
shards, portable serial shards, macOS stock Bash
- heavy (Herdr only): 20m step tripwire on the family run so always()
cleanup still runs, under a 75m job-level last-resort backstop
The workflow's header comment states the policy and points at
docs/fm-test-portable-shards.md "Timeouts", which now owns it, and each
job names its tier beside timeout-minutes. tests/fm-ci-workflow.test.sh
asserts the policy against the parsed workflow instead of the old
per-job minute values: every job joins exactly one tier, exactly three
distinct job-level values exist, the fast tier stays within 5-10
minutes, the normal budget stays at least double the modeled parallel
lane sum reported by fm-test-run.sh --check-coverage, and the Herdr step
tripwire stays below its job backstop with an always() cleanup after it.
Concurrency supersession, shard counts, lane membership, and fail-fast
settings are unchanged.
* no-mistakes(review): Decouple the normal timeout from packing estimates
* no-mistakes(review): Assert Herdr teardown follows the family run
* no-mistakes(review): Pin Herdr family-run timeout to 20 minutes
* no-mistakes(review): Ignore comments when identifying Herdr steps
* no-mistakes(review): Identify Herdr steps by declarative ids
* no-mistakes(document): Clarify authoritative three-tier timeout policy
* fix(bin): keep supervisor status closes from waking the same home (#4895)
* fix(bin): keep supervisor status closes from waking the same home
A drain that already folded OPEN DECISIONS has presented those bytes even
when the watcher has no matching seen marker. Treat that fold, and the
presentation cursor, as known so the bookkeeping close stays quiet while
later worker lines still signal.
* no-mistakes(review): Keep folded worker failures waking past supervisor closes
* no-mistakes(review): Wake on unlisted folded worker lines; batch multi-key closes
* no-mistakes(review): Stop folded worker resolved lines from counting as already read
* no-mistakes(document): Correct self-announced close marker contract in docs
* fix(bin): stop labeling Herdr as experimental (#4972)
* Stop steering operators away from Herdr
* no-mistakes(review): Neutralize remaining Herdr opt-out documentation wording
* fix(bin): treat a live no-mistakes run as current after rebase (#4973)
* fix(bin): treat a live no-mistakes run as current after rebase
A running run on the task's branch is authoritative regardless of head.
Matching only the local head made a rebased in-flight run look failed.
* no-mistakes(review): restrict coarse live-any-head to foreign-branch answers
* no-mistakes(review): reject gate-parked runs from the executing predicate
* no-mistakes(review): hoist gate-marker patterns into single run-lib owner
* no-mistakes(review): require live daemon for head-free run binding
* no-mistakes(review): require answered daemon-down before unbinding live runs
* no-mistakes(review): extend daemon guard to anchored continuation routes
* no-mistakes(review): delete live-any-head; restore dead-daemon verdict
* no-mistakes(review): keep parked gates parked; name dead daemon everywhere
* no-mistakes(review): set dead-daemon verdict instead of emitting early
* no-mistakes(review): align selected route with legacy dead-daemon handling
* no-mistakes(review): drop unproven-record binds; narrow coarse gate reading
* no-mistakes(review): narrow header, drop vestigial guard, retarget tests
* no-mistakes(review): revert coarse gate override; require answered-down probe
* no-mistakes(review): cache one daemon probe; stop duplicating run id
* no-mistakes(review): restrict coarse dead-daemon verdict to moved-off rows
* no-mistakes(review): delete coarse dead-daemon extension and gate note
* no-mistakes(review): delete remaining coarse dead-daemon block and stale docs
* no-mistakes(document): document rebase-safe live-run bind and unverified-record verdict
* fix(bin): prevent long worker launch command truncation (#4994)
* fix(bin): stage the launch command in a private file and type a short source line
A long launch line typed while the fresh pane shell is still busy waits in the
terminal's canonical line buffer, which drops input past about 1,024 bytes on
macOS, so the pane was left at an unfinished command with no agent running.
fm-spawn now writes the assembled command to the task's own temp root under
umask 077 and types only a short line that sources it.
Refs #4559
* fix(bin): keep the per-task temp root private before staging the launch command
The root lives at a predictable path under /tmp and now holds the whole launch
command. Create it with mode 0700, refuse one that already exists as anything but
a directory owned by this user that nobody else can write, and tighten an owned
one, so no other local user can plant or swap the staged file.
Refs #4559
* fix(bin): enforce private staged launch file mode
* test(spawn): cover long staged Claude launches
* no-mistakes(review): Namespace launch files and prove truncation staging
* no-mistakes(review): Use immutable per-spawn launch filenames
* no-mistakes(document): Document staged launch delivery safeguards
* no-mistakes(ci): Updated eight behavior tests/fakes to execute or inspect immutable staged launch files instead of expecting inline launch commands. This restores Muse, secondmate lifecycle/restart, remote trace/parent binding, compact-adviser, and Orca coverage. All affected tests, dispatch-profile regression, fixture tests, syntax checks, ShellCheck, and git diff checks pass
---------
Co-authored-by: Vytautas Stankus <svycka@gmail.com>
* test: authorize isolated Herdr lab validation (#4998)
* Add isolated Herdr runbook to test instructions
* no-mistakes(review): Drop substring matching from test.instructions contract
* no-mistakes(review): Assert commands.test key absence in YAML
* Drop unit-first sentence and instructions contract test
Captain-scoped follow-up on the Herdr-lab test.instructions ship:
keep the lab safety runbook only, and leave the no-mistakes contract
test focused on commands.test absence.
* docs(vision): accept vendor-semantics and 9k AGENTS ceiling (#4873) (#5001)
* docs(vision): accept vendor-semantics and 9k contract-ceiling amendments (#4873)
Replace the pixels-of-today's-UI rule with a quarantined, version-pinned
surface-adapter exception recorded as standing debt. Cap the always-loaded
contract at 9,000 words and require prune-or-trigger before a crossing change
lands.
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
* docs(vision): restore accepted three-sentence vendor-semantics form (#4873)
Replace the compressed paraphrase with the issue's accepted wording:
a named quarantined version-pinned adapter, expected to break, recorded
as standing debt that never hardens into a shared contract.
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
* feat(bin): defer the wedge escalation for a lane parked at a supervisor-owed gate (#4974)
* fix(watch): recheck a gate awaiting a human instead of wedge-escalating it
A lane whose validation run is parked at a gate waiting on a human
decision is correctly quiet, but nothing in its status line says so: the
evidence is the pipeline's own gate state rather than anything the worker
wrote. The wedge timer read that silence as a suspected wedge and climbed
the escalation ladder for as long as the wait lasted, and each escalation
cost a supervising turn. The landed declared-wait consult does not reach
it, because a live ordinary crewmate never reports a declared pause, and
raising FM_STALE_ESCALATE_SECS would delay genuine wedge detection for
every lane by the same amount.
The threshold now reads a second, independent record when the status line
accounts for nothing: whether the crew's current state is a gate whose
answer is owed by a human. That is minted only from the gate's own
findings table, by a row whose `action` column is exactly `ask-user`,
located by position out of the table header the way nm_gate_step_row
already reads its row - never searched for over the run payload, where a
finding's free-text description or a branch name satisfies a search just
as well. A gate awaiting the CREWMATE's own answer keeps the unchanged
escalation schedule, reason and demand-deep-inspection wording, because a
crewmate that goes quiet before answering its own gate is exactly the
wedge the ladder exists to catch.
Each kind of wait now carries the human it is on, the action that clears
it, and whether that human is the captain as data alongside the verdict,
rather than as wording chosen per branch where the recheck is written, so
the deferral cannot word one kind of wait as another and a new kind
cannot ship without deciding all of them. A parked gate has no written
record of when its wait began, so its recheck publishes no wait age at
all rather than one read from the quiet window this deferral resets on
every pass, which would report the same small number for a gate of any
age. Like every other captain-facing recheck here it is absorbed in
silence while the away-posture record exists, arming no throttle, so the
recheck is owed in full the moment the record is archived.
The consult runs only in the at-threshold branch that was about to
escalate, beside the worktree walk already there, and only for lanes
whose status line explained nothing.
Closes #3055
* no-mistakes(review): require an unanswered decision before deferring a parked gate
* no-mistakes(review): reset the away-silenced timer, fail-safe findings parse, US-joined wait records
* test(watch): pass the pane hash wedge_timer_check now takes
Upstream gave wedge_timer_check a sixth <pane-hash> argument for its
dead-record probe. The malformed-wait-record rounds drive the real function
directly, so they pass one, and stub fm_backend_agent_state to a live agent so
the probe that runs after a refused deferral keeps the unchanged ladder rather
than reading a backend the child shell has none of.
* no-mistakes(review): Bind parked-gate wait to its run, owe it firstmate
* no-mistakes(document): correct wait-kind count, crew-state reader scope, gate-key coupling
* feat(watch): make the parked-gate wait deferral opt-in
The wedge timer deferring a lane parked at a validation gate is new
supervision behaviour rather than a restored one, and it decides which
lanes give up the escalation ladder, so it now ships as a default-off
per-home option instead of changing every home on upgrade.
config/wedge-defer-parked-gate arms it. The flag is read before the
decision fold, so an unconfigured home spends no fold or current-state
read, writes no record, and keeps the unchanged escalation schedule,
reasons and demand-deep-inspection wording; a test counts the reader
calls in both directions to pin that.
It is not inherited by secondmate homes: each home supervises its own
crew and owns that trade separately, the same reason
config/turnend-churn-absorb is home-local.
The away-posture absorb returns to leaving the idle timer alone, which
it had restarted only because the costly consult could reach it. A
parked-gate wait is owed to the supervisor rather than the captain, so
it never enters that branch, and the recheck owed on return is again
owed in full the moment the record is archived.
* test(watch): pin that the away-silenced hold leaves the idle timer alone
The absorb no longer restarts the timer, so the recheck owed on return is
owed in full rather than a cadence into the return. Nothing asserted
that, so a restart could be reintroduced silently.
* no-mistakes(review): document away-silence rationale, pin captured gate component
* no-mistakes(test): anchor gate row scan to the braced findings header
* no-mistakes(document): pin same-block gate row invariant in crew-state comment
* fix(bin): reclaim a task whose herdr endpoint was destroyed (#5007)
* fix(control): let the owning seat reclaim a task whose endpoint is gone
A destroyed pane or workspace made `missing` a terminal state. Relaunch
accepted only `dead` and said to stop the agent first; exit refused
`missing` and said to reconcile the task first; there is no reconcile
verb. Each command named the other as its prerequisite, so a task whose
terminal went away could not be reclaimed by anything, and a no-mistakes
approval it was parked on had no seat left to answer it.
`missing` is agent-free a fortiori: there is no endpoint, so there is no
agent in it. Widen the existing guards rather than add a verb.
- fm-spawn --relaunch accepts a positively proven `missing` and creates
one fresh endpoint in the recorded worktree; the record it already
republishes rebinds the task to it. A `dead` endpoint is still adopted
in place.
- fm-control exit reports `endpoint-gone` instead of dying, so the
relaunch transaction's stop step no longer dead-ends, and re-resolves
the endpoint from the record before verifying the replacement.
The duplicate-agent refusal is untouched: both verdicts come from the
same recovery-grade classifier, which claims `missing` only from positive
absence, so `alive`, `ambiguous`, and `unreadable` all still refuse. The
backends' own create paths refuse a live same-labeled endpoint as a
second independent guard. The worktree, its branch, commits, uncommitted
changes, armed poll and registration, record rows, and status log are all
untouched - a reclaim is a recovery, never a teardown.
A secondmate is excluded: its gone-endpoint recovery already has one
owner in the session-start liveness sweep, so relaunch refuses and names
it rather than becoming a second path to the same outcome.
Tests reproduce both halves of the deadlock, the reclaim succeeding,
unlanded work surviving it, and the refusals that still hold.
* no-mistakes(review): prove endpoint absence per backend before reclaim rebinds
* no-mistakes(review): give exit and relaunch one absence proof; pin herdr rebind session
* no-mistakes(review): narrow endpoint reclaim to herdr; tmux refuses honestly
* no-mistakes(review): stop refusals and docs asserting unestablished causes
* no-mistakes(review): stop herdr fixture helper losing tmp-root registration
* no-mistakes(review): document workspace drift and absence-probe server residue
* no-mistakes(review): correct rebind limitation to its one reachable case
* no-mistakes(review): stop claiming reclaim leaves instructions untouched
* no-mistakes(document): scope fm-control-lib purity claim, note reclaim coverage
* no-mistakes(rebase): read the staged launch file in the herdr fixture
Rebasing onto main picked up #4994, which stages a long worker launch
command into a script and delivers the short `. '<path>'` line instead of
the literal command. The tmux fake and tests/fixtures.sh were updated for
that; the herdr fake this branch adds was written before it and still
keyed "an agent now exists on this pane" off the literal
`encode launch-brief` text, so after the rebase it never marked the
rebound pane live and the reclaim's alive-wait read `dead`.
Dereference the staged file first, exactly as the tmux fake above does.
Test-fixture only; no production path changes.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* no-mistakes(document): note reclaim placement in herdr and scripts inventories
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat(bin): stamp status events with their emission time (#3764)
* test(status): reproduce missing event emission time
* wip(status): preserve optional event emission time
* test(status): document indirect clock stub invocation
* no-mistakes(review): Preserve historical status bytes during reply recovery
* no-mistakes(test): Fix timestamped status assertions and remote fixture dependencies
* no-mistakes(review): Preserve captain regex overrides for timestamped status events
* no-mistakes(document): Clarify status event timing and publication contracts
* no-mistakes(lint): Quote literal done to satisfy ShellCheck
* no-mistakes(ci): Captain, updated .github/workflows/ci.yml to expect 19 snapshot tests instead of 18, matching the PR’s added regression. Reproduced the failure before the fix. Stock Bash 3.2.57 verification passed: parse sweep, 19 snapshot tests, 53 Bearings tests, and the public-followup regression. Workflow lint and diff checks passed
* no-mistakes(test): Preserve terminal notifications with malformed timestamp tags
* no-mistakes(test): Stamp Rovo spawn failures with emission time
* no-mistakes(document): Verify status event documentation
* no-mistakes(lint): Fix ShellCheck quoting in status emission-time tests
* no-mistakes(ci): Captain, fixed four lifecycle assertions to accept emission timestamps while preserving publication and retry checks. Reproduced the CI failure before the fix. The lifecycle suite now passes with six Beads capability skips; syntax, targeted ShellCheck, and diff checks passed
* no-mistakes(ci): Captain, fixed malformed timestamp colons hiding actionable events using shared normalization. Original bytes and unknown ages are preserved. Regression reproduced before the fix; classifier and remote-reply suites, targeted lint, syntax, and diff checks passed
* no-mistakes(review): Stamp remote escalations at call sites, drop new flag
* no-mistakes(review): Accept stamped escalation and close lines in test assertions
* no-mistakes(review): Restore reserved-key answered-note guard for stamped closes
* test(status): accept optional emission time in PR-provenance assertions
The #4148 provenance test landed on main with exact unstamped greps.
Parent-channel lines from this branch carry [at=<epoch>], so strip only
that tag before the same exact match. No production change.
* no-mistakes(review): Accept stamped ready signal in PR fallback scrape
* no-mistakes(review): Drop relay flag, stamp parent events at call sites
* no-mistakes(review): Stamp worker terminal-signal instructions, revert fm-on fixture
* no-mistakes(review): Accept optional stamp in live cmux drift guard
* no-mistakes(review): Restore original test invocation order in two suites
* no-mistakes(review): Strip only well-formed numeric status time tags
* no-mistakes(document): Drop stale unstamped PR-ready line spelling from channel doc
* no-mistakes(review): Stamp agy spawn-failure status lines with event time
* fix(bin): normalize status event times in-shell and freeze the budget test clock
Two paths made a status event's emission time cost more than it should.
The captain-relevance fallback piped every line through awk to drop a
well-formed `[at=<epoch>]` tag before matching, so a supervisor sweep paid a
fork per line just to prepare a regex match. Shell parameter expansion does the
same strip with no fork, and the retry-dedup scan now reuses that one helper
instead of carrying a second copy of the rule in awk. The copies had already
drifted: the shell side stripped tags from lines with no colon, which the awk
rule left whole, so a colonless line could be mistaken for one already
recorded. One definition, checked against the awk rule it replaces over the
edge cases and a 4000-line fuzz.
tests/fm-contributions.test.sh froze its fixture clock only in exhaust mode. In
hang mode the poll set DEADLINE to the real now plus a one-second budget, and
when the second ticked before the first forge call the loop broke without ever
calling gh: forge/calls was never written and the assertion failed reading a
missing file. Freezing the clock in both modes removes the dependence on wall
time; the bounded call is still cut by the real timeout, so the observation the
test asserts still starts.
Emission time stays optional on new status records, and legacy or malformed
lines keep an unknown age.
* no-mistakes(review): Stamp ask-user escalation line and fix Kimi status assertion
* no-mistakes(document): Drop stale unstamped done-line spelling from watcher docs
* test: fold emission-time snapshot coverage into the fixture case
Drop the incidental ci.yml 18-to-19 count hunk so the PR no longer
touches workflows. Keep every emission-time assertion by folding it
into test_fixture_snapshot_json.
* no-mistakes(review): replace brief date substitution with epoch placeholder; drop emitted_at_epoch
* no-mistakes(review): align untimed normalizer with epoch parser; tolerate placeholder stamp in PR scrape
* no-mistakes(review): strip undelimited at-tags; correct brief stamp header
* no-mistakes(review): normalize stamps at both captain-regex sites; restore mtime freshness
* no-mistakes(review): strip colon-bearing stamps for relevance; fix headers and test oracles
* no-mistakes(review): narrow escalation match to stamp tolerance; pin note verb
* no-mistakes(review): read note and key past colon-bearing stamps
* test(status): keep inactive reconcile assertions stamp-tolerant
These two oracles were made stamp-tolerant while resolving one of the
branch's merges from main. The rebase drops merge commits, so that
adaptation was lost and both assertions went back to matching an exact
substring that a stamped line no longer contains: the tag lands before
the colon, so "failed [key=k]: ..." is now "failed [key=k] [at=N]: ...".
Strip a well-formed tag before matching, as the branch's other oracles do.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* no-mistakes(review): unstamp fold colon tests; reserve stamp width in cap
* no-mistakes(document): correct stale unstamped status-line spellings in docs
* no-mistakes(document): quote brief-test literals for lint; correct stamp-helper contract comments
* no-mistakes(ci): rename subshell-local epoch in delivery-race stub
The serialization test overrides fm_pending_reply_mark_delivered inside a
(..) subshell. Its `epoch` local collided with the same name in
status_line_at_epoch/status_stamp_line, which this branch added and this
suite now calls at top level, so ShellCheck 0.11.0 reported SC2030 and
failed Lint 2. The stub already prefixes its other locals with `pending_`
for the same reason; `epoch` was the leftover.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(bin): unify Lavish host and disconnect handling (#5060)
* fix: ship clean Lavish host fixes
* no-mistakes(review): Fix Lavish classifications and fail-closed host loading
* no-mistakes(review): Restore Lavish host state across retries and launches
* no-mistakes(review): Preserve destination Lavish host when configuration is absent
* no-mistakes(document): Document Lavish status and host guarantees
* feat: act on captain's away words during AFK supervision (#5076)
* feat(afk): make the captain's away words the whole mandate
Retire the clause fields, verb list, never-set scan, refused records, and
the per-task merge-grant list from the away-posture record. The record is
now version 2: the captain's words verbatim plus expected return, spend
cap, and reach line; a version 1 record still validates, reads, and
archives so a live away window is never broken by the upgrade.
The supervision branch reads the words at the tail of every wake and acts
on them by its own judgment through the guarded scripts under standing
authority, never by analogy, holding for the return on doubt, and opens
each such outcome summary with "per your away instructions:" so the
return brief can render the words beside the session's account. While the
record exists any green merge runs under away authority (ledger tag
"away"); red merges, --allow-red, asynchronous and queued merges, and
local-only landing stay refused. The branch may file a backlog item the
words explicitly call for before dispatching it under the spend cap.
Tests drive fm-afk-contract.sh, fm-afk-launch.sh, fm-afk-return.sh, and
fm-pr-merge.sh as commands: version 2 written, version 1 read, retired
flags and subcommands refused by name, green merges landing under the
record, red and waived-red refused, the record lock still closing the
authority-read window, and the Pi away tail carrying the words.
* no-mistakes(review): carry the away read-back to the session verbatim
* no-mistakes(review): match the exact away-action marker in the return brief
* no-mistakes(review): refuse a words block truncated by a damaged line
* no-mistakes(document): Refresh away-role contract documentation
* fix(bin): render the remote charter's steering-inbox path host-local (#5049)
* fix(bin): render the remote charter's steering-inbox path host-local
A freshly provisioned remote secondmate read a parent-home absolute
steering-inbox path in its charter - a location that exists on no route -
and spent its first turn discovering the gap and filing a blocked
decision for what was a render defect. The seed's remote-copy rewrite now
maps the inbox to the route's host-local parent-route inbox, exactly as
it already maps the reply-log path, so every mention - bare path, listing,
and handled/ acknowledgement - lands host-local.
Both rewrites also become plain assignments, because a quoted substitution
nested inside a double-quoted printf argument leaks literal quotes into
the replacement text on stock macOS bash. The lifecycle suite pins the
corrected render both directions against the real seed, provisioning,
and delivery route, sharing one fixture value between the render truth
and the delivery truth.
Closes #5012
* no-mistakes(document): document remote charter's host-local steering inbox
* feat: route Lavish feedback directly to owning workers (#5099)
* feat(procevent): route worker-owned Lavish rounds
* no-mistakes(review): drop duplicate artifact field from task-owned registration
* no-mistakes(review): post worker reply once, fix ring label, keep re-arm atomic
* no-mistakes(review): keep worker board owned until terminal round acknowledged
* no-mistakes(review): refuse every retirement of an open worker-owned round
* no-mistakes(review): use real lavish reply flag, isolate reply generations
* no-mistakes(review): drop .posted marker for best-effort reply posting
* no-mistakes(review): consume staged reply after listener setup, refuse orphaned captures
* no-mistakes(review): require a reachable owner, redeliver open rounds, roll back failed re-arms
* no-mistakes(review): re-arm only to acknowledge an open round
* no-mistakes(review): conclude only a still-open terminal round
* no-mistakes(review): record the acknowledgement before retiring the board
* no-mistakes(review): retain the registration across a conclude, qualify terminal docs
* no-mistakes(document): Document worker-owned Lavish round lifecycle
* fix(bin): fit pull observation within the contribution poll budget (#5107)
* fix(bin): reserve contribution observation budget
* no-mistakes(review): Strengthen slow-read regression test to exceed the poll budget
* feat(bin): add idempotent inbox capture, replies, receipts, and readiness JSON (#5103)
* feat(bin): add idempotent inbox orders, receipts, replies, and readiness
Let a caller supply a request id when publishing a captain inbox note so a
retry returns the original note instead of creating a second one, including
across the crash window between save and wake announcement. Separate saved
from announced so a failed wake is repairable without enqueueing again.
Add bounded receipts JSON with omission disclosure, a durable primary reply
against a note id, and a read-only readiness projection that can say
unknown instead of inferring liveness from a lock file.
* no-mistakes(review): fix(bin): honest inbox announce, reply cursor, and readiness verdict
* fix(bin): resolve ready from lock-holder ancestry; drop lock status --json
Remove the extra JSON surface from fm-lock.sh so its human status still
always exits zero. Have the readiness projection classify the inspected
home from the lock-holder pid via fm-harness.sh ancestry, with an explicit
FM_SUPERVISION_MODEL still winning and an unknown model when there is no
holder. Prove the yes path when that ancestry names a known harness.
* no-mistakes(review): Harden inbox announce, receipts reads, and reply sequence cursor
* no-mistakes(document): Note read-only lock inspection in scripts inventory
* no-mistakes(lint): Pass missing id argument to malformed-reply test printf
---------
Co-authored-by: cliflacata-svg <304148223+cliflacata-svg@users.noreply.github.com>
* fix(bin): stop harness footer rows below a composer from reading as pending text (#5118)
* fix(composer): stop a harness footer row from reading as a composer holding text
A harness draws its own furniture below the composer - a user statusLine, a
permission-mode hint - and the cursorless "bottom-most shape wins" rule looks
exactly there. `→` (U+2192) is Cursor's prompt glyph but ordinary text
everywhere else, so a statusLine opening with `→` was selected as a bare
composer, swallowed the hint row beneath it as wrapped input, and answered
`pending` on a visibly empty pane. `fm_task_inbox_ring` defers on exactly that
verdict, and `bin/fm-watch.sh`'s re-ring calls the same function, so the first
doorbell and every retry were skipped and the worker never saw the steer.
Measured live on 2026-09-20: three of five Claude Code 2.1.236 worker panes on
Herdr 0.8.0 had genuinely empty composers and every one of them was refused.
A separator pair that closed over a bare agent-glyph row is a proven composer
container, so the contiguous non-blank rows below its closing rule are that
composer's footer and are no longer composer candidates. The demotion is bounded
by all three of its own preconditions: a blank row ends the zone, a pair that
closed over no glyph row demotes nothing, and a shape with no separator pair at
all (Cursor's half-block rules) is untouched. Real unsubmitted text in that same
composer, including a stray SGR mouse report left by a click in the pane, still
reads `pending`.
Pinned by two portable regressions and by a new cursorless arm on the live
composer-matrix guard, which re-reads each harness's already-proven-idle pane
the way every non-tmux backend reads it and fails naming the harness and
version when that read is `pending`.
* no-mistakes(review): make composer footer-zone demotion shape-independent
* no-mistakes(review): make footer-zone demotion refuse-only and drop rescan
* no-mistakes(lint): quote probe-absent sentinel to clear ShellCheck SC2100
---------
Co-authored-by: Koen Muller <koen@catapult.nl>
* feat(bin): append optional home-local include to briefs (#5115)
Co-authored-by: guanchengh-lgtm <271917158+guanchengh-lgtm@users.noreply.github.com>
* fix(bin): report a branch with no validation run as absent instead of an unreadable runs table (#5114)
* fix(bin): stop misreading a no-run branch as an unreadable runs table
Defect: when `no-mistakes axi status`'s overview is truncated (a task's
own branch has zero rows among the shown ones), fm_nm_select_run's
Python fallback derived the repo identity for its direct SQLite query
from a `repo: <path>` line it expected in the overview text. The real
CLI never emits that line, truncated or not (see the genuine capture at
tests/captures/no-mistakes-v1.70.1/overview.toon, which has only
`count:`/`runs[...]:`), so the lookup always failed and reported
"unreadable runs table" for a task that simply has no run on its
branch. On a fleet with many concurrent runs, every idle-branch task
hits the truncated-overview path routinely, so this fired every few
minutes and drowned genuine unreadable/blocked verdicts in noise.
Fix: derive the repo identity from the task worktree path instead,
which is exactly the value `no-mistakes` records as a repo's
`working_path` (confirmed against the existing capped-overview test
fixtures, which already register repos by worktree path). A worktree
path that is not absolute cannot be matched and still reads as
unreadable rather than being guessed at. Also raise the reader's
SQLite busy timeout from 1s to 30s so ordinary lock contention on a
busy fleet cannot masquerade as an unreadable database.
Safety: every other verdict byte-for-byte unchanged - the repo lookup
still requires exactly one matching row (a genuinely corrupt or
mismatched repos table still reports unreadable, per the existing
`repo` failure-mode test), the branch query and row validation are
untouched, and a zero-row result for the branch still flows through
the same recursive re-parse that already turns an empty `runs[0]{...}`
table into `absent`. Added a regression test
(test_capped_overview_without_repo_line_and_no_runs_reports_absent)
that reproduces the real overview shape - capped, zero rows for the
task's branch, no `repo: ` line - and asserts the crew state falls
through to the pane/busy verdict instead of reporting unknown or
"unreadable". Full fm-crew-state.test.sh suite passes unchanged
otherwise.
* fix: recovered same-branch inventory awk misreads empty result as unreadable
fm_nm_select_run's deep SQLite reader rebuilds a `count:`/`runs[...]:`
overview and re-runs it through the same awk selection pass. When that
rebuilt inventory has zero rows for the branch, the row-matching loop never
executes, so its counters (`seen`) stay at awk's uninitialized empty string
while `expected` and `shown` are plain strings parsed from the header text.
Comparing an uninitialized value against a non-numeric string uses string
comparison, so "" != "0" is true, and the END block takes the "unreadable
runs table" branch instead of falling through to the correct "absent"
verdict for a branch with genuinely zero runs.
Coerce the affected END comparisons with `+0` so they are always numeric,
matching seen/expected/shown/total regardless of whether awk classified
them as strings or numeric strings. A truncated or genuinely malformed
inventory still differs numerically and still reports unreadable.
* no-mistakes(review): bound capped-overview inventory reader and canonicalize worktree lookup
* no-mistakes(review): match recorded repo path first, tolerate duplicate spellings
* no-mistakes(review): revert repo lookup to exact working_path match
* no-mistakes(document): note state-db inventory read under crew-state nm timeout
* fix(bin): require a non-draft pull request before a PR-based done report (#5141)
* fix(bin): require a no…
KaranSantra
added a commit
to KaranSantra/firstmate
that referenced
this pull request
Sep 23, 2026
) * fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586) * fix(watch): honour a declared wait before wedge-escalating a quiet pane wedge_timer_check escalated on elapsed idle time alone. Nothing asked whether the worker had already said why its pane was quiet, so a lane that declared a bounded external wait climbed the escalation ladder for as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every repeat carried demand-deep-inspection - which by its own wording forbids re-absorbing on the run-step or pane state, so the supervisor could not use the evidence that was there either. The generated brief promises that declaring `paused:` buys the long recheck cadence instead of a wedge, but the timer was still reachable while that declaration stood: a crew that declares a wait and then has an active run or busy pane attributed to it is handed to the timer as provably-working. The declaration is what the worker said about its own silence, so it now outranks a liveness verdict that only says something is running. The consult runs in the at-threshold branch that was about to escalate, beside the worktree walk already there, and costs one status-line read. Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS recheck the declared-wait absorber already uses, so the wait is still rechecked and cannot rot invisibly. Which verb declared it decides the wording, because the two block on different people: a `paused:` wait is owed by an external dependency and asks the reader to confirm it still holds, while a `captain-held:` transfer is owed by the captain reading the recheck and asks them to answer or release the hold. A hold is not rechecked at all while the away-posture record exists, as on every other captain-held path, and that absorb arms no throttle so the recheck is owed in full on return. A declared clearing time that has already passed stops counting, and a lane that never declared one keeps the identical escalation schedule, reason, count and demand-deep-inspection wording, so detection and its worst-case time are unchanged. The deferral restarts the idle timer rather than cancelling it, so a lane that stops waiting escalates again within one threshold. A lane quiet because its own validation run is parked at a gate awaiting a human decision is deliberately out of scope: reading that state needs a signal carrying who the wait is on and what clears it, rather than one inferred from a parked verdict that also covers gates awaiting the crewmate itself. Tests pin both directions for each case and were each confirmed to fail with the consult removed. * no-mistakes(document): docs: honour declared waits in stale-escalation docs * fix(bin): report verified PR state for passed runs (#4624) * fix(bin): derive passed PR state from PR record A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists. For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim. Fixes #4607 * no-mistakes(review): Add bounded GitLab merge-request state reads * no-mistakes(review): Preserve network-free inactive crew-state scans * no-mistakes(document): Document PR record readers in shared library * fix: restore published contribution follow-up (Fixes #4469) (#4627) * fix: restore published contribution follow-up (Fixes #4469) * fix(review): Fix contribution freshness and merge actor routing * fix(review): Restore issue triage and scope contribution follow-up * fix(test): test: assert one wake per contribution signal * fix(document): Document contribution follow-up * fix: restore truthful terminal delivery evidence * fix(review): Disclose unsupported contributions and deduplicate watcher wakes * fix(review): Preserve unmeasured unsupported contributions across Bearings * fix(review): Deduplicate shared contribution wakes and isolate diagnostics * fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower * fix(bin): make remote report transfers explicit and fail-open (#4658) * fix(bin): make a remote-reply document gap self-clearing and re-attemptable A remote mate's undelivered document raised a keyed `blocked` decision that nothing could ever resolve, and any `data/*.md` substring in any mirrored line was an unconditional fetch instruction. A mate announcing a report it had not written yet therefore manufactured a permanent, factually false blocker, and its own explanation of the false alarm manufactured more. The reader has no permanence vocabulary: a report still being written refuses exactly like a path that will never exist. So an undelivered document is now a durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`, re-attempted on the next delta and on the channel's own quiet poll, and retired with a matching `resolved` line naming the local copy once it arrives. The cursor still advances and no delta stalls on one bad pointer. Only a structured `report=data/....md` pointer now offers a document, so a path merely mentioned in prose - including one under another home's mirror tree, which is provably not that mate's to serve - is never fetched. Offers are deduplicated across the whole delta, the escalation names each missing document once and carries the reader's own reason instead of discarding it, and a strictly increasing notice ordinal keeps a later escalation from being swallowed as duplicate bytes. A mirrored line still lands once whichever pointer form it was first written under. * no-mistakes(review): Require structured pointer token boundaries * no-mistakes(review): Unify boundary-safe pointer extraction and rewriting * fix(bin): identify a mirrored line independently of its delivery state Two defects in the boundary-safe pointer work. The at-most-once check compared only the all-remote and all-local renderings of a line, so it could not recognize a mixed one. A line offering two documents where only the first was deliverable mirrored as local-plus-remote; once the second arrived, a cursor-loss whole-log recapture rendered the same line all-local, matched neither alternate, and mirrored a second time. A line's identity is now the canonical form every boundary-valid pointer would take once delivered, derived by the same parser that does extraction and rewriting, so it no longer depends on which documents happened to be deliverable at the time. The pointer map was passed to awk through the process environment. A delta may carry up to the configured 1 MiB bound, and an expanded map of delivered pointers can exceed the platform's exec argument limit, so awk would fail to start; because no caller checked, the empty result would have been appended as blank lines while the cursor advanced past dropped status content. The map now travels in a file, and every call site checks the exit status and stops the ingest rather than committing a delta it could not render. Both passes now run once per stream instead of twice per line. * no-mistakes(review): Abort ingest when document pointer extraction fails * no-mistakes(review): Exclude structured cross-home pointers from document transfer * fix(bin): fail open on an undeliverable remote document instead of tracking it Narrow the remote-reply document fix to the scope the diagnosis actually requires, as decided after measuring a simpler alternative. A document the reader cannot deliver now fails open. The mate's line is mirrored with its own pointer, the cursor advances, and one unkeyed note carries the reader's reason. A note never enters the open-decision fold, so it cannot stand open the way the original keyed block did - which removes the never-clearing false blocker by construction rather than by resolving it. That makes the durable self-clearing obligation unnecessary, so it goes: the per-mate pending-documents record, its notice ordinal and resolved announcements, and the poll-side retry. Canonical line identity goes too, and with it a way to silently drop a genuine status line; mirroring is back to at-most-once on exact bytes. The cross-home exclusion goes as well: under fail-open a cross-home report= either fails harmlessly or is a nested remote report this mate genuinely holds, which is now relayed again. Kept: fetching only on a structured report= pointer, the boundary-correct parser, the file-based rewrite map, and checked extraction and rewrite exit status. The parser now scans behind a sentinel byte so a rejected candidate can no longer give the text right after it a false leading boundary. The reported incident is covered end to end: a report path announced in prose before it exists raises no decision, and the report still arrives through the ledger publisher's structured offer once written. * no-mistakes(review): Preserve source-line identity across remote reply replays * no-mistakes(document): Document remote reply transfer and replay semantics * no-mistakes(lint): Fix staging truncation lint checks * fix(calm): preserve substantive mid-turn responses (#4655) * Preserve substantive Calm mid-turn text * no-mistakes(review): Distinguish newline-preserved replies from short narration * no-mistakes(document): Document Calm mid-turn preservation boundaries * no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check` * fix(bin): preserve PR merge polls across volume remounts (#4656) * fix(bin): re-record PR poll identity after a volume device renumber (Fixes #4260) A volume remount can renumber the state filesystem's st_dev while every inode and byte stays the same; APFS does this across a reboot. A poll registration records its sidecar and check as device:inode, so every poll armed before the remount failed strict validation and the watcher refused all of them as unauthenticated state checks until each was re-armed by hand. There are two device comparisons. fm_pr_private_file_valid compares a live file's device with the state directory's device read in the same invocation: it refuses a file that is not on the state directory's own filesystem and already survives a renumber, so it is unchanged. The registration's recorded identity versus the live identity (from #556, reused by the #932 retirement receipt) binds the registration to the exact files published in its own transaction; its device part is what breaks. When strict capture fails, the watcher now proves the device is the only difference: every other artifact check passes (template bytes, both hashes, private mode, single link, live device, metadata), both recorded identities name one device, and each recorded inode equals its live inode. Only then, under the task's control lock, does it rewrite the two identity lines, repeating the whole proof and comparing the registration's file identity and bytes just before the rename, and then capture strictly again. A swapped, altered, re-moded, relinked, split-device, or foreign-device artifact still fails a proof and is still refused, and a pending retirement receipt blocks the rewrite. Reproduction: on macOS a poll armed on an APFS disk image that was detached and re-attached behind another image moved st_dev 16777239 -> 16777243 with inodes, bytes, mode, and link count unchanged; the real watcher refused it on main and reports its merge with this change. The portable regression test rewrites a real registration's recorded device and drives the watcher. Not changed here: the status presentation cursor keys rows by its own device:inode identity in bin/fm-classify-lib.sh, a different helper that needs its own fix; a retirement receipt left by a reboot between its publication and removal still names the old device and stays refused; custom check trust binds only a content hash and is unaffected. * fix(review): Serialize PR poll publication writers * fix(review): Bound PR poll publication lock scope * fix(bin): keep contribution records when the poll budget runs out (follow-up to #4627) (#4661) A budget that expires partway through an observation no longer records an error or prints the unavailable wake; the URL keeps its prior record and is observed first next poll. forge() flags budget exhaustion at the point it refuses, or when a read is killed at the budget's own deadline, so a genuine forge failure still records the error and wakes. Each distinct URL is now observed once per poll and applied to every owning task. * fix(bin): clear parent pending-replies on local secondmate retirement (#4680) * fix(bin): clear parent pending-replies on local secondmate retirement Local secondmate teardown left resolved parent pending-reply records behind after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced retirement while any reply for that id is still unresolved, and delete every matching record plus its delivery confirmation after a successful local or remote retirement, matching the remote cleanup path. * no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup * no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung * no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern * no-mistakes(review): Pending-replies Basename und corr_id abgleichen * no-mistakes(document): Clarify forced retirement pending-reply cleanup --------- Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc> * fix(bin): accept Orca's composite worktree id when tearing down a task (#4677) * fix(bin): accept Orca's composite worktree id at teardown Teardown refused every Orca-backed task because the endpoint validator checked orca_worktree_id with the simple-atom rule meant for tmux-style window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca returns that id as `<orca id>::<absolute worktree path>`, so the colon and slashes in every real value made validation fail and finished Orca tasks could never be cleaned up. Validate the field as the composite it is: both halves of the first `::` split present, the path half absolute, and no embedded newline, carriage return, or tab. The terminal field keeps the atom check, which is correct for it, and no other backend's validation changes. The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca never returns, which is why the suite passed a check the real value fails. They now carry the composite form, so the tests exercise the real value. * no-mistakes(document): name Orca's repo id in the composite worktree id * no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points * feat(bin): add opt-in typed dispatch resolution (#4692) * feat(bin): add opt-in typed dispatch resolution through typesafe.ai Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or scout profile from a written brief with typesafe.ai's System One model: one Choice question over the rules' `when` texts, then the confidence floor, the rule's `approval` and `floor`, each profile's `provider` and `floor`, one quota-axi snapshot, and the spendPriority argmax all in code. It is off unless TYPESAFE_API_KEY is in the environment or the home's gitignored .env; off means one stderr line, exit 0, and no network call, so firstmate dispatches exactly as before. The key reaches curl on a file descriptor, never argv. Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates the four new optional dispatch fields. Document the schema, the operator contract, the AGENTS.md intake step, and the live and benchmark evidence. * no-mistakes(review): Harden typed dispatch resolution and quota bounds * no-mistakes(review): Validate dispatch floors and ranking evidence * no-mistakes(review): Tighten dispatch response and floor evidence * no-mistakes(review): Neutralize none matching and resolve defaults locally * no-mistakes(review): Preserve providerless profiles outside typed resolution * no-mistakes(review): Validate response usage and reject duplicate profiles * no-mistakes(review): Escalate unverifiable floors and validate probabilities * no-mistakes(review): Validate probability mass and unknown profile floors * no-mistakes(review): Simplify resolver interface and preserve fallback routing * no-mistakes(review): Fix constants and rank partial quota evidence * no-mistakes(review): Add authoritative provider mapping and enforce explicit providers * no-mistakes(review): Declare provider for documented Pi profile * no-mistakes(review): Validate provider identifiers and support Gemini dispatch * no-mistakes(review): Strictly anchor provider identifiers * no-mistakes(review): Validate selectors and preserve fallback candidate evidence * no-mistakes(review): Gate typed validation and harden resolver evidence * no-mistakes(review): Preserve opt-in routing and harden candidate evidence * no-mistakes(review): Prioritize known exhaustion over quota uncertainty * no-mistakes(review): Isolate API secrets and preserve no-key diagnostics * no-mistakes(review): Fallback safely when dispatch rules are absent * no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets * no-mistakes(document): Document typed dispatch safety and fallback behavior * fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753) * test: reproduce buried status declarations in shared readers * fix: share status event reads and preserve open blockers * fix: retain terminal scout and ship status declarations * no-mistakes(review): Fix status chronology, legacy completions, and reader performance * no-mistakes(review): Share terminal decision reconciliation across fleet snapshots * no-mistakes(review): Unify terminal supersession across cached folds and consumers * no-mistakes(review): Filter per-key status history while preserving terminal chronology * no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells * no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses * no-mistakes(document): Document latest-event status read and kind-scoped fold cursor * no-mistakes(lint): Quote literal done in test for-lists for SC1010 * ci: expect 19 snapshot/fleet-view tests This branch adds a fleet-snapshot regression, so the stock macOS Bash lane's hardcoded guard of 18 'ok - ' lines fails on the new count. Bump the guard and its message to 19. * no-mistakes(review): Restore multiline child outcome reporting * no-mistakes(review): Select ledger terminal events through bounded shared reader * no-mistakes(review): Report newest open decision instead of preferring blocked * no-mistakes(review): Require colon before ship/scout terminal supersession in fold * no-mistakes(review): Gate socket-down override on latest event; drop lock matrix * no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions * no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold * no-mistakes(test): Update fleet-view expectations to newest-open-decision rule * no-mistakes(document): Align status-read docs with fold-resolved crew state * no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers * no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree * test: fold terminal-cleanup snapshot coverage into the completed-scout case Keep the ship/scout/secondmate supersession assertions without adding a nineteenth top-level fleet-view test, so CI can stay at the upstream suite count. * no-mistakes(document): Clarify socket-down override expiry in architecture doc * ci: retrigger flaky contribution check * fix(bin): launch codex crewmates with codex's hook layer disabled (#4689) * fix(spawn): launch codex crewmates with codex's hook layer disabled A freshly launched Codex worker never reached its instructions. Codex stopped it on an interactive "Hooks need review" modal whose selection sits on "Review hooks", which is neither trusting nor declining. Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow navigation, so the selection cannot be moved, and pre-accepting the prompt by writing Codex's own trust store would record an operator consent that was never given. The hooks are the machine's own ~/.codex/hooks.json plus any project's .codex/hooks.json. A crewmate needs neither: its turn-end signal is the -c notify= program on the same launch, and Firstmate's project hooks are primary-session infrastructure that stands down in a child worktree. Crewmate and scout launches now pass --disable hooks. That is the opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted hooks; disabling the feature runs none of them and leaves the operator's ~/.codex untouched. An unknown feature name is a hard Codex error, so a release that drops the flag fails the launch loudly instead of silently restoring the modal. A secondmate is a primary in its own home and keeps the project hooks its turn-end guard and session-start digest ride on. Verified on codex-cli 0.151.0: the modal is gone and the turn-end notification still lands. This unblocks the second review that every finished pull request is supposed to get. Fixes kunchenguid/firstmate#4673 * no-mistakes(review): Fix contradictory hook count in Codex verification record * fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710) * fix(bin): settle terminal contributions and wake once per read-failure episode A contribution whose last good observation is merged or closed is final: poll no longer re-reads it, projection keeps it fresh, and a stale error recorded beside it is cleared once. A genuine forge-read failure on an open contribution still records its error on every cycle but prints the unavailable wake only when it starts a failure episode; a successful read ends the episode. Open PRs linked from done tasks keep being observed. The false unavailable beside a complete observation was budget exhaustion mid-observation, already fixed by #4661. * fix(review): Settle terminal contribution owners * fix(review): Deduplicate shared contribution failure episodes * fix(test): Preserve settled terminal contribution records * fix: select authoritative no-mistakes runs (#4476) * fix(crew-state): select authoritative validation runs by identity Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row. Refs: https://github.com/kunchenguid/firstmate/issues/3215 * fix(review): Resolve same-branch run identities beyond capped history * fix(review): Fix run-selection compatibility, races, and worker-state fallbacks * fix(review): Limit run validation to the requested branch * fix(test): Anchor AXI fixtures and document remaining live evidence gaps * fix(document): Clarify run selection documentation and capture ownership * fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation * fix: distinguish captain outcomes from no-op updates (#4738) * fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work MAIN answered a supervision-branch outcome for completed captain-requested work (implementation done, PR ready for review and merge approval) with "Captain, shipshape.", reading section 9's no-action reply as covering it and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no captain-facing response is owed". Section 9 now limits the shipshape reply to true no-ops (idle re-read, empty heartbeat, consequence-free acknowledgement) and requires a short outcome response naming what finished and what word is needed whenever requested work finishes or a result needs the captain's word, even when a transcript entry already shows the substance. The Pi protocol's re-emit rule now says it bounds repetition only, and carries a worked example of the ready-for-review outcome whose correct processing turn a shipshape reply fails. No executable contract evaluates the content of MAIN's captain-facing reply, so the regression is the protocol example in the owner doc rather than a text-match test. * no-mistakes(document): Clarify captain-facing outcomes versus no-ops * docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line The document step condensed the Pi protocol's re-emit rule and dropped the worked example of a finished, ready-for-review outcome whose correct processing turn a "Captain, shipshape." reply fails. That example is the contract's regression: no executable contract evaluates the content of MAIN's captain-facing reply, so the owner doc's example is the test case. Restore it directly under the re-emit rule, prefixed as a regression example that is kept verbatim and never condensed or summarized away. * no-mistakes(review): Clarify captain outcome and decision-word requirements * no-mistakes(document): Clarify captain-facing completion outcomes * docs(pi): require the PR URL in the visible captain-facing outcome reply Captain review on the regression example: drop the sample reply string and say only that the ready-for-review outcome requires relaying a captain-facing outcome response, not just "Captain, shipshape.". Fold in the visible-PR-handoff failure seen this session: after the branch outcome reporting this fix green, MAIN's visible reply was only "Awaiting your merge call." with no PR URL, leaning on the dim anchor. Section 9's URL rule now also covers a review or merge ask and names the visible reply as where the URL goes, sourced from the ready status, pr= metadata, or the supervision branch's summary and never left to a transcript entry. The Pi protocol adds the same-way failure and places the captain-facing text in the final visible assistant reply after the fm_branch_processed call, because Calm hides assistant text emitted in the same step as a tool call as a working note. Investigation verdict, evidence in the PR comment: no recent PR caused the handoff failure; Pi has hidden same-step pre-tool assistant text since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and #4658 touched only remote report transfer. * no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs * no-mistakes(document): Clarify captain-facing supervision outcomes * docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule The consolidated section 9 URL rule narrowed its trigger to a review or merge ask, dropping the "whenever a PR is mentioned" catch-all from #3648 that keeps every PR URL copied from a durable record and never assembled from memory. Restore that trigger as a union with the review or merge ask so the one consolidated rule covers both. * fix(bin): let non-owner Claude Stops exit safely (#4777) * Fix foreign-owner turn-end supervision loop * no-mistakes(review): Scope foreign-owner safe exit to Claude guard * no-mistakes(document): Document Claude foreign-owner safe exit * fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778) Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty indexed array as an unbound variable and aborts the shell. In signal_turnend_panes_churned() the missing_keys loop was reachable with an empty array whenever every churned key already held a fresh .churn-since-* marker (a second churning turn-end inside an open deferral window), so each watcher cycle died about half a minute in and supervision restarted endlessly. The created_keys rollback loops had the same latent crash on their error paths. Audit of bin/ for the same pattern found one more confirmed-reachable case: remote_handoff's noncanonical-body scan iterates to_move, which is empty when a retried remote handoff finds every key already staged in the outbox. All other "${arr[@]}" sites are either count-guarded, guaranteed non-empty by construction, or unreachable while empty. Guard the three reachable expansions with the repo's existing "${arr[@]+...}" idiom. New regression test drives a real watcher through the all-marked churn path; the macos-stock-bash CI lane runs it under real /bin/bash 3.2 via FM_TEST_ONLY. * Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783) The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: require complete captain-facing final responses (#4779) * docs: require complete final responses across harnesses * no-mistakes(document): Document complete final replies for Grok Bot * docs: point Grok replies to the shared contract owner * no-mistakes(review): Clarify final recap without batching decision asks * fix: preserve substantive mid-turn text in Pi Calm (#4788) * fix(calm): preserve substantive Pi mid-turn text * no-mistakes(review): Preserve substantive Pi Calm text per block * no-mistakes(test): Cover shared Calm preservation boundaries behaviorally * no-mistakes(document): Consolidate Calm preservation documentation * fix: harden mail checks and rebalance full-coverage CI (#4800) * Improve CI reliability and rebalance full-coverage validation * no-mistakes(document): Clarify lint partition documentation * fix(bin): answer Kimi 2.0.0 folder-trust dialog during spawn (#4799) * Handle Kimi workspace trust dialog * no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers * no-mistakes(review): Gate Kimi ready on any trust marker and clean captures * no-mistakes(review): Read visible pane for Kimi trust and ready gates * no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate * no-mistakes(review): Harden Kimi viewport capture and trust dialog detection * no-mistakes(document): Document Kimi spawn refusal on cmux and Orca * fix(bin): report a dead-agent record once instead of escalating forever (#4775) * fix(bin): report a record whose agent is gone once instead of escalating forever The wedge escalation path never asked whether there was still an agent to be wedged. A wedge is something stuck that might recover, so re-alarming it earns its cost; an agent that is gone never moves again, its pane never churns, the idle timer never resets, and the escalate path clears its own timer and re-arms with nothing bounding the count. Observed on a live fleet: two finished lanes reached 226 and 203 consecutive escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400 notifications a day from two lanes with no agent running at all. On one, fm-control.sh exit answered already-stopped and fm-crew-state.sh read "failed - run failed". Closing the Herdr pane did not stop it either: with the pane genuinely gone and herdr pane read returning pane_not_found, the count kept climbing, because the poll is driven by the record's window= line rather than by the pane. The cost is not the repetition but that it drowns the alarms that matter. fm_backend_agent_state already separates a thinking agent from a gone one at process level. In the branch that was about to escalate, read it once and treat only its two recovery-grade verdicts - dead (endpoint present, no agent in it) and missing (endpoint authoritatively absent) - as proof, reporting that record once and not re-escalating it while it stays that way. Every other verdict, including alive, ambiguous, unreadable, unverified, and a read that failed outright, keeps the identical schedule, reason, and escalation count, so a genuinely wedged live agent is unaffected. The probe costs at most one backend read per window per threshold, the same budget the declared-wait consult and the worktree write probe already take. The report decides nothing about the record's fate: both lanes still held unlanded work and teardown refusing them was correct, so retiring, relaunching, or cleaning up stays with the supervisor. The once-only marker is owned entirely by that function and is dropped by the same read the moment the endpoint stops reading gone, so a replacement launched into the same window escalates normally and its own later death is reported again. Related, and not closed by this: #4412, #4482, #4316. Tests drive the real watcher against a record whose endpoint does not exist and pin both directions: dead and missing report once and never advance the count across later thresholds, while alive, ambiguous, and unreadable endpoints keep escalating with the identical reason and a climbing count. * fix(bin): bind the once-only dead report to the pane it reported Review of the parent commit found a reachable sequence where a later death in the same window lost its promised report. The marker was keyed on the verdict string alone and dropped only when a threshold probe read a non-gone verdict, but probes run only at thresholds: a replacement launched into the same window that dies without ever being probed alive - it crashes at startup, or works and then crashes - was absorbed by the previous death's marker. The pane's first sight yielded only the generic stale wake and every later threshold matched the stale marker, so the second death never got the detailed once-report that both the function's own comment and docs/architecture.md promise. Record the verdict together with the pane hash it was reported for, and absorb a repeat only while both still match. A replacement churns the pane, which resets the stale suppressor, wedge timer, and escalation count while no reset site touches this marker, so the pane half is what tells the second death apart from the first. The live-probe drop stays as it was. Clearing the marker at those reset sites instead would re-open unbounded re-alarming for a dead pane whose display ever ticks, which is the exact defect the parent commit exists to close. The noise bound is unchanged: an unchanged dead pane still absorbs on every later threshold and never advances the escalation count, and every verdict short of proof still escalates exactly as before. * no-mistakes(review): Key the dead-record once-marker on the busy incarnation token * no-mistakes(document): Document dead-record escalation cap in stale-pane config entry * no-mistakes(document): Add busy-state inventory line to AGENTS.md * no-mistakes(document): Document dead-record probe on busy-turn-bound wedge path * fix(bin): create captain-hold rows when Beads requires due (#4854) Captain holds have no due semantics and are a hold kind, not a Beads issue type. The create path now waives due.required and maps to native type task. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: disable compact adviser for spawned agents (#4877) * feat(bin): launch every spawned agent with the compact adviser disabled Every crewmate, scout, and secondmate Firstmate launches now starts with COMPACT_ADVISER_DISABLE=1, on a fresh spawn and on a relaunch alike, so an unattended session never activates the compact adviser. The value is unconditional: no configuration file gates it and there is no override, unlike the trace carrier beside it. Three carriers deliver it, because no single one covers every launch shape. The pane shell receives an export beside GOTMPDIR, so the agent's own children inherit it too. The launch command carries an explicit assignment, prepended outermost so it wins over any ambient value the pane already held. The cleared launch environment sets it again at the `env -i` boundary and keeps COMPACT_ADVISER_DISABLE in the fixed operational floor, which is what preserves the switch when config/launch-env-allowlist empties the environment, and what delivers it on a remote host that never had the value. bin/fm-control.sh relaunch, the bootstrap secondmate relaunch, and the remote secondmate transport all rebuild their launch through bin/fm-spawn.sh, so they inherit the same floor. The captain's own primary session is untouched. The two new suites drive the real spawn and then execute the launch command the pane actually received, with the harness replaced by a probe that prints its own environment, rather than matching script text. They cover ship and secondmate launches with the allowlist absent and enabled, the pane export and its ordering, fm-control.sh relaunch, and the full parent to remote-host chain. * no-mistakes(review): Export compact-adviser disable across compound launches * no-mistakes(document): Document spawned-agent compact-adviser environment guarantee * fix(bin): preserve Claude lock ownership after helper recycling (#4894) * fix(bin): let a background Claude session keep owning its session lock Session-lock ownership was decided by process ancestry alone. Under an unattended Claude session the model loop runs in a transient bg-spare bridged to the front-end by a shared daemon; when that bridge is recycled the contiguous claude-named ancestry from a hook to the recorded owner breaks while the owner pid stays alive, so the Stop auto-arm stood down as a foreign live owner, the turn-end guard ended every turn with its read-only diagnostic, and fm-lock.sh refused - a self-sustaining outage until restart. Ownership is now ancestry membership OR a trusted same-session id, never id-first: - fm-session-lock-lib.sh accepts CLAUDE_CODE_SESSION_ID only when CLAUDE_PID is a Claude-shaped member of the current contiguous run, compares it against the id recorded in state/.lock-session, and requires the recorded pid to still be a live harness. No id, no sidecar, an untrusted id, a different id, or a dead recorded pid leaves the ancestry verdict unchanged. Ids are never read from ps argv. - fm-lock.sh accepts a same-session holder at both refusal sites, writes, refreshes, and clears the sidecar only under its claim lock (including the early already-mine exit, skipped only while the deferred startup sweep leases that lock), keeps it byte-identical across a same-session confirmation, records CLAUDE_PID on lock line 1 for a session with a trusted id so a shared daemon or front-end that outlives the session never keeps a dead session's lock alive, never rewrites a live line 1 on a same-session confirmation, and names the recorded id in the live-owner refusal. - The .lock line-1 format is unchanged, so every reader that takes the whole first line as the pid keeps working; the guard's foreign-owner exit is unchanged and inherits the fix through the shared predicate. Tests: the ancestry suite drives the ancestry and id signals apart in a deterministic process table (asserting the divergence) and runs a real orphaned front-end/daemon/pty-host/spare tree through six phases with the real lock, auto-arm, and guard scripts; the foreign-owner repro keeps its negative control and adds a same-id positive control. Disclosure: no live unattended Claude background session ran on the verifying machine. The topology is documented by the real process listings in #3902, #2314, #3398, and #4066; coverage is the structural predicate plus the executable fixtures, not a live pass. Residual: bin/fm-sessionstart-nudge.sh keeps its own private ancestry walk (it only decides whether to print a nudge) and may nudge on a resume in the recycled case. Out of scope, deliberately: no structured lock format, no guard budget changes, no daemon-identity rejection, no fork lineage. * no-mistakes(review): Wait for claim lock; revert failed sidecars * no-mistakes(review): Revalidate ownership after wait; restore sidecars * no-mistakes(review): Roll back sidecar by publication phase * no-mistakes(review): Restore sidecar only if lock line is unchanged * no-mistakes(review): Trust session ids without a spelling allowlist * no-mistakes(review): Disarm sidecar rollback before backup cleanup * no-mistakes(document): Updated session-lock ownership documentation * feat: park main under the away posture on Pi (#4889) * feat: park main under the away posture on Pi While the away-posture record exists on a Pi primary, the supervision branch takes every actionable wake, no processing turn opens on main, captain rows accumulate for the return brief, and main's standing authority relocates to the branch through the existing guarded scripts. - lib/fm-branch-dispatch.ts: read the record at every routing decision; while it exists claim check, decision-owned, and heartbeat rows too, keeping the two broken-queue vetoes; expose checkSeqs so a claimed check row lifts task scoping. - fm-primary-pi-watch.ts: offer every actionable row under the record; a declined wake and every watcher-failure alarm still reach main. - fm-branch-supervision.ts: drop the legacy .afk decline; append a fixed POSTURE: AWAY tail carrying the record's read-back verbatim per wake; open no processing request while the record exists, re-checked immediately before a request would open and at every run boundary; present the accumulated rows at the first run boundary after archive. - fm-lease-lib.sh: fm_lease_forbid_branch passes the branch for opted-in actions only while fm-afk-contract.sh validate succeeds on a confirmed live record; PR merge, fresh spawn, and decision answer opt in, local landing never does. - fm-send.sh: a --resolve-key naming an open needs-decision or captain-held task is a decision answer and meets the partition; blocked: keys stay steering. - fm-spawn.sh: enforce the record's spend cap for a fresh ordinary spawn by either actor; relaunches and secondmates exempt. - fm-branch-prompt.sh: fixed Postures section and the verbatim ask-user-authority policy; the prefix stays byte-stable. - fm-afk-return.sh: count what the away session handled from the store. - docs, afk skill, AGENTS.md stub: main parked on Pi, green merge gate absolute while away. - tests: watcher and branch extension suites, fleet-record, merge, and decision-answer suites cover the relocation, the vetoes, the tail, the parked processing turn, the cancellation, the re-presentation, and the spend cap; dated live-guard evidence recorded. * no-mistakes(review): Refuse branch merge after preflight archive race * no-mistakes(review): Fix away wake, spawn, and processing races * no-mistakes(review): Suppress parked processing; narrow away-only rejection * no-mistakes(review): Abort dedicated processing; gate branch spawn once * no-mistakes(review): Stamp away-only on the dispatch offer * no-mistakes(review): Treat invalid away records as spend-cap absence * no-mistakes(review): Drop spawn test hook; abort processing-opened runs * no-mistakes(review): Bind abort to opening prompt; cap-read absence * no-mistakes(review): Limit away branch spawn to queued work only * no-mistakes(document): Correct AFK posture documentation * ci: standardize workflow timeouts into three tiers (#4910) * ci: simplify CI job timeouts to a three-tier policy Replace the scattered per-job timeout values (10m parallel, 25m lint, 30m serial, 10m macOS) with three readable tiers, each a hang tripwire with headroom rather than a packing estimate: - fast (5m): coverage guard, repo invariants, timing aggregate - normal (30m, one shared budget): lint partitions, portable parallel shards, portable serial shards, macOS stock Bash - heavy (Herdr only): 20m step tripwire on the family run so always() cleanup still runs, under a 75m job-level last-resort backstop The workflow's header comment states the policy and points at docs/fm-test-portable-shards.md "Timeouts", which now owns it, and each job names its tier beside timeout-minutes. tests/fm-ci-workflow.test.sh asserts the policy against the parsed workflow instead of the old per-job minute values: every job joins exactly one tier, exactly three distinct job-level values exist, the fast tier stays within 5-10 minutes, the normal budget stays at least double the modeled parallel lane sum reported by fm-test-run.sh --check-coverage, and the Herdr step tripwire stays below its job backstop with an always() cleanup after it. Concurrency supersession, shard counts, lane membership, and fail-fast settings are unchanged. * no-mistakes(review): Decouple the normal timeout from packing estimates * no-mistakes(review): Assert Herdr teardown follows the family run * no-mistakes(review): Pin Herdr family-run timeout to 20 minutes * no-mistakes(review): Ignore comments when identifying Herdr steps * no-mistakes(review): Identify Herdr steps by declarative ids * no-mistakes(document): Clarify authoritative three-tier timeout policy * fix(bin): keep supervisor status closes from waking the same home (#4895) * fix(bin): keep supervisor status closes from waking the same home A drain that already folded OPEN DECISIONS has presented those bytes even when the watcher has no matching seen marker. Treat that fold, and the presentation cursor, as known so the bookkeeping close stays quiet while later worker lines still signal. * no-mistakes(review): Keep folded worker failures waking past supervisor closes * no-mistakes(review): Wake on unlisted folded worker lines; batch multi-key closes * no-mistakes(review): Stop folded worker resolved lines from counting as already read * no-mistakes(document): Correct self-announced close marker contract in docs * fix(bin): stop labeling Herdr as experimental (#4972) * Stop steering operators away from Herdr * no-mistakes(review): Neutralize remaining Herdr opt-out documentation wording * fix(bin): treat a live no-mistakes run as current after rebase (#4973) * fix(bin): treat a live no-mistakes run as current after rebase A running run on the task's branch is authoritative regardless of head. Matching only the local head made a rebased in-flight run look failed. * no-mistakes(review): restrict coarse live-any-head to foreign-branch answers * no-mistakes(review): reject gate-parked runs from the executing predicate * no-mistakes(review): hoist gate-marker patterns into single run-lib owner * no-mistakes(review): require live daemon for head-free run binding * no-mistakes(review): require answered daemon-down before unbinding live runs * no-mistakes(review): extend daemon guard to anchored continuation routes * no-mistakes(review): delete live-any-head; restore dead-daemon verdict * no-mistakes(review): keep parked gates parked; name dead daemon everywhere * no-mistakes(review): set dead-daemon verdict instead of emitting early * no-mistakes(review): align selected route with legacy dead-daemon handling * no-mistakes(review): drop unproven-record binds; narrow coarse gate reading * no-mistakes(review): narrow header, drop vestigial guard, retarget tests * no-mistakes(review): revert coarse gate override; require answered-down probe * no-mistakes(review): cache one daemon probe; stop duplicating run id * no-mistakes(review): restrict coarse dead-daemon verdict to moved-off rows * no-mistakes(review): delete coarse dead-daemon extension and gate note * no-mistakes(review): delete remaining coarse dead-daemon block and stale docs * no-mistakes(document): document rebase-safe live-run bind and unverified-record verdict * fix(bin): prevent long worker launch command truncation (#4994) * fix(bin): stage the launch command in a private file and type a short source line A long launch line typed while the fresh pane shell is still busy waits in the terminal's canonical line buffer, which drops input past about 1,024 bytes on macOS, so the pane was left at an unfinished command with no agent running. fm-spawn now writes the assembled command to the task's own temp root under umask 077 and types only a short line that sources it. Refs #4559 * fix(bin): keep the per-task temp root private before staging the launch command The root lives at a predictable path under /tmp and now holds the whole launch command. Create it with mode 0700, refuse one that already exists as anything but a directory owned by this user that nobody else can write, and tighten an owned one, so no other local user can plant or swap the staged file. Refs #4559 * fix(bin): enforce private staged launch file mode * test(spawn): cover long staged Claude launches * no-mistakes(review): Namespace launch files and prove truncation staging * no-mistakes(review): Use immutable per-spawn launch filenames * no-mistakes(document): Document staged launch delivery safeguards * no-mistakes(ci): Updated eight behavior tests/fakes to execute or inspect immutable staged launch files instead of expecting inline launch commands. This restores Muse, secondmate lifecycle/restart, remote trace/parent binding, compact-adviser, and Orca coverage. All affected tests, dispatch-profile regression, fixture tests, syntax checks, ShellCheck, and git diff checks pass --------- Co-authored-by: Vytautas Stankus <svycka@gmail.com> * test: authorize isolated Herdr lab validation (#4998) * Add isolated Herdr runbook to test instructions * no-mistakes(review): Drop substring matching from test.instructions contract * no-mistakes(review): Assert commands.test key absence in YAML * Drop unit-first sentence and instructions contract test Captain-scoped follow-up on the Herdr-lab test.instructions ship: keep the lab safety runbook only, and leave the no-mistakes contract test focused on commands.test absence. * docs(vision): accept vendor-semantics and 9k AGENTS ceiling (#4873) (#5001) * docs(vision): accept vendor-semantics and 9k contract-ceiling amendments (#4873) Replace the pixels-of-today's-UI rule with a quarantined, version-pinned surface-adapter exception recorded as standing debt. Cap the always-loaded contract at 9,000 words and require prune-or-trigger before a crossing change lands. Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com> * docs(vision): restore accepted three-sentence vendor-semantics form (#4873) Replace the compressed paraphrase with the issue's accepted wording: a named quarantined version-pinned adapter, expected to break, recorded as standing debt that never hardens into a shared contract. Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com> --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com> * feat(bin): defer the wedge escalation for a lane parked at a supervisor-owed gate (#4974) * fix(watch): recheck a gate awaiting a human instead of wedge-escalating it A lane whose validation run is parked at a gate waiting on a human decision is correctly quiet, but nothing in its status line says so: the evidence is the pipeline's own gate state rather than anything the worker wrote. The wedge timer read that silence as a suspected wedge and climbed the escalation ladder for as long as the wait lasted, and each escalation cost a supervising turn. The landed declared-wait consult does not reach it, because a live ordinary crewmate never reports a declared pause, and raising FM_STALE_ESCALATE_SECS would delay genuine wedge detection for every lane by the same amount. The threshold now reads a second, independent record when the status line accounts for nothing: whether the crew's current state is a gate whose answer is owed by a human. That is minted only from the gate's own findings table, by a row whose `action` column is exactly `ask-user`, located by position out of the table header the way nm_gate_step_row already reads its row - never searched for over the run payload, where a finding's free-text description or a branch name satisfies a search just as well. A gate awaiting the CREWMATE's own answer keeps the unchanged escalation schedule, reason and demand-deep-inspection wording, because a crewmate that goes quiet before answering its own gate is exactly the wedge the ladder exists to catch. Each kind of wait now carries the human it is on, the action that clears it, and whether that human is the captain as data alongside the verdict, rather than as wording chosen per branch where the recheck is written, so the deferral cannot word one kind of wait as another and a new kind cannot ship without deciding all of them. A parked gate has no written record of when its wait began, so its recheck publishes no wait age at all rather than one read from the quiet window this deferral resets on every pass, which would report the same small number for a gate of any age. Like every other captain-facing recheck here it is absorbed in silence while the away-posture record exists, arming no throttle, so the recheck is owed in full the moment the record is archived. The consult runs only in the at-threshold branch that was about to escalate, beside the worktree walk already there, and only for lanes whose status line explained nothing. Closes #3055 * no-mistakes(review): require an unanswered decision before deferring a parked gate * no-mistakes(review): reset the away-silenced timer, fail-safe findings parse, US-joined wait records * test(watch): pass the pane hash wedge_timer_check now takes Upstream gave wedge_timer_check a sixth <pane-hash> argument for its dead-record probe. The malformed-wait-record rounds drive the real function directly, so they pass one, and stub fm_backend_agent_state to a live agent so the probe that runs after a refused deferral keeps the unchanged ladder rather than reading a backend the child shell has none of. * no-mistakes(review): Bind parked-gate wait to its run, owe it firstmate * no-mistakes(document): correct wait-kind count, crew-state reader scope, gate-key coupling * feat(watch): make the parked-gate wait deferral opt-in The wedge timer deferring a lane parked at a validation gate is new supervision behaviour rather than a restored one, and it decides which lanes give up the escalation ladder, so it now ships as a default-off per-home option instead of changing every home on upgrade. config/wedge-defer-parked-gate arms it. The flag is read before the decision fold, so an unconfigured home spends no fold or current-state read, writes no record, and keeps the unchanged escalation schedule, reasons and demand-deep-inspection wording; a test counts the reader calls in both directions to pin that. It is not inherited by secondmate homes: each home supervises its own crew and owns that trade separately, the same reason config/turnend-churn-absorb is home-local. The away-posture absorb returns to leaving the idle timer alone, which it had restarted only because the costly consult could reach it. A parked-gate wait is owed to the supervisor rather than the captain, so it never enters that branch, and the recheck owed on return is again owed in full the moment the record is archived. * test(watch): pin that the away-silenced hold leaves the idle timer alone The absorb no longer restarts the timer, so the recheck owed on return is owed in full rather than a cadence into the return. Nothing asserted that, so a restart could be reintroduced silently. * no-mistakes(review): document away-silence rationale, pin captured gate component * no-mistakes(test): anchor gate row scan to the braced findings header * no-mistakes(document): pin same-block gate row invariant in crew-state comment * fix(bin): reclaim a task whose herdr endpoint was destroyed (#5007) * fix(control): let the owning seat reclaim a task whose endpoint is gone A destroyed pane or workspace made `missing` a terminal state. Relaunch accepted only `dead` and said to stop the agent first; exit refused `missing` and said to reconcile the task first; there is no reconcile verb. Each command named the other as its prerequisite, so a task whose terminal went away could not be reclaimed by anything, and a no-mistakes approval it was parked on had no seat left to answer it. `missing` is agent-free a fortiori: there is no endpoint, so there is no agent in it. Widen the existing guards rather than add a verb. - fm-spawn --relaunch accepts a positively proven `missing` and creates one fresh endpoint in the recorded worktree; the record it already republishes rebinds the task to it. A `dead` endpoint is still adopted in place. - fm-control exit reports `endpoint-gone` instead of dying, so the relaunch transaction's stop step no longer dead-ends, and re-resolves the endpoint from the record before verifying the replacement. The duplicate-agent refusal is untouched: both verdicts come from the same recovery-grade classifier, which claims `missing` only from positive absence, so `alive`, `ambiguous`, and `unreadable` all still refuse. The backends' own create paths refuse a live same-labeled endpoint as a second independent guard. The worktree, its branch, commits, uncommitted changes, armed poll and registration, record rows, and status log are all untouched - a reclaim is a recovery, never a teardown. A secondmate is excluded: its gone-endpoint recovery already has one owner in the session-start liveness sweep, so relaunch refuses and names it rather than becoming a second path to the same outcome. Tests reproduce both halves of the deadlock, the reclaim succeeding, unlanded work surviving it, and the refusals that still hold. * no-mistakes(review): prove endpoint absence per backend before reclaim rebinds * no-mistakes(review): give exit and relaunch one absence proof; pin herdr rebind session * no-mistakes(review): narrow endpoint reclaim to herdr; tmux refuses honestly * no-mistakes(review): stop refusals and docs asserting unestablished causes * no-mistakes(review): stop herdr fixture helper losing tmp-root registration * no-mistakes(review): document workspace drift and absence-probe server residue * no-mistakes(review): correct rebind limitation to its one reachable case * no-mistakes(review): stop claiming reclaim leaves instructions untouched * no-mistakes(document): scope fm-control-lib purity claim, note reclaim coverage * no-mistakes(rebase): read the staged launch file in the herdr fixture Rebasing onto main picked up #4994, which stages a long worker launch command into a script and delivers the short `. '<path>'` line instead of the literal command. The tmux fake and tests/fixtures.sh were updated for that; the herdr fake this branch adds was written before it and still keyed "an agent now exists on this pane" off the literal `encode launch-brief` text, so after the rebase it never marked the rebound pane live and the reclaim's alive-wait read `dead`. Dereference the staged file first, exactly as the tmux fake above does. Test-fixture only; no production path changes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * no-mistakes(document): note reclaim placement in herdr and scripts inventories --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(bin): stamp status events with their emission time (#3764) * test(status): reproduce missing event emission time * wip(status): preserve optional event emission time * test(status): document indirect clock stub invocation * no-mistakes(review): Preserve historical status bytes during reply recovery * no-mistakes(test): Fix timestamped status assertions and remote fixture dependencies * no-mistakes(review): Preserve captain regex overrides for timestamped status events * no-mistakes(document): Clarify status event timing and publication contracts * no-mistakes(lint): Quote literal done to satisfy ShellCheck * no-mistakes(ci): Captain, updated .github/workflows/ci.yml to expect 19 snapshot tests instead of 18, matching the PR’s added regression. Reproduced the failure before the fix. Stock Bash 3.2.57 verification passed: parse sweep, 19 snapshot tests, 53 Bearings tests, and the public-followup regression. Workflow lint and diff checks passed * no-mistakes(test): Preserve terminal notifications with malformed timestamp tags * no-mistakes(test): Stamp Rovo spawn failures with emission time * no-mistakes(document): Verify status event documentation * no-mistakes(lint): Fix ShellCheck quoting in status emission-time tests * no-mistakes(ci): Captain, fixed four lifecycle assertions to accept emission timestamps while preserving publication and retry checks. Reproduced the CI failure before the fix. The lifecycle suite now passes with six Beads capability skips; syntax, targeted ShellCheck, and diff checks passed * no-mistakes(ci): Captain, fixed malformed timestamp colons hiding actionable events using shared normalization. Original bytes and unknown ages are preserved. Regression reproduced before the fix; classifier and remote-reply suites, targeted lint, syntax, and diff checks passed * no-mistakes(review): Stamp remote escalations at call sites, drop new flag * no-mistakes(review): Accept stamped escalation and close lines in test assertions * no-mistakes(review): Restore reserved-key answered-note guard for stamped closes * test(status): accept optional emission time in PR-provenance assertions The #4148 provenance test landed on main with exact unstamped greps. Parent-channel lines from this branch carry [at=<epoch>], so strip only that tag before the same exact match. No production change. * no-mistakes(review): Accept stamped ready signal in PR fallback scrape * no-mistakes(review): Drop relay flag, stamp parent events at call sites * no-mistakes(review): Stamp worker terminal-signal instructions, revert fm-on fixture * no-mistakes(review): Accept optional stamp in live cmux drift guard * no-mistakes(review): Restore original test invocation order in two suites * no-mistakes(review): Strip only well-formed numeric status time tags * no-mistakes(document): Drop stale unstamped PR-ready line spelling from channel doc * no-mistakes(review): Stamp agy spawn-failure status lines with event time * fix(bin): normalize status event times in-shell and freeze the budget test clock Two paths made a status event's emission time cost more than it should. The captain-relevance fallback piped every line through awk to drop a well-formed `[at=<epoch>]` tag before matching, so a supervisor sweep paid a fork per line just to prepare a regex match. Shell parameter expansion does the same strip with no fork, and the retry-dedup scan now reuses that one helper instead o…
dsantosg1103
added a commit
to dsantosg1103/firstmate
that referenced
this pull request
Sep 24, 2026
* fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586)
* fix(watch): honour a declared wait before wedge-escalating a quiet pane
wedge_timer_check escalated on elapsed idle time alone. Nothing asked
whether the worker had already said why its pane was quiet, so a lane
that declared a bounded external wait climbed the escalation ladder for
as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every
repeat carried demand-deep-inspection - which by its own wording forbids
re-absorbing on the run-step or pane state, so the supervisor could not
use the evidence that was there either.
The generated brief promises that declaring `paused:` buys the long
recheck cadence instead of a wedge, but the timer was still reachable
while that declaration stood: a crew that declares a wait and then has an
active run or busy pane attributed to it is handed to the timer as
provably-working. The declaration is what the worker said about its own
silence, so it now outranks a liveness verdict that only says something
is running.
The consult runs in the at-threshold branch that was about to escalate,
beside the worktree walk already there, and costs one status-line read.
Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS
recheck the declared-wait absorber already uses, so the wait is still
rechecked and cannot rot invisibly. Which verb declared it decides the
wording, because the two block on different people: a `paused:` wait is
owed by an external dependency and asks the reader to confirm it still
holds, while a `captain-held:` transfer is owed by the captain reading
the recheck and asks them to answer or release the hold. A hold is not
rechecked at all while the away-posture record exists, as on every other
captain-held path, and that absorb arms no throttle so the recheck is
owed in full on return.
A declared clearing time that has already passed stops counting, and a
lane that never declared one keeps the identical escalation schedule,
reason, count and demand-deep-inspection wording, so detection and its
worst-case time are unchanged. The deferral restarts the idle timer
rather than cancelling it, so a lane that stops waiting escalates again
within one threshold.
A lane quiet because its own validation run is parked at a gate awaiting
a human decision is deliberately out of scope: reading that state needs a
signal carrying who the wait is on and what clears it, rather than one
inferred from a parked verdict that also covers gates awaiting the
crewmate itself.
Tests pin both directions for each case and were each confirmed to fail
with the consult removed.
* no-mistakes(document): docs: honour declared waits in stale-escalation docs
* fix(bin): report verified PR state for passed runs (#4624)
* fix(bin): derive passed PR state from PR record
A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists.
For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim.
Fixes #4607
* no-mistakes(review): Add bounded GitLab merge-request state reads
* no-mistakes(review): Preserve network-free inactive crew-state scans
* no-mistakes(document): Document PR record readers in shared library
* fix: restore published contribution follow-up (Fixes #4469) (#4627)
* fix: restore published contribution follow-up (Fixes #4469)
* fix(review): Fix contribution freshness and merge actor routing
* fix(review): Restore issue triage and scope contribution follow-up
* fix(test): test: assert one wake per contribution signal
* fix(document): Document contribution follow-up
* fix: restore truthful terminal delivery evidence
* fix(review): Disclose unsupported contributions and deduplicate watcher wakes
* fix(review): Preserve unmeasured unsupported contributions across Bearings
* fix(review): Deduplicate shared contribution wakes and isolate diagnostics
* fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower
* fix(bin): make remote report transfers explicit and fail-open (#4658)
* fix(bin): make a remote-reply document gap self-clearing and re-attemptable
A remote mate's undelivered document raised a keyed `blocked` decision that
nothing could ever resolve, and any `data/*.md` substring in any mirrored line
was an unconditional fetch instruction. A mate announcing a report it had not
written yet therefore manufactured a permanent, factually false blocker, and
its own explanation of the false alarm manufactured more.
The reader has no permanence vocabulary: a report still being written refuses
exactly like a path that will never exist. So an undelivered document is now a
durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`,
re-attempted on the next delta and on the channel's own quiet poll, and retired
with a matching `resolved` line naming the local copy once it arrives. The
cursor still advances and no delta stalls on one bad pointer.
Only a structured `report=data/....md` pointer now offers a document, so a path
merely mentioned in prose - including one under another home's mirror tree,
which is provably not that mate's to serve - is never fetched. Offers are
deduplicated across the whole delta, the escalation names each missing document
once and carries the reader's own reason instead of discarding it, and a
strictly increasing notice ordinal keeps a later escalation from being
swallowed as duplicate bytes. A mirrored line still lands once whichever
pointer form it was first written under.
* no-mistakes(review): Require structured pointer token boundaries
* no-mistakes(review): Unify boundary-safe pointer extraction and rewriting
* fix(bin): identify a mirrored line independently of its delivery state
Two defects in the boundary-safe pointer work.
The at-most-once check compared only the all-remote and all-local renderings
of a line, so it could not recognize a mixed one. A line offering two documents
where only the first was deliverable mirrored as local-plus-remote; once the
second arrived, a cursor-loss whole-log recapture rendered the same line
all-local, matched neither alternate, and mirrored a second time. A line's
identity is now the canonical form every boundary-valid pointer would take once
delivered, derived by the same parser that does extraction and rewriting, so it
no longer depends on which documents happened to be deliverable at the time.
The pointer map was passed to awk through the process environment. A delta may
carry up to the configured 1 MiB bound, and an expanded map of delivered
pointers can exceed the platform's exec argument limit, so awk would fail to
start; because no caller checked, the empty result would have been appended as
blank lines while the cursor advanced past dropped status content. The map now
travels in a file, and every call site checks the exit status and stops the
ingest rather than committing a delta it could not render.
Both passes now run once per stream instead of twice per line.
* no-mistakes(review): Abort ingest when document pointer extraction fails
* no-mistakes(review): Exclude structured cross-home pointers from document transfer
* fix(bin): fail open on an undeliverable remote document instead of tracking it
Narrow the remote-reply document fix to the scope the diagnosis actually
requires, as decided after measuring a simpler alternative.
A document the reader cannot deliver now fails open. The mate's line is
mirrored with its own pointer, the cursor advances, and one unkeyed note
carries the reader's reason. A note never enters the open-decision fold, so it
cannot stand open the way the original keyed block did - which removes the
never-clearing false blocker by construction rather than by resolving it.
That makes the durable self-clearing obligation unnecessary, so it goes: the
per-mate pending-documents record, its notice ordinal and resolved
announcements, and the poll-side retry. Canonical line identity goes too, and
with it a way to silently drop a genuine status line; mirroring is back to
at-most-once on exact bytes. The cross-home exclusion goes as well: under
fail-open a cross-home report= either fails harmlessly or is a nested remote
report this mate genuinely holds, which is now relayed again.
Kept: fetching only on a structured report= pointer, the boundary-correct
parser, the file-based rewrite map, and checked extraction and rewrite exit
status. The parser now scans behind a sentinel byte so a rejected candidate can
no longer give the text right after it a false leading boundary.
The reported incident is covered end to end: a report path announced in prose
before it exists raises no decision, and the report still arrives through the
ledger publisher's structured offer once written.
* no-mistakes(review): Preserve source-line identity across remote reply replays
* no-mistakes(document): Document remote reply transfer and replay semantics
* no-mistakes(lint): Fix staging truncation lint checks
* fix(calm): preserve substantive mid-turn responses (#4655)
* Preserve substantive Calm mid-turn text
* no-mistakes(review): Distinguish newline-preserved replies from short narration
* no-mistakes(document): Document Calm mid-turn preservation boundaries
* no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`
* fix(bin): preserve PR merge polls across volume remounts (#4656)
* fix(bin): re-record PR poll identity after a volume device renumber (Fixes #4260)
A volume remount can renumber the state filesystem's st_dev while every
inode and byte stays the same; APFS does this across a reboot. A poll
registration records its sidecar and check as device:inode, so every poll
armed before the remount failed strict validation and the watcher refused
all of them as unauthenticated state checks until each was re-armed by hand.
There are two device comparisons. fm_pr_private_file_valid compares a live
file's device with the state directory's device read in the same invocation:
it refuses a file that is not on the state directory's own filesystem and
already survives a renumber, so it is unchanged. The registration's recorded
identity versus the live identity (from #556, reused by the #932 retirement
receipt) binds the registration to the exact files published in its own
transaction; its device part is what breaks.
When strict capture fails, the watcher now proves the device is the only
difference: every other artifact check passes (template bytes, both hashes,
private mode, single link, live device, metadata), both recorded identities
name one device, and each recorded inode equals its live inode. Only then,
under the task's control lock, does it rewrite the two identity lines,
repeating the whole proof and comparing the registration's file identity and
bytes just before the rename, and then capture strictly again. A swapped,
altered, re-moded, relinked, split-device, or foreign-device artifact still
fails a proof and is still refused, and a pending retirement receipt blocks
the rewrite.
Reproduction: on macOS a poll armed on an APFS disk image that was detached
and re-attached behind another image moved st_dev 16777239 -> 16777243 with
inodes, bytes, mode, and link count unchanged; the real watcher refused it on
main and reports its merge with this change. The portable regression test
rewrites a real registration's recorded device and drives the watcher.
Not changed here: the status presentation cursor keys rows by its own
device:inode identity in bin/fm-classify-lib.sh, a different helper that
needs its own fix; a retirement receipt left by a reboot between its
publication and removal still names the old device and stays refused; custom
check trust binds only a content hash and is unaffected.
* fix(review): Serialize PR poll publication writers
* fix(review): Bound PR poll publication lock scope
* fix(bin): keep contribution records when the poll budget runs out (follow-up to #4627) (#4661)
A budget that expires partway through an observation no longer records an
error or prints the unavailable wake; the URL keeps its prior record and is
observed first next poll. forge() flags budget exhaustion at the point it
refuses, or when a read is killed at the budget's own deadline, so a genuine
forge failure still records the error and wakes. Each distinct URL is now
observed once per poll and applied to every owning task.
* fix(bin): clear parent pending-replies on local secondmate retirement (#4680)
* fix(bin): clear parent pending-replies on local secondmate retirement
Local secondmate teardown left resolved parent pending-reply records behind
after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced
retirement while any reply for that id is still unresolved, and delete every
matching record plus its delivery confirmation after a successful local or
remote retirement, matching the remote cleanup path.
* no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup
* no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung
* no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern
* no-mistakes(review): Pending-replies Basename und corr_id abgleichen
* no-mistakes(document): Clarify forced retirement pending-reply cleanup
---------
Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>
* fix(bin): accept Orca's composite worktree id when tearing down a task (#4677)
* fix(bin): accept Orca's composite worktree id at teardown
Teardown refused every Orca-backed task because the endpoint validator
checked orca_worktree_id with the simple-atom rule meant for tmux-style
window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca
returns that id as `<orca id>::<absolute worktree path>`, so the colon and
slashes in every real value made validation fail and finished Orca tasks
could never be cleaned up.
Validate the field as the composite it is: both halves of the first `::`
split present, the path half absolute, and no embedded newline, carriage
return, or tab. The terminal field keeps the atom check, which is correct
for it, and no other backend's validation changes.
The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca
never returns, which is why the suite passed a check the real value fails.
They now carry the composite form, so the tests exercise the real value.
* no-mistakes(document): name Orca's repo id in the composite worktree id
* no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points
* feat(bin): add opt-in typed dispatch resolution (#4692)
* feat(bin): add opt-in typed dispatch resolution through typesafe.ai
Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or
scout profile from a written brief with typesafe.ai's System One model:
one Choice question over the rules' `when` texts, then the confidence
floor, the rule's `approval` and `floor`, each profile's `provider` and
`floor`, one quota-axi snapshot, and the spendPriority argmax all in code.
It is off unless TYPESAFE_API_KEY is in the environment or the home's
gitignored .env; off means one stderr line, exit 0, and no network call,
so firstmate dispatches exactly as before. The key reaches curl on a file
descriptor, never argv.
Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and
the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new
tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates
the four new optional dispatch fields. Document the schema, the operator
contract, the AGENTS.md intake step, and the live and benchmark evidence.
* no-mistakes(review): Harden typed dispatch resolution and quota bounds
* no-mistakes(review): Validate dispatch floors and ranking evidence
* no-mistakes(review): Tighten dispatch response and floor evidence
* no-mistakes(review): Neutralize none matching and resolve defaults locally
* no-mistakes(review): Preserve providerless profiles outside typed resolution
* no-mistakes(review): Validate response usage and reject duplicate profiles
* no-mistakes(review): Escalate unverifiable floors and validate probabilities
* no-mistakes(review): Validate probability mass and unknown profile floors
* no-mistakes(review): Simplify resolver interface and preserve fallback routing
* no-mistakes(review): Fix constants and rank partial quota evidence
* no-mistakes(review): Add authoritative provider mapping and enforce explicit providers
* no-mistakes(review): Declare provider for documented Pi profile
* no-mistakes(review): Validate provider identifiers and support Gemini dispatch
* no-mistakes(review): Strictly anchor provider identifiers
* no-mistakes(review): Validate selectors and preserve fallback candidate evidence
* no-mistakes(review): Gate typed validation and harden resolver evidence
* no-mistakes(review): Preserve opt-in routing and harden candidate evidence
* no-mistakes(review): Prioritize known exhaustion over quota uncertainty
* no-mistakes(review): Isolate API secrets and preserve no-key diagnostics
* no-mistakes(review): Fallback safely when dispatch rules are absent
* no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets
* no-mistakes(document): Document typed dispatch safety and fallback behavior
* fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753)
* test: reproduce buried status declarations in shared readers
* fix: share status event reads and preserve open blockers
* fix: retain terminal scout and ship status declarations
* no-mistakes(review): Fix status chronology, legacy completions, and reader performance
* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots
* no-mistakes(review): Unify terminal supersession across cached folds and consumers
* no-mistakes(review): Filter per-key status history while preserving terminal chronology
* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells
* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses
* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor
* no-mistakes(lint): Quote literal done in test for-lists for SC1010
* ci: expect 19 snapshot/fleet-view tests
This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.
* no-mistakes(review): Restore multiline child outcome reporting
* no-mistakes(review): Select ledger terminal events through bounded shared reader
* no-mistakes(review): Report newest open decision instead of preferring blocked
* no-mistakes(review): Require colon before ship/scout terminal supersession in fold
* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix
* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions
* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold
* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule
* no-mistakes(document): Align status-read docs with fold-resolved crew state
* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers
* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree
* test: fold terminal-cleanup snapshot coverage into the completed-scout case
Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.
* no-mistakes(document): Clarify socket-down override expiry in architecture doc
* ci: retrigger flaky contribution check
* fix(bin): launch codex crewmates with codex's hook layer disabled (#4689)
* fix(spawn): launch codex crewmates with codex's hook layer disabled
A freshly launched Codex worker never reached its instructions. Codex
stopped it on an interactive "Hooks need review" modal whose selection
sits on "Review hooks", which is neither trusting nor declining.
Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow
navigation, so the selection cannot be moved, and pre-accepting the
prompt by writing Codex's own trust store would record an operator
consent that was never given.
The hooks are the machine's own ~/.codex/hooks.json plus any project's
.codex/hooks.json. A crewmate needs neither: its turn-end signal is the
-c notify= program on the same launch, and Firstmate's project hooks are
primary-session infrastructure that stands down in a child worktree.
Crewmate and scout launches now pass --disable hooks. That is the
opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted
hooks; disabling the feature runs none of them and leaves the operator's
~/.codex untouched. An unknown feature name is a hard Codex error, so a
release that drops the flag fails the launch loudly instead of silently
restoring the modal. A secondmate is a primary in its own home and keeps
the project hooks its turn-end guard and session-start digest ride on.
Verified on codex-cli 0.151.0: the modal is gone and the turn-end
notification still lands.
This unblocks the second review that every finished pull request is supposed to get.
Fixes kunchenguid/firstmate#4673
* no-mistakes(review): Fix contradictory hook count in Codex verification record
* fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710)
* fix(bin): settle terminal contributions and wake once per read-failure episode
A contribution whose last good observation is merged or closed is final:
poll no longer re-reads it, projection keeps it fresh, and a stale error
recorded beside it is cleared once. A genuine forge-read failure on an open
contribution still records its error on every cycle but prints the
unavailable wake only when it starts a failure episode; a successful read
ends the episode. Open PRs linked from done tasks keep being observed.
The false unavailable beside a complete observation was budget exhaustion
mid-observation, already fixed by #4661.
* fix(review): Settle terminal contribution owners
* fix(review): Deduplicate shared contribution failure episodes
* fix(test): Preserve settled terminal contribution records
* fix: select authoritative no-mistakes runs (#4476)
* fix(crew-state): select authoritative validation runs by identity
Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row.
Refs: https://github.com/kunchenguid/firstmate/issues/3215
* fix(review): Resolve same-branch run identities beyond capped history
* fix(review): Fix run-selection compatibility, races, and worker-state fallbacks
* fix(review): Limit run validation to the requested branch
* fix(test): Anchor AXI fixtures and document remaining live evidence gaps
* fix(document): Clarify run selection documentation and capture ownership
* fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation
* fix: distinguish captain outcomes from no-op updates (#4738)
* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work
MAIN answered a supervision-branch outcome for completed captain-requested
work (implementation done, PR ready for review and merge approval) with
"Captain, shipshape.", reading section 9's no-action reply as covering it
and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no
captain-facing response is owed".
Section 9 now limits the shipshape reply to true no-ops (idle re-read,
empty heartbeat, consequence-free acknowledgement) and requires a short
outcome response naming what finished and what word is needed whenever
requested work finishes or a result needs the captain's word, even when a
transcript entry already shows the substance. The Pi protocol's re-emit
rule now says it bounds repetition only, and carries a worked example of
the ready-for-review outcome whose correct processing turn a shipshape
reply fails.
No executable contract evaluates the content of MAIN's captain-facing
reply, so the regression is the protocol example in the owner doc rather
than a text-match test.
* no-mistakes(document): Clarify captain-facing outcomes versus no-ops
* docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line
The document step condensed the Pi protocol's re-emit rule and dropped the
worked example of a finished, ready-for-review outcome whose correct
processing turn a "Captain, shipshape." reply fails. That example is the
contract's regression: no executable contract evaluates the content of
MAIN's captain-facing reply, so the owner doc's example is the test case.
Restore it directly under the re-emit rule, prefixed as a regression
example that is kept verbatim and never condensed or summarized away.
* no-mistakes(review): Clarify captain outcome and decision-word requirements
* no-mistakes(document): Clarify captain-facing completion outcomes
* docs(pi): require the PR URL in the visible captain-facing outcome reply
Captain review on the regression example: drop the sample reply string
and say only that the ready-for-review outcome requires relaying a
captain-facing outcome response, not just "Captain, shipshape.".
Fold in the visible-PR-handoff failure seen this session: after the
branch outcome reporting this fix green, MAIN's visible reply was only
"Awaiting your merge call." with no PR URL, leaning on the dim anchor.
Section 9's URL rule now also covers a review or merge ask and names the
visible reply as where the URL goes, sourced from the ready status, pr=
metadata, or the supervision branch's summary and never left to a
transcript entry. The Pi protocol adds the same-way failure and places
the captain-facing text in the final visible assistant reply after the
fm_branch_processed call, because Calm hides assistant text emitted in
the same step as a tool call as a working note.
Investigation verdict, evidence in the PR comment: no recent PR caused
the handoff failure; Pi has hidden same-step pre-tool assistant text
since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and
#4658 touched only remote report transfer.
* no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs
* no-mistakes(document): Clarify captain-facing supervision outcomes
* docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule
The consolidated section 9 URL rule narrowed its trigger to a review or
merge ask, dropping the "whenever a PR is mentioned" catch-all from
#3648 that keeps every PR URL copied from a durable record and never
assembled from memory. Restore that trigger as a union with the review
or merge ask so the one consolidated rule covers both.
* fix(bin): let non-owner Claude Stops exit safely (#4777)
* Fix foreign-owner turn-end supervision loop
* no-mistakes(review): Scope foreign-owner safe exit to Claude guard
* no-mistakes(document): Document Claude foreign-owner safe exit
* fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778)
Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty
indexed array as an unbound variable and aborts the shell. In
signal_turnend_panes_churned() the missing_keys loop was reachable with
an empty array whenever every churned key already held a fresh
.churn-since-* marker (a second churning turn-end inside an open
deferral window), so each watcher cycle died about half a minute in and
supervision restarted endlessly. The created_keys rollback loops had the
same latent crash on their error paths.
Audit of bin/ for the same pattern found one more confirmed-reachable
case: remote_handoff's noncanonical-body scan iterates to_move, which is
empty when a retried remote handoff finds every key already staged in
the outbox. All other "${arr[@]}" sites are either count-guarded,
guaranteed non-empty by construction, or unreachable while empty.
Guard the three reachable expansions with the repo's existing
"${arr[@]+...}" idiom. New regression test drives a real watcher
through the all-marked churn path; the macos-stock-bash CI lane runs it
under real /bin/bash 3.2 via FM_TEST_ONLY.
* Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783)
The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: require complete captain-facing final responses (#4779)
* docs: require complete final responses across harnesses
* no-mistakes(document): Document complete final replies for Grok Bot
* docs: point Grok replies to the shared contract owner
* no-mistakes(review): Clarify final recap without batching decision asks
* fix: preserve substantive mid-turn text in Pi Calm (#4788)
* fix(calm): preserve substantive Pi mid-turn text
* no-mistakes(review): Preserve substantive Pi Calm text per block
* no-mistakes(test): Cover shared Calm preservation boundaries behaviorally
* no-mistakes(document): Consolidate Calm preservation documentation
* fix: harden mail checks and rebalance full-coverage CI (#4800)
* Improve CI reliability and rebalance full-coverage validation
* no-mistakes(document): Clarify lint partition documentation
* fix(bin): answer Kimi 2.0.0 folder-trust dialog during spawn (#4799)
* Handle Kimi workspace trust dialog
* no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers
* no-mistakes(review): Gate Kimi ready on any trust marker and clean captures
* no-mistakes(review): Read visible pane for Kimi trust and ready gates
* no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate
* no-mistakes(review): Harden Kimi viewport capture and trust dialog detection
* no-mistakes(document): Document Kimi spawn refusal on cmux and Orca
* fix(bin): report a dead-agent record once instead of escalating forever (#4775)
* fix(bin): report a record whose agent is gone once instead of escalating forever
The wedge escalation path never asked whether there was still an agent to be
wedged. A wedge is something stuck that might recover, so re-alarming it earns
its cost; an agent that is gone never moves again, its pane never churns, the
idle timer never resets, and the escalate path clears its own timer and re-arms
with nothing bounding the count.
Observed on a live fleet: two finished lanes reached 226 and 203 consecutive
escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400
notifications a day from two lanes with no agent running at all. On one,
fm-control.sh exit answered already-stopped and fm-crew-state.sh read
"failed - run failed". Closing the Herdr pane did not stop it either: with the
pane genuinely gone and herdr pane read returning pane_not_found, the count kept
climbing, because the poll is driven by the record's window= line rather than by
the pane. The cost is not the repetition but that it drowns the alarms that
matter.
fm_backend_agent_state already separates a thinking agent from a gone one at
process level. In the branch that was about to escalate, read it once and treat
only its two recovery-grade verdicts - dead (endpoint present, no agent in it)
and missing (endpoint authoritatively absent) - as proof, reporting that record
once and not re-escalating it while it stays that way. Every other verdict,
including alive, ambiguous, unreadable, unverified, and a read that failed
outright, keeps the identical schedule, reason, and escalation count, so a
genuinely wedged live agent is unaffected. The probe costs at most one backend
read per window per threshold, the same budget the declared-wait consult and the
worktree write probe already take.
The report decides nothing about the record's fate: both lanes still held
unlanded work and teardown refusing them was correct, so retiring, relaunching,
or cleaning up stays with the supervisor. The once-only marker is owned entirely
by that function and is dropped by the same read the moment the endpoint stops
reading gone, so a replacement launched into the same window escalates normally
and its own later death is reported again.
Related, and not closed by this: #4412, #4482, #4316.
Tests drive the real watcher against a record whose endpoint does not exist and
pin both directions: dead and missing report once and never advance the count
across later thresholds, while alive, ambiguous, and unreadable endpoints keep
escalating with the identical reason and a climbing count.
* fix(bin): bind the once-only dead report to the pane it reported
Review of the parent commit found a reachable sequence where a later death in
the same window lost its promised report. The marker was keyed on the verdict
string alone and dropped only when a threshold probe read a non-gone verdict,
but probes run only at thresholds: a replacement launched into the same window
that dies without ever being probed alive - it crashes at startup, or works and
then crashes - was absorbed by the previous death's marker. The pane's first
sight yielded only the generic stale wake and every later threshold matched the
stale marker, so the second death never got the detailed once-report that both
the function's own comment and docs/architecture.md promise.
Record the verdict together with the pane hash it was reported for, and absorb a
repeat only while both still match. A replacement churns the pane, which resets
the stale suppressor, wedge timer, and escalation count while no reset site
touches this marker, so the pane half is what tells the second death apart from
the first. The live-probe drop stays as it was.
Clearing the marker at those reset sites instead would re-open unbounded
re-alarming for a dead pane whose display ever ticks, which is the exact defect
the parent commit exists to close.
The noise bound is unchanged: an unchanged dead pane still absorbs on every
later threshold and never advances the escalation count, and every verdict short
of proof still escalates exactly as before.
* no-mistakes(review): Key the dead-record once-marker on the busy incarnation token
* no-mistakes(document): Document dead-record escalation cap in stale-pane config entry
* no-mistakes(document): Add busy-state inventory line to AGENTS.md
* no-mistakes(document): Document dead-record probe on busy-turn-bound wedge path
* fix(bin): create captain-hold rows when Beads requires due (#4854)
Captain holds have no due semantics and are a hold kind, not a Beads issue
type. The create path now waives due.required and maps to native type task.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: disable compact adviser for spawned agents (#4877)
* feat(bin): launch every spawned agent with the compact adviser disabled
Every crewmate, scout, and secondmate Firstmate launches now starts with
COMPACT_ADVISER_DISABLE=1, on a fresh spawn and on a relaunch alike, so an
unattended session never activates the compact adviser.
The value is unconditional: no configuration file gates it and there is no
override, unlike the trace carrier beside it.
Three carriers deliver it, because no single one covers every launch shape.
The pane shell receives an export beside GOTMPDIR, so the agent's own children
inherit it too.
The launch command carries an explicit assignment, prepended outermost so it
wins over any ambient value the pane already held.
The cleared launch environment sets it again at the `env -i` boundary and keeps
COMPACT_ADVISER_DISABLE in the fixed operational floor, which is what preserves
the switch when config/launch-env-allowlist empties the environment, and what
delivers it on a remote host that never had the value.
bin/fm-control.sh relaunch, the bootstrap secondmate relaunch, and the remote
secondmate transport all rebuild their launch through bin/fm-spawn.sh, so they
inherit the same floor.
The captain's own primary session is untouched.
The two new suites drive the real spawn and then execute the launch command the
pane actually received, with the harness replaced by a probe that prints its own
environment, rather than matching script text.
They cover ship and secondmate launches with the allowlist absent and enabled,
the pane export and its ordering, fm-control.sh relaunch, and the full parent to
remote-host chain.
* no-mistakes(review): Export compact-adviser disable across compound launches
* no-mistakes(document): Document spawned-agent compact-adviser environment guarantee
* fix(bin): preserve Claude lock ownership after helper recycling (#4894)
* fix(bin): let a background Claude session keep owning its session lock
Session-lock ownership was decided by process ancestry alone. Under an
unattended Claude session the model loop runs in a transient bg-spare
bridged to the front-end by a shared daemon; when that bridge is
recycled the contiguous claude-named ancestry from a hook to the
recorded owner breaks while the owner pid stays alive, so the Stop
auto-arm stood down as a foreign live owner, the turn-end guard ended
every turn with its read-only diagnostic, and fm-lock.sh refused - a
self-sustaining outage until restart.
Ownership is now ancestry membership OR a trusted same-session id,
never id-first:
- fm-session-lock-lib.sh accepts CLAUDE_CODE_SESSION_ID only when
CLAUDE_PID is a Claude-shaped member of the current contiguous run,
compares it against the id recorded in state/.lock-session, and
requires the recorded pid to still be a live harness. No id, no
sidecar, an untrusted id, a different id, or a dead recorded pid
leaves the ancestry verdict unchanged. Ids are never read from ps
argv.
- fm-lock.sh accepts a same-session holder at both refusal sites,
writes, refreshes, and clears the sidecar only under its claim lock
(including the early already-mine exit, skipped only while the
deferred startup sweep leases that lock), keeps it byte-identical
across a same-session confirmation, records CLAUDE_PID on lock line 1
for a session with a trusted id so a shared daemon or front-end that
outlives the session never keeps a dead session's lock alive, never
rewrites a live line 1 on a same-session confirmation, and names the
recorded id in the live-owner refusal.
- The .lock line-1 format is unchanged, so every reader that takes the
whole first line as the pid keeps working; the guard's foreign-owner
exit is unchanged and inherits the fix through the shared predicate.
Tests: the ancestry suite drives the ancestry and id signals apart in a
deterministic process table (asserting the divergence) and runs a real
orphaned front-end/daemon/pty-host/spare tree through six phases with
the real lock, auto-arm, and guard scripts; the foreign-owner repro
keeps its negative control and adds a same-id positive control.
Disclosure: no live unattended Claude background session ran on the
verifying machine. The topology is documented by the real process
listings in #3902, #2314, #3398, and #4066; coverage is the structural
predicate plus the executable fixtures, not a live pass.
Residual: bin/fm-sessionstart-nudge.sh keeps its own private ancestry
walk (it only decides whether to print a nudge) and may nudge on a
resume in the recycled case.
Out of scope, deliberately: no structured lock format, no guard budget
changes, no daemon-identity rejection, no fork lineage.
* no-mistakes(review): Wait for claim lock; revert failed sidecars
* no-mistakes(review): Revalidate ownership after wait; restore sidecars
* no-mistakes(review): Roll back sidecar by publication phase
* no-mistakes(review): Restore sidecar only if lock line is unchanged
* no-mistakes(review): Trust session ids without a spelling allowlist
* no-mistakes(review): Disarm sidecar rollback before backup cleanup
* no-mistakes(document): Updated session-lock ownership documentation
* feat: park main under the away posture on Pi (#4889)
* feat: park main under the away posture on Pi
While the away-posture record exists on a Pi primary, the supervision branch
takes every actionable wake, no processing turn opens on main, captain rows
accumulate for the return brief, and main's standing authority relocates to
the branch through the existing guarded scripts.
- lib/fm-branch-dispatch.ts: read the record at every routing decision; while
it exists claim check, decision-owned, and heartbeat rows too, keeping the
two broken-queue vetoes; expose checkSeqs so a claimed check row lifts task
scoping.
- fm-primary-pi-watch.ts: offer every actionable row under the record; a
declined wake and every watcher-failure alarm still reach main.
- fm-branch-supervision.ts: drop the legacy .afk decline; append a fixed
POSTURE: AWAY tail carrying the record's read-back verbatim per wake; open no
processing request while the record exists, re-checked immediately before a
request would open and at every run boundary; present the accumulated rows
at the first run boundary after archive.
- fm-lease-lib.sh: fm_lease_forbid_branch passes the branch for opted-in
actions only while fm-afk-contract.sh validate succeeds on a confirmed live
record; PR merge, fresh spawn, and decision answer opt in, local landing
never does.
- fm-send.sh: a --resolve-key naming an open needs-decision or captain-held
task is a decision answer and meets the partition; blocked: keys stay
steering.
- fm-spawn.sh: enforce the record's spend cap for a fresh ordinary spawn by
either actor; relaunches and secondmates exempt.
- fm-branch-prompt.sh: fixed Postures section and the verbatim
ask-user-authority policy; the prefix stays byte-stable.
- fm-afk-return.sh: count what the away session handled from the store.
- docs, afk skill, AGENTS.md stub: main parked on Pi, green merge gate
absolute while away.
- tests: watcher and branch extension suites, fleet-record, merge, and
decision-answer suites cover the relocation, the vetoes, the tail, the
parked processing turn, the cancellation, the re-presentation, and the
spend cap; dated live-guard evidence recorded.
* no-mistakes(review): Refuse branch merge after preflight archive race
* no-mistakes(review): Fix away wake, spawn, and processing races
* no-mistakes(review): Suppress parked processing; narrow away-only rejection
* no-mistakes(review): Abort dedicated processing; gate branch spawn once
* no-mistakes(review): Stamp away-only on the dispatch offer
* no-mistakes(review): Treat invalid away records as spend-cap absence
* no-mistakes(review): Drop spawn test hook; abort processing-opened runs
* no-mistakes(review): Bind abort to opening prompt; cap-read absence
* no-mistakes(review): Limit away branch spawn to queued work only
* no-mistakes(document): Correct AFK posture documentation
* ci: standardize workflow timeouts into three tiers (#4910)
* ci: simplify CI job timeouts to a three-tier policy
Replace the scattered per-job timeout values (10m parallel, 25m lint, 30m
serial, 10m macOS) with three readable tiers, each a hang tripwire with
headroom rather than a packing estimate:
- fast (5m): coverage guard, repo invariants, timing aggregate
- normal (30m, one shared budget): lint partitions, portable parallel
shards, portable serial shards, macOS stock Bash
- heavy (Herdr only): 20m step tripwire on the family run so always()
cleanup still runs, under a 75m job-level last-resort backstop
The workflow's header comment states the policy and points at
docs/fm-test-portable-shards.md "Timeouts", which now owns it, and each
job names its tier beside timeout-minutes. tests/fm-ci-workflow.test.sh
asserts the policy against the parsed workflow instead of the old
per-job minute values: every job joins exactly one tier, exactly three
distinct job-level values exist, the fast tier stays within 5-10
minutes, the normal budget stays at least double the modeled parallel
lane sum reported by fm-test-run.sh --check-coverage, and the Herdr step
tripwire stays below its job backstop with an always() cleanup after it.
Concurrency supersession, shard counts, lane membership, and fail-fast
settings are unchanged.
* no-mistakes(review): Decouple the normal timeout from packing estimates
* no-mistakes(review): Assert Herdr teardown follows the family run
* no-mistakes(review): Pin Herdr family-run timeout to 20 minutes
* no-mistakes(review): Ignore comments when identifying Herdr steps
* no-mistakes(review): Identify Herdr steps by declarative ids
* no-mistakes(document): Clarify authoritative three-tier timeout policy
* fix(bin): keep supervisor status closes from waking the same home (#4895)
* fix(bin): keep supervisor status closes from waking the same home
A drain that already folded OPEN DECISIONS has presented those bytes even
when the watcher has no matching seen marker. Treat that fold, and the
presentation cursor, as known so the bookkeeping close stays quiet while
later worker lines still signal.
* no-mistakes(review): Keep folded worker failures waking past supervisor closes
* no-mistakes(review): Wake on unlisted folded worker lines; batch multi-key closes
* no-mistakes(review): Stop folded worker resolved lines from counting as already read
* no-mistakes(document): Correct self-announced close marker contract in docs
* fix(bin): stop labeling Herdr as experimental (#4972)
* Stop steering operators away from Herdr
* no-mistakes(review): Neutralize remaining Herdr opt-out documentation wording
* fix(bin): treat a live no-mistakes run as current after rebase (#4973)
* fix(bin): treat a live no-mistakes run as current after rebase
A running run on the task's branch is authoritative regardless of head.
Matching only the local head made a rebased in-flight run look failed.
* no-mistakes(review): restrict coarse live-any-head to foreign-branch answers
* no-mistakes(review): reject gate-parked runs from the executing predicate
* no-mistakes(review): hoist gate-marker patterns into single run-lib owner
* no-mistakes(review): require live daemon for head-free run binding
* no-mistakes(review): require answered daemon-down before unbinding live runs
* no-mistakes(review): extend daemon guard to anchored continuation routes
* no-mistakes(review): delete live-any-head; restore dead-daemon verdict
* no-mistakes(review): keep parked gates parked; name dead daemon everywhere
* no-mistakes(review): set dead-daemon verdict instead of emitting early
* no-mistakes(review): align selected route with legacy dead-daemon handling
* no-mistakes(review): drop unproven-record binds; narrow coarse gate reading
* no-mistakes(review): narrow header, drop vestigial guard, retarget tests
* no-mistakes(review): revert coarse gate override; require answered-down probe
* no-mistakes(review): cache one daemon probe; stop duplicating run id
* no-mistakes(review): restrict coarse dead-daemon verdict to moved-off rows
* no-mistakes(review): delete coarse dead-daemon extension and gate note
* no-mistakes(review): delete remaining coarse dead-daemon block and stale docs
* no-mistakes(document): document rebase-safe live-run bind and unverified-record verdict
* fix(bin): prevent long worker launch command truncation (#4994)
* fix(bin): stage the launch command in a private file and type a short source line
A long launch line typed while the fresh pane shell is still busy waits in the
terminal's canonical line buffer, which drops input past about 1,024 bytes on
macOS, so the pane was left at an unfinished command with no agent running.
fm-spawn now writes the assembled command to the task's own temp root under
umask 077 and types only a short line that sources it.
Refs #4559
* fix(bin): keep the per-task temp root private before staging the launch command
The root lives at a predictable path under /tmp and now holds the whole launch
command. Create it with mode 0700, refuse one that already exists as anything but
a directory owned by this user that nobody else can write, and tighten an owned
one, so no other local user can plant or swap the staged file.
Refs #4559
* fix(bin): enforce private staged launch file mode
* test(spawn): cover long staged Claude launches
* no-mistakes(review): Namespace launch files and prove truncation staging
* no-mistakes(review): Use immutable per-spawn launch filenames
* no-mistakes(document): Document staged launch delivery safeguards
* no-mistakes(ci): Updated eight behavior tests/fakes to execute or inspect immutable staged launch files instead of expecting inline launch commands. This restores Muse, secondmate lifecycle/restart, remote trace/parent binding, compact-adviser, and Orca coverage. All affected tests, dispatch-profile regression, fixture tests, syntax checks, ShellCheck, and git diff checks pass
---------
Co-authored-by: Vytautas Stankus <svycka@gmail.com>
* test: authorize isolated Herdr lab validation (#4998)
* Add isolated Herdr runbook to test instructions
* no-mistakes(review): Drop substring matching from test.instructions contract
* no-mistakes(review): Assert commands.test key absence in YAML
* Drop unit-first sentence and instructions contract test
Captain-scoped follow-up on the Herdr-lab test.instructions ship:
keep the lab safety runbook only, and leave the no-mistakes contract
test focused on commands.test absence.
* docs(vision): accept vendor-semantics and 9k AGENTS ceiling (#4873) (#5001)
* docs(vision): accept vendor-semantics and 9k contract-ceiling amendments (#4873)
Replace the pixels-of-today's-UI rule with a quarantined, version-pinned
surface-adapter exception recorded as standing debt. Cap the always-loaded
contract at 9,000 words and require prune-or-trigger before a crossing change
lands.
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
* docs(vision): restore accepted three-sentence vendor-semantics form (#4873)
Replace the compressed paraphrase with the issue's accepted wording:
a named quarantined version-pinned adapter, expected to break, recorded
as standing debt that never hardens into a shared contract.
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
* feat(bin): defer the wedge escalation for a lane parked at a supervisor-owed gate (#4974)
* fix(watch): recheck a gate awaiting a human instead of wedge-escalating it
A lane whose validation run is parked at a gate waiting on a human
decision is correctly quiet, but nothing in its status line says so: the
evidence is the pipeline's own gate state rather than anything the worker
wrote. The wedge timer read that silence as a suspected wedge and climbed
the escalation ladder for as long as the wait lasted, and each escalation
cost a supervising turn. The landed declared-wait consult does not reach
it, because a live ordinary crewmate never reports a declared pause, and
raising FM_STALE_ESCALATE_SECS would delay genuine wedge detection for
every lane by the same amount.
The threshold now reads a second, independent record when the status line
accounts for nothing: whether the crew's current state is a gate whose
answer is owed by a human. That is minted only from the gate's own
findings table, by a row whose `action` column is exactly `ask-user`,
located by position out of the table header the way nm_gate_step_row
already reads its row - never searched for over the run payload, where a
finding's free-text description or a branch name satisfies a search just
as well. A gate awaiting the CREWMATE's own answer keeps the unchanged
escalation schedule, reason and demand-deep-inspection wording, because a
crewmate that goes quiet before answering its own gate is exactly the
wedge the ladder exists to catch.
Each kind of wait now carries the human it is on, the action that clears
it, and whether that human is the captain as data alongside the verdict,
rather than as wording chosen per branch where the recheck is written, so
the deferral cannot word one kind of wait as another and a new kind
cannot ship without deciding all of them. A parked gate has no written
record of when its wait began, so its recheck publishes no wait age at
all rather than one read from the quiet window this deferral resets on
every pass, which would report the same small number for a gate of any
age. Like every other captain-facing recheck here it is absorbed in
silence while the away-posture record exists, arming no throttle, so the
recheck is owed in full the moment the record is archived.
The consult runs only in the at-threshold branch that was about to
escalate, beside the worktree walk already there, and only for lanes
whose status line explained nothing.
Closes #3055
* no-mistakes(review): require an unanswered decision before deferring a parked gate
* no-mistakes(review): reset the away-silenced timer, fail-safe findings parse, US-joined wait records
* test(watch): pass the pane hash wedge_timer_check now takes
Upstream gave wedge_timer_check a sixth <pane-hash> argument for its
dead-record probe. The malformed-wait-record rounds drive the real function
directly, so they pass one, and stub fm_backend_agent_state to a live agent so
the probe that runs after a refused deferral keeps the unchanged ladder rather
than reading a backend the child shell has none of.
* no-mistakes(review): Bind parked-gate wait to its run, owe it firstmate
* no-mistakes(document): correct wait-kind count, crew-state reader scope, gate-key coupling
* feat(watch): make the parked-gate wait deferral opt-in
The wedge timer deferring a lane parked at a validation gate is new
supervision behaviour rather than a restored one, and it decides which
lanes give up the escalation ladder, so it now ships as a default-off
per-home option instead of changing every home on upgrade.
config/wedge-defer-parked-gate arms it. The flag is read before the
decision fold, so an unconfigured home spends no fold or current-state
read, writes no record, and keeps the unchanged escalation schedule,
reasons and demand-deep-inspection wording; a test counts the reader
calls in both directions to pin that.
It is not inherited by secondmate homes: each home supervises its own
crew and owns that trade separately, the same reason
config/turnend-churn-absorb is home-local.
The away-posture absorb returns to leaving the idle timer alone, which
it had restarted only because the costly consult could reach it. A
parked-gate wait is owed to the supervisor rather than the captain, so
it never enters that branch, and the recheck owed on return is again
owed in full the moment the record is archived.
* test(watch): pin that the away-silenced hold leaves the idle timer alone
The absorb no longer restarts the timer, so the recheck owed on return is
owed in full rather than a cadence into the return. Nothing asserted
that, so a restart could be reintroduced silently.
* no-mistakes(review): document away-silence rationale, pin captured gate component
* no-mistakes(test): anchor gate row scan to the braced findings header
* no-mistakes(document): pin same-block gate row invariant in crew-state comment
* fix(bin): reclaim a task whose herdr endpoint was destroyed (#5007)
* fix(control): let the owning seat reclaim a task whose endpoint is gone
A destroyed pane or workspace made `missing` a terminal state. Relaunch
accepted only `dead` and said to stop the agent first; exit refused
`missing` and said to reconcile the task first; there is no reconcile
verb. Each command named the other as its prerequisite, so a task whose
terminal went away could not be reclaimed by anything, and a no-mistakes
approval it was parked on had no seat left to answer it.
`missing` is agent-free a fortiori: there is no endpoint, so there is no
agent in it. Widen the existing guards rather than add a verb.
- fm-spawn --relaunch accepts a positively proven `missing` and creates
one fresh endpoint in the recorded worktree; the record it already
republishes rebinds the task to it. A `dead` endpoint is still adopted
in place.
- fm-control exit reports `endpoint-gone` instead of dying, so the
relaunch transaction's stop step no longer dead-ends, and re-resolves
the endpoint from the record before verifying the replacement.
The duplicate-agent refusal is untouched: both verdicts come from the
same recovery-grade classifier, which claims `missing` only from positive
absence, so `alive`, `ambiguous`, and `unreadable` all still refuse. The
backends' own create paths refuse a live same-labeled endpoint as a
second independent guard. The worktree, its branch, commits, uncommitted
changes, armed poll and registration, record rows, and status log are all
untouched - a reclaim is a recovery, never a teardown.
A secondmate is excluded: its gone-endpoint recovery already has one
owner in the session-start liveness sweep, so relaunch refuses and names
it rather than becoming a second path to the same outcome.
Tests reproduce both halves of the deadlock, the reclaim succeeding,
unlanded work surviving it, and the refusals that still hold.
* no-mistakes(review): prove endpoint absence per backend before reclaim rebinds
* no-mistakes(review): give exit and relaunch one absence proof; pin herdr rebind session
* no-mistakes(review): narrow endpoint reclaim to herdr; tmux refuses honestly
* no-mistakes(review): stop refusals and docs asserting unestablished causes
* no-mistakes(review): stop herdr fixture helper losing tmp-root registration
* no-mistakes(review): document workspace drift and absence-probe server residue
* no-mistakes(review): correct rebind limitation to its one reachable case
* no-mistakes(review): stop claiming reclaim leaves instructions untouched
* no-mistakes(document): scope fm-control-lib purity claim, note reclaim coverage
* no-mistakes(rebase): read the staged launch file in the herdr fixture
Rebasing onto main picked up #4994, which stages a long worker launch
command into a script and delivers the short `. '<path>'` line instead of
the literal command. The tmux fake and tests/fixtures.sh were updated for
that; the herdr fake this branch adds was written before it and still
keyed "an agent now exists on this pane" off the literal
`encode launch-brief` text, so after the rebase it never marked the
rebound pane live and the reclaim's alive-wait read `dead`.
Dereference the staged file first, exactly as the tmux fake above does.
Test-fixture only; no production path changes.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* no-mistakes(document): note reclaim placement in herdr and scripts inventories
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat(bin): stamp status events with their emission time (#3764)
* test(status): reproduce missing event emission time
* wip(status): preserve optional event emission time
* test(status): document indirect clock stub invocation
* no-mistakes(review): Preserve historical status bytes during reply recovery
* no-mistakes(test): Fix timestamped status assertions and remote fixture dependencies
* no-mistakes(review): Preserve captain regex overrides for timestamped status events
* no-mistakes(document): Clarify status event timing and publication contracts
* no-mistakes(lint): Quote literal done to satisfy ShellCheck
* no-mistakes(ci): Captain, updated .github/workflows/ci.yml to expect 19 snapshot tests instead of 18, matching the PR’s added regression. Reproduced the failure before the fix. Stock Bash 3.2.57 verification passed: parse sweep, 19 snapshot tests, 53 Bearings tests, and the public-followup regression. Workflow lint and diff checks passed
* no-mistakes(test): Preserve terminal notifications with malformed timestamp tags
* no-mistakes(test): Stamp Rovo spawn failures with emission time
* no-mistakes(document): Verify status event documentation
* no-mistakes(lint): Fix ShellCheck quoting in status emission-time tests
* no-mistakes(ci): Captain, fixed four lifecycle assertions to accept emission timestamps while preserving publication and retry checks. Reproduced the CI failure before the fix. The lifecycle suite now passes with six Beads capability skips; syntax, targeted ShellCheck, and diff checks passed
* no-mistakes(ci): Captain, fixed malformed timestamp colons hiding actionable events using shared normalization. Original bytes and unknown ages are preserved. Regression reproduced before the fix; classifier and remote-reply suites, targeted lint, syntax, and diff checks passed
* no-mistakes(review): Stamp remote escalations at call sites, drop new flag
* no-mistakes(review): Accept stamped escalation and close lines in test assertions
* no-mistakes(review): Restore reserved-key answered-note guard for stamped closes
* test(status): accept optional emission time in PR-provenance assertions
The #4148 provenance test landed on main with exact unstamped greps.
Parent-channel lines from this branch carry [at=<epoch>], so strip only
that tag before the same exact match. No production change.
* no-mistakes(review): Accept stamped ready signal in PR fallback scrape
* no-mistakes(review): Drop relay flag, stamp parent events at call sites
* no-mistakes(review): Stamp worker terminal-signal instructions, revert fm-on fixture
* no-mistakes(review): Accept optional stamp in live cmux drift guard
* no-mistakes(review): Restore original test invocation order in two suites
* no-mistakes(review): Strip only well-formed numeric status time tags
* no-mistakes(document): Drop stale unstamped PR-ready line spelling from channel doc
* no-mistakes(review): Stamp agy spawn-failure status lines with event time
* fix(bin): normalize status event times in-shell and freeze the budget test clock
Two paths made a status event's emission time cost more than it should.
The captain-relevance fallback piped every line through awk to drop a
well-formed `[at=<epoch>]` tag before matching, so a supervisor sweep paid a
fork per line just to prepare a regex match. Shell parameter expansion does the
same strip with no fork, and the retry-dedup scan now reuses that one helper
instead of carrying a …
jorguez96
added a commit
to jorguez96/firstmate
that referenced
this pull request
Sep 25, 2026
…ts resolved (#19) * fix(bin): support process events under symlinked homes (#3484) * fix(bin): resolve process-event state roots before validating them The process-event module validated the caller's spelling of a home's state root instead of the directory it operates on: it required the supplied path to equal its own lexical normalization, which rejects any path reached through a symlinked ancestor. On macOS both /tmp and $TMPDIR are symlinks, so an operator home under either could never claim a source. Reconcile still reported the runner started, while the detached runner died writing "cannot claim source" to the discarded stderr, and the source silently never fired. Resolve the state root to its physical directory once, then apply the existing private-directory validation to that resolved directory and derive every path, recorded claim identity, and later confinement check from it. This keeps the confinement contract for the directory actually operated on rather than only for callers that already spelled it physically, and removes the window where an ancestor symlink could be repointed between check and use. Homes already spelled physically behave identically. This was the single cause of both deterministic macOS failures in tests/fm-procevent.test.sh ("reconcile never claimed the registered source") and tests/fm-procevent-when.test.sh ("the winning concurrent arm did not produce an outcome"). The new case pins the behavior with an explicit symlinked-ancestor home, so it fails without the fix on any platform rather than only where the temp root happens to be a symlink. * fix(bin): pin the external capture staging boundary to its physical path The extension capture path pinned its registry staging boundary by comparing `pwd -P` against the caller-spelled registry directory, so a home reached through a symlinked ancestor still refused to start an extension-backed source after the state root itself resolved correctly. That left such a home half working: built-in sources ran while external ones failed. The staging preparer now prints the physical registry directory it validated, matching the inbox and reservation preparers beside it, and the start path pins on that returned path. The new end-to-end case drives the shipped file-signal package from a symlinked home spelling. * no-mistakes(review): Propagate canonical process-event state roots * no-mistakes(review): Propagate canonical state to process-event adapters * no-mistakes(document): Document physical process-event state roots * fix(pi): deliver captain outcomes as deterministic transcript entries (#3312) * fix(pi): persist captain outcomes visibly * no-mistakes(review): Recover captain outcomes after cold-start lock acquisition * no-mistakes(document): Document cold-start captain-outcome recovery * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes(review): Prove immediate Pi captain-outcome transcript delivery * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * fix(pi): process captain outcomes through a sequence-keyed turn PR #3312 made every captain-facing supervision outcome a durable, exact-once visible transcript entry with the read cursor advancing only after that entry exists. That is the display half of the delivery contract. Left alone it turns a probabilistic silent loss into a deterministic one: the captain sees an anchor line, and firstmate never acts, because nothing opens a turn and nothing records whether main ever processed the outcome. The 2026-08-31 timeline showed the two shapes this must survive on the previous hidden-turn path: seven delivered decision outcomes each answered by an empty assistant message (cursor advanced, no retry, unanswered for close to three hours), and two answered by an unrelated prior reply. Both happened because delivery advanced the cursor at enqueue and accepted whatever the next assistant message was. Add the processing half on top of the persistence half: - bin/fm-branch-outcome.sh keeps a processed marker separate from the read cursor (`unprocessed`, `mark-processed --through`, `processed-init`). It only advances through an explicit sequence-bound acknowledgement, never past the read cursor and never backwards; an absent marker reads as zero and `processed-init` migrates delivered history once so an upgraded home is not re-presented its past. - After the visible entry for a captain outcome exists, the extension hands every still-unprocessed captain row to main as one hidden, typed `fm-branch-process` request listing each `[seq N] task: summary`, opening exactly one main turn. Main closes it only by calling the new `fm_branch_processed` tool with the highest sequence listed. An unrelated, empty, or paraphrased answer leaves the sequence open, and the same request is presented again at the end of the next main run and at session start. The first two presentations of a sequence set open a turn of their own; after that the request rides the captain's next prompt so an ignored request cannot loop, and a session replacement resets that budget. Routine outcomes stay turn-free. - The regressions cover exactly those incident shapes against the real store scripts: an empty answer and an unrelated prior answer neither advance the marker nor stop re-presentation, the acknowledgement is refused beyond the read cursor and outside lock ownership, a partial acknowledgement keeps the newer sequence open, and #3312's own assertions now forbid an unkeyed turn rather than any turn. The store suite pins the marker's bounds and the migration; the real-SDK guard for appendEntry persistence and model exclusion is unchanged. Docs move the protocol from "no model turn" to "one sequence-keyed processing turn closed only by its acknowledgement", and the verification record carries the dated run against Pi 0.84.4. * no-mistakes(review): Harden outcome listing and sequence-bound acknowledgements * no-mistakes(review): Harden outcome state validation and request pacing * no-mistakes(review): Reject unsafe sidecars and unterminated outcome stores * no-mistakes(review): Validate canonical mark-read cursor state * no-mistakes(review): Guard cursor advancement against corrupt processed state * no-mistakes(review): Bind acknowledgements to active processing requests * no-mistakes(review): Reset pacing when processing sequence membership changes * no-mistakes(review): Enforce silent outcome invariants at storage boundary * no-mistakes(document): Document hardened captain outcome processing contracts --------- Co-authored-by: kunchenguid <kun@kunchenguid.com> * feat: add bounded concurrent Bearings ledger collection (#3481) * feat: bound Bearings remote ledger collection * no-mistakes(review): Clarify default remote-ledger collection behavior * no-mistakes(review): Detach reconcile delivery from watcher loop * no-mistakes(review): Enforce bounded snapshot and request captures * no-mistakes(review): Bound legacy summary capture before parsing * no-mistakes(review): Bound primary remote ledger captures * no-mistakes(document): Correct snapshot and reconcile documentation * no-mistakes(lint): Fix ShellCheck quoting in bounded collector * no-mistakes(ci): Fixed all three CI failures: updated the macOS Bearings assertion to 44 tests, made the home-summary test deterministic and aligned with default ledger consumption, and increased the asynchronous reconcile retirement wait for loaded CI. Verified both focused suites, all 44 Bearings tests, ShellCheck, actionlint, Bash parsing, and git diff checks * test: await reconcile request retirement * no-mistakes(review): Avoid empty reconcile queue process churn * no-mistakes(review): Read ledger summaries from immutable snapshots * no-mistakes(review): Reject multi-document home ledger streams * no-mistakes(review): Coalesce durable reconcile requests per target * no-mistakes(review): Unify reconcile keys and reject snapshot streams * no-mistakes(review): Key reconcile requests by stable target ID * no-mistakes(document): Document per-target reconcile request coalescing * no-mistakes(lint): Remove unused snapshot summary file variable * no-mistakes(ci): Adjusted the concurrent collector regression’s end-to-end timing ceiling to account for stock macOS process/jq overhead outside the three-second remote collection budget, while remaining below the 15-second serial-read floor. Verified with stock /bin/bash 3.2: all 44 Bearings tests pass; bash syntax and git diff checks pass * no-mistakes(ci): Fixed legacy summary validation to require exactly one top-level JSON document and added behavioral regression coverage. Stabilized CI by conditionally waiting longer for durable reconcile delivery and synchronously stopping the fm-on worker tree before fixture cleanup. Removed a redundant flaky healthy-path timing assertion; the wedged-reader test still proves concurrent bounded collection. Verified fm-bearings-snapshot, fm-secondmate-reconcile, and fm-on tests, plus project ShellCheck, bash syntax, and git diff checks * ci: rebalance portable serial test shards (#3489) * fix(ci): rebalance the portable serial shards on measured durations The "Behavior portable serial 3" shard ran 17-20 minutes against its 20-minute job cap and intermittently timed out seconds after a passing test, on branches and on main alike. Shards are packed longest-processing-time from per-script duration hints, and those hints were last measured on 2026-08-21 at 116 scripts. The lane has since grown to 139 scripts and from ~42 to ~63 minutes: 17 scripts had no hint at all and fell back to the 20 s default, and several existing hints were low by 2-5x (fm-watch-triage 142 s hinted vs 263 s measured, fm-public-followup 36 s vs 197 s). The partition therefore looked perfectly balanced in hint space, 734.6 s per shard, while really running 11.5, 13.6, 18.8 and 16.5 minutes. Script-count balance, which is what the tests asserted, stayed normal throughout and hid it. Refresh the hints from the timing artifacts of three green runs, taking the slowest measurement of each script so the balance holds on a slow runner, and split the lane across five shards instead of four. Replayed against those runs' real per-script durations the worst shard is now 12.54 minutes, 63% of the unchanged 20-minute cap, and the serial lane's wall clock drops from ~20 to ~12.5 minutes. Bound the drift that caused this rather than relying on the hints being refreshed by hand: the coverage guard now reports the unmeasured share as serial_unhinted= and refuses past PORTABLE_SERIAL_MAX_UNHINTED_PERCENT, which leaves room for newly added tests while making a stale table fail the guard instead of silently pushing one shard into its cap. No test changes what it asserts and no test stops running; only the partition across shards changes. * no-mistakes(document): Clarify conservative shard timing aggregate * fix(pi): fall back on incomplete supervision branch prompts (#3491) * fix(pi): fall back after settled branch errors * no-mistakes(review): Detect provider errors across prompt compaction * no-mistakes(review): Preserve in-flight branch state across selection changes * fix(pi): re-probe supervision branch after cooldown (#3497) * fix(pi): recover supervision branch after cooldown * no-mistakes(review): Defer branch recovery until prompt settlement * no-mistakes(document): Clarify supervision cooldown recovery contract * fix(bin): remove legacy remote snapshot reads (#3501) * refactor: remove legacy remote summary reads * no-mistakes(document): Document ledger-only snapshot reads * no-mistakes(ci): Fixed the snapshot test fixture so ledger refreshes use the same fake executable PATH as the snapshot consumer. This preserves observable endpoint freshness after removing legacy summary computation. Verified stock Bash parsing and all 44 Bearings tests pass under /bin/bash; git diff checks pass * no-mistakes(ci): Fixed the CI-only snapshot fixture failure by ensuring the bounded-ledger refresh uses its fake tmux backend. This removes host tmux availability as a source of nondeterminism. Verified all 44 Bearings tests pass, Bash syntax passes, and git diff checks are clean * no-mistakes(ci): Fixed CI nondeterminism in the Bearings fixture: all local ledger refreshes now use the fixture’s fake tmux backend when available, instead of depending on host tmux state. Verified stock /bin/bash syntax, git diff checks, and all 44 Bearings tests with a deliberately failing host tmux * fix(pi): preserve watcher continuity across session replacement (#3498) * fix(pi): rearm watcher after session replacement * no-mistakes(review): Queue actionable closes across Pi session replacement * no-mistakes(review): Stop replacement arm when handoff persistence fails * no-mistakes(review): Preserve actionable wakes through branch and late child races * no-mistakes(review): Surface late handoff failures without crashing Pi * no-mistakes(review): Coordinate replacement delivery settlement and unique handoff tokens * no-mistakes(review): Retry stale deliveries and release settled claims * no-mistakes(review): Distinguish branch settlement and retry handoff cleanup * no-mistakes(review): Deduplicate persistent handoff cleanup alerts * no-mistakes(review): Acknowledge watcher follow-ups only when consumed * no-mistakes(review): Persist idle follow-ups until agent consumption * no-mistakes(review): Preserve pending outcomes when handoff persistence fails * no-mistakes(review): Arm replacement before awaiting prior delivery settlement * no-mistakes(review): Adopt pending handoffs after lock reclamation * no-mistakes(review): Prevent stale generations from adopting replacement handoffs * no-mistakes(review): Scope replacement handoffs by watcher state * no-mistakes(document): Clarify replacement handoff documentation * no-mistakes(ci): Fixed the failing branch-extension tests to model the new settlement-promise contract. Failure cases now assert that delivery ownership returns to the watcher instead of expecting direct extension fallback. Verified the updated branch suite, Pi watcher suite, shell syntax, and diff checks * no-mistakes(review): Update branch settlement tests and preserve chunked outcomes * no-mistakes(document): Document watcher-owned replacement handoffs * no-mistakes(document): Verify replacement handoff documentation * test(pi): cover watcher-owned branch fallback * no-mistakes(document): Refresh watcher-owned fallback documentation * fix(bin): resurface task statuses missed by wake handling (#3495) * fix(bin): resurface terminal statuses lost after branch handling * test(watch): canonicalize process-event fixture homes * no-mistakes(review): Index branch outcomes by causal status position * no-mistakes(review): Recover outcome indexes and deduplicate resurfaced statuses * no-mistakes(review): Handle legacy ambiguity and oversized status diagnostics * no-mistakes(review): Keep unclassifiable oversized statuses silent * no-mistakes(document): Document lost-wake outcome backstop * no-mistakes(document): Update outcome backstop documentation * no-mistakes(ci): Fixed CI regressions in wake-drain: parseable reserved-key decisions can no longer bypass the durable decision-fold guard, and status output is prepared and receipt-committed before presentation to prevent repeated one-shot outcomes after later failures. Added a behavioral regression for receipt commit failure and retry. Targeted backstop, correlation-token, decision-cursor, open-decision, unread-status, syntax, and diff checks pass locally. Shard-4 failures appeared unrelated/flaky; the network-parallel test passed locally * no-mistakes(ci): Fixed the Greptile P1 data-loss issue by committing presentation receipts only after prepared output reaches stdout. Added behavioral coverage proving output failure leaves the backstop retryable and receipt failure may duplicate but never lose a presentation. Relevant wake-drain suites and syntax/diff checks pass. The shard-4 Pi extension failure is unrelated to this PR and did not warrant changes * no-mistakes(ci): Stabilized tests/fm-bootstrap-network-parallel.test.sh by replacing scheduler-sensitive equal-sleep timing with bounded synchronization between mocked fetch and remote probes. This preserves detection of real serialization while avoiding false failures under CI load. Verified with five consecutive test runs, bash syntax validation, ShellCheck, and git diff checks. The separate Pi stock-rendering failure reproduces locally but is unrelated environment/version drift * no-mistakes(ci): Fixed Behavior portable serial 4 by adding fm-classify-lib.sh and fm-timeout-lib.sh to the broken-root Pi test fixture; fm-branch-outcome.sh now depends on them. Verified the full Pi branch-extension suite with real-Pi checks skipped, the wake-drain outcome-backstop suite, Bash syntax, and git diff checks. Greptile findings are already addressed at HEAD; the no-mistakes attestation failure is external head-SHA state * fix(bin): collect follow-up results from remote work homes (#3503) * fix(bin): deliver typed terminal results from remote work homes A public commitment whose work is bound to a REMOTE secondmate home could never receive its typed terminal result. `fm-public-followup.sh brief` printed an emit command carrying this home's own absolute path and this checkout's own script path, neither of which exists on the machine the worker runs on, so the worker had nothing it could write to that the owning home would ever read - and `consume` kept finding nothing while the promise stayed open. The brief is now route-aware: for a remote work home it prints that route's own code root and home with `--stage-in`, so the typed event is staged in the home where the work actually runs, and the closing paragraph names the owning home as the one on the other machine instead of pointing at the path above it. The owning home collects those staged results over the same SSH route it reaches that secondmate on, because the transport only runs outbound: `consume` pulls them into its own inbox and reconciles them exactly as it reconciles a local report. Collection is non-destructive until the result is durably held, so a dropped connection cannot lose a terminal result, and a route that could not be reached is named in `consume`'s output with the promise left open rather than reported as an empty inbox. A local work home is untouched: the brief still prints `--home` with this home and this checkout's script, and the event still lands directly in this home's typed terminal-result inbox. This is the emit-side counterpart of the retire/clear fix in #3479 and reuses the remote-route resolution that landed with it. Reconciling a loop bound to a remote route now reaches that route, so the existing remote cases drive `consume` through the same faked transport their other steps already use. * no-mistakes(review): Fail loudly on unresolved routes and invalid staging homes * no-mistakes(review): Fail collection when remote outbox is unreadable * no-mistakes(review): Surface reassigned remote routes during empty collection * no-mistakes(review): Fail remote collection on invalid registrations * no-mistakes(review): Reject unsafe registration entries during remote collection * no-mistakes(review): Restore healthy empty remote collection behavior * no-mistakes(review): Skip remote collection for delivered registrations * no-mistakes(review): Skip delivered registrations before route validation * no-mistakes(document): Document remote follow-up collection semantics * fix(bin): exclude secondmates from home-summary validity (#3504) * fix(bin): exclude secondmates from home-summary child inventory kind=secondmate meta records never have backlog rows, so counting them in unowned_children or terminal_in_flight made a clean main home look invalid once earlier ledger checks passed. * no-mistakes(review): Cover terminal secondmate in-flight exclusion * no-mistakes(ci): Updated the stock macOS Bash CI snapshot expectation from 15 to 16 tests. Verified all 16 snapshot/fleet-view tests pass under Bash 3.2.57 and `git diff --check` succeeds * fix(bin): self-heal outcome indexes on first drain (#3509) * fix(bin): self-heal status-outcome indexes on every drain Missing ready markers were skipping the lost-wake backstop on non-Pi homes because only the Pi branch ran processed-init. Drain now rebuilds those indexes under the outcome lock and fails closed only on a real store fault. * no-mistakes(review): Guard held-lock initialization and fail marker writes * no-mistakes(document): Document cross-harness outcome-index self-healing * fix(bearings): keep active children underway during captain holds (#3505) * fix(bearings): keep active children underway beside a captain hold Project each readable home's active children into Underway independently of the home-level captain-decision classification so a hold no longer hides live work. * no-mistakes(review): Preserve Underway repos and disclose child truncation * no-mistakes(review): Fall back to task project for Underway repos * no-mistakes(ci): Updated the stock macOS Bash CI assertion from 44 to 45 Bearings tests, matching the newly added behavioral regression. Verified all 45 tests pass under /bin/bash, Bash syntax checks pass, and git diff validation is clean * fix(pi): settle watcher delivery on Pi accepting the follow-up (#3513) * fix(pi): settle watcher delivery on Pi accepting the follow-up A follow-up queued while main is streaming joins the running run without ever raising before_agent_start, so waiting on that event before clearing the successor pipeline (#3498) stalled every later actionable close: no successor started, no wake was delivered or offered to the branch, and the turn-end guard woke main to re-arm by hand after every close. The pipeline now settles once Pi accepts the follow-up. Consumption is observed at before_agent_start for an idle main and at the user message_start for a streaming main, and decides only what a replacement session (/new, /resume, /fork, reload) replays. An exhausted restoration delivers its typed failure without launching an arm past the retry bound, which the stall had hidden. The replacement-coordinator map is typed so the strict no-emit typecheck passes again. Tests: the doubles no longer raise before_agent_start for a streaming send, a portable regression drives two actionable closes while main streams and proves the successor chain plus consumption-scoped replay, and a credential-free real-SDK probe pins Pi's event contract for both the streaming and the idle follow-up. Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a * fix(pi): retry a verified successor that fails during wake delivery A verified successor can exit while the wake it was started for is still being delivered, most plausibly during a branch turn that holds the settlement for minutes. Its failure close arrived while the pipeline's single-flight guard was set, so the close handler skipped the retry, and the pipeline's end no longer launched an arm, which left the live generation with no watcher and no retry timer. The close handler now records that failure when the child had reported readiness and was not retired by the restoration itself, and the pipeline runs the ordinary bounded, lock-checked retry for it once the delivery settles. A restoration started for a later pending supersedes it, and an exhausted restoration still hands repair to main without a further arm. The regression holds a branch settlement open while the verified successor exits with a failure and proves one retry watcher starts after the settlement releases, none while it is held. Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a * fix(bin): bound repeat stale wakes for parked workers (#3532) * fix(bin): bound repeat stale wakes for a parked but live worker A worker parked on a declared wait - `paused:` for an external or pipeline wait, or a verified `captain-held` transfer - kept waking firstmate far inside FM_PAUSE_RESURFACE_SECS. Observed as five consecutive alarms on one captain-held worker and dozens across a day on a pipeline wait, and reported upstream as four wakes in 75 minutes against a 3600s window. pause_state_class deliberately answers `none` for a still-live agent even under a declared wait, so a worker genuinely waiting on a decision is never silenced. That classification is correct and is left alone; it routes every parked but live worker through surface_nonterminal_stale on first sight of each distinct stale hash, and an idle parked pane still churns its hash on a clock or a token counter without changing what is being waited on. Two places let that churn re-alarm: - surface_nonterminal_stale queued the wake BEFORE consulting whether a wait was declared, then wrote `.paused-resurfaced-<key>` - the very throttle that should have suppressed it. The throttle was never read on this path and was advanced by the wake it should have prevented. - The hash-change path cleared that throttle through clear_pause_tracking whenever the classification came back `none`, so each tick also bought the same declared wait a fresh window. Fixing only the first site changes nothing. Read the throttle before anything is queued and advance it only on a wake that really fires, and on the hash-change path reset only the per-hash bookkeeping while the declaration still stands, via a clear_stale_hash_tracking split so neither half of clear_pause_tracking is duplicated. The throttle is keyed to the declaration, not to the pane. First sight still wakes, so an inconclusive state is still inspected, and the window's end still re-surfaces once, so a forgotten wait cannot rot invisibly - noise traded for a bounded cadence, never for silence. The wake identity stays the plain `stale: <win>` the away-mode handoff depends on. Tests cover both observed forms and were confirmed to fail against three deliberate breaks: each site reverted on its own, and a re-surface that never fires again. * fix(document): Clarify declared-wait wake cadence documentation * fix(ci): Captain, fixed the stale-throttle inheritance: cadence markers now bind to the current wait declaration, so replacement paused and captain-held waits each emit their first plain `stale:` wake. Added behavioral coverage for both forms. Bite proof failed as expected when identity matching was removed, then passed after restoration. Full watcher triage suite, `bin/fm-lint.sh`, syntax checks, and diff checks pass. Changes remain uncommitted for the outer executor * fix(ci): Captain, fixed the confirmed Greptile finding. `resurface_absorbed` now applies a throttle only when its stored declaration scope matches the current wait, so replacement `paused:` and `captain-held` waits surface immediately without changing classification. Added executable coverage for both absorbed forms. Bite proof failed before the fix at the intended assertion; afterward the full watcher triage suite, `bin/fm-lint.sh`, shell syntax checks, and `git diff --check` passed * fix(bin): accept the away-mode daemon as the turn-end supervision owner (#3567) * fix(turnend): accept the away-mode daemon as the supervision owner While state/.afk exists the away-mode daemon owns supervision and runs bin/fm-watch.sh one-shot: the watcher exits on every wake and the daemon starts its replacement. The turn-end guard tested for a live watcher process holding the watch lock at that instant, so a turn boundary that landed in the hand-off blocked with "TURN WOULD END BLIND" while supervision was completely healthy, costing a full handling turn each time. Reproduced with the real daemon wrapping the real watcher and the real guard sampling the same home: 6 of 40 samples blocked, every one of them with the daemon alive and the beacon 2-3 seconds old, and a new watcher pid on each cycle. After the fix the same reproduction blocks 0 of 40, and killing the daemon and its watcher (away mode still on, beacon still fresh) blocks again. The guard now accepts a live, identity-matched daemon holding this home as proof of supervision while away mode is active. The identity match is the same discipline the watcher lock uses, so a recycled pid or a lock left by a killed daemon proves nothing. The fresh-beacon half of the predicate is unchanged: a daemon that stops restarting its watcher still blocks once the beacon passes grace, a home with no supervisor blocks exactly as before, and with away mode off the strict watcher predicate is untouched. The predicate reads only durable state, so it behaves identically for every primary harness and runtime backend. * no-mistakes(document): clarify away-mode daemon supervision proof and test coverage * no-mistakes(document): generalize stale turn-end predicate summary in architecture.md * fix(backlog): omit --file from row probes for non-markdown backends (#3582) * fix(backlog): omit markdown file for beads probes * no-mistakes(document): Narrow backlog addressing doc to mutations for backend-aware probes * no-mistakes(ci): Fixed the Greptile P2 review comment (the only failing check) on tests/fm-backlog-atomicity.test.sh. The comment correctly noted that an exported TASKS_AXI_BACKEND environment variable would inherit into the spawned scripts and, because fm_tasks_axi_backend gives it top precedence, override each test case's .tasks.toml backend fixture — making the backend-specific argv assertions fail for environmental reasons. Fix: unset TASKS_AXI_BACKEND in the test harness right after sourcing tests/lib.sh, with a comment explaining why, so every case deterministically exercises its declared backend (4 lines added; no production code touched). Verified: reproduced the leak before the fix (TASKS_AXI_BACKEND=beads made the markdown dispatch case fail with 'beads show failed', exactly the reported failure mode); after the fix the full suite passes (0 failures, exit 0) both with and without TASKS_AXI_BACKEND=beads exported. The added lines are shellcheck-clean (the only shellcheck note, SC1091 on the lib.sh source line, pre-exists this change) * fix(bin): classify progress updates on requested work as routine (#3589) The supervision branch's verdict rule escalated every outcome that answered a captain request, so "the work started" and "still working" notes reached the captain with nothing to look at. The rule now keeps a finished result of requested work captain-facing, even when healthy, and treats start or still-working updates that bring no new artifact, finding, or decision as routine. The captain list for review-ready PRs, ask-user findings, exhausted blockers, credentials, and destructive or security-sensitive cases is unchanged, as are the unsolicited-routine, silent-fleet-review, and doubt-chooses-captain rules. The fm_branch_report tool description and the two docs that restated the old unconditional rule now point at the prompt's "Verdict: routine or captain" section as the one owner instead of carrying a second copy. * fix(bin): preserve captain calls during teardown (#3595) * fix(bin): never close a captain call during cleanup A scout that held its own work item for the captain, which is what captain-hold-lifecycle prefers ("hold the work item the question gates"), was closed by bin/fm-teardown.sh's automatic backlog transition. The completion gate passed, cleanup ran, and the captain's question moved to Done with no recorded answer: the one thing the policy says must never happen. `tasks-axi done` closes a held row silently, and nothing in teardown asked whether the row was the captain's own call. bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when the task is still an open captain call, 1 when it is not, 2 when that cannot be established. It reads the row through the transition library's backend-aware probe, so it addresses the same backlog teardown does; the script's other commands now address the configured data directory the same way instead of FM_HOME, which also fixes captain holds in a home with a relocated data directory. Teardown asks `open` before any destructive step and refuses on 2. On 0 only the close changes: after cleanup and still under the task's own lock, the row gets one "Deliverable of the finished work" line at the end of its body and returns to Queued through `tasks-axi reopen`, keeping its hold, so it lands in Captain's Call instead of reading as work under way. --force does not lift this: it authorizes discarding unlanded work, never the captain's question. The deliverable goes into the body because `tasks-axi update --report` rewrites the title of a row that is not Done. The crash window reuses the pending-close record teardown already stages: a `mode=retain` line makes the existing replay record the deliverable and reopen instead of closing, with the same validator, stale-generation check, cleanup-incomplete marking, and non-blocking bootstrap lock as an ordinary close. A retained row the captain answered first simply retires the record. No parallel record type, recovery command, or second bootstrap loop is introduced. Regressions run the real executables: the captain-held scout survives cleanup queued, held, with its deliverable and on the board, only `answer` closes it, --force keeps it open, and an ordinary scout still closes with its report; an interrupted cleanup leaves the row untouched and the next session start retains it; a relocated backlog keeps the retention in its one configured file; and a ship row whose hold cannot be read refuses cleanup before anything destructive. Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np * no-mistakes(review): Serialize captain holds and fix backend-aware listing * no-mistakes(document): Update captain-call retention documentation * no-mistakes(document): Fix relocated captain-hold backlog diagnostics * fix(bin): deliver secondmate outcomes to the parent channel (#3592) * fix(bin): deliver every secondmate outcome on the parent channel from the recording scripts A secondmate's captain-facing outcomes could miss: the mate model addressed the captain in its own unread chat instead of appending to the parent channel, and a PR-ready report, a finding, a decision, a blocker, and a failure all depended on that one remembered append. Make delivery structural, so the parent channel never depends on the model: - bin/fm-parent-channel-lib.sh is the one owner of channel resolution and exact-line append-once; the merge outcome path and the inactive-outcome scan now publish through it instead of two private copies. - bin/fm-inactive-reconcile.sh gains a ledger-first path that runs on every watcher poll in a secondmate home: a direct child's whole terminal done or failed line is delivered at once with its note, recorded PR, mode, merge posture, and scout report pointer, keyed and receipted so it is delivered once, and the inactive path yields to it. `report <task-id>` runs the same delivery for a caller holding the child's meta lock. - bin/fm-pr-check.sh publishes the PR-ready line with the canonical URL at registration. - bin/fm-captain-hold.sh publishes a hold and its answer, keyed by task id and resolution-record count, with no new persisted state. - bin/fm-teardown.sh delivers the child's final line before removing its record and refuses, retaining every record, while the channel cannot be written. - The charter opens with the parent-channel rule and confines the mate's own appends to judgement; AGENTS.md carries the carve-out at the persona address rule and the escalation list. docs/secondmate-parent-channel.md records the design and its coverage, and docs/verification/secondmate-parent-channel.md records the live run with real tmux panes and both real watchers delivering every line with no model. Supersedes #3569. * no-mistakes(review): Fix parent outcome retries and reconciliation locking * no-mistakes(review): Prevent busy children from starving ledger delivery * no-mistakes(review): Correct ledger metadata and hold occurrence handling * no-mistakes(review): Disambiguate ledger outcomes and normalize hold reasons * no-mistakes(review): Close ledger races and preserve teardown records * no-mistakes(document): Correct parent-channel receipt and scanner documentation * no-mistakes(lint): Quote done arguments for ShellCheck compliance * no-mistakes(ci): Fixed both CI failures. Updated GOTMP teardown fixtures for the new final-outcome reporter and isolated them from host tmux state. Updated the PR security assertion to distinguish the accepted PR-ready line from duplicate merge outcomes. Verified with both failing test suites, bash syntax checks, and git diff checks * no-mistakes(ci): Fixed Greptile’s duplicate-delivery race in bin/fm-inactive-reconcile.sh. Ledger events now claim matching already-delivered inactive receipts using the prior status fingerprint, preventing duplicate parent reports while preserving later same-state completions. Added behavioral regression coverage. Verified inactive-reconcile tests, project lint, documentation audience checks, syntax, and diff checks. Teardown tests passed relevant cases before the documented pre-existing herdr-preflight-missing-adapter failure * fix(bin): sync remote second mates to primary commit (#3599) * fix(bin): sync remote second-mate homes to the parent primary commit Session start and remote launch pointed a remote second-mate home at whatever Firstmate copy its own host kept, so a home that had already advanced past that copy refused as a non-fast-forward and every other home stopped at the host's older commit while the primary ran ahead. The parent now resolves ITS primary default-branch commit with the existing helper and hands that commit to the host on both paths. Because a remote home is a standalone clone, the host imports that one commit before advancing - already present, else from that host's Firstmate copy without moving it, else from the home's own origin - and then runs the SAME ff_target guards a local home gets, so dirty, diverged, feature-branch, and unresolvable targets skip untouched and the ancestry rules keep one owner. An unimportable target now names /updatefirstmate instead of failing opaquely, and a host still running an older Firstmate copy is reported the same way rather than echoing a bare refusal. The host-local launch leg no longer re-runs its own secondmate sync, so the spawn it drives cannot re-target that host's copy after the parent has already converged the home. /updatefirstmate is unchanged: it still refreshes the remote code root from that host's origin and then syncs the home to that refreshed copy, which is what the sync call with no target commit means. * no-mistakes(document): Document primary-targeted remote secondmate synchronization * fix(bin): separate captain intent from firstmate specs (#3597) * fix(bin): split brief task into captain intent and firstmate spec Keep no-mistakes --intent as the captain's ask plus later captain words, not the build spec or worker tradeoffs. * fix(bin): stop task-subsection copies at the next heading Promotion was swallowing the scout Setup contract into Firstmate spec, and pre-subsection briefs lost their # Task body. * no-mistakes(review): Validate brief content and preserve nested specifications * no-mistakes(review): Scope placeholder validation to scaffold-only subsection bodies * no-mistakes(review): Ignore fenced subsection headings during brief validation * no-mistakes(review): Preserve captain intent across scout promotion * no-mistakes(review): Enforce safe intent boundaries for legacy promotions * no-mistakes(review): Allow marked legacy intent and reject empty promotions * no-mistakes(review): Scope task parsing and overlay legacy intent contracts * no-mistakes(review): Overlay current intent contract for all no-mistakes spawns * no-mistakes(review): Preserve later captain clarifications in intent overlays * no-mistakes(document): Document brief intent enforcement and ownership * no-mistakes(ci): Updated spawn-related test fixtures to use valid Captain intent and Firstmate spec subsections, corrected launch-path expectations to launch-brief.md, and resolved ShellCheck quoting findings. Verified with fm-lint.sh and 15 affected behavior tests, including real Herdr tests; all passed * no-mistakes(ci): Updated stale spawn/promotion fixtures in the Muse, Orca, secondmate-harness, and public-followup suites to provide valid Captain's intent and Firstmate spec subsections. Verified full Orca and secondmate-harness suites, targeted public-followup promotion behavior, Bash syntax, diff checks, and fm-lint * fix: start a fresh supervision branch for every main session (#3600) * fix(pi): start a new supervision branch conversation per main session The supervision branch reopened one recorded conversation forever, so every main session start reloaded the current generated prompt and then weeks of accumulated thread, where a superseded rule could still outweigh today's. The branch conversation is now scoped to one main session: the session generation owns the recorded conversation, so a cold start, /new, /resume, /fork, or a reload always builds a new one, while a rebuild inside one session (a model or effort change) still continues that session's own conversation. The dialog mirror re-anchors with it. Its durable cursor records what the previous branch conversation received, so a /resume or reload - which keeps main's own session file - would otherwise leave the new branch blind to dialog main itself still has. The reset is bounded by the current main session, and the cursor keeps advancing incrementally within it. The durable outcome store and its processed marker are untouched, so unacknowledged captain-facing outcomes still re-present on the new main session. * no-mistakes(document): Document fresh Pi supervision conversations * no-mistakes(ci): Fixed the flaky concurrent inbox failure. Lock acquisition now retries when a competing lock disappears between a failed claim and inspection. Added a behavioral regression covering that race. Verified the full inbox test four times, project lint, and git diff checks * feat: restart second mates after instruction updates (#3614) * feat(update): restart second mates whose instructions changed /updatefirstmate pulled new bytes onto disk and then asked each advanced second mate to re-read them. A running agent holds AGENTS.md and every loaded skill frozen from launch and no verified harness offers a reload, so that steer could not reach a loaded skill at all and left the mate holding two contradictory copies of its own job description. An eligible mate is now restarted instead, in the same home and endpoint, through the existing transactional relaunch. The restart is gated on the mate first writing down the open work it holds only in conversation - the open-record half of /stow, never its memory sweeps - so an unregistered captain call is flushed before the conversation is spent. Anything that leaves the reload unprovable falls back to the old re-read message and is reported as exactly that, never as a clean reload. Remote mates take the same path: fm-remote-secondmate-control.sh gains a relaunch verb whose host-local leg runs that same control plane, since the mate is an ordinary local secondmate from its host's point of view. The primary resolves the profile and passes it explicitly, because config/secondmate-harness is not inherited and the file on that host belongs to a different home. fm-update.sh now splits its advanced live mates into a restart set and a nudge residual, and both sets require a changed instruction surface, which also closes the over-nudge against the session-start sweep. Restart is stricter still: a bin/-only advance reloads itself on the next call, so it never costs a conversation. Colocated tests cover the gating, the persist-then-restart order, the task-subset persist request, each unsafe fallback, the remote hop, and the remote sync's new instruction-surface report. * no-mistakes(review): Fix restart correlation, concurrent waits, and lifecycle reporting * no-mistakes(review): Parallelize relaunches and classify replacement incarnations * no-mistakes(review): Gate restart actions on live agent state * no-mistakes(review): Handle failed restart workers without hanging * no-mistakes(review): Nudge legacy remotes and preserve persist recovery * no-mistakes(review): Document one-time secondmate restart rollout * no-mistakes(review): Honor arrived replies and refresh remote profiles * no-mistakes(review): Revert remote parent profile reconciliation * no-mistakes(review): Reset remote profile defaults and honor published results * no-mistakes(review): Preserve fallback nudges for unverifiable secondmates * no-mistakes(document): Document second-mate restart update flow * no-mistakes(lint): Fix ShellCheck warnings in restart scripts * perf: accelerate local validation with bounded concurrency (#3644) * perf(tests): route gate verification through the bounded concurrent runner Local validation was the pipeline's dominant cost: across 67 recorded no-mistakes agent sessions on this repo, 99.3% of command execution was `bash tests/*.test.sh`, run strictly one script at a time, and 2% of those calls were killed by an agent-guessed timeout and paid for twice. Three changes, each measured: - `.no-mistakes.yaml` pins `commands.test` to `bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner already owns changed-file selection, bounded concurrency, the refusal of unproven scripts, and a generous automatic per-script bound, so the gate's baseline is neither a serial chain nor a guessed timeout. It stays intent-targeted - the Test step still runs its evidence agent on top - and excludes the live-Herdr family the required Herdr lane owns. - `bin/fm-test-run.sh` gives a plain list of script paths the same bounded automatic scheduler and automatic bound that `--changed` gets. Naming several subjects is how a verification round asks for exactly those scripts. The curated selections are untouched: `--lane` still composes CI shards whose serial lane must stay serial, `--family` is what the required Herdr lane runs, and `--all` stays a deliberate complete regression. - `pr-forge` is admitted to the concurrent-safe family registry on two consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those, and records `secondmate` and `session-bootstrap` as refused with the exact script and reason each failed on, so the refusals are actionable rather than silent. Measured on this host, 0 failures on both sides: verification round, 4 scripts 448s chained -> 231s through the runner (-48%) pr-forge family 409.2s at 1 worker -> 237.9s at 4 (1.72x) watcher-wake-lock family 1311.1s at 1 worker -> 539.3s at 4 (2.43x) A fourth lever was implemented and then removed because the measurement refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made `fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s unchanged, back to back. Those sleeps are not overhead added to the clock - they are how a test waits for a subject moving on fm-watch.sh's own one-second cadence - so sampling less often only delays detection. It also broke `fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a settled condition. CONTRIBUTING.md records that result so the experiment is not repeated. * no-mistakes(review): Separate concurrent runs by isolation proof family * no-mistakes(review): Limit automatic timeouts to changed-file validation * no-mistakes(document): Clarify validation concurrency documentation * fix: copy PR URLs from durable records (#3648) * fix: copy PR URLs from records or abstain, never assemble them Supervision reported a plausible but dead PR link three times because its prompt demanded a full https:// URL at a moment when only a PR number was observable, so the model assembled an owner/repository from memory, and the PR check then accepted that URL and wrote it into the task record, after which the model kept defending its own tool-endorsed guess over the worker's real link. Three changes close that chain without any live forge lookup, so private forges are treated exactly like public ones: - bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy or abstain" section requires a URL to be copied verbatim from a durable record (the done: PR <url> status line, pr= metadata, or the backlog note), forbids assembling owner, repository, host, or number from memory, and has the branch report only the identifier it actually holds when no record names the URL yet, leaving the PR check unarmed until the worker's ready line arrives. AGENTS.md section 7 and 9 carry the same copy-or-abstain rule for main in place of the bare full-URL mandate. - Worker briefs (bin/fm-brief.sh, ship and scout rules) require the full https:// URL wherever a PR is mentioned - status line, terminal, or summary - never a bare "PR 108", so the link is in view as early as the number is. - bin/fm-pr-check.sh refuses, offline and before any side effect, a URL that the task's own done lines contradict, printing both spellings; a log naming no URL still records the argument as before. fm_pr_status_ready_urls in bin/fm-pr-lib.sh owns reading those lines. The refusal also reaches bin/fm-pr-merge.sh, so nothing merges under a contradicted URL. Tests cover the offline refusal with zero side effects, the recorded spelling being accepted, markdown-wrapped and punctuated URLs, working lines not counting, the merge wrapper propagation, a self-hosted merge request with no forge call, the prompt carrying the rule, and the brief carrying the worker rule. * no-mistakes(review): Remove stale PR URL enforcement * no-mistakes(ci): Removed backlog notes as an accepted PR identity source. PR URLs may now be copied only from the task’s `done: PR <url>` status or canonical `pr=` metadata; otherwise supervision reports only the known identifier and leaves PR checking unarmed. Updated related guidance/docs and verified with branch-supervision tests, brief tests, ShellCheck, and `git diff --check` * fix(bin): disable Claude feedback drafts for fleet launches (#3661) * fix(bin): disable Claude's feedback-draft flow for fleet-launched agents Scope --settings '{"feedbackDrafts":"off"}' to every Firstmate-launched Claude crewmate and secondmate, so /bug and /feedback never queue or submit a bug report on the captain's behalf. feedbackDrafts is the documented settings key (Claude Code changelog 2.1.247); the per-launch CLI flag never touches the captain's global settings.json. Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3 * no-mistakes(review): Prevent managed settings from re-enabling Claude feedback drafts * no-mistakes(document): Fix Claude feedback documentation formatting * fix(bin): layer both feedback-draft controls for defense in depth The prior --settings-only fix can be overridden by a managed Claude settings policy (feedbackDrafts precedence). Keep CLAUDE_CODE_SEND_FEEDBACK=0 alongside --settings '{"feedbackDrafts":"off"}': either control alone disables the SendFeedback tool, so a managed override of one still leaves the other in force. Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3 * no-mistakes(document): Document Claude feedback-draft suppression ownership * feat(tests): run three more validation families concurrently (#3662) * perf(tests): admit three more families to concurrent validation The three families that `docs/fm-test-isolation-proof.md` recorded as refused were not refused for concurrency. Each blocker was a test that decided a property by wall clock, or a script filed where it cannot run. Fixing those three things admits all three families and recovers 28.6 minutes of local validation with no assertion removed or weakened. - `tests/fm-backlog-handoff.test.sh` injected its pre-move crash by killing the handoff, sleeping a fixed second, then delegating the move to the real binary. Nothing ever killed the fake, so on a host slow enough for the case's next assertions to take longer than a second, the orphan woke and completed the very move the case requires left undone, and recovery then failed with `Task "pre-move-crash" not found in this backlog`. Watching the two backlogs during the injected crash showed exactly that, the item moving one second after the crash. All four crash injections in the file now go through a new `fm_fake_crash_injector` shim that signals the target and returns only once it is observably gone, and the pre-move fake never delegates the move at all. - `tests/fm-session-start.test.sh` proved the startup digest does not block on a slow current-state read by timing the whole digest against a fixed eight-second sleep, which a loaded host exceeds without the property being violated. It now holds that read open until the case releases it and asserts, the moment the digest returns, that the read has not finished. A digest that waited would wait indefinitely rather than for an interval a slow host can out-run, so the assertion is stronger than the bound it replaces. Its scan budget moves to the maximum, because the old value left two seconds of margin over the fixed sleep and measured the host rather than the deadline that `tests/fm-inactive-reconcile.test.sh` owns. - `fm-backend-herdr-focus-flash-e2e` was filed in the family map's catch-all, which put it in the portable serial lane, where Linux CI gate-skips it: that real-Herdr regression was running nowhere. It moves to `real-herdr-gated` and the required Herdr lane. `fm-claude-stop-autoarm-live-e2e` gate-skips on its opt-in variable and moves to `live-harness-optin`. The 28 remaining ungrouped scripts become an enumerated `standalone` family instead of admitting `unclassified` itself. `unclassified` is the family map's `*)` arm, so admitting it would silently grant concurrency to every test added afterwards, which is exactly the population with no proof. A new test still lands in `unclassified` and stays serial, and `tests/fm-test-run.test.sh` covers that split behaviorally. Each family passes two consecutive four-worker proofs with zero failures. On the production runner, `secondmate` goes 1233.1s to 453.4s, `session-bootstrap` 756.4s to 286.4s, and `standalone` 724.6s to 261.1s: 2.71x overall and 1713.2s recovered. The whole suite runs 177 scripts in 52.6 minutes of wall clock against 121 minutes of summed script time. * no-mistakes(document): Refresh concurrent validation and shard documentation * no-mistakes(ci): Fixed the real-Herdr focus-flash E2E race exposed by reclassification. Part C now starts its persistent child atomically via `pane run` and verifies stable child identity through Herdr’s public `process-info` interface, avoiding the racy send-text/send-keys sequence and platform-specific `ps` matching. Verified with bash syntax checking, ShellCheck, git diff checks, and the complete E2E test on Herdr 0.8.2 * feat: structure no-mistakes ask-user escalations (#3670) * feat(brief): structure no-mistakes ask-user escalation as event + snapshot file Crewmates escalating a no-mistakes ask-user gate now report one status event naming every finding id plus a snapshot file holding the gate's axi finding records verbatim (id, severity, file, line, description, authority), using the same shape even for a single finding. The status line never paraphrases. The format is defined once in fm-dod-lib.sh and rendered into both the scout and ship rule 6 in fm-brief.sh, so a promoted scout - whose rule 6 fm-promote.sh preserves unchanged - gets the identical contract as a freshly-spawned no-mistakes ship worker. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PpiWaDerbYavTLPPtEjQei * no-mistakes(review): Preserve ask-user escalation output contract * no-mistakes(review): Align escalation format test expectation * no-mistakes(review): Scope ask-user escalation instructions correctly * no-mistakes(review): Remove ask-user from generic decision rules --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> * fix(bin): require self-sufficient no-mistakes intent (#3671) * fix(bin): require a self-sufficient no-mistakes intent A no-mistakes worker's --intent is only as useful as the string it passes. PR #3604 shipped with an intent that was only "do 1, 2, 3, 7 from the report": the real contract lived in a private scout report and never reached --intent, so nobody holding that string plus the codebase could have derived the specification. This is pure instruction at the contract's one owner; no spawn-side or promotion-side check is added. - bin/fm-dod-lib.sh: the generated no-mistakes Definition of done now states that the --intent string must be self-sufficient (the string plus the codebase reconstructs roughly the same specification) and tells the worker to write the substance of any report, decision, or PR the captain's intent refers to into --intent rather than the pointer, while Firstmate build instructions and the worker's own decisions still stay out. The spawn-time overlay points back at that rule so its "supersedes" wording cannot cancel it, and the header's owner statement carries the rule. - AGENTS.md section 11 and bin/fm-brief.sh's header ask Firstmate to include the substance of referenced material when filling ## Captain's intent, and section 11 points at the owner of the rule. - tests/fm-brief.test.sh and tests/fm-task-delivery.test.sh assert the rendered brief and launch contract carry the rule. Claude-Session: https://claude.ai/code/session_01YMhEe42q7BAAoN6RxNuzim * no-mistakes(document): Replace incident-specific intent test commentary * fix: accelerate local Bearings snapshot composition (#3499) * Speed local fleet snapshot composition * no-mistakes(review): Stabilize task inventory during concurrent snapshot composition * no-mistakes(document): Document local snapshot observation concurrency * no-mistakes(ci): Fixed CI failures by making empty task manifests compatible with stock macOS Bash 3.2, snapshotting task metadata before concurrent observations to prevent generation drift, strengthening the behavioral race regression, and updating the stock-Bash Bearings test count to 45. Verified fleet snapshot tests (15), Bearings tests (45), workflow lint tests, project lint, Bash 3.2 parsing, and diff checks * no-mistakes(ci): Fixed the Linux CI failure caused by passing large backlog/task JSON through jq command-line arguments, which exceeded the per-argument size limit. Both inventory projections now stream large JSON inputs through stdin. Verified with fm-bearings-snapshot.test.sh (45 tests), fm-fleet-snapshot-view.test.sh (15 tests), Bash syntax, and git diff checks * no-mistakes(ci): Fixed concurrent task teardown during metadata capture: vanished metadata is now omitted while genuine copy failures remain fatal. Added a deterministic public Bearings regression test and updated CI’s expected test count. Verified with the full Bearings suite, workflow-lint suite, Bash syntax checks, and git diff checks * no-mistakes(ci): Fixed PR-caused CI and review issues: streamed large fleet JSON through jq stdin to avoid Linux argument limits, kept crew-state reads bound to captured metadata generations, and strengthened the behavioral race test. Bearings (46 tests), fleet snapshot (15 tests), crew-state, backend, lint, Bash syntax, and diff checks pass locally. Serial shard 5’s unrelated task-inbox segmentation fault appears infrastructural/flaky * no-mistakes(ci): Fixed endpoint-state generation crossing by validating captured spawn_gen before and after local endpoint probes, falling back to exact metadata identity for legacy tasks. Stale probe results now become unknown instead of false unhealthy state. Added a behavioral relaunch-race regression test. Verified the full Bearings snapshot suite, shellcheck, bash syntax, and git diff checks * fix(snapshot): keep live observations generation-coherent * no-mistakes(review): Keep secondmate observations generation-bound without copying reports * no-mistakes(document): Document generation-coherent snapshot observations * test(bearings): measure local read overlap instead of wall-clock budget The large-local-snapshot regression asserted that a whole snapshot composed in under five seconds. That bound measures how loaded the host is, not whether the per-task reads actually overlap, so it failed intermittently on a contended machine: one run in six on a box at load 16-20, landing exactly on the five second boundary. Time a serialized run and a concurrent run of the same workload instead and require the concurrent one to save at least two seconds. Both runs pay the same composition overhead, so the difference isolates the overlap this change delivers. Five one-second reads serialize into five seconds and overlap into about one, and re-serializing the reads collapses the saving to roughly zero, so the assertion still fails loudly if the concurrency regresses. Also bump the pinned Bearings test count to 48, since rebasing onto the current default branch picked up its captain-hold test. * no-mistakes(review): Restore JSON-derived decision flags * no-mistakes(review): Unify status-derived snapshot observations * no-mistakes(ci): Updated the stock macOS Bash CI check’s Bearings test count from 48 to 49. Verified the full Bearings suite passes and emits exactly 49 TAP successes; git diff checks pass * fix: prevent stale supervision wake loops (#3672) * fix(bin): stop the supervision branch's stale-ack and ghost-report loops Clean-slate implementation of the four authorized recommendations from the supervision-ghost-retrigger analysis (items 1, 2, 3, and 7), in their minimal form, superseding PR #3604: - fm_branch_report refuses a task the wake being handled never named. The extension fixes the reportable task set from the eligible rows before each prompt (signal and stale rows resolve to their tasks, a heartbeat allows any task with a live record, fleet is always allowed), so a report typed from memory about a task whose records teardown already removed is never stored or delivered. - An acknowledgement that consumes nothing says "nothing was acknowledged through N" and prints the exact --ack-through / --recovery-generation command for the current presented wake, instead of "re-run the drain", which re-fed the same stale acknowledgement in a loop. - bin/fm-guard.sh no longer tells the branch actor to drain queued wakes while it is handling them; it names the granted rows instead. - Teardown removes state/.<task>.branch-outcome-index for ordinary tasks and descendants; the index rebuild and the append-side index write both skip a task with neither a live record nor a status log, so the branch's report of a teardown it just performed is stored without recreating the index. No new locking, no spawn-generation binding, and no retired-task refusal: the branch can still report the outcome of a task it just tore down, and the teardown test now proves that path end to end. * fix(bin): narrow the branch report scope and guard silence to the minimal form Apply the four review decisions on the clean-slate branch: - A signal or stale prompt may report only the tasks its own rows resolve to; fleet is refused there too. A heartbeat review is not scoped by task at all, so the extension no longer tracks live task records and refuses nothing by task id during a fleet review. - The outcome-index rebuild no longer skips retired tasks; the append-side skip alone keeps a torn-down task's index from being recreated. - bin/fm-guard.sh keeps the queued-wakes warning silent for the branch actor instead of printing a replacement note. * no-mistakes(document): Align supervision docs with scoped wake handling * fix(bin): avoid fleet snapshot argument limits (#3677) * Fix fleet snapshot large JSON transport * no-mistakes(review): Captain: file-back fleet snapshot transport safely * no-mistakes(review): Captain: file-back parent summary aggregation * no-mistakes(ci): Rebased the PR's three commits onto f4d7875824ecc5e274b4bb896f10c1e1f207b7e4 and resolved the fleet snapshot conflict while preserving the base's task-observation lifecycle. Fixed Greptile's valid finding by recursively removing the private mktemp transport directory, so future transport files cannot cause cleanup to fail. Verified with tests/fm-home-summary-refresh.test.sh, bin/fm-lint.sh, git diff --check, and ancestry checks. All passed; the fix remains as an uncommitted worktree change for the outer executor * fix(bin): attribute active runs with unfetched pipeline heads (#3681) * fix(bin): rec…
mituso89
pushed a commit
to mituso89/firstmate
that referenced
this pull request
Sep 26, 2026
* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work MAIN answered a supervision-branch outcome for completed captain-requested work (implementation done, PR ready for review and merge approval) with "Captain, shipshape.", reading section 9's no-action reply as covering it and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no captain-facing response is owed". Section 9 now limits the shipshape reply to true no-ops (idle re-read, empty heartbeat, consequence-free acknowledgement) and requires a short outcome response naming what finished and what word is needed whenever requested work finishes or a result needs the captain's word, even when a transcript entry already shows the substance. The Pi protocol's re-emit rule now says it bounds repetition only, and carries a worked example of the ready-for-review outcome whose correct processing turn a shipshape reply fails. No executable contract evaluates the content of MAIN's captain-facing reply, so the regression is the protocol example in the owner doc rather than a text-match test. * no-mistakes(document): Clarify captain-facing outcomes versus no-ops * docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line The document step condensed the Pi protocol's re-emit rule and dropped the worked example of a finished, ready-for-review outcome whose correct processing turn a "Captain, shipshape." reply fails. That example is the contract's regression: no executable contract evaluates the content of MAIN's captain-facing reply, so the owner doc's example is the test case. Restore it directly under the re-emit rule, prefixed as a regression example that is kept verbatim and never condensed or summarized away. * no-mistakes(review): Clarify captain outcome and decision-word requirements * no-mistakes(document): Clarify captain-facing completion outcomes * docs(pi): require the PR URL in the visible captain-facing outcome reply Captain review on the regression example: drop the sample reply string and say only that the ready-for-review outcome requires relaying a captain-facing outcome response, not just "Captain, shipshape.". Fold in the visible-PR-handoff failure seen this session: after the branch outcome reporting this fix green, MAIN's visible reply was only "Awaiting your merge call." with no PR URL, leaning on the dim anchor. Section 9's URL rule now also covers a review or merge ask and names the visible reply as where the URL goes, sourced from the ready status, pr= metadata, or the supervision branch's summary and never left to a transcript entry. The Pi protocol adds the same-way failure and places the captain-facing text in the final visible assistant reply after the fm_branch_processed call, because Calm hides assistant text emitted in the same step as a tool call as a working note. Investigation verdict, evidence in the PR comment: no recent PR caused the handoff failure; Pi has hidden same-step pre-tool assistant text since kunchenguid#2339 (2026-08-13), kunchenguid#4655 changed only the Claude Code mod, and kunchenguid#4658 touched only remote report transfer. * no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs * no-mistakes(document): Clarify captain-facing supervision outcomes * docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule The consolidated section 9 URL rule narrowed its trigger to a review or merge ask, dropping the "whenever a PR is mentioned" catch-all from kunchenguid#3648 that keeps every PR URL copied from a durable record and never assembled from memory. Restore that trigger as a union with the review or merge ask so the one consolidated rule covers both.
rub-a-dub-dub
added a commit
to rub-a-dub-dub/firstmate
that referenced
this pull request
Sep 27, 2026
…gences (#29) * fix(bin): treat Claude Code's default external-imports flags as never asked, not declined (#4387) * fix(bin): read Claude Code's default external-imports flags as never asked, not declined (#4378) fm-claude-trust.sh refused the whole trust registration whenever the project-root entry carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code writes that value only on an explicit "No, disable". Claude Code's default project entry carries Approved and WarningShown both false before the dialog is ever shown, so every such project refused every spawn. Only Approved === false with WarningShown === true — the pair the dialog writes on a decline — now counts as a decline. false/false behaves like an absent flag: trust is registered and no import consent is manufactured. New case test_project_root_entry_default_import_flags_are_not_a_decline fails on b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31, bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * no-mistakes(review): Correct harness doc's external-imports decline predicate --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(bin): keep operator-address labels out of no-mistakes intent (#4445) * fix(brief): keep operator address out of composed intent Teach raw-word authoring for intent sections and mid-task relays, with a neutral [captain] provenance marker for legacy mixed tasks. Keep headings and contract prose outside the serialized intent body. The legacy selector already excluded the old speaker labels from its output; preserve that read compatibility. The reproduced leak comes from adding labels inside a modern intent body, not from the legacy selector. Do not scrub actual request content. Add exact serialized-input and generated-contract regressions, retaining refusal of unmarked legacy tasks and coverage of scout promotion. Fixes https://github.com/kunchenguid/firstmate/issues/3882 * no-mistakes(review): Refuse operator-address lines in Captain's intent body * no-mistakes(document): Document operator-address refusal in intent contract comments * fix: classify OpenCode ellipsis hint as idle (#4451) * fix(composer): recognize Grok 1.0.5's oversized titled bottom border as a proven empty composer (#4455) * fix(composer): accept Grok title overhang * no-mistakes(review): summary: named Grok overhang constant, doc caveat, restored tmux typed-title coverage * fix(bin): translate Stop hook timeout signals into durable auto-arm failure (#4474) * fix(bin): recover Claude auto-arm after timeout * no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list * fix(spawn): establish Claude task channel authority (#4464) * fix(spawn): establish Claude task channel authority * no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference * fix(bin): refuse fm-control.sh exit when the composer holds unproven or pending text (#4458) * fix: guard relaunch exit against pending input * no-mistakes(review): Verifying test run in progress * no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard * no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now * fix(spawn): establish crewmate identity first (#4481) * fix(bin): reconcile redundant secondmate divergence during updates (#4460) * fix: reconcile diverged secondmate updates * no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md * no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md * feat: enable gpt-5.6-luna max reasoning for crew dispatch (#4497) * fix(dispatch): support Codex Luna max effort * no-mistakes(review): use portable CODEX_HOME path in codex effort reference * feat(calm): render smooth Unicode swell with asymmetric two-color sail (#4498) * feat(calm): render smooth Unicode swell * feat(calm): make sails asymmetric * feat(calm): use quarter sail glyph * no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer * no-mistakes(document): docs: sync calm wave phase doc comment * no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。 * fix(bin): supersede stale scout delivery text in brief.md on promotion (#4491) * fix: supersede scout delivery brief on promotion * fix: preserve ship safety contract after promotion * no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch * fix(bin): make captain holds work on hosts with an older JSON::PP, and stop cleanup dropping accents from a held body (#4471) * fix(bin): let captain holds work on hosts with an older JSON::PP Holding a task for the captain, and the cleanup that keeps a captain-held row open, both fail outright on any host whose JSON::PP defaults allow_nonref off - 2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`, but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older library rejects that whole value with "must be object or array". The consequence is fleet-wide on such a host, not one broken command: a worker there cannot formally record a decision for the captain at all. It can only mention the decision in passing in a status line, where it can be missed - which is how a real decision goes unrecorded. The hold reports that the task lost its hold-set stamp; the cleanup cannot return the row to Queued. Both call sites now ask for allow_nonref explicitly rather than inheriting whatever the installed library defaults to. The second one is worth naming: its `/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string case that fails, so the guard selects for the failing input rather than protecting against it. The regression case forces the older default back off for every perl the commands spawn, then drives both paths - holding a task that carries a body, and tearing down a captain-held row whose deliverable must still be appended. It also probes that the simulation genuinely rejects a bare scalar, so the case cannot pass vacuously on a lenient host. Each half was verified failing on its own unfixed call site with that site's real error message. Suites: fm-captain-hold-lifecycle 51 cases, fm-backlog-atomicity 99 cases, 0 failures. Verification limit: the mechanism is reproduced and tested, but neither fix is verified against a real JSON::PP 2.27202 host, because none is in the loop. This laptop runs 4.06, where the bug does not manifest. `bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a brace-delimited object before decoding, so allow_nonref never applies. * fix(bin): stop cleanup silently dropping accented characters from a held body Cleanup rewrites a captain-held row's body to append the finished work's deliverable, and the decoder it reads that body with printed decoded characters to a stream with no `:raw` layer. A character at or below U+00FF then came out as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the accent. `fm_backlog_retain` writes that body straight back through `--body-file`, and nothing reported an error - the character was simply gone from a row still waiting on the captain. The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus `utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already used. Review of the parent commit found this on one of the lines that commit already changed. It predates that change. The test asserts bytes rather than decoded strings, because comparing strings cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any character above U+00FF makes perl print the whole string as UTF-8, so one body carrying both an accent and an em dash passes even unfixed and proves nothing. Verified failing before the fix on the accented row, passing after. Suites: fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures. * no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc * no-mistakes(review): drop whole-file UTF-8 check from retained-body test * no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc * no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc * fix(bin): read codex 0.154's idle braille starfield rows as composer furniture (#4532) * fix(composer): read codex 0.154's idle starfield and status footer as furniture codex-cli 0.154.0 animates a braille "starfield" around its idle composer: on the row above the bold `›` prompt row, on the `›` row behind the SGR-2 dim `Ask Codex to do anything` placeholder, and on the row below it, then draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`). The cells are truecolor greys on both sides of the ghost luminance ceiling, so the brighter ones survive ghost stripping, and the rows below the glyph carry no structural edge. The shared classifier selected the bare `›` shape, extended its wrap region over the two rows beneath the glyph, read the survivors and the footer as wrapped typed input, and answered `pending`; the steering doorbell defers on exactly that verdict, so no doorbell ever reached an idle codex 0.154 pane. bin/fm-composer-lib.sh now recognises that furniture by shape, declared once next to the idle placeholders and reached from the two wrap-region boundary points: - a row whose non-whitespace content is entirely braille cells (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it never counts as wrapped typed content and bounds a bare composer's wrap region; braille behind the glyph row's content is stripped before the emptiness decision when nothing else follows the glyph; a row mixing braille with other text stays typed content; - the codex status footer bounds the wrap region exactly as omp's status row does, anchored on the effort token, a spaced middle dot, and a `~` or `/` path cell, so a typed `fix · tests` stays composer input; - `^Ask Codex to do anything$` joins the verified idle-placeholder set; the ghost strip remains what proves that row empty, and the bare-row rule that bright placeholder text is real input is unchanged. Unchanged: the strict blank-row rule, the styled=0 degradation (a plain cmux/orca capture of this screen still reads `unknown`, never `pending`), FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape. tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte with the divergence (letters in place of the starfield read `pending`) and the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh is the default-on live guard (token-free, skips explicitly without codex or tmux) that launches the installed codex idle and asserts `empty` through both the tmux and the cursorless styled reads, naming codex --version on failure. docs/verification/runtime-backends.md records the dated Herdr evidence: `pending` before, `empty` after, on the captured screen. * no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry --------- Co-authored-by: Todd Billings <todd@usdvcapital.com> * fix(bin): refuse empty text steers in fm-send (#4259) * fix(bin): refuse empty text steers in fm-send A marked secondmate request sent with an empty message delivered only marker and correlation bytes and minted a pending-reply expectation the parent could never see resolved, stalling the fleet with no loud error (#4255). Fail closed on an empty or whitespace-only message on the text path, mirroring the existing --resolve-key refusal. * chore: retain ambient Pi-lens autoformat as its own commit Formatting-only edits produced by ambient Pi-lens autoformat during the msg-loss investigation, kept separate from the behavioural change in c23acba6 so the fix stays reviewable on its own. AGENTS.md is deliberately excluded: its only autoformat edit stripped the trailing space from the documented FM_OPERATIONAL_PREFIX value, which bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11 records as permanent compatibility. Documenting that constant without its trailing space makes the doc wrong about the contract, so that one line was restored rather than retained. * fix(calm): paint the working ship one yellow over all-blue water (#4554) On rose-pine-moon the two-color water (cyan crests over blue troughs) read as a pink stripe over aqua, the yellow left sail and mast clashed with the red right sail, and the hull carried a blue interior run. Every water cell is now blue so the swell reads through glyph height alone, and both sail halves, the mast, and the whole hull are one yellow run. Geometry, cadence, animation, direction flip, resize clamping, and the narrow fallback are unchanged. Update the unit and real-TUI color assertions to the new palette and the Calm docs that described the old one. * fix(bin): stop aging a second mate's active turn from its launch (#4270) * fix(watch): stop aging a second mate's active turn from its launch The parent watcher's second-mate wake-loop stall check exempts a mate that is demonstrably inside an active turn, but secondmate_in_active_turn asked busy_turn_over_age first and returned "not in a turn" whenever that said the bound was crossed. busy_turn_over_age ages from state/<task>.turn-ended, falling back to state/<task>.meta. A second mate's turns end in its own home, so the parent never gets a turn-ended mark for it and the fallback ages the mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago was therefore permanently "over age", the busy pane was never consulted, and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false wake-loop stall. The gate now bounds the busy exemption by <idle> - how long the queue's drain position has not moved - which is evidence this home actually holds. A busy mate stays exempt while the queue has been frozen for less than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so the bound that stops a busy pane from proving liveness forever is kept rather than removed. busy_turn_over_age is untouched; its remaining callers are the ordinary crew busy-pane bound. The regression pins the case that actually broke: a mate whose launch record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn must not escalate, while the same mate with its queue frozen past the bound still publishes exactly one notification. The existing coverage only exercised a freshly launched mate, which passes either way. Reaching that alert now costs a pane capture inside the gate, so the three checkpoints in this suite that assert an alert move from a 1s to a 4s bound - the value the neighbouring active-turn cases already use. The bound is a ceiling, not a wait: the checkpoint returns on the first actionable wake. On a loaded machine a 1s bound missed the alert repeatedly; at 4s it did not miss in 20 runs under the same load. * no-mistakes(review): scope the second-mate active-turn regression test's coverage claim * no-mistakes(document): fix stale second-mate active-turn comments in fm-watch * feat(bin): add read-only PR blocker and reviewer discovery commands (#4278) * feat(bin): add read-only PR blocker and reviewer-discovery commands Two focused, opt-in commands that read GitHub and never write to it. fm-pr-state.sh reports what still blocks one pull request from the author's side: a closed or merged state, draft state, unknown or conflicting mergeability, absent or failing required checks, and a blocking CHANGES_REQUESTED decision explained by each reviewer's latest verdict, marked STALE when it was left at a superseded head. A pull request that only awaits an approval is not reported as blocked, and advisory checks are omitted. Every reading is taken against one exact head; a push that lands mid-read invalidates the whole result rather than mixing two snapshots. fm-pr-reviewers.sh suggests reviewers from the most recent commits to the pull request's exact changed paths, counting each commit once, resolving handles through GitHub's own commit author.login mapping, and excluding the author and Bot accounts. Both stay read-only: no review request, no approval, no merge. Unresolved review-thread state is left unreported because the REST API does not expose it and unattended commands may not use GraphQL. Closes #3731 * no-mistakes(review): accept only PR URLs and stop at terminal state * no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate * no-mistakes(review): stop attributing readings to unverified heads * no-mistakes(review): narrow readiness contract to checks that have reported * no-mistakes(review): read the pull request once, drop the head guard * no-mistakes(document): scope pr-forge isolation proof to its measured members * no-mistakes(document): record uncovered pr-forge members and their pending proof * docs(isolation-proof): re-prove pr-forge at its full membership tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the pr-forge family in this branch, and script_allows_concurrency grants four workers by family membership alone, so both ran concurrently on a proof measured before they existed. Re-proved the family at all eight members: two consecutive runs, 0 failures, each begun with the one-minute load average below 6.0 so the result measures isolation rather than contention. A third run taken between them is disclosed rather than recorded, because it started while the previous run's workers were still decaying. The new durations are not comparable with the six-member measurement above them, so they are not presented as evidence about the two new members, and that record's 1.72x four-worker figure is left as a statement about its own run rather than restated as current. * no-mistakes(review): disclose gh error-text coupling at its matching site and tests * fix(bin): teach validation-round pauses in generated briefs (#2752) * fix(bin): teach validation-round pauses in briefs * no-mistakes(document): Point classifier comments to authoritative pause examples * docs(readme): add star history chart (#4558) * fix(bin): refuse teardown when a task's endpoint close fails (#4510) * fix(teardown): refuse a cleanup whose endpoint close failed bin/fm-teardown.sh discarded both the exit status and the stderr of every fm_backend_kill call, so a close that genuinely failed was indistinguishable from one that succeeded. Teardown continued past it, deleted the task's durable records, returned its worktree, and reported the cleanup as completed. The deleted metadata is the only record of which endpoint belongs to the task, so such a close did not merely leave a stray session behind, it stranded one: nothing was left on disk naming it. The adapters could not carry that signal either. Driven against the real code, every backend arm returned 0 for a genuine failure exactly as it did for an already-exited endpoint, so there was nothing for the four call sites to propagate even once they stopped swallowing it. The tmux arm now resolves a close that did not succeed against the window's exact recorded identity, since kill-window fails the same way for a window that is gone and one that is still there. The Orca arm reports a close its missing CLI never attempted. Both stay silent for an endpoint that is already legitimately gone, and the remaining arms are unchanged: their close-command timing cannot be established without the real Zellij, Orca, and cmux binaries, and a gate that refused ordinary cleanup of an already-exited session would be worse than the defect. docs/verification/runtime-backends.md records what each backend can prove. A reported close failure now reaches teardown's existing retain-and-stop refusal before the records naming the endpoint are removed, matching where the Herdr confirmed-gone gates already sit for the same hazard, and the retained records let a rerun finish once the close works. * no-mistakes(review): refuse unreadable tmux close re-read; honor --force override * no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close * no-mistakes(document): document endpoint-close refusal in its backend and retirement owners * no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree * feat(calm): add flag-gated Claude Code Calm mode (#4565) * feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that brings Calm to Claude Code: the sailboat replaces the stock working row through a Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration, and canonically classified operational user rows draw at zero height. /calm is registered by the hooks module itself and toggles the same per-home config/calm preference the Pi extension uses, so one choice applies on either harness; rows redraw retroactively on toggle and stay hidden across claude --continue. The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS flag is on. Nothing sets that flag in any settings file, and the plugin carries no command file, skill, agent, or classic hook, so it is a complete no-op while the flag is off. The trusted project auto-loads it through an .agents/skills symlink, the only path Claude Code scans for project plugins. Extract the working-ship geometry, bounce track, cadences, and freeze/resume state into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module imports from outside the plugin folder) and have the Pi widget paint that core's frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify operational rows through a port of bin/fm-operational-input.sh's classify command guarded by a corpus parity test against the shell owner. Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering, Raster packing, policy, classifier parity), the mod's own claude plugin test suites behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op, the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272. Docs: record the version-scoped Claude Code evidence and the three bounded gaps in docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md, and make the shared preference, layout, and contributor notes harness-neutral. * no-mistakes(review): Preserve colliding final replies and strengthen parser parity * no-mistakes(review): Preserve final replies and strengthen canonical parity checks * no-mistakes(review): Require exact function-hooks opt-in before Calm activation * no-mistakes(review): Clarify Calm module loading and activation boundaries * no-mistakes(review): Reset Calm presentation state across session starts * no-mistakes(document): Refresh Calm session lifecycle documentation * feat(calm): paint the Claude Code working ship in Claude's own theme colors The captain picked the "Claude native" palette for the Claude Code mod's Raster: every water cell takes the spinner blue of the active theme family (#93a5ff dark, #5769f7 light) and the whole boat takes the Claude orange of the stock spinner (#d77757), one water color and one boat color. The family follows the `theme` setting's prefix, read at load through $.config.list and re-read on a config.set of that row, with `auto` and custom themes falling back to the dark set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte. Rename the shared sprite's color classes from hue names to `water` and `boat`, since each harness now maps them to its own colors; geometry, motion, cadence, and the activation gate are untouched. Tests cover both palettes' packing and the family rule under Node, and the plugin kit drives every theme value, a theme change mid-session, the Calm-off pass-through, and inertness of the menu read while the flag is off. The docs describe the Claude Code colors and record the guard passing on 2.1.273. * no-mistakes(review): Use light palette for unresolved Claude themes * no-mistakes(document): Refresh Claude Calm verification evidence * fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586) * fix(watch): honour a declared wait before wedge-escalating a quiet pane wedge_timer_check escalated on elapsed idle time alone. Nothing asked whether the worker had already said why its pane was quiet, so a lane that declared a bounded external wait climbed the escalation ladder for as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every repeat carried demand-deep-inspection - which by its own wording forbids re-absorbing on the run-step or pane state, so the supervisor could not use the evidence that was there either. The generated brief promises that declaring `paused:` buys the long recheck cadence instead of a wedge, but the timer was still reachable while that declaration stood: a crew that declares a wait and then has an active run or busy pane attributed to it is handed to the timer as provably-working. The declaration is what the worker said about its own silence, so it now outranks a liveness verdict that only says something is running. The consult runs in the at-threshold branch that was about to escalate, beside the worktree walk already there, and costs one status-line read. Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS recheck the declared-wait absorber already uses, so the wait is still rechecked and cannot rot invisibly. Which verb declared it decides the wording, because the two block on different people: a `paused:` wait is owed by an external dependency and asks the reader to confirm it still holds, while a `captain-held:` transfer is owed by the captain reading the recheck and asks them to answer or release the hold. A hold is not rechecked at all while the away-posture record exists, as on every other captain-held path, and that absorb arms no throttle so the recheck is owed in full on return. A declared clearing time that has already passed stops counting, and a lane that never declared one keeps the identical escalation schedule, reason, count and demand-deep-inspection wording, so detection and its worst-case time are unchanged. The deferral restarts the idle timer rather than cancelling it, so a lane that stops waiting escalates again within one threshold. A lane quiet because its own validation run is parked at a gate awaiting a human decision is deliberately out of scope: reading that state needs a signal carrying who the wait is on and what clears it, rather than one inferred from a parked verdict that also covers gates awaiting the crewmate itself. Tests pin both directions for each case and were each confirmed to fail with the consult removed. * no-mistakes(document): docs: honour declared waits in stale-escalation docs * fix(bin): report verified PR state for passed runs (#4624) * fix(bin): derive passed PR state from PR record A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists. For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim. Fixes #4607 * no-mistakes(review): Add bounded GitLab merge-request state reads * no-mistakes(review): Preserve network-free inactive crew-state scans * no-mistakes(document): Document PR record readers in shared library * fix: restore published contribution follow-up (Fixes #4469) (#4627) * fix: restore published contribution follow-up (Fixes #4469) * fix(review): Fix contribution freshness and merge actor routing * fix(review): Restore issue triage and scope contribution follow-up * fix(test): test: assert one wake per contribution signal * fix(document): Document contribution follow-up * fix: restore truthful terminal delivery evidence * fix(review): Disclose unsupported contributions and deduplicate watcher wakes * fix(review): Preserve unmeasured unsupported contributions across Bearings * fix(review): Deduplicate shared contribution wakes and isolate diagnostics * fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower * fix(bin): make remote report transfers explicit and fail-open (#4658) * fix(bin): make a remote-reply document gap self-clearing and re-attemptable A remote mate's undelivered document raised a keyed `blocked` decision that nothing could ever resolve, and any `data/*.md` substring in any mirrored line was an unconditional fetch instruction. A mate announcing a report it had not written yet therefore manufactured a permanent, factually false blocker, and its own explanation of the false alarm manufactured more. The reader has no permanence vocabulary: a report still being written refuses exactly like a path that will never exist. So an undelivered document is now a durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`, re-attempted on the next delta and on the channel's own quiet poll, and retired with a matching `resolved` line naming the local copy once it arrives. The cursor still advances and no delta stalls on one bad pointer. Only a structured `report=data/....md` pointer now offers a document, so a path merely mentioned in prose - including one under another home's mirror tree, which is provably not that mate's to serve - is never fetched. Offers are deduplicated across the whole delta, the escalation names each missing document once and carries the reader's own reason instead of discarding it, and a strictly increasing notice ordinal keeps a later escalation from being swallowed as duplicate bytes. A mirrored line still lands once whichever pointer form it was first written under. * no-mistakes(review): Require structured pointer token boundaries * no-mistakes(review): Unify boundary-safe pointer extraction and rewriting * fix(bin): identify a mirrored line independently of its delivery state Two defects in the boundary-safe pointer work. The at-most-once check compared only the all-remote and all-local renderings of a line, so it could not recognize a mixed one. A line offering two documents where only the first was deliverable mirrored as local-plus-remote; once the second arrived, a cursor-loss whole-log recapture rendered the same line all-local, matched neither alternate, and mirrored a second time. A line's identity is now the canonical form every boundary-valid pointer would take once delivered, derived by the same parser that does extraction and rewriting, so it no longer depends on which documents happened to be deliverable at the time. The pointer map was passed to awk through the process environment. A delta may carry up to the configured 1 MiB bound, and an expanded map of delivered pointers can exceed the platform's exec argument limit, so awk would fail to start; because no caller checked, the empty result would have been appended as blank lines while the cursor advanced past dropped status content. The map now travels in a file, and every call site checks the exit status and stops the ingest rather than committing a delta it could not render. Both passes now run once per stream instead of twice per line. * no-mistakes(review): Abort ingest when document pointer extraction fails * no-mistakes(review): Exclude structured cross-home pointers from document transfer * fix(bin): fail open on an undeliverable remote document instead of tracking it Narrow the remote-reply document fix to the scope the diagnosis actually requires, as decided after measuring a simpler alternative. A document the reader cannot deliver now fails open. The mate's line is mirrored with its own pointer, the cursor advances, and one unkeyed note carries the reader's reason. A note never enters the open-decision fold, so it cannot stand open the way the original keyed block did - which removes the never-clearing false blocker by construction rather than by resolving it. That makes the durable self-clearing obligation unnecessary, so it goes: the per-mate pending-documents record, its notice ordinal and resolved announcements, and the poll-side retry. Canonical line identity goes too, and with it a way to silently drop a genuine status line; mirroring is back to at-most-once on exact bytes. The cross-home exclusion goes as well: under fail-open a cross-home report= either fails harmlessly or is a nested remote report this mate genuinely holds, which is now relayed again. Kept: fetching only on a structured report= pointer, the boundary-correct parser, the file-based rewrite map, and checked extraction and rewrite exit status. The parser now scans behind a sentinel byte so a rejected candidate can no longer give the text right after it a false leading boundary. The reported incident is covered end to end: a report path announced in prose before it exists raises no decision, and the report still arrives through the ledger publisher's structured offer once written. * no-mistakes(review): Preserve source-line identity across remote reply replays * no-mistakes(document): Document remote reply transfer and replay semantics * no-mistakes(lint): Fix staging truncation lint checks * fix(calm): preserve substantive mid-turn responses (#4655) * Preserve substantive Calm mid-turn text * no-mistakes(review): Distinguish newline-preserved replies from short narration * no-mistakes(document): Document Calm mid-turn preservation boundaries * no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check` * fix(bin): preserve PR merge polls across volume remounts (#4656) * fix(bin): re-record PR poll identity after a volume device renumber (Fixes #4260) A volume remount can renumber the state filesystem's st_dev while every inode and byte stays the same; APFS does this across a reboot. A poll registration records its sidecar and check as device:inode, so every poll armed before the remount failed strict validation and the watcher refused all of them as unauthenticated state checks until each was re-armed by hand. There are two device comparisons. fm_pr_private_file_valid compares a live file's device with the state directory's device read in the same invocation: it refuses a file that is not on the state directory's own filesystem and already survives a renumber, so it is unchanged. The registration's recorded identity versus the live identity (from #556, reused by the #932 retirement receipt) binds the registration to the exact files published in its own transaction; its device part is what breaks. When strict capture fails, the watcher now proves the device is the only difference: every other artifact check passes (template bytes, both hashes, private mode, single link, live device, metadata), both recorded identities name one device, and each recorded inode equals its live inode. Only then, under the task's control lock, does it rewrite the two identity lines, repeating the whole proof and comparing the registration's file identity and bytes just before the rename, and then capture strictly again. A swapped, altered, re-moded, relinked, split-device, or foreign-device artifact still fails a proof and is still refused, and a pending retirement receipt blocks the rewrite. Reproduction: on macOS a poll armed on an APFS disk image that was detached and re-attached behind another image moved st_dev 16777239 -> 16777243 with inodes, bytes, mode, and link count unchanged; the real watcher refused it on main and reports its merge with this change. The portable regression test rewrites a real registration's recorded device and drives the watcher. Not changed here: the status presentation cursor keys rows by its own device:inode identity in bin/fm-classify-lib.sh, a different helper that needs its own fix; a retirement receipt left by a reboot between its publication and removal still names the old device and stays refused; custom check trust binds only a content hash and is unaffected. * fix(review): Serialize PR poll publication writers * fix(review): Bound PR poll publication lock scope * fix(bin): keep contribution records when the poll budget runs out (follow-up to #4627) (#4661) A budget that expires partway through an observation no longer records an error or prints the unavailable wake; the URL keeps its prior record and is observed first next poll. forge() flags budget exhaustion at the point it refuses, or when a read is killed at the budget's own deadline, so a genuine forge failure still records the error and wakes. Each distinct URL is now observed once per poll and applied to every owning task. * fix(bin): clear parent pending-replies on local secondmate retirement (#4680) * fix(bin): clear parent pending-replies on local secondmate retirement Local secondmate teardown left resolved parent pending-reply records behind after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced retirement while any reply for that id is still unresolved, and delete every matching record plus its delivery confirmation after a successful local or remote retirement, matching the remote cleanup path. * no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup * no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung * no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern * no-mistakes(review): Pending-replies Basename und corr_id abgleichen * no-mistakes(document): Clarify forced retirement pending-reply cleanup --------- Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc> * fix(bin): accept Orca's composite worktree id when tearing down a task (#4677) * fix(bin): accept Orca's composite worktree id at teardown Teardown refused every Orca-backed task because the endpoint validator checked orca_worktree_id with the simple-atom rule meant for tmux-style window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca returns that id as `<orca id>::<absolute worktree path>`, so the colon and slashes in every real value made validation fail and finished Orca tasks could never be cleaned up. Validate the field as the composite it is: both halves of the first `::` split present, the path half absolute, and no embedded newline, carriage return, or tab. The terminal field keeps the atom check, which is correct for it, and no other backend's validation changes. The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca never returns, which is why the suite passed a check the real value fails. They now carry the composite form, so the tests exercise the real value. * no-mistakes(document): name Orca's repo id in the composite worktree id * no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points * feat(bin): add opt-in typed dispatch resolution (#4692) * feat(bin): add opt-in typed dispatch resolution through typesafe.ai Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or scout profile from a written brief with typesafe.ai's System One model: one Choice question over the rules' `when` texts, then the confidence floor, the rule's `approval` and `floor`, each profile's `provider` and `floor`, one quota-axi snapshot, and the spendPriority argmax all in code. It is off unless TYPESAFE_API_KEY is in the environment or the home's gitignored .env; off means one stderr line, exit 0, and no network call, so firstmate dispatches exactly as before. The key reaches curl on a file descriptor, never argv. Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates the four new optional dispatch fields. Document the schema, the operator contract, the AGENTS.md intake step, and the live and benchmark evidence. * no-mistakes(review): Harden typed dispatch resolution and quota bounds * no-mistakes(review): Validate dispatch floors and ranking evidence * no-mistakes(review): Tighten dispatch response and floor evidence * no-mistakes(review): Neutralize none matching and resolve defaults locally * no-mistakes(review): Preserve providerless profiles outside typed resolution * no-mistakes(review): Validate response usage and reject duplicate profiles * no-mistakes(review): Escalate unverifiable floors and validate probabilities * no-mistakes(review): Validate probability mass and unknown profile floors * no-mistakes(review): Simplify resolver interface and preserve fallback routing * no-mistakes(review): Fix constants and rank partial quota evidence * no-mistakes(review): Add authoritative provider mapping and enforce explicit providers * no-mistakes(review): Declare provider for documented Pi profile * no-mistakes(review): Validate provider identifiers and support Gemini dispatch * no-mistakes(review): Strictly anchor provider identifiers * no-mistakes(review): Validate selectors and preserve fallback candidate evidence * no-mistakes(review): Gate typed validation and harden resolver evidence * no-mistakes(review): Preserve opt-in routing and harden candidate evidence * no-mistakes(review): Prioritize known exhaustion over quota uncertainty * no-mistakes(review): Isolate API secrets and preserve no-key diagnostics * no-mistakes(review): Fallback safely when dispatch rules are absent * no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets * no-mistakes(document): Document typed dispatch safety and fallback behavior * fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753) * test: reproduce buried status declarations in shared readers * fix: share status event reads and preserve open blockers * fix: retain terminal scout and ship status declarations * no-mistakes(review): Fix status chronology, legacy completions, and reader performance * no-mistakes(review): Share terminal decision reconciliation across fleet snapshots * no-mistakes(review): Unify terminal supersession across cached folds and consumers * no-mistakes(review): Filter per-key status history while preserving terminal chronology * no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells * no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses * no-mistakes(document): Document latest-event status read and kind-scoped fold cursor * no-mistakes(lint): Quote literal done in test for-lists for SC1010 * ci: expect 19 snapshot/fleet-view tests This branch adds a fleet-snapshot regression, so the stock macOS Bash lane's hardcoded guard of 18 'ok - ' lines fails on the new count. Bump the guard and its message to 19. * no-mistakes(review): Restore multiline child outcome reporting * no-mistakes(review): Select ledger terminal events through bounded shared reader * no-mistakes(review): Report newest open decision instead of preferring blocked * no-mistakes(review): Require colon before ship/scout terminal supersession in fold * no-mistakes(review): Gate socket-down override on latest event; drop lock matrix * no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions * no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold * no-mistakes(test): Update fleet-view expectations to newest-open-decision rule * no-mistakes(document): Align status-read docs with fold-resolved crew state * no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers * no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree * test: fold terminal-cleanup snapshot coverage into the completed-scout case Keep the ship/scout/secondmate supersession assertions without adding a nineteenth top-level fleet-view test, so CI can stay at the upstream suite count. * no-mistakes(document): Clarify socket-down override expiry in architecture doc * ci: retrigger flaky contribution check * fix(bin): launch codex crewmates with codex's hook layer disabled (#4689) * fix(spawn): launch codex crewmates with codex's hook layer disabled A freshly launched Codex worker never reached its instructions. Codex stopped it on an interactive "Hooks need review" modal whose selection sits on "Review hooks", which is neither trusting nor declining. Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow navigation, so the selection cannot be moved, and pre-accepting the prompt by writing Codex's own trust store would record an operator consent that was never given. The hooks are the machine's own ~/.codex/hooks.json plus any project's .codex/hooks.json. A crewmate needs neither: its turn-end signal is the -c notify= program on the same launch, and Firstmate's project hooks are primary-session infrastructure that stands down in a child worktree. Crewmate and scout launches now pass --disable hooks. That is the opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted hooks; disabling the feature runs none of them and leaves the operator's ~/.codex untouched. An unknown feature name is a hard Codex error, so a release that drops the flag fails the launch loudly instead of silently restoring the modal. A secondmate is a primary in its own home and keeps the project hooks its turn-end guard and session-start digest ride on. Verified on codex-cli 0.151.0: the modal is gone and the turn-end notification still lands. This unblocks the second review that every finished pull request is supposed to get. Fixes kunchenguid/firstmate#4673 * no-mistakes(review): Fix contradictory hook count in Codex verification record * fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710) * fix(bin): settle terminal contributions and wake once per read-failure episode A contribution whose last good observation is merged or closed is final: poll no longer re-reads it, projection keeps it fresh, and a stale error recorded beside it is cleared once. A genuine forge-read failure on an open contribution still records its error on every cycle but prints the unavailable wake only when it starts a failure episode; a successful read ends the episode. Open PRs linked from done tasks keep being observed. The false unavailable beside a complete observation was budget exhaustion mid-observation, already fixed by #4661. * fix(review): Settle terminal contribution owners * fix(review): Deduplicate shared contribution failure episodes * fix(test): Preserve settled terminal contribution records * fix: select authoritative no-mistakes runs (#4476) * fix(crew-state): select authoritative validation runs by identity Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row. Refs: https://github.com/kunchenguid/firstmate/issues/3215 * fix(review): Resolve same-branch run identities beyond capped history * fix(review): Fix run-selection compatibility, races, and worker-state fallbacks * fix(review): Limit run validation to the requested branch * fix(test): Anchor AXI fixtures and document remaining live evidence gaps * fix(document): Clarify run selection documentation and capture ownership * fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation * fix: distinguish captain outcomes from no-op updates (#4738) * fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work MAIN answered a supervision-branch outcome for completed captain-requested work (implementation done, PR ready for review and merge approval) with "Captain, shipshape.", reading section 9's no-action reply as covering it and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no captain-facing response is owed". Section 9 now limits the shipshape reply to true no-ops (idle re-read, empty heartbeat, consequence-free acknowledgement) and requires a short outcome response naming what finished and what word is needed whenever requested work finishes or a result needs the captain's word, even when a transcript entry already shows the substance. The Pi protocol's re-emit rule now says it bounds repetition only, and carries a worked example of the ready-for-review outcome whose correct processing turn a shipshape reply fails. No executable contract evaluates the content of MAIN's captain-facing reply, so the regression is the protocol example in the owner doc rather than a text-match test. * no-mistakes(document): Clarify captain-facing outcomes versus no-ops * docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line The document step condensed the Pi protocol's re-emit rule and dropped the worked example of a finished, ready-for-review outcome whose correct processing turn a "Captain, shipshape." reply fails. That example is the contract's regression: no executable contract evaluates the content of MAIN's captain-facing reply, so the owner doc's example is the test case. Restore it directly under the re-emit rule, prefixed as a regression example that is kept verbatim and never condensed or summarized away. * no-mistakes(review): Clarify captain outcome and decision-word requirements * no-mistakes(document): Clarify captain-facing completion outcomes * docs(pi): require the PR URL in the visible captain-facing outcome reply Captain review on the regression example: drop the sample reply string and say only that the ready-for-review outcome requires relaying a captain-facing outcome response, not just "Captain, shipshape.". Fold in the visible-PR-handoff failure seen this session: after the branch outcome reporting this fix green, MAIN's visible reply was only "Awaiting your merge call." with no PR URL, leaning on the dim anchor. Section 9's URL rule now also covers a review or merge ask and names the visible reply as where the URL goes, sourced from the ready status, pr= metadata, or the supervision branch's summary and never left to a transcript entry. The Pi protocol adds the same-way failure and places the captain-facing text in the final visible assistant reply after the fm_branch_processed call, because Calm hides assistant text emitted in the same step as a tool call as a working note. Investigation verdict, evidence in the PR comment: no recent PR caused the handoff failure; Pi has hidden same-step pre-tool assistant text since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and #4658 touched only remote report transfer. * no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs * no-mistakes(document): Clarify captain-facing supervision outcomes * docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule The consolidated section 9 URL rule narrowed its trigger to a review or merge ask, dropping the "whenever a PR is mentioned" catch-all from #3648 that keeps every PR URL copied from a durable record and never assembled from memory. Restore that trigger as a union with the review or merge ask so the one consolidated rule covers both. * fix(bin): let non-owner Claude Stops exit safely (#4777) * Fix foreign-owner turn-end supervision loop * no-mistakes(review): Scope foreign-owner safe exit to Claude guard * no-mistakes(document): Document Claude foreign-owner safe exit * fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778) Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty indexed array as an unbound variable and aborts the shell. In signal_turnend_panes_churned() the missing_keys loop was reachable with an empty array whenever every churned key already held a fresh .churn-since-* marker (a second churning turn-end inside an open deferral window), so each watcher cycle died about half a minute in and supervision restarted endlessly. The created_keys rollback loops had the same latent crash on their error paths. Audit of bin/ for the same pattern found one more confirmed-reachable case: remote_handoff's noncanonical-body scan iterates to_move, which is empty when a retried remote handoff finds every key already staged in the outbox. All other "${arr[@]}" sites are either count-guarded, guaranteed non-empty by construction, or unreachable while empty. Guard the three reachable expansions with the repo's existing "${arr[@]+...}" idiom. New regression test drives a real watcher through the all-marked churn path; the macos-stock-bash CI lane runs it under real /bin/bash 3.2 via FM_TEST_ONLY. * Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783) The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: require complete captain-facing final responses (#4779) * docs: require complete final responses across harnesses * no-mistakes(document): Document complete final replies for Grok Bot * docs: point Grok replies to the shared contract owner * no-mistakes(review): Clarify final recap without batching decision asks * fix: preserve substantive mid-turn text in Pi Calm (#4788) * fix(calm): preserve substantive Pi mid-turn text * no-mistakes(review): Preserve substantive Pi Calm text per block * no-mistakes(test): Cover shared Calm preservation boundaries behaviorally * no-mistakes(document): Consolidate Calm preservation documentation * fix: harden mail checks and rebalance full-coverage CI (#4800) * Improve CI reliability and rebalance full-coverage validation * no-mistakes(document): Clarify lint partition documentation * fix(bin): answer Kimi 2.0.0 folder-trust dialog during spawn (#4799) * Handle Kimi workspace trust dialog * no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers * no-mistakes(review): Gate Kimi ready on any trust marker and clean captures * no-mistakes(review): Read visible pane for Kimi trust and ready gates * no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate * no-mistakes(review): Harden Kimi viewport capture and trust dialog detection * no-mistakes(document): Document Kimi spawn refusal on cmux and Orca * fix(bin): report a dead-agent record once instead of escalating forever (#4775) * fix(bin): report a record whose agent is gone once instead of escalating forever The wedge escalation path never asked whether there was still an agent to be wedged. A wedge is something stuck that might recover, so re-alarming it earns its cost; an agent that is gone never moves again, its pane never churns, the idle timer never resets, and the escalate path clears its own timer and re-arms with nothing bounding the count. Observed on a live fleet: two finished lanes reached 226 and 203 consecutive escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400 notifications a day from two lanes with no agent running at all. On one, fm-control.sh exit answered already-stopped and fm-crew-state.sh read "failed - run failed". Closing the Herdr pane did not stop it either: with the pane genuinely gone and herdr pane read returning pane_not_found, the count kept climbing, because the poll is driven by the record's window= line rather than by the pane. The cost is not the repetition but that it drowns the alarms that matter. fm_backend_agent_state already separates a thinking agent from a gone one at process level. In the branch that was about to escalate, read it once and treat only its two recovery-grade verdicts - dead (endpoint present, no agent in it) and missing (endpoint authoritatively absent) - as proof, reporting that record once and not re-escalating it while it stays that way. Every other verdict, including alive, ambiguous, unreadable, unverified, and a read that failed outright, keeps the identical schedule, reason, and escalation count, so a genuinely wedged live agent is unaffected. The probe costs at most one backend read per window per threshold, the same budget the declared-wait consult and the worktree write probe already take. The report decides nothing about the record's fate: both lanes still held unlanded work and teardown refusing them was correct, so retiring, relaunching, or cleaning up stays with the supervisor. The once-only marker is owned entirely by that function and is dropped by the same read the moment the endpoint stops reading gone, so a replacement launched into the same window escalates normally and its own later death is reported again. Related, and not closed by this: #4412, #4482, #4316. Tests drive the real watcher against a record whose endpoint does not exist and pin both directions: dead and missing report once and never advance the count across later thresholds, while alive, ambiguous, and unreadable endpoints keep escalating with the identical reason and a climbing count. * fix(bin): bind the once-only dead report to the pane it reported Review of the parent commit found a reachable sequence where a later death in the same window lost its promised report. The marker was keyed on the verdict string alone and dropped only when a threshold probe read a non-gone verdict, but probes run only at thresholds: a replacement launched into the same window that dies without ever being probed alive - it crashes at startup, or works and then crashes - was absorbed by the previous death's marker. The pane's first sight yielded only the generic stale wake and every later threshold matched the stale marker, so the second death never got the detailed once-report that both the function's own comment and docs/architecture.md promise. Record the verdict together with the pane hash it was reported for, and absorb a repeat only while both still match. A replacement churns the pane, which resets the stale suppressor, wedge timer, and escalation count while no reset site touches this marker, so the pane half is what tells the second death apart from the first. The live-probe drop stays as it was. Clearing the marker at those reset sites instead would re-open unbounded re-alarming for a dead pane whose display ever ticks, which is the exact defect the pa…
gk-io-dev
added a commit
to gk-io-dev/firstmate
that referenced
this pull request
Sep 30, 2026
…bility fixes (#3) * feat(bin): add opt-in typed dispatch resolution (#4692) * feat(bin): add opt-in typed dispatch resolution through typesafe.ai Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or scout profile from a written brief with typesafe.ai's System One model: one Choice question over the rules' `when` texts, then the confidence floor, the rule's `approval` and `floor`, each profile's `provider` and `floor`, one quota-axi snapshot, and the spendPriority argmax all in code. It is off unless TYPESAFE_API_KEY is in the environment or the home's gitignored .env; off means one stderr line, exit 0, and no network call, so firstmate dispatches exactly as before. The key reaches curl on a file descriptor, never argv. Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates the four new optional dispatch fields. Document the schema, the operator contract, the AGENTS.md intake step, and the live and benchmark evidence. * no-mistakes(review): Harden typed dispatch resolution and quota bounds * no-mistakes(review): Validate dispatch floors and ranking evidence * no-mistakes(review): Tighten dispatch response and floor evidence * no-mistakes(review): Neutralize none matching and resolve defaults locally * no-mistakes(review): Preserve providerless profiles outside typed resolution * no-mistakes(review): Validate response usage and reject duplicate profiles * no-mistakes(review): Escalate unverifiable floors and validate probabilities * no-mistakes(review): Validate probability mass and unknown profile floors * no-mistakes(review): Simplify resolver interface and preserve fallback routing * no-mistakes(review): Fix constants and rank partial quota evidence * no-mistakes(review): Add authoritative provider mapping and enforce explicit providers * no-mistakes(review): Declare provider for documented Pi profile * no-mistakes(review): Validate provider identifiers and support Gemini dispatch * no-mistakes(review): Strictly anchor provider identifiers * no-mistakes(review): Validate selectors and preserve fallback candidate evidence * no-mistakes(review): Gate typed validation and harden resolver evidence * no-mistakes(review): Preserve opt-in routing and harden candidate evidence * no-mistakes(review): Prioritize known exhaustion over quota uncertainty * no-mistakes(review): Isolate API secrets and preserve no-key diagnostics * no-mistakes(review): Fallback safely when dispatch rules are absent * no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets * no-mistakes(document): Document typed dispatch safety and fallback behavior * fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753) * test: reproduce buried status declarations in shared readers * fix: share status event reads and preserve open blockers * fix: retain terminal scout and ship status declarations * no-mistakes(review): Fix status chronology, legacy completions, and reader performance * no-mistakes(review): Share terminal decision reconciliation across fleet snapshots * no-mistakes(review): Unify terminal supersession across cached folds and consumers * no-mistakes(review): Filter per-key status history while preserving terminal chronology * no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells * no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses * no-mistakes(document): Document latest-event status read and kind-scoped fold cursor * no-mistakes(lint): Quote literal done in test for-lists for SC1010 * ci: expect 19 snapshot/fleet-view tests This branch adds a fleet-snapshot regression, so the stock macOS Bash lane's hardcoded guard of 18 'ok - ' lines fails on the new count. Bump the guard and its message to 19. * no-mistakes(review): Restore multiline child outcome reporting * no-mistakes(review): Select ledger terminal events through bounded shared reader * no-mistakes(review): Report newest open decision instead of preferring blocked * no-mistakes(review): Require colon before ship/scout terminal supersession in fold * no-mistakes(review): Gate socket-down override on latest event; drop lock matrix * no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions * no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold * no-mistakes(test): Update fleet-view expectations to newest-open-decision rule * no-mistakes(document): Align status-read docs with fold-resolved crew state * no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers * no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree * test: fold terminal-cleanup snapshot coverage into the completed-scout case Keep the ship/scout/secondmate supersession assertions without adding a nineteenth top-level fleet-view test, so CI can stay at the upstream suite count. * no-mistakes(document): Clarify socket-down override expiry in architecture doc * ci: retrigger flaky contribution check * fix(bin): launch codex crewmates with codex's hook layer disabled (#4689) * fix(spawn): launch codex crewmates with codex's hook layer disabled A freshly launched Codex worker never reached its instructions. Codex stopped it on an interactive "Hooks need review" modal whose selection sits on "Review hooks", which is neither trusting nor declining. Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow navigation, so the selection cannot be moved, and pre-accepting the prompt by writing Codex's own trust store would record an operator consent that was never given. The hooks are the machine's own ~/.codex/hooks.json plus any project's .codex/hooks.json. A crewmate needs neither: its turn-end signal is the -c notify= program on the same launch, and Firstmate's project hooks are primary-session infrastructure that stands down in a child worktree. Crewmate and scout launches now pass --disable hooks. That is the opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted hooks; disabling the feature runs none of them and leaves the operator's ~/.codex untouched. An unknown feature name is a hard Codex error, so a release that drops the flag fails the launch loudly instead of silently restoring the modal. A secondmate is a primary in its own home and keeps the project hooks its turn-end guard and session-start digest ride on. Verified on codex-cli 0.151.0: the modal is gone and the turn-end notification still lands. This unblocks the second review that every finished pull request is supposed to get. Fixes kunchenguid/firstmate#4673 * no-mistakes(review): Fix contradictory hook count in Codex verification record * fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710) * fix(bin): settle terminal contributions and wake once per read-failure episode A contribution whose last good observation is merged or closed is final: poll no longer re-reads it, projection keeps it fresh, and a stale error recorded beside it is cleared once. A genuine forge-read failure on an open contribution still records its error on every cycle but prints the unavailable wake only when it starts a failure episode; a successful read ends the episode. Open PRs linked from done tasks keep being observed. The false unavailable beside a complete observation was budget exhaustion mid-observation, already fixed by #4661. * fix(review): Settle terminal contribution owners * fix(review): Deduplicate shared contribution failure episodes * fix(test): Preserve settled terminal contribution records * fix: select authoritative no-mistakes runs (#4476) * fix(crew-state): select authoritative validation runs by identity Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row. Refs: https://github.com/kunchenguid/firstmate/issues/3215 * fix(review): Resolve same-branch run identities beyond capped history * fix(review): Fix run-selection compatibility, races, and worker-state fallbacks * fix(review): Limit run validation to the requested branch * fix(test): Anchor AXI fixtures and document remaining live evidence gaps * fix(document): Clarify run selection documentation and capture ownership * fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation * fix: distinguish captain outcomes from no-op updates (#4738) * fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work MAIN answered a supervision-branch outcome for completed captain-requested work (implementation done, PR ready for review and merge approval) with "Captain, shipshape.", reading section 9's no-action reply as covering it and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no captain-facing response is owed". Section 9 now limits the shipshape reply to true no-ops (idle re-read, empty heartbeat, consequence-free acknowledgement) and requires a short outcome response naming what finished and what word is needed whenever requested work finishes or a result needs the captain's word, even when a transcript entry already shows the substance. The Pi protocol's re-emit rule now says it bounds repetition only, and carries a worked example of the ready-for-review outcome whose correct processing turn a shipshape reply fails. No executable contract evaluates the content of MAIN's captain-facing reply, so the regression is the protocol example in the owner doc rather than a text-match test. * no-mistakes(document): Clarify captain-facing outcomes versus no-ops * docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line The document step condensed the Pi protocol's re-emit rule and dropped the worked example of a finished, ready-for-review outcome whose correct processing turn a "Captain, shipshape." reply fails. That example is the contract's regression: no executable contract evaluates the content of MAIN's captain-facing reply, so the owner doc's example is the test case. Restore it directly under the re-emit rule, prefixed as a regression example that is kept verbatim and never condensed or summarized away. * no-mistakes(review): Clarify captain outcome and decision-word requirements * no-mistakes(document): Clarify captain-facing completion outcomes * docs(pi): require the PR URL in the visible captain-facing outcome reply Captain review on the regression example: drop the sample reply string and say only that the ready-for-review outcome requires relaying a captain-facing outcome response, not just "Captain, shipshape.". Fold in the visible-PR-handoff failure seen this session: after the branch outcome reporting this fix green, MAIN's visible reply was only "Awaiting your merge call." with no PR URL, leaning on the dim anchor. Section 9's URL rule now also covers a review or merge ask and names the visible reply as where the URL goes, sourced from the ready status, pr= metadata, or the supervision branch's summary and never left to a transcript entry. The Pi protocol adds the same-way failure and places the captain-facing text in the final visible assistant reply after the fm_branch_processed call, because Calm hides assistant text emitted in the same step as a tool call as a working note. Investigation verdict, evidence in the PR comment: no recent PR caused the handoff failure; Pi has hidden same-step pre-tool assistant text since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and #4658 touched only remote report transfer. * no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs * no-mistakes(document): Clarify captain-facing supervision outcomes * docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule The consolidated section 9 URL rule narrowed its trigger to a review or merge ask, dropping the "whenever a PR is mentioned" catch-all from #3648 that keeps every PR URL copied from a durable record and never assembled from memory. Restore that trigger as a union with the review or merge ask so the one consolidated rule covers both. * fix(bin): let non-owner Claude Stops exit safely (#4777) * Fix foreign-owner turn-end supervision loop * no-mistakes(review): Scope foreign-owner safe exit to Claude guard * no-mistakes(document): Document Claude foreign-owner safe exit * fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778) Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty indexed array as an unbound variable and aborts the shell. In signal_turnend_panes_churned() the missing_keys loop was reachable with an empty array whenever every churned key already held a fresh .churn-since-* marker (a second churning turn-end inside an open deferral window), so each watcher cycle died about half a minute in and supervision restarted endlessly. The created_keys rollback loops had the same latent crash on their error paths. Audit of bin/ for the same pattern found one more confirmed-reachable case: remote_handoff's noncanonical-body scan iterates to_move, which is empty when a retried remote handoff finds every key already staged in the outbox. All other "${arr[@]}" sites are either count-guarded, guaranteed non-empty by construction, or unreachable while empty. Guard the three reachable expansions with the repo's existing "${arr[@]+...}" idiom. New regression test drives a real watcher through the all-marked churn path; the macos-stock-bash CI lane runs it under real /bin/bash 3.2 via FM_TEST_ONLY. * Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783) The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: require complete captain-facing final responses (#4779) * docs: require complete final responses across harnesses * no-mistakes(document): Document complete final replies for Grok Bot * docs: point Grok replies to the shared contract owner * no-mistakes(review): Clarify final recap without batching decision asks * fix: preserve substantive mid-turn text in Pi Calm (#4788) * fix(calm): preserve substantive Pi mid-turn text * no-mistakes(review): Preserve substantive Pi Calm text per block * no-mistakes(test): Cover shared Calm preservation boundaries behaviorally * no-mistakes(document): Consolidate Calm preservation documentation * fix: harden mail checks and rebalance full-coverage CI (#4800) * Improve CI reliability and rebalance full-coverage validation * no-mistakes(document): Clarify lint partition documentation * fix(bin): answer Kimi 2.0.0 folder-trust dialog during spawn (#4799) * Handle Kimi workspace trust dialog * no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers * no-mistakes(review): Gate Kimi ready on any trust marker and clean captures * no-mistakes(review): Read visible pane for Kimi trust and ready gates * no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate * no-mistakes(review): Harden Kimi viewport capture and trust dialog detection * no-mistakes(document): Document Kimi spawn refusal on cmux and Orca * fix(bin): report a dead-agent record once instead of escalating forever (#4775) * fix(bin): report a record whose agent is gone once instead of escalating forever The wedge escalation path never asked whether there was still an agent to be wedged. A wedge is something stuck that might recover, so re-alarming it earns its cost; an agent that is gone never moves again, its pane never churns, the idle timer never resets, and the escalate path clears its own timer and re-arms with nothing bounding the count. Observed on a live fleet: two finished lanes reached 226 and 203 consecutive escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400 notifications a day from two lanes with no agent running at all. On one, fm-control.sh exit answered already-stopped and fm-crew-state.sh read "failed - run failed". Closing the Herdr pane did not stop it either: with the pane genuinely gone and herdr pane read returning pane_not_found, the count kept climbing, because the poll is driven by the record's window= line rather than by the pane. The cost is not the repetition but that it drowns the alarms that matter. fm_backend_agent_state already separates a thinking agent from a gone one at process level. In the branch that was about to escalate, read it once and treat only its two recovery-grade verdicts - dead (endpoint present, no agent in it) and missing (endpoint authoritatively absent) - as proof, reporting that record once and not re-escalating it while it stays that way. Every other verdict, including alive, ambiguous, unreadable, unverified, and a read that failed outright, keeps the identical schedule, reason, and escalation count, so a genuinely wedged live agent is unaffected. The probe costs at most one backend read per window per threshold, the same budget the declared-wait consult and the worktree write probe already take. The report decides nothing about the record's fate: both lanes still held unlanded work and teardown refusing them was correct, so retiring, relaunching, or cleaning up stays with the supervisor. The once-only marker is owned entirely by that function and is dropped by the same read the moment the endpoint stops reading gone, so a replacement launched into the same window escalates normally and its own later death is reported again. Related, and not closed by this: #4412, #4482, #4316. Tests drive the real watcher against a record whose endpoint does not exist and pin both directions: dead and missing report once and never advance the count across later thresholds, while alive, ambiguous, and unreadable endpoints keep escalating with the identical reason and a climbing count. * fix(bin): bind the once-only dead report to the pane it reported Review of the parent commit found a reachable sequence where a later death in the same window lost its promised report. The marker was keyed on the verdict string alone and dropped only when a threshold probe read a non-gone verdict, but probes run only at thresholds: a replacement launched into the same window that dies without ever being probed alive - it crashes at startup, or works and then crashes - was absorbed by the previous death's marker. The pane's first sight yielded only the generic stale wake and every later threshold matched the stale marker, so the second death never got the detailed once-report that both the function's own comment and docs/architecture.md promise. Record the verdict together with the pane hash it was reported for, and absorb a repeat only while both still match. A replacement churns the pane, which resets the stale suppressor, wedge timer, and escalation count while no reset site touches this marker, so the pane half is what tells the second death apart from the first. The live-probe drop stays as it was. Clearing the marker at those reset sites instead would re-open unbounded re-alarming for a dead pane whose display ever ticks, which is the exact defect the parent commit exists to close. The noise bound is unchanged: an unchanged dead pane still absorbs on every later threshold and never advances the escalation count, and every verdict short of proof still escalates exactly as before. * no-mistakes(review): Key the dead-record once-marker on the busy incarnation token * no-mistakes(document): Document dead-record escalation cap in stale-pane config entry * no-mistakes(document): Add busy-state inventory line to AGENTS.md * no-mistakes(document): Document dead-record probe on busy-turn-bound wedge path * fix(bin): create captain-hold rows when Beads requires due (#4854) Captain holds have no due semantics and are a hold kind, not a Beads issue type. The create path now waives due.required and maps to native type task. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: disable compact adviser for spawned agents (#4877) * feat(bin): launch every spawned agent with the compact adviser disabled Every crewmate, scout, and secondmate Firstmate launches now starts with COMPACT_ADVISER_DISABLE=1, on a fresh spawn and on a relaunch alike, so an unattended session never activates the compact adviser. The value is unconditional: no configuration file gates it and there is no override, unlike the trace carrier beside it. Three carriers deliver it, because no single one covers every launch shape. The pane shell receives an export beside GOTMPDIR, so the agent's own children inherit it too. The launch command carries an explicit assignment, prepended outermost so it wins over any ambient value the pane already held. The cleared launch environment sets it again at the `env -i` boundary and keeps COMPACT_ADVISER_DISABLE in the fixed operational floor, which is what preserves the switch when config/launch-env-allowlist empties the environment, and what delivers it on a remote host that never had the value. bin/fm-control.sh relaunch, the bootstrap secondmate relaunch, and the remote secondmate transport all rebuild their launch through bin/fm-spawn.sh, so they inherit the same floor. The captain's own primary session is untouched. The two new suites drive the real spawn and then execute the launch command the pane actually received, with the harness replaced by a probe that prints its own environment, rather than matching script text. They cover ship and secondmate launches with the allowlist absent and enabled, the pane export and its ordering, fm-control.sh relaunch, and the full parent to remote-host chain. * no-mistakes(review): Export compact-adviser disable across compound launches * no-mistakes(document): Document spawned-agent compact-adviser environment guarantee * fix(bin): preserve Claude lock ownership after helper recycling (#4894) * fix(bin): let a background Claude session keep owning its session lock Session-lock ownership was decided by process ancestry alone. Under an unattended Claude session the model loop runs in a transient bg-spare bridged to the front-end by a shared daemon; when that bridge is recycled the contiguous claude-named ancestry from a hook to the recorded owner breaks while the owner pid stays alive, so the Stop auto-arm stood down as a foreign live owner, the turn-end guard ended every turn with its read-only diagnostic, and fm-lock.sh refused - a self-sustaining outage until restart. Ownership is now ancestry membership OR a trusted same-session id, never id-first: - fm-session-lock-lib.sh accepts CLAUDE_CODE_SESSION_ID only when CLAUDE_PID is a Claude-shaped member of the current contiguous run, compares it against the id recorded in state/.lock-session, and requires the recorded pid to still be a live harness. No id, no sidecar, an untrusted id, a different id, or a dead recorded pid leaves the ancestry verdict unchanged. Ids are never read from ps argv. - fm-lock.sh accepts a same-session holder at both refusal sites, writes, refreshes, and clears the sidecar only under its claim lock (including the early already-mine exit, skipped only while the deferred startup sweep leases that lock), keeps it byte-identical across a same-session confirmation, records CLAUDE_PID on lock line 1 for a session with a trusted id so a shared daemon or front-end that outlives the session never keeps a dead session's lock alive, never rewrites a live line 1 on a same-session confirmation, and names the recorded id in the live-owner refusal. - The .lock line-1 format is unchanged, so every reader that takes the whole first line as the pid keeps working; the guard's foreign-owner exit is unchanged and inherits the fix through the shared predicate. Tests: the ancestry suite drives the ancestry and id signals apart in a deterministic process table (asserting the divergence) and runs a real orphaned front-end/daemon/pty-host/spare tree through six phases with the real lock, auto-arm, and guard scripts; the foreign-owner repro keeps its negative control and adds a same-id positive control. Disclosure: no live unattended Claude background session ran on the verifying machine. The topology is documented by the real process listings in #3902, #2314, #3398, and #4066; coverage is the structural predicate plus the executable fixtures, not a live pass. Residual: bin/fm-sessionstart-nudge.sh keeps its own private ancestry walk (it only decides whether to print a nudge) and may nudge on a resume in the recycled case. Out of scope, deliberately: no structured lock format, no guard budget changes, no daemon-identity rejection, no fork lineage. * no-mistakes(review): Wait for claim lock; revert failed sidecars * no-mistakes(review): Revalidate ownership after wait; restore sidecars * no-mistakes(review): Roll back sidecar by publication phase * no-mistakes(review): Restore sidecar only if lock line is unchanged * no-mistakes(review): Trust session ids without a spelling allowlist * no-mistakes(review): Disarm sidecar rollback before backup cleanup * no-mistakes(document): Updated session-lock ownership documentation * feat: park main under the away posture on Pi (#4889) * feat: park main under the away posture on Pi While the away-posture record exists on a Pi primary, the supervision branch takes every actionable wake, no processing turn opens on main, captain rows accumulate for the return brief, and main's standing authority relocates to the branch through the existing guarded scripts. - lib/fm-branch-dispatch.ts: read the record at every routing decision; while it exists claim check, decision-owned, and heartbeat rows too, keeping the two broken-queue vetoes; expose checkSeqs so a claimed check row lifts task scoping. - fm-primary-pi-watch.ts: offer every actionable row under the record; a declined wake and every watcher-failure alarm still reach main. - fm-branch-supervision.ts: drop the legacy .afk decline; append a fixed POSTURE: AWAY tail carrying the record's read-back verbatim per wake; open no processing request while the record exists, re-checked immediately before a request would open and at every run boundary; present the accumulated rows at the first run boundary after archive. - fm-lease-lib.sh: fm_lease_forbid_branch passes the branch for opted-in actions only while fm-afk-contract.sh validate succeeds on a confirmed live record; PR merge, fresh spawn, and decision answer opt in, local landing never does. - fm-send.sh: a --resolve-key naming an open needs-decision or captain-held task is a decision answer and meets the partition; blocked: keys stay steering. - fm-spawn.sh: enforce the record's spend cap for a fresh ordinary spawn by either actor; relaunches and secondmates exempt. - fm-branch-prompt.sh: fixed Postures section and the verbatim ask-user-authority policy; the prefix stays byte-stable. - fm-afk-return.sh: count what the away session handled from the store. - docs, afk skill, AGENTS.md stub: main parked on Pi, green merge gate absolute while away. - tests: watcher and branch extension suites, fleet-record, merge, and decision-answer suites cover the relocation, the vetoes, the tail, the parked processing turn, the cancellation, the re-presentation, and the spend cap; dated live-guard evidence recorded. * no-mistakes(review): Refuse branch merge after preflight archive race * no-mistakes(review): Fix away wake, spawn, and processing races * no-mistakes(review): Suppress parked processing; narrow away-only rejection * no-mistakes(review): Abort dedicated processing; gate branch spawn once * no-mistakes(review): Stamp away-only on the dispatch offer * no-mistakes(review): Treat invalid away records as spend-cap absence * no-mistakes(review): Drop spawn test hook; abort processing-opened runs * no-mistakes(review): Bind abort to opening prompt; cap-read absence * no-mistakes(review): Limit away branch spawn to queued work only * no-mistakes(document): Correct AFK posture documentation * ci: standardize workflow timeouts into three tiers (#4910) * ci: simplify CI job timeouts to a three-tier policy Replace the scattered per-job timeout values (10m parallel, 25m lint, 30m serial, 10m macOS) with three readable tiers, each a hang tripwire with headroom rather than a packing estimate: - fast (5m): coverage guard, repo invariants, timing aggregate - normal (30m, one shared budget): lint partitions, portable parallel shards, portable serial shards, macOS stock Bash - heavy (Herdr only): 20m step tripwire on the family run so always() cleanup still runs, under a 75m job-level last-resort backstop The workflow's header comment states the policy and points at docs/fm-test-portable-shards.md "Timeouts", which now owns it, and each job names its tier beside timeout-minutes. tests/fm-ci-workflow.test.sh asserts the policy against the parsed workflow instead of the old per-job minute values: every job joins exactly one tier, exactly three distinct job-level values exist, the fast tier stays within 5-10 minutes, the normal budget stays at least double the modeled parallel lane sum reported by fm-test-run.sh --check-coverage, and the Herdr step tripwire stays below its job backstop with an always() cleanup after it. Concurrency supersession, shard counts, lane membership, and fail-fast settings are unchanged. * no-mistakes(review): Decouple the normal timeout from packing estimates * no-mistakes(review): Assert Herdr teardown follows the family run * no-mistakes(review): Pin Herdr family-run timeout to 20 minutes * no-mistakes(review): Ignore comments when identifying Herdr steps * no-mistakes(review): Identify Herdr steps by declarative ids * no-mistakes(document): Clarify authoritative three-tier timeout policy * fix(bin): keep supervisor status closes from waking the same home (#4895) * fix(bin): keep supervisor status closes from waking the same home A drain that already folded OPEN DECISIONS has presented those bytes even when the watcher has no matching seen marker. Treat that fold, and the presentation cursor, as known so the bookkeeping close stays quiet while later worker lines still signal. * no-mistakes(review): Keep folded worker failures waking past supervisor closes * no-mistakes(review): Wake on unlisted folded worker lines; batch multi-key closes * no-mistakes(review): Stop folded worker resolved lines from counting as already read * no-mistakes(document): Correct self-announced close marker contract in docs * fix(bin): stop labeling Herdr as experimental (#4972) * Stop steering operators away from Herdr * no-mistakes(review): Neutralize remaining Herdr opt-out documentation wording * fix(bin): treat a live no-mistakes run as current after rebase (#4973) * fix(bin): treat a live no-mistakes run as current after rebase A running run on the task's branch is authoritative regardless of head. Matching only the local head made a rebased in-flight run look failed. * no-mistakes(review): restrict coarse live-any-head to foreign-branch answers * no-mistakes(review): reject gate-parked runs from the executing predicate * no-mistakes(review): hoist gate-marker patterns into single run-lib owner * no-mistakes(review): require live daemon for head-free run binding * no-mistakes(review): require answered daemon-down before unbinding live runs * no-mistakes(review): extend daemon guard to anchored continuation routes * no-mistakes(review): delete live-any-head; restore dead-daemon verdict * no-mistakes(review): keep parked gates parked; name dead daemon everywhere * no-mistakes(review): set dead-daemon verdict instead of emitting early * no-mistakes(review): align selected route with legacy dead-daemon handling * no-mistakes(review): drop unproven-record binds; narrow coarse gate reading * no-mistakes(review): narrow header, drop vestigial guard, retarget tests * no-mistakes(review): revert coarse gate override; require answered-down probe * no-mistakes(review): cache one daemon probe; stop duplicating run id * no-mistakes(review): restrict coarse dead-daemon verdict to moved-off rows * no-mistakes(review): delete coarse dead-daemon extension and gate note * no-mistakes(review): delete remaining coarse dead-daemon block and stale docs * no-mistakes(document): document rebase-safe live-run bind and unverified-record verdict * fix(bin): prevent long worker launch command truncation (#4994) * fix(bin): stage the launch command in a private file and type a short source line A long launch line typed while the fresh pane shell is still busy waits in the terminal's canonical line buffer, which drops input past about 1,024 bytes on macOS, so the pane was left at an unfinished command with no agent running. fm-spawn now writes the assembled command to the task's own temp root under umask 077 and types only a short line that sources it. Refs #4559 * fix(bin): keep the per-task temp root private before staging the launch command The root lives at a predictable path under /tmp and now holds the whole launch command. Create it with mode 0700, refuse one that already exists as anything but a directory owned by this user that nobody else can write, and tighten an owned one, so no other local user can plant or swap the staged file. Refs #4559 * fix(bin): enforce private staged launch file mode * test(spawn): cover long staged Claude launches * no-mistakes(review): Namespace launch files and prove truncation staging * no-mistakes(review): Use immutable per-spawn launch filenames * no-mistakes(document): Document staged launch delivery safeguards * no-mistakes(ci): Updated eight behavior tests/fakes to execute or inspect immutable staged launch files instead of expecting inline launch commands. This restores Muse, secondmate lifecycle/restart, remote trace/parent binding, compact-adviser, and Orca coverage. All affected tests, dispatch-profile regression, fixture tests, syntax checks, ShellCheck, and git diff checks pass --------- Co-authored-by: Vytautas Stankus <svycka@gmail.com> * test: authorize isolated Herdr lab validation (#4998) * Add isolated Herdr runbook to test instructions * no-mistakes(review): Drop substring matching from test.instructions contract * no-mistakes(review): Assert commands.test key absence in YAML * Drop unit-first sentence and instructions contract test Captain-scoped follow-up on the Herdr-lab test.instructions ship: keep the lab safety runbook only, and leave the no-mistakes contract test focused on commands.test absence. * docs(vision): accept vendor-semantics and 9k AGENTS ceiling (#4873) (#5001) * docs(vision): accept vendor-semantics and 9k contract-ceiling amendments (#4873) Replace the pixels-of-today's-UI rule with a quarantined, version-pinned surface-adapter exception recorded as standing debt. Cap the always-loaded contract at 9,000 words and require prune-or-trigger before a crossing change lands. Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com> * docs(vision): restore accepted three-sentence vendor-semantics form (#4873) Replace the compressed paraphrase with the issue's accepted wording: a named quarantined version-pinned adapter, expected to break, recorded as standing debt that never hardens into a shared contract. Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com> --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com> * feat(bin): defer the wedge escalation for a lane parked at a supervisor-owed gate (#4974) * fix(watch): recheck a gate awaiting a human instead of wedge-escalating it A lane whose validation run is parked at a gate waiting on a human decision is correctly quiet, but nothing in its status line says so: the evidence is the pipeline's own gate state rather than anything the worker wrote. The wedge timer read that silence as a suspected wedge and climbed the escalation ladder for as long as the wait lasted, and each escalation cost a supervising turn. The landed declared-wait consult does not reach it, because a live ordinary crewmate never reports a declared pause, and raising FM_STALE_ESCALATE_SECS would delay genuine wedge detection for every lane by the same amount. The threshold now reads a second, independent record when the status line accounts for nothing: whether the crew's current state is a gate whose answer is owed by a human. That is minted only from the gate's own findings table, by a row whose `action` column is exactly `ask-user`, located by position out of the table header the way nm_gate_step_row already reads its row - never searched for over the run payload, where a finding's free-text description or a branch name satisfies a search just as well. A gate awaiting the CREWMATE's own answer keeps the unchanged escalation schedule, reason and demand-deep-inspection wording, because a crewmate that goes quiet before answering its own gate is exactly the wedge the ladder exists to catch. Each kind of wait now carries the human it is on, the action that clears it, and whether that human is the captain as data alongside the verdict, rather than as wording chosen per branch where the recheck is written, so the deferral cannot word one kind of wait as another and a new kind cannot ship without deciding all of them. A parked gate has no written record of when its wait began, so its recheck publishes no wait age at all rather than one read from the quiet window this deferral resets on every pass, which would report the same small number for a gate of any age. Like every other captain-facing recheck here it is absorbed in silence while the away-posture record exists, arming no throttle, so the recheck is owed in full the moment the record is archived. The consult runs only in the at-threshold branch that was about to escalate, beside the worktree walk already there, and only for lanes whose status line explained nothing. Closes #3055 * no-mistakes(review): require an unanswered decision before deferring a parked gate * no-mistakes(review): reset the away-silenced timer, fail-safe findings parse, US-joined wait records * test(watch): pass the pane hash wedge_timer_check now takes Upstream gave wedge_timer_check a sixth <pane-hash> argument for its dead-record probe. The malformed-wait-record rounds drive the real function directly, so they pass one, and stub fm_backend_agent_state to a live agent so the probe that runs after a refused deferral keeps the unchanged ladder rather than reading a backend the child shell has none of. * no-mistakes(review): Bind parked-gate wait to its run, owe it firstmate * no-mistakes(document): correct wait-kind count, crew-state reader scope, gate-key coupling * feat(watch): make the parked-gate wait deferral opt-in The wedge timer deferring a lane parked at a validation gate is new supervision behaviour rather than a restored one, and it decides which lanes give up the escalation ladder, so it now ships as a default-off per-home option instead of changing every home on upgrade. config/wedge-defer-parked-gate arms it. The flag is read before the decision fold, so an unconfigured home spends no fold or current-state read, writes no record, and keeps the unchanged escalation schedule, reasons and demand-deep-inspection wording; a test counts the reader calls in both directions to pin that. It is not inherited by secondmate homes: each home supervises its own crew and owns that trade separately, the same reason config/turnend-churn-absorb is home-local. The away-posture absorb returns to leaving the idle timer alone, which it had restarted only because the costly consult could reach it. A parked-gate wait is owed to the supervisor rather than the captain, so it never enters that branch, and the recheck owed on return is again owed in full the moment the record is archived. * test(watch): pin that the away-silenced hold leaves the idle timer alone The absorb no longer restarts the timer, so the recheck owed on return is owed in full rather than a cadence into the return. Nothing asserted that, so a restart could be reintroduced silently. * no-mistakes(review): document away-silence rationale, pin captured gate component * no-mistakes(test): anchor gate row scan to the braced findings header * no-mistakes(document): pin same-block gate row invariant in crew-state comment * fix(bin): reclaim a task whose herdr endpoint was destroyed (#5007) * fix(control): let the owning seat reclaim a task whose endpoint is gone A destroyed pane or workspace made `missing` a terminal state. Relaunch accepted only `dead` and said to stop the agent first; exit refused `missing` and said to reconcile the task first; there is no reconcile verb. Each command named the other as its prerequisite, so a task whose terminal went away could not be reclaimed by anything, and a no-mistakes approval it was parked on had no seat left to answer it. `missing` is agent-free a fortiori: there is no endpoint, so there is no agent in it. Widen the existing guards rather than add a verb. - fm-spawn --relaunch accepts a positively proven `missing` and creates one fresh endpoint in the recorded worktree; the record it already republishes rebinds the task to it. A `dead` endpoint is still adopted in place. - fm-control exit reports `endpoint-gone` instead of dying, so the relaunch transaction's stop step no longer dead-ends, and re-resolves the endpoint from the record before verifying the replacement. The duplicate-agent refusal is untouched: both verdicts come from the same recovery-grade classifier, which claims `missing` only from positive absence, so `alive`, `ambiguous`, and `unreadable` all still refuse. The backends' own create paths refuse a live same-labeled endpoint as a second independent guard. The worktree, its branch, commits, uncommitted changes, armed poll and registration, record rows, and status log are all untouched - a reclaim is a recovery, never a teardown. A secondmate is excluded: its gone-endpoint recovery already has one owner in the session-start liveness sweep, so relaunch refuses and names it rather than becoming a second path to the same outcome. Tests reproduce both halves of the deadlock, the reclaim succeeding, unlanded work surviving it, and the refusals that still hold. * no-mistakes(review): prove endpoint absence per backend before reclaim rebinds * no-mistakes(review): give exit and relaunch one absence proof; pin herdr rebind session * no-mistakes(review): narrow endpoint reclaim to herdr; tmux refuses honestly * no-mistakes(review): stop refusals and docs asserting unestablished causes * no-mistakes(review): stop herdr fixture helper losing tmp-root registration * no-mistakes(review): document workspace drift and absence-probe server residue * no-mistakes(review): correct rebind limitation to its one reachable case * no-mistakes(review): stop claiming reclaim leaves instructions untouched * no-mistakes(document): scope fm-control-lib purity claim, note reclaim coverage * no-mistakes(rebase): read the staged launch file in the herdr fixture Rebasing onto main picked up #4994, which stages a long worker launch command into a script and delivers the short `. '<path>'` line instead of the literal command. The tmux fake and tests/fixtures.sh were updated for that; the herdr fake this branch adds was written before it and still keyed "an agent now exists on this pane" off the literal `encode launch-brief` text, so after the rebase it never marked the rebound pane live and the reclaim's alive-wait read `dead`. Dereference the staged file first, exactly as the tmux fake above does. Test-fixture only; no production path changes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * no-mistakes(document): note reclaim placement in herdr and scripts inventories --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(bin): stamp status events with their emission time (#3764) * test(status): reproduce missing event emission time * wip(status): preserve optional event emission time * test(status): document indirect clock stub invocation * no-mistakes(review): Preserve historical status bytes during reply recovery * no-mistakes(test): Fix timestamped status assertions and remote fixture dependencies * no-mistakes(review): Preserve captain regex overrides for timestamped status events * no-mistakes(document): Clarify status event timing and publication contracts * no-mistakes(lint): Quote literal done to satisfy ShellCheck * no-mistakes(ci): Captain, updated .github/workflows/ci.yml to expect 19 snapshot tests instead of 18, matching the PR’s added regression. Reproduced the failure before the fix. Stock Bash 3.2.57 verification passed: parse sweep, 19 snapshot tests, 53 Bearings tests, and the public-followup regression. Workflow lint and diff checks passed * no-mistakes(test): Preserve terminal notifications with malformed timestamp tags * no-mistakes(test): Stamp Rovo spawn failures with emission time * no-mistakes(document): Verify status event documentation * no-mistakes(lint): Fix ShellCheck quoting in status emission-time tests * no-mistakes(ci): Captain, fixed four lifecycle assertions to accept emission timestamps while preserving publication and retry checks. Reproduced the CI failure before the fix. The lifecycle suite now passes with six Beads capability skips; syntax, targeted ShellCheck, and diff checks passed * no-mistakes(ci): Captain, fixed malformed timestamp colons hiding actionable events using shared normalization. Original bytes and unknown ages are preserved. Regression reproduced before the fix; classifier and remote-reply suites, targeted lint, syntax, and diff checks passed * no-mistakes(review): Stamp remote escalations at call sites, drop new flag * no-mistakes(review): Accept stamped escalation and close lines in test assertions * no-mistakes(review): Restore reserved-key answered-note guard for stamped closes * test(status): accept optional emission time in PR-provenance assertions The #4148 provenance test landed on main with exact unstamped greps. Parent-channel lines from this branch carry [at=<epoch>], so strip only that tag before the same exact match. No production change. * no-mistakes(review): Accept stamped ready signal in PR fallback scrape * no-mistakes(review): Drop relay flag, stamp parent events at call sites * no-mistakes(review): Stamp worker terminal-signal instructions, revert fm-on fixture * no-mistakes(review): Accept optional stamp in live cmux drift guard * no-mistakes(review): Restore original test invocation order in two suites * no-mistakes(review): Strip only well-formed numeric status time tags * no-mistakes(document): Drop stale unstamped PR-ready line spelling from channel doc * no-mistakes(review): Stamp agy spawn-failure status lines with event time * fix(bin): normalize status event times in-shell and freeze the budget test clock Two paths made a status event's emission time cost more than it should. The captain-relevance fallback piped every line through awk to drop a well-formed `[at=<epoch>]` tag before matching, so a supervisor sweep paid a fork per line just to prepare a regex match. Shell parameter expansion does the same strip with no fork, and the retry-dedup scan now reuses that one helper instead of carrying a second copy of the rule in awk. The copies had already drifted: the shell side stripped tags from lines with no colon, which the awk rule left whole, so a colonless line could be mistaken for one already recorded. One definition, checked against the awk rule it replaces over the edge cases and a 4000-line fuzz. tests/fm-contributions.test.sh froze its fixture clock only in exhaust mode. In hang mode the poll set DEADLINE to the real now plus a one-second budget, and when the second ticked before the first forge call the loop broke without ever calling gh: forge/calls was never written and the assertion failed reading a missing file. Freezing the clock in both modes removes the dependence on wall time; the bounded call is still cut by the real timeout, so the observation the test asserts still starts. Emission time stays optional on new status records, and legacy or malformed lines keep an unknown age. * no-mistakes(review): Stamp ask-user escalation line and fix Kimi status assertion * no-mistakes(document): Drop stale unstamped done-line spelling from watcher docs * test: fold emission-time snapshot coverage into the fixture case Drop the incidental ci.yml 18-to-19 count hunk so the PR no longer touches workflows. Keep every emission-time assertion by folding it into test_fixture_snapshot_json. * no-mistakes(review): replace brief date substitution with epoch placeholder; drop emitted_at_epoch * no-mistakes(review): align untimed normalizer with epoch parser; tolerate placeholder stamp in PR scrape * no-mistakes(review): strip undelimited at-tags; correct brief stamp header * no-mistakes(review): normalize stamps at both captain-regex sites; restore mtime freshness * no-mistakes(review): strip colon-bearing stamps for relevance; fix headers and test oracles * no-mistakes(review): narrow escalation match to stamp tolerance; pin note verb * no-mistakes(review): read note and key past colon-bearing stamps * test(status): keep inactive reconcile assertions stamp-tolerant These two oracles were made stamp-tolerant while resolving one of the branch's merges from main. The rebase drops merge commits, so that adaptation was lost and both assertions went back to matching an exact substring that a stamped line no longer contains: the tag lands before the colon, so "failed [key=k]: ..." is now "failed [key=k] [at=N]: ...". Strip a well-formed tag before matching, as the branch's other oracles do. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * no-mistakes(review): unstamp fold colon tests; reserve stamp width in cap * no-mistakes(document): correct stale unstamped status-line spellings in docs * no-mistakes(document): quote brief-test literals for lint; correct stamp-helper contract comments * no-mistakes(ci): rename subshell-local epoch in delivery-race stub The serialization test overrides fm_pending_reply_mark_delivered inside a (..) subshell. Its `epoch` local collided with the same name in status_line_at_epoch/status_stamp_line, which this branch added and this suite now calls at top level, so ShellCheck 0.11.0 reported SC2030 and failed Lint 2. The stub already prefixes its other locals with `pending_` for the same reason; `epoch` was the leftover. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(bin): unify Lavish host and disconnect handling (#5060) * fix: ship clean Lavish host fixes * no-mistakes(review): Fix Lavish classifications and fail-closed host loading * no-mistakes(review): Restore Lavish host state across retries and launches * no-mistakes(review): Preserve destination Lavish host when configuration is absent * no-mistakes(document): Document Lavish status and host guarantees * feat: act on captain's away words during AFK supervision (#5076) * feat(afk): make the captain's away words the whole mandate Retire the clause fields, verb list, never-set scan, refused records, and the per-task merge-grant list from the away-posture record. The record is now version 2: the captain's words verbatim plus expected return, spend cap, and reach line; a version 1 record still validates, reads, and archives so a live away window is never broken by the upgrade. The supervision branch reads the words at the tail of every wake and acts on them by its own judgment through the guarded scripts under standing authority, never by analogy, holding for the return on doubt, and opens each such outcome summary with "per your away instructions:" so the return brief can render the words beside the session's account. While the record exists any green merge runs under away authority (ledger tag "away"); red merges, --allow-red, asynchronous and queued merges, and local-only landing stay refused. The branch may file a backlog item the words explicitly call for before dispatching it under the spend cap. Tests drive fm-afk-contract.sh, fm-afk-launch.sh, fm-afk-return.sh, and fm-pr-merge.sh as commands: version 2 written, version 1 read, retired flags and subcommands refused by name, green merges landing under the record, red and waived-red refused, the record lock still closing the authority-read window, and the Pi away tail carrying the words. * no-mistakes(review): carry the away read-back to the session verbatim * no-mistakes(review): match the exact away-action marker in the return brief * no-mistakes(review): refuse a words block truncated by a damaged line * no-mistakes(document): Refresh away-role contract documentation * fix(bin): render the remote charter's steering-inbox path host-local (#5049) * fix(bin): render the remote charter's steering-inbox path host-local A freshly provisioned remote secondmate read a parent-home absolute steering-inbox path in its charter - a location that exists on no route - and spent its first turn discovering the gap and filing a blocked decision for what was a render defect. The seed's remote-copy rewrite now maps the inbox to the route's host-local parent-route inbox, exactly as it already maps the reply-log path, so every mention - bare path, listing, and handled/ acknowledgement - lands host-local. Both rewrites also become plain assignments, because a quoted substitution nested inside a double-quoted printf argument leaks literal quotes into the replacement text on stock macOS bash. The lifecycle suite pins the corrected render both directions against the real seed, provisioning, and delivery route, sharing one fixture value between the render truth and the delivery truth. Closes #5012 * no-mistakes(document): document remote charter's host-local steering inbox * feat: route Lavish feedback directly to owning workers (#5099) * feat(procevent): route worker-owned Lavish rounds * no-mistakes(review): drop duplicate artifact field from task-owned registration * no-mistakes(review): post worker reply once, fix ring label, keep re-arm atomic * no-mistakes(review): keep worker board owned until terminal round acknowledged * no-mistakes(review): refuse every retirement of an open worker-owned round * no-mistakes(review): use real lavish reply flag, isolate reply generations * no-mistakes(review): drop .posted marker for best-effort reply posting * no-mistakes(review): consume staged reply after listener setup, refuse orphaned captures * no-mistakes(review): require a reachable owner, redeliver open rounds, roll back failed re-arms * no-mistakes(review): re-arm only to acknowledge an open round * no-mistakes(review): conclude only a still-open terminal round * no-mistakes(review): record the acknowledgement before retiring the board * no-mistakes(review): retain the registration across a conclude, qualify terminal docs * no-mistakes(document): Document worker-owned Lavish round lifecycle * fix(bin): fit pull observation within the contribution poll budget (#5107) * fix(bin): reserve contribution observation budget * no-mistakes(review): Strengthen slow-read regression test to exceed the poll budget * feat(bin): add idempotent inbox capture, replies, receipts, and readiness JSON (#5103) * feat(bin): add idempotent inbox orders, receipts, replies, and readiness Let a caller supply a request id when publishing a captain inbox note so a retry returns the original note instead of creating a second one, including across the crash window between save and wake announcement. Separate saved from announced so a failed wake is repairable without enqueueing again. Add bounded receipts JSON with omission disclosure, a durable primary reply against a note id, and a read-only readiness projection that can say unknown instead of inferring liveness from a lock file. * no-mistakes(review): fix(bin): honest inbox announce, reply cursor, and readiness verdict * fix(bin): resolve ready from lock-holder ancestry; drop lock status --json Remove the extra JSON surface from fm-lock.sh so its human status still always exits zero. Have the readiness projection classify the inspected home from the lock-holder pid via fm-harness.sh ancestry, with an explicit FM_SUPERVISION_MODEL still winning and an unknown model when there is no holder. Prove the yes path when that ancestry names a known harness. * no-mistakes(review): Harden inbox announce, receipts reads, and reply sequence cursor * no-mistakes(document): Note read-only lock inspection in scripts inventory * no-mistakes(lint): Pass missing id argument to malformed-reply test printf --------- Co-authored-by: cliflacata-svg <304148223+cliflacata-svg@users.noreply.github.com> * fix(bin): stop harness footer rows below a composer from reading as pending text (#5118) * fix(composer): stop a harness footer row from reading as a composer holding text A harness draws its own furniture below the composer - a user statusLine, a permission-mode hint - and the cursorless "bottom-most shape wins" rule looks exactly there. `→` (U+2192) is Cursor's prompt glyph but ordinary text everywhere else, so a statusLine opening with `→` was selected as a bare composer, swallowed the hint row beneath it as wrapped input, and answered `pending` on a visibly empty pane. `fm_task_inbox_ring` defers on exactly that verdict, and `bin/fm-watch.sh`'s re-ring calls the same function, so the first doorbell and every retry were skipped and the worker never saw the steer. Measured live on 2026-09-20: three of five Claude Code 2.1.236 worker panes on Herdr 0.8.0 had genuinely empty composers and every one of them was refused. A separator pair that closed over a bare agent-glyph row is a proven composer container, so the contiguous non-blank rows below its closing rule are that composer's footer and are no longer composer candidates. The demotion is bounded by all three of its own preconditions: a blank row ends the zone, a pair that closed over no glyph row demotes nothing, and a shape with no separator pair at all (Cursor's half-block rules) is untouched. Real unsubmitted text in that same composer, including a stray SGR mouse report left by a click in the pane, still reads `pending`. Pinned by two portable regressions and by a new cursorless arm on the live composer-matrix guard, which re-reads each harness's already-proven-idle pane the way every non-tmux backend reads it and fails naming the harness and version when that read is `pending`. * no-mistakes(review): make composer footer-zone demotion shape-independent * no-mistakes(review): make footer-zone demotion refuse-only and drop rescan * no-mistakes(lint): quote probe-absent sentinel to clear ShellCheck SC2100 --------- Co-authored-by: Koen Muller <koen@catapult.nl> * feat(bin): append optional home-local include to briefs (#5115) Co-authored-by: guanchengh-lgtm <271917158+guanchengh-lgtm@users.noreply.github.com> * fix(bin): report a branch with no validation run as absent instead of an unreadable runs table (#5114) * fix(bin): stop misreading a no-run branch as an unreadable runs table Defect: when `no-mistakes axi status`'s overview is truncated (a task's own branch has zero rows among the shown ones), fm_nm_select_run's Python fallback derived the repo identity for its direct SQLite query from a `repo: <path>` line it expected in the overview text. The real CLI never emits that line, truncated or not (see the genuine capture at tests/captures/no-mistakes-v1.70.1/overview.toon, which has only `count:`/`runs[...]:`), so the lookup always failed and reported "unreadable runs table" for a task that simply has no run on its branch. On a fleet with many concurrent runs, every idle-branch task hits the truncated-overview path routinely, so this fired every few minutes and drowned genuine unreadable/blocked verdicts in noise. Fix: derive the repo identity from the task worktree path instead, which is exactly the value `no-mistakes` records as a repo's `working_path` (confirmed against the existing capped-overview test fixtures, which already register repos by worktree path). A worktree path that is not absolute cannot be matched and still reads as unreadable rather than being guessed at. Also raise the reader's SQLite busy timeout from 1s to 30s so ordinary lock contention on a busy fleet cannot masquerade as an unreadable database. Safety: every other verdict byte-for-byte unchanged - the repo lookup still requires exactly one matching row (a genuinely corrupt or mismatched repos table still reports unreadable, per the existing `repo` failure-mode test), the branch query and row validation are untouched, and a zero-row result for the branch still flows through the same recursive re-parse that already turns an empty `runs[0]{...}` table into `absent`. Added a regression test (test_capped_overview_without_repo_line_and_no_runs_reports_absent) that reproduces the real overview shape - capped, zero rows for the task's branch, no `repo: ` line - and asserts the crew state falls through to the pane/busy verdict instead of reporting unknown or "unreadable". Full fm-crew-state.test.sh suite passes unchanged otherwise. * fix: recovered same-branch inventory awk misreads empty result as unreadable fm_nm_select_run's deep SQLite reader rebuilds a `count:`/`runs[...]:` overview and re-runs it through the same awk selection pass. When that rebuilt inventory has zero rows for the branch, the row-matching loop never executes, so its counters (`seen`) stay at awk's uninitialized empty string while `expected` and `shown` are plain strings parsed from the header text. Comparing an uninitialized value against a non-numeric string uses string comparison, so "" != "0" is true, and the END block takes the "unreadable runs table" branch instead of falling through to the correct "absent" verdict for a branch with genuinely zero runs. Coerce the affected END comparisons with `+0` so they are always numeric, matching seen/expected/shown/total regardless of whether awk classified them as strings or numeric strings. A truncated or genuinely malformed inventory still differs numerically and still reports unreadable. * no-mistakes(review): bound capped-overview inventory reader and canonicalize worktree lookup * no-mistakes(review): match recorded repo path first, tolerate duplicate spellings * no-mistakes(review): revert repo lookup to exact working_path match * no-mistakes(document): note state-db inventory read under crew-state nm timeout --------- Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.c…
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Abandon 3637. Have the crewmate who did the investigation start a clean slate PR. I wonder whether the fix should be that workers shouldn't report something like "PR 108". We should instruct it to always send full URL PRs whenever a PR is mentioned. Also, stop mandating a full URL - adjust the instruction there to something like: send full PR URL if you have one in the records, otherwise only report whatever identifier you do possess. Ask whether what I am proposing would be a sufficient fix.
Closed: #3637
Investigation: data/fm-supervision-url-hallucination-cause-s1/report.md
What Changed
Risk Assessment
✅ Low: The change is limited to aligned worker and supervisor instructions plus behavioral checks of generated prompt output, with the previously unrequested enforcement fully removed.
Testing
Completed 1 recorded test check.
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
🔧 **Review** - 1 issue found → auto-fixed ✅
bin/fm-pr-check.sh:58- The new status-log refusal path exceeds the stated intent, which only requires worker/supervisor copy-or-abstain instructions. It also does not provide the claimed durable protection: status logs are append-only, andfm_pr_status_ready_urlsreturns every historical done-line URL, so afterdone: PR <wrong-url>followed by the diagnostic’s recommended corrected done line, the old URL still matches and can be recorded or merged. Remove this unrequested enforcement component and retain the required instruction changes; ask the user before adding a deterministic enforcement policy with explicit correction/supersession semantics.🔧 Fix: Remove stale PR URL enforcement
✅ Re-checked - no issues remain.
bin/fm-test-run.sh --changed --exclude-family real-herdr-gated✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.