fix(bin): preserve captain calls during teardown - #3595
Merged
Merged
Conversation
A scout that held its own work item for the captain, which is what
captain-hold-lifecycle prefers ("hold the work item the question gates"),
was closed by bin/fm-teardown.sh's automatic backlog transition. The
completion gate passed, cleanup ran, and the captain's question moved to
Done with no recorded answer: the one thing the policy says must never
happen. `tasks-axi done` closes a held row silently, and nothing in
teardown asked whether the row was the captain's own call.
bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when
the task is still an open captain call, 1 when it is not, 2 when that
cannot be established. It reads the row through the transition library's
backend-aware probe, so it addresses the same backlog teardown does; the
script's other commands now address the configured data directory the
same way instead of FM_HOME, which also fixes captain holds in a home
with a relocated data directory.
Teardown asks `open` before any destructive step and refuses on 2. On 0
only the close changes: after cleanup and still under the task's own
lock, the row gets one "Deliverable of the finished work" line at the end
of its body and returns to Queued through `tasks-axi reopen`, keeping its
hold, so it lands in Captain's Call instead of reading as work under way.
--force does not lift this: it authorizes discarding unlanded work, never
the captain's question. The deliverable goes into the body because
`tasks-axi update --report` rewrites the title of a row that is not Done.
The crash window reuses the pending-close record teardown already stages:
a `mode=retain` line makes the existing replay record the deliverable and
reopen instead of closing, with the same validator, stale-generation
check, cleanup-incomplete marking, and non-blocking bootstrap lock as an
ordinary close. A retained row the captain answered first simply retires
the record. No parallel record type, recovery command, or second bootstrap
loop is introduced.
Regressions run the real executables: the captain-held scout survives
cleanup queued, held, with its deliverable and on the board, only
`answer` closes it, --force keeps it open, and an ordinary scout still
closes with its report; an interrupted cleanup leaves the row untouched
and the next session start retains it; a relocated backlog keeps the
retention in its one configured file; and a ship row whose hold cannot be
read refuses cleanup before anything destructive.
Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np
Confidence Score: 5/5The PR appears safe to merge because the captain-call state check, destructive teardown, and recovery transition are serialized and fail closed when the hold state cannot be determined. The changed lifecycle preserves open captain calls through normal and interrupted teardown, avoids reopening calls already answered by the captain, and retains the existing ordinary close behavior without an accepted correctness or security defect. Reviews (1): Last reviewed commit: "no-mistakes(document): Fix relocated cap..." | Re-trigger Greptile |
Marcos-Leyva
added a commit
to Marcos-Leyva/firstmate
that referenced
this pull request
Sep 3, 2026
) * fix: surface inbound Relay media to responding agents (kunchenguid#3442) * fix: surface inbound Relay attachments to the responding agent A Discord support thread's screenshots were never seen by the agent handling the mention. The relay delivered them and the poll stashed them: the reporter's images arrived on the `thread_starter` entry of `in_reply_to_chain` while the mention's own media list was empty. The gap was in the responder's playbook, which enumerated a fixed field list (`request_id`, `text`, `in_reply_to`, `in_reply_to_chain`) and so made every other field, attachments included, invisible. Fix it where the gap is, in prose: - Read the complete payload object rather than a fixed field list, so media and later relay fields are never skipped again. - Fetch and view attached media with the agent's own tools, on the mention and on every chain entry, and call out the common shape where only the thread starter carries the screenshots. - Restrict those fetches to known-good platform media hosts over https (Discord: cdn.discordapp.com, media.discordapp.net, images-ext-1.discordapp.net, images-ext-2.discordapp.net; X: pbs.twimg.com, video.twimg.com), report a blocked host instead of working around it, and treat everything fetched as untrusted public input on the same terms as the surrounding thread text. The poll stays out of it and downloads nothing, so no third-party bytes are pulled on the polling path. The new test pins the contract the playbook depends on: a mention in the incident's shape, with an empty top-level media list and screenshots on the thread starter, must reach the inbox with the payload intact and its media URLs unfetched. * no-mistakes(review): Preserve media authority and enforce poll-only fetching * no-mistakes(document): Clarify Relay attachment safety prose * fix(bin): defer inactive reconciliation during startup (kunchenguid#3480) * Defer inactive startup reconciliation * no-mistakes(review): Queue deferred inactive reconciliation diagnostics durably * no-mistakes(review): Require worker phases to cover startup requests * no-mistakes(review): Make diagnostic wakes safely acknowledgeable * no-mistakes(document): Document deferred startup phase coverage * fix(bin): bound wake drain presentation lock waits (kunchenguid#3475) * fix: bound status presentation lock waits * no-mistakes(review): Distinguish malformed presentation locks from live contention * no-mistakes(review): Bound no-ack drain queue lock acquisition * no-mistakes(document): Document bounded presentation-lock drain behavior * no-mistakes(lint): Annotate bounded lock output global * no-mistakes(ci): Added deterministic regression coverage for successful bounded-lock acquisition after live contention, verifying helper-to-caller PID ownership handoff and caller release. Verified with bash syntax checks, git diff checks, and the full fm-wake-queue test suite * fix(bin): retire public follow-ups in remote homes (kunchenguid#3479) * fix(relay): close a public loop whose work lives in a remote secondmate home A public-followup loop bound to a REMOTE secondmate could never be closed. `clear_public_followup_link` (bin/fm-public-followup.sh:701) required an absolute recorded `work_home_path` for a `secondmate:*` work home, but a remote route has no local path on this machine, so registration records that field empty (bin/fm-public-followup.sh:291). Every close ran that clear first, so `retire` died with "could not clear the legacy X link ... retained for reconciliation" forever, and `deliver` posted the public reply and then stranded the loop at `posted`. `--force` never covered that step. The clear now goes to the remote home over that route's SSH transport, running `fm-x-followup.sh --clear <work-id>` through `bin/fm-on.sh`. The route is decided from `data/secondmates.md` before any local path is consulted, so a same-named local directory can never stand in for a remote home, and registrations already on disk retire without needing a new field. `fm-on.sh` passes ssh's status through, so 255 stays the established "delivered but completion unknown" result this codebase already reconciles: the close is refused, the registration and the remote link are left exactly as they were, and the message names the unknown completion instead of claiming a definite failure. Local secondmate and `main` work homes are untouched, and `--force` still governs only the unresolved-obligation refusal. Three regression cases drive a remote route end to end, faking only the ssh binary at the FM_SSH_BIN seam and then running the real remote entrypoint against a local checkout, so the clear that must reach the remote home actually happens there. * no-mistakes(review): Guard remote link clears by request identity * no-mistakes(review): Fail guarded clears on unreadable remote state * no-mistakes(review): Reject guarded clears on non-writable remote state * no-mistakes(review): Allow no-link retirement in non-writable remote state * no-mistakes(document): Correct public-followup verification guarantee count * no-mistakes(ci): Fixed the guarded link-clear race by ensuring absence is decided under the metadata lock whenever publication is possible. Added a behavioral concurrency regression test. Verified with fm-x-mode and fm-public-followup suites, Bash syntax checks, diff checks, and bin/fm-lint.sh * no-mistakes(ci): Fixed the guarded link-clear race by refusing an unlocked absence decision when a publisher already owns the metadata lock in a non-writable directory. Added a behavioral concurrency regression test. Verified with fm-x-mode, fm-public-followup, syntax/diff checks, and fm-lint * no-mistakes(ci): Fixed the guarded-clear race by refusing all guarded clears when the metadata parent is non-writable, including apparent link absence. Added a behavioral regression with a publisher waiting to create the lock, updated remote-retirement expectations and verification docs. Passed fm-x-mode, fm-public-followup, fm-lint, documentation audience, Bash syntax, and diff checks * fix(relay): bound the guarded remote link clear so it refuses instead of hanging The guarded clear checks that the remote state directory is writable before taking the metadata lock, but that check cannot close the window: the parent can turn non-writable between the check and lock creation, and a lock held by a live holder is indistinguishable from that at the acquire. `fm_lock_acquire_wait` is an unbounded `while ! try; do sleep 0.1; done`, so either case retried forever and `deliver` or `retire` wedged with nothing reported, instead of returning the retained-for-reconciliation refusal the guard exists to produce. This path runs unattended over the secondmate transport, where a wedge is worse than either outcome the guard defines. The guarded clear now acquires through `fm_lock_acquire_wait_bounded` (FMX_LINK_CLEAR_LOCK_TIMEOUT, default 10 seconds) and refuses on timeout through the existing failure path. Unguarded local callers keep the ordinary unbounded wait, so local behavior is unchanged. The bounded primitive's header no longer claims presentation-only scope, since this is a second authorized caller; nothing else in the shared lock infrastructure changed. The regression holds the metadata lock with a genuinely live process while leaving the state directory writable, so the refusal can only come from the bound and never from the writability precondition. Against the unbounded wait it does not terminate at all; with the bound it refuses, retains the registration, writes no receipt, and leaves the remote link untouched. * no-mistakes(review): Harden lock-timeout regression with independent deadline * no-mistakes(review): Restore no-op guarded clears on read-only state * no-mistakes(document): Clarify remote public-followup cleanup contract * fix(bin): support process events under symlinked homes (kunchenguid#3484) * fix(bin): resolve process-event state roots before validating them The process-event module validated the caller's spelling of a home's state root instead of the directory it operates on: it required the supplied path to equal its own lexical normalization, which rejects any path reached through a symlinked ancestor. On macOS both /tmp and $TMPDIR are symlinks, so an operator home under either could never claim a source. Reconcile still reported the runner started, while the detached runner died writing "cannot claim source" to the discarded stderr, and the source silently never fired. Resolve the state root to its physical directory once, then apply the existing private-directory validation to that resolved directory and derive every path, recorded claim identity, and later confinement check from it. This keeps the confinement contract for the directory actually operated on rather than only for callers that already spelled it physically, and removes the window where an ancestor symlink could be repointed between check and use. Homes already spelled physically behave identically. This was the single cause of both deterministic macOS failures in tests/fm-procevent.test.sh ("reconcile never claimed the registered source") and tests/fm-procevent-when.test.sh ("the winning concurrent arm did not produce an outcome"). The new case pins the behavior with an explicit symlinked-ancestor home, so it fails without the fix on any platform rather than only where the temp root happens to be a symlink. * fix(bin): pin the external capture staging boundary to its physical path The extension capture path pinned its registry staging boundary by comparing `pwd -P` against the caller-spelled registry directory, so a home reached through a symlinked ancestor still refused to start an extension-backed source after the state root itself resolved correctly. That left such a home half working: built-in sources ran while external ones failed. The staging preparer now prints the physical registry directory it validated, matching the inbox and reservation preparers beside it, and the start path pins on that returned path. The new end-to-end case drives the shipped file-signal package from a symlinked home spelling. * no-mistakes(review): Propagate canonical process-event state roots * no-mistakes(review): Propagate canonical state to process-event adapters * no-mistakes(document): Document physical process-event state roots * fix(pi): deliver captain outcomes as deterministic transcript entries (kunchenguid#3312) * fix(pi): persist captain outcomes visibly * no-mistakes(review): Recover captain outcomes after cold-start lock acquisition * no-mistakes(document): Document cold-start captain-outcome recovery * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes(review): Prove immediate Pi captain-outcome transcript delivery * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * fix(pi): process captain outcomes through a sequence-keyed turn PR kunchenguid#3312 made every captain-facing supervision outcome a durable, exact-once visible transcript entry with the read cursor advancing only after that entry exists. That is the display half of the delivery contract. Left alone it turns a probabilistic silent loss into a deterministic one: the captain sees an anchor line, and firstmate never acts, because nothing opens a turn and nothing records whether main ever processed the outcome. The 2026-08-31 timeline showed the two shapes this must survive on the previous hidden-turn path: seven delivered decision outcomes each answered by an empty assistant message (cursor advanced, no retry, unanswered for close to three hours), and two answered by an unrelated prior reply. Both happened because delivery advanced the cursor at enqueue and accepted whatever the next assistant message was. Add the processing half on top of the persistence half: - bin/fm-branch-outcome.sh keeps a processed marker separate from the read cursor (`unprocessed`, `mark-processed --through`, `processed-init`). It only advances through an explicit sequence-bound acknowledgement, never past the read cursor and never backwards; an absent marker reads as zero and `processed-init` migrates delivered history once so an upgraded home is not re-presented its past. - After the visible entry for a captain outcome exists, the extension hands every still-unprocessed captain row to main as one hidden, typed `fm-branch-process` request listing each `[seq N] task: summary`, opening exactly one main turn. Main closes it only by calling the new `fm_branch_processed` tool with the highest sequence listed. An unrelated, empty, or paraphrased answer leaves the sequence open, and the same request is presented again at the end of the next main run and at session start. The first two presentations of a sequence set open a turn of their own; after that the request rides the captain's next prompt so an ignored request cannot loop, and a session replacement resets that budget. Routine outcomes stay turn-free. - The regressions cover exactly those incident shapes against the real store scripts: an empty answer and an unrelated prior answer neither advance the marker nor stop re-presentation, the acknowledgement is refused beyond the read cursor and outside lock ownership, a partial acknowledgement keeps the newer sequence open, and kunchenguid#3312's own assertions now forbid an unkeyed turn rather than any turn. The store suite pins the marker's bounds and the migration; the real-SDK guard for appendEntry persistence and model exclusion is unchanged. Docs move the protocol from "no model turn" to "one sequence-keyed processing turn closed only by its acknowledgement", and the verification record carries the dated run against Pi 0.84.4. * no-mistakes(review): Harden outcome listing and sequence-bound acknowledgements * no-mistakes(review): Harden outcome state validation and request pacing * no-mistakes(review): Reject unsafe sidecars and unterminated outcome stores * no-mistakes(review): Validate canonical mark-read cursor state * no-mistakes(review): Guard cursor advancement against corrupt processed state * no-mistakes(review): Bind acknowledgements to active processing requests * no-mistakes(review): Reset pacing when processing sequence membership changes * no-mistakes(review): Enforce silent outcome invariants at storage boundary * no-mistakes(document): Document hardened captain outcome processing contracts --------- Co-authored-by: kunchenguid <kun@kunchenguid.com> * feat: add bounded concurrent Bearings ledger collection (kunchenguid#3481) * feat: bound Bearings remote ledger collection * no-mistakes(review): Clarify default remote-ledger collection behavior * no-mistakes(review): Detach reconcile delivery from watcher loop * no-mistakes(review): Enforce bounded snapshot and request captures * no-mistakes(review): Bound legacy summary capture before parsing * no-mistakes(review): Bound primary remote ledger captures * no-mistakes(document): Correct snapshot and reconcile documentation * no-mistakes(lint): Fix ShellCheck quoting in bounded collector * no-mistakes(ci): Fixed all three CI failures: updated the macOS Bearings assertion to 44 tests, made the home-summary test deterministic and aligned with default ledger consumption, and increased the asynchronous reconcile retirement wait for loaded CI. Verified both focused suites, all 44 Bearings tests, ShellCheck, actionlint, Bash parsing, and git diff checks * test: await reconcile request retirement * no-mistakes(review): Avoid empty reconcile queue process churn * no-mistakes(review): Read ledger summaries from immutable snapshots * no-mistakes(review): Reject multi-document home ledger streams * no-mistakes(review): Coalesce durable reconcile requests per target * no-mistakes(review): Unify reconcile keys and reject snapshot streams * no-mistakes(review): Key reconcile requests by stable target ID * no-mistakes(document): Document per-target reconcile request coalescing * no-mistakes(lint): Remove unused snapshot summary file variable * no-mistakes(ci): Adjusted the concurrent collector regression’s end-to-end timing ceiling to account for stock macOS process/jq overhead outside the three-second remote collection budget, while remaining below the 15-second serial-read floor. Verified with stock /bin/bash 3.2: all 44 Bearings tests pass; bash syntax and git diff checks pass * no-mistakes(ci): Fixed legacy summary validation to require exactly one top-level JSON document and added behavioral regression coverage. Stabilized CI by conditionally waiting longer for durable reconcile delivery and synchronously stopping the fm-on worker tree before fixture cleanup. Removed a redundant flaky healthy-path timing assertion; the wedged-reader test still proves concurrent bounded collection. Verified fm-bearings-snapshot, fm-secondmate-reconcile, and fm-on tests, plus project ShellCheck, bash syntax, and git diff checks * ci: rebalance portable serial test shards (kunchenguid#3489) * fix(ci): rebalance the portable serial shards on measured durations The "Behavior portable serial 3" shard ran 17-20 minutes against its 20-minute job cap and intermittently timed out seconds after a passing test, on branches and on main alike. Shards are packed longest-processing-time from per-script duration hints, and those hints were last measured on 2026-08-21 at 116 scripts. The lane has since grown to 139 scripts and from ~42 to ~63 minutes: 17 scripts had no hint at all and fell back to the 20 s default, and several existing hints were low by 2-5x (fm-watch-triage 142 s hinted vs 263 s measured, fm-public-followup 36 s vs 197 s). The partition therefore looked perfectly balanced in hint space, 734.6 s per shard, while really running 11.5, 13.6, 18.8 and 16.5 minutes. Script-count balance, which is what the tests asserted, stayed normal throughout and hid it. Refresh the hints from the timing artifacts of three green runs, taking the slowest measurement of each script so the balance holds on a slow runner, and split the lane across five shards instead of four. Replayed against those runs' real per-script durations the worst shard is now 12.54 minutes, 63% of the unchanged 20-minute cap, and the serial lane's wall clock drops from ~20 to ~12.5 minutes. Bound the drift that caused this rather than relying on the hints being refreshed by hand: the coverage guard now reports the unmeasured share as serial_unhinted= and refuses past PORTABLE_SERIAL_MAX_UNHINTED_PERCENT, which leaves room for newly added tests while making a stale table fail the guard instead of silently pushing one shard into its cap. No test changes what it asserts and no test stops running; only the partition across shards changes. * no-mistakes(document): Clarify conservative shard timing aggregate * fix(pi): fall back on incomplete supervision branch prompts (kunchenguid#3491) * fix(pi): fall back after settled branch errors * no-mistakes(review): Detect provider errors across prompt compaction * no-mistakes(review): Preserve in-flight branch state across selection changes * fix(pi): re-probe supervision branch after cooldown (kunchenguid#3497) * fix(pi): recover supervision branch after cooldown * no-mistakes(review): Defer branch recovery until prompt settlement * no-mistakes(document): Clarify supervision cooldown recovery contract * fix(bin): remove legacy remote snapshot reads (kunchenguid#3501) * refactor: remove legacy remote summary reads * no-mistakes(document): Document ledger-only snapshot reads * no-mistakes(ci): Fixed the snapshot test fixture so ledger refreshes use the same fake executable PATH as the snapshot consumer. This preserves observable endpoint freshness after removing legacy summary computation. Verified stock Bash parsing and all 44 Bearings tests pass under /bin/bash; git diff checks pass * no-mistakes(ci): Fixed the CI-only snapshot fixture failure by ensuring the bounded-ledger refresh uses its fake tmux backend. This removes host tmux availability as a source of nondeterminism. Verified all 44 Bearings tests pass, Bash syntax passes, and git diff checks are clean * no-mistakes(ci): Fixed CI nondeterminism in the Bearings fixture: all local ledger refreshes now use the fixture’s fake tmux backend when available, instead of depending on host tmux state. Verified stock /bin/bash syntax, git diff checks, and all 44 Bearings tests with a deliberately failing host tmux * fix(pi): preserve watcher continuity across session replacement (kunchenguid#3498) * fix(pi): rearm watcher after session replacement * no-mistakes(review): Queue actionable closes across Pi session replacement * no-mistakes(review): Stop replacement arm when handoff persistence fails * no-mistakes(review): Preserve actionable wakes through branch and late child races * no-mistakes(review): Surface late handoff failures without crashing Pi * no-mistakes(review): Coordinate replacement delivery settlement and unique handoff tokens * no-mistakes(review): Retry stale deliveries and release settled claims * no-mistakes(review): Distinguish branch settlement and retry handoff cleanup * no-mistakes(review): Deduplicate persistent handoff cleanup alerts * no-mistakes(review): Acknowledge watcher follow-ups only when consumed * no-mistakes(review): Persist idle follow-ups until agent consumption * no-mistakes(review): Preserve pending outcomes when handoff persistence fails * no-mistakes(review): Arm replacement before awaiting prior delivery settlement * no-mistakes(review): Adopt pending handoffs after lock reclamation * no-mistakes(review): Prevent stale generations from adopting replacement handoffs * no-mistakes(review): Scope replacement handoffs by watcher state * no-mistakes(document): Clarify replacement handoff documentation * no-mistakes(ci): Fixed the failing branch-extension tests to model the new settlement-promise contract. Failure cases now assert that delivery ownership returns to the watcher instead of expecting direct extension fallback. Verified the updated branch suite, Pi watcher suite, shell syntax, and diff checks * no-mistakes(review): Update branch settlement tests and preserve chunked outcomes * no-mistakes(document): Document watcher-owned replacement handoffs * no-mistakes(document): Verify replacement handoff documentation * test(pi): cover watcher-owned branch fallback * no-mistakes(document): Refresh watcher-owned fallback documentation * fix(bin): resurface task statuses missed by wake handling (kunchenguid#3495) * fix(bin): resurface terminal statuses lost after branch handling * test(watch): canonicalize process-event fixture homes * no-mistakes(review): Index branch outcomes by causal status position * no-mistakes(review): Recover outcome indexes and deduplicate resurfaced statuses * no-mistakes(review): Handle legacy ambiguity and oversized status diagnostics * no-mistakes(review): Keep unclassifiable oversized statuses silent * no-mistakes(document): Document lost-wake outcome backstop * no-mistakes(document): Update outcome backstop documentation * no-mistakes(ci): Fixed CI regressions in wake-drain: parseable reserved-key decisions can no longer bypass the durable decision-fold guard, and status output is prepared and receipt-committed before presentation to prevent repeated one-shot outcomes after later failures. Added a behavioral regression for receipt commit failure and retry. Targeted backstop, correlation-token, decision-cursor, open-decision, unread-status, syntax, and diff checks pass locally. Shard-4 failures appeared unrelated/flaky; the network-parallel test passed locally * no-mistakes(ci): Fixed the Greptile P1 data-loss issue by committing presentation receipts only after prepared output reaches stdout. Added behavioral coverage proving output failure leaves the backstop retryable and receipt failure may duplicate but never lose a presentation. Relevant wake-drain suites and syntax/diff checks pass. The shard-4 Pi extension failure is unrelated to this PR and did not warrant changes * no-mistakes(ci): Stabilized tests/fm-bootstrap-network-parallel.test.sh by replacing scheduler-sensitive equal-sleep timing with bounded synchronization between mocked fetch and remote probes. This preserves detection of real serialization while avoiding false failures under CI load. Verified with five consecutive test runs, bash syntax validation, ShellCheck, and git diff checks. The separate Pi stock-rendering failure reproduces locally but is unrelated environment/version drift * no-mistakes(ci): Fixed Behavior portable serial 4 by adding fm-classify-lib.sh and fm-timeout-lib.sh to the broken-root Pi test fixture; fm-branch-outcome.sh now depends on them. Verified the full Pi branch-extension suite with real-Pi checks skipped, the wake-drain outcome-backstop suite, Bash syntax, and git diff checks. Greptile findings are already addressed at HEAD; the no-mistakes attestation failure is external head-SHA state * fix(bin): collect follow-up results from remote work homes (kunchenguid#3503) * fix(bin): deliver typed terminal results from remote work homes A public commitment whose work is bound to a REMOTE secondmate home could never receive its typed terminal result. `fm-public-followup.sh brief` printed an emit command carrying this home's own absolute path and this checkout's own script path, neither of which exists on the machine the worker runs on, so the worker had nothing it could write to that the owning home would ever read - and `consume` kept finding nothing while the promise stayed open. The brief is now route-aware: for a remote work home it prints that route's own code root and home with `--stage-in`, so the typed event is staged in the home where the work actually runs, and the closing paragraph names the owning home as the one on the other machine instead of pointing at the path above it. The owning home collects those staged results over the same SSH route it reaches that secondmate on, because the transport only runs outbound: `consume` pulls them into its own inbox and reconciles them exactly as it reconciles a local report. Collection is non-destructive until the result is durably held, so a dropped connection cannot lose a terminal result, and a route that could not be reached is named in `consume`'s output with the promise left open rather than reported as an empty inbox. A local work home is untouched: the brief still prints `--home` with this home and this checkout's script, and the event still lands directly in this home's typed terminal-result inbox. This is the emit-side counterpart of the retire/clear fix in kunchenguid#3479 and reuses the remote-route resolution that landed with it. Reconciling a loop bound to a remote route now reaches that route, so the existing remote cases drive `consume` through the same faked transport their other steps already use. * no-mistakes(review): Fail loudly on unresolved routes and invalid staging homes * no-mistakes(review): Fail collection when remote outbox is unreadable * no-mistakes(review): Surface reassigned remote routes during empty collection * no-mistakes(review): Fail remote collection on invalid registrations * no-mistakes(review): Reject unsafe registration entries during remote collection * no-mistakes(review): Restore healthy empty remote collection behavior * no-mistakes(review): Skip remote collection for delivered registrations * no-mistakes(review): Skip delivered registrations before route validation * no-mistakes(document): Document remote follow-up collection semantics * fix(bin): exclude secondmates from home-summary validity (kunchenguid#3504) * fix(bin): exclude secondmates from home-summary child inventory kind=secondmate meta records never have backlog rows, so counting them in unowned_children or terminal_in_flight made a clean main home look invalid once earlier ledger checks passed. * no-mistakes(review): Cover terminal secondmate in-flight exclusion * no-mistakes(ci): Updated the stock macOS Bash CI snapshot expectation from 15 to 16 tests. Verified all 16 snapshot/fleet-view tests pass under Bash 3.2.57 and `git diff --check` succeeds * fix(bin): self-heal outcome indexes on first drain (kunchenguid#3509) * fix(bin): self-heal status-outcome indexes on every drain Missing ready markers were skipping the lost-wake backstop on non-Pi homes because only the Pi branch ran processed-init. Drain now rebuilds those indexes under the outcome lock and fails closed only on a real store fault. * no-mistakes(review): Guard held-lock initialization and fail marker writes * no-mistakes(document): Document cross-harness outcome-index self-healing * fix(bearings): keep active children underway during captain holds (kunchenguid#3505) * fix(bearings): keep active children underway beside a captain hold Project each readable home's active children into Underway independently of the home-level captain-decision classification so a hold no longer hides live work. * no-mistakes(review): Preserve Underway repos and disclose child truncation * no-mistakes(review): Fall back to task project for Underway repos * no-mistakes(ci): Updated the stock macOS Bash CI assertion from 44 to 45 Bearings tests, matching the newly added behavioral regression. Verified all 45 tests pass under /bin/bash, Bash syntax checks pass, and git diff validation is clean * fix(pi): settle watcher delivery on Pi accepting the follow-up (kunchenguid#3513) * fix(pi): settle watcher delivery on Pi accepting the follow-up A follow-up queued while main is streaming joins the running run without ever raising before_agent_start, so waiting on that event before clearing the successor pipeline (kunchenguid#3498) stalled every later actionable close: no successor started, no wake was delivered or offered to the branch, and the turn-end guard woke main to re-arm by hand after every close. The pipeline now settles once Pi accepts the follow-up. Consumption is observed at before_agent_start for an idle main and at the user message_start for a streaming main, and decides only what a replacement session (/new, /resume, /fork, reload) replays. An exhausted restoration delivers its typed failure without launching an arm past the retry bound, which the stall had hidden. The replacement-coordinator map is typed so the strict no-emit typecheck passes again. Tests: the doubles no longer raise before_agent_start for a streaming send, a portable regression drives two actionable closes while main streams and proves the successor chain plus consumption-scoped replay, and a credential-free real-SDK probe pins Pi's event contract for both the streaming and the idle follow-up. Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a * fix(pi): retry a verified successor that fails during wake delivery A verified successor can exit while the wake it was started for is still being delivered, most plausibly during a branch turn that holds the settlement for minutes. Its failure close arrived while the pipeline's single-flight guard was set, so the close handler skipped the retry, and the pipeline's end no longer launched an arm, which left the live generation with no watcher and no retry timer. The close handler now records that failure when the child had reported readiness and was not retired by the restoration itself, and the pipeline runs the ordinary bounded, lock-checked retry for it once the delivery settles. A restoration started for a later pending supersedes it, and an exhausted restoration still hands repair to main without a further arm. The regression holds a branch settlement open while the verified successor exits with a failure and proves one retry watcher starts after the settlement releases, none while it is held. Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a * fix(bin): bound repeat stale wakes for parked workers (kunchenguid#3532) * fix(bin): bound repeat stale wakes for a parked but live worker A worker parked on a declared wait - `paused:` for an external or pipeline wait, or a verified `captain-held` transfer - kept waking firstmate far inside FM_PAUSE_RESURFACE_SECS. Observed as five consecutive alarms on one captain-held worker and dozens across a day on a pipeline wait, and reported upstream as four wakes in 75 minutes against a 3600s window. pause_state_class deliberately answers `none` for a still-live agent even under a declared wait, so a worker genuinely waiting on a decision is never silenced. That classification is correct and is left alone; it routes every parked but live worker through surface_nonterminal_stale on first sight of each distinct stale hash, and an idle parked pane still churns its hash on a clock or a token counter without changing what is being waited on. Two places let that churn re-alarm: - surface_nonterminal_stale queued the wake BEFORE consulting whether a wait was declared, then wrote `.paused-resurfaced-<key>` - the very throttle that should have suppressed it. The throttle was never read on this path and was advanced by the wake it should have prevented. - The hash-change path cleared that throttle through clear_pause_tracking whenever the classification came back `none`, so each tick also bought the same declared wait a fresh window. Fixing only the first site changes nothing. Read the throttle before anything is queued and advance it only on a wake that really fires, and on the hash-change path reset only the per-hash bookkeeping while the declaration still stands, via a clear_stale_hash_tracking split so neither half of clear_pause_tracking is duplicated. The throttle is keyed to the declaration, not to the pane. First sight still wakes, so an inconclusive state is still inspected, and the window's end still re-surfaces once, so a forgotten wait cannot rot invisibly - noise traded for a bounded cadence, never for silence. The wake identity stays the plain `stale: <win>` the away-mode handoff depends on. Tests cover both observed forms and were confirmed to fail against three deliberate breaks: each site reverted on its own, and a re-surface that never fires again. * fix(document): Clarify declared-wait wake cadence documentation * fix(ci): Captain, fixed the stale-throttle inheritance: cadence markers now bind to the current wait declaration, so replacement paused and captain-held waits each emit their first plain `stale:` wake. Added behavioral coverage for both forms. Bite proof failed as expected when identity matching was removed, then passed after restoration. Full watcher triage suite, `bin/fm-lint.sh`, syntax checks, and diff checks pass. Changes remain uncommitted for the outer executor * fix(ci): Captain, fixed the confirmed Greptile finding. `resurface_absorbed` now applies a throttle only when its stored declaration scope matches the current wait, so replacement `paused:` and `captain-held` waits surface immediately without changing classification. Added executable coverage for both absorbed forms. Bite proof failed before the fix at the intended assertion; afterward the full watcher triage suite, `bin/fm-lint.sh`, shell syntax checks, and `git diff --check` passed * fix(bin): accept the away-mode daemon as the turn-end supervision owner (kunchenguid#3567) * fix(turnend): accept the away-mode daemon as the supervision owner While state/.afk exists the away-mode daemon owns supervision and runs bin/fm-watch.sh one-shot: the watcher exits on every wake and the daemon starts its replacement. The turn-end guard tested for a live watcher process holding the watch lock at that instant, so a turn boundary that landed in the hand-off blocked with "TURN WOULD END BLIND" while supervision was completely healthy, costing a full handling turn each time. Reproduced with the real daemon wrapping the real watcher and the real guard sampling the same home: 6 of 40 samples blocked, every one of them with the daemon alive and the beacon 2-3 seconds old, and a new watcher pid on each cycle. After the fix the same reproduction blocks 0 of 40, and killing the daemon and its watcher (away mode still on, beacon still fresh) blocks again. The guard now accepts a live, identity-matched daemon holding this home as proof of supervision while away mode is active. The identity match is the same discipline the watcher lock uses, so a recycled pid or a lock left by a killed daemon proves nothing. The fresh-beacon half of the predicate is unchanged: a daemon that stops restarting its watcher still blocks once the beacon passes grace, a home with no supervisor blocks exactly as before, and with away mode off the strict watcher predicate is untouched. The predicate reads only durable state, so it behaves identically for every primary harness and runtime backend. * no-mistakes(document): clarify away-mode daemon supervision proof and test coverage * no-mistakes(document): generalize stale turn-end predicate summary in architecture.md * fix(backlog): omit --file from row probes for non-markdown backends (kunchenguid#3582) * fix(backlog): omit markdown file for beads probes * no-mistakes(document): Narrow backlog addressing doc to mutations for backend-aware probes * no-mistakes(ci): Fixed the Greptile P2 review comment (the only failing check) on tests/fm-backlog-atomicity.test.sh. The comment correctly noted that an exported TASKS_AXI_BACKEND environment variable would inherit into the spawned scripts and, because fm_tasks_axi_backend gives it top precedence, override each test case's .tasks.toml backend fixture — making the backend-specific argv assertions fail for environmental reasons. Fix: unset TASKS_AXI_BACKEND in the test harness right after sourcing tests/lib.sh, with a comment explaining why, so every case deterministically exercises its declared backend (4 lines added; no production code touched). Verified: reproduced the leak before the fix (TASKS_AXI_BACKEND=beads made the markdown dispatch case fail with 'beads show failed', exactly the reported failure mode); after the fix the full suite passes (0 failures, exit 0) both with and without TASKS_AXI_BACKEND=beads exported. The added lines are shellcheck-clean (the only shellcheck note, SC1091 on the lib.sh source line, pre-exists this change) * fix(bin): classify progress updates on requested work as routine (kunchenguid#3589) The supervision branch's verdict rule escalated every outcome that answered a captain request, so "the work started" and "still working" notes reached the captain with nothing to look at. The rule now keeps a finished result of requested work captain-facing, even when healthy, and treats start or still-working updates that bring no new artifact, finding, or decision as routine. The captain list for review-ready PRs, ask-user findings, exhausted blockers, credentials, and destructive or security-sensitive cases is unchanged, as are the unsolicited-routine, silent-fleet-review, and doubt-chooses-captain rules. The fm_branch_report tool description and the two docs that restated the old unconditional rule now point at the prompt's "Verdict: routine or captain" section as the one owner instead of carrying a second copy. * fix(bin): preserve captain calls during teardown (kunchenguid#3595) * fix(bin): never close a captain call during cleanup A scout that held its own work item for the captain, which is what captain-hold-lifecycle prefers ("hold the work item the question gates"), was closed by bin/fm-teardown.sh's automatic backlog transition. The completion gate passed, cleanup ran, and the captain's question moved to Done with no recorded answer: the one thing the policy says must never happen. `tasks-axi done` closes a held row silently, and nothing in teardown asked whether the row was the captain's own call. bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when the task is still an open captain call, 1 when it is not, 2 when that cannot be established. It reads the row through the transition library's backend-aware probe, so it addresses the same backlog teardown does; the script's other commands now address the configured data directory the same way instead of FM_HOME, which also fixes captain holds in a home with a relocated data directory. Teardown asks `open` before any destructive step and refuses on 2. On 0 only the close changes: after cleanup and still under the task's own lock, the row gets one "Deliverable of the finished work" line at the end of its body and returns to Queued through `tasks-axi reopen`, keeping its hold, so it lands in Captain's Call instead of reading as work under way. --force does not lift this: it authorizes discarding unlanded work, never the captain's question. The deliverable goes into the body because `tasks-axi update --report` rewrites the title of a row that is not Done. The crash window reuses the pending-close record teardown already stages: a `mode=retain` line makes the existing replay record the deliverable and reopen instead of closing, with the same validator, stale-generation check, cleanup-incomplete marking, and non-blocking bootstrap lock as an ordinary close. A retained row the captain answered first simply retires the record. No parallel record type, recovery command, or second bootstrap loop is introduced. Regressions run the real executables: the captain-held scout survives cleanup queued, held, with its deliverable and on the board, only `answer` closes it, --force keeps it open, and an ordinary scout still closes with its report; an interrupted cleanup leaves the row untouched and the next session start retains it; a relocated backlog keeps the retention in its one configured file; and a ship row whose hold cannot be read refuses cleanup before anything destructive. Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np * no-mistakes(review): Serialize captain holds and fix backend-aware listing * no-mistakes(document): Update captain-call retention documentation * no-mistakes(document): Fix relocated captain-hold backlog diagnostics * fix(bin): deliver secondmate outcomes to the parent channel (kunchenguid#3592) * fix(bin): deliver every secondmate outcome on the parent channel from the recording scripts A secondmate's captain-facing outcomes could miss: the mate model addressed the captain in its own unread chat instead of appending to the parent channel, and a PR-ready report, a finding, a decision, a blocker, and a failure all depended on that one remembered append. Make delivery structural, so the parent channel never depends on the model: - bin/fm-parent-channel-lib.sh is the one owner of channel resolution and exact-line append-once; the merge outcome path and the inactive-outcome scan now publish through it instead of two private copies. - bin/fm-inactive-reconcile.sh gains a ledger-first path that runs on every watcher poll in a secondmate home: a direct child's whole terminal done or failed line is delivered at once with its note, recorded PR, mode, merge posture, and scout report pointer, keyed and receipted so it is delivered once, and the inactive path yields to it. `report <task-id>` runs the same delivery for a caller holding the child's meta lock. - bin/fm-pr-check.sh publishes the PR-ready line with the canonical URL at registration. - bin/fm-captain-hold.sh publishes a hold and its answer, keyed by task id and resolution-record count, with no new persisted state. - bin/fm-teardown.sh delivers the child's final line before removing its record and refuses, retaining every record, while the channel cannot be written. - The charter opens with the parent-channel rule and confines the mate's own appends to judgement; AGENTS.md carries the carve-out at the persona address rule and the escalation list. docs/secondmate-parent-channel.md records the design and its coverage, and docs/verification/secondmate-parent-channel.md records the live run with real tmux panes and both real watchers delivering every line with no model. Supersedes kunchenguid#3569. * no-mistakes(review): Fix parent outcome retries and reconciliation locking * no-mistakes(review): Prevent busy children from starving ledger delivery * no-mistakes(review): Correct ledger metadata and hold occurrence handling * no-mistakes(review): Disambiguate ledger outcomes and normalize hold reasons * no-mistakes(review): Close ledger races and preserve teardown records * no-mistakes(document): Correct parent-channel receipt and scanner documentation * no-mistakes(lint): Quote done arguments for ShellCheck compliance * no-mistakes(ci): Fixed both CI failures. Updated GOTMP teardown fixtures for the new final-outcome reporter and isolated them from host tmux state. Updated the PR security assertion to distinguish the accepted PR-ready line from duplicate merge outcomes. Verified with both failing test suites, bash syntax checks, and git diff checks * no-mistakes(ci): Fixed Greptile’s duplicate-delivery race in bin/fm-inactive-reconcile.sh. Ledger events now claim matching already-delivered inactive receipts using the prior status fingerprint, preventing duplicate parent reports while preserving later same-state completions. Added behavioral regression coverage. Verified inactive-reconcile tests, project lint, documentation audience checks, syntax, and diff checks. Teardown tests passed relevant cases before the documented pre-existing herdr-preflight-missing-adapter failure * fix(bin): sync remote second mates to primary commit (kunchenguid#3599) * fix(bin): sync remote second-mate homes to the parent primary commit Session start and remote launch pointed a remote second-mate home at whatever Firstmate copy its own host kept, so a home that had already advanced past that copy refused as a non-fast-forward and every other home stopped at the host's older commit while the primary ran ahead. The parent now resolves ITS primary default-branch commit with the existing helper and hands that commit to the host on both paths. Because a remote home is a standalone clone, the host imports that one commit before advancing - already present, else from that host's Firstmate copy without moving it, else from the home's own origin - and then runs the SAME ff_target guards a local home gets, so dirty, diverged, feature-branch, and unresolvable targets skip untouched and the ancestry rules keep one owner. An unimportable target now names /updatefirstmate instead of failing opaquely, and a host still running an older Firstmate copy is reported the same way rather than echoing a bare refusal. The host-local launch leg no longer re-runs its own secondmate sync, so the spawn it drives cannot re-target that host's copy after the parent has already converged the home. /updatefirstmate is unchanged: it still refreshes the remote code root from that host's origin and then syncs the home to that refreshed copy, which is what the sync call with no target commit means. * no-mistakes(document): Document primary-targeted remote secondmate synchronization * fix(bin): separate captain intent from firstmate specs (kunchenguid#3597) * fix(bin): split brief task into captain intent and firstmate spec Keep no-mistakes --intent as the captain's ask plus later captain words, not the build spec or worker tradeoffs. * fix(bin): stop task-subsection copies at the next heading Promotion was swallowing the scout Setup contract into Firstmate spec, and pre-subsection briefs lost their # Task body. * no-mistakes(review): Validate brief content and preserve nested specifications * no-mistakes(review): Scope placeholder validation to scaffold-only subsection bodies * no-mistakes(review): Ignore fenced subsection headings during brief validation * no-mistakes(review): Preserve captain intent across scout promotion * no-mistakes(review): Enforce safe intent boundaries for legacy promotions * no-mistakes(review): Allow marked legacy intent and reject empty promotions * no-mistakes(review): Scope task parsing and overlay legacy intent contracts * no-mistakes(review): Overlay current intent contract for all no-mistakes spawns * no-mistakes(review): Preserve later captain clarifications in intent overlays * no-mistakes(document): Document brief intent enforcement and ownership * no-mistakes(ci): Updated spawn-related test fixtures to use valid Captain intent and Firstmate spec subsections, corrected launch-path expectations to launch-brief.md, and resolved ShellCheck quoting findings. Verified with fm-lint.sh and 15 affected behavior tests, including real Herdr tests; all passed * no-mistakes(ci): Updated stale spawn/promotion fixtures in the Muse, Orca, secondmate-harness, and public-followup suites to provide valid Captain's intent and Firstmate spec subsections. Verified full Orca and secondmate-harness suites, targeted public-followup promotion behavior, Bash syntax, diff checks, and fm-lint * fix: start a fresh supervision branch for every main session (kunchenguid#3600) * fix(pi): start a new supervision branch conversation per main session The supervision branch reopened one recorded conversation forever, so every main session start reloaded the current generated prompt and then weeks of accumulated thread, where a superseded rule could still outweigh today's. The branch conversation is now scoped to one main session: the session generation owns the recorded conversation, so a cold start, /new, /resume, /fork, or a reload always builds a new one, while a rebuild inside one session (a model or effort change) still continues that session's own conversation. The dialog mirror re-anchors with it. Its durable cursor records what the previous branch conversation received, so a /resume or reload - which keeps main's own session file - would otherwise leave the new branch blind to dialog main itself still has. The reset is bounded by the current main session, and the cursor keeps advancing incrementally within it. The durable outcome store and its processed marker are untouched, so unacknowledged captain-facing outcomes still re-present on the new main session. * no-mistakes(document): Document fresh Pi supervision conversations * no-mistakes(ci): Fixed the flaky concurrent inbox failure. Lock acquisition now retries when a competing lock disappears between a failed claim and inspection. Added a behavioral regression covering that race. Verified the full inbox test four times, project lint, and git diff checks * no-mistakes(document): Add process-group isolation and per-script timeout to fm-test-run.sh docs --------- Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Co-authored-by: FocalFactotum <305704917+FocalFactotum@users.noreply.github.com> Co-authored-by: kunchenguid <kun@kunchenguid.com> Co-authored-by: Mickaël Rémond <mremond@process-one.net> Co-authored-by: Joel Le <143022894+krakns@users.noreply.github.com> Co-authored-by: Jon Roosevelt <rooseveltadvisors@gmail.com>
RooseveltAdvisors
pushed a commit
to RooseveltAdvisors/firstmate
that referenced
this pull request
Sep 3, 2026
* fix(bin): never close a captain call during cleanup
A scout that held its own work item for the captain, which is what
captain-hold-lifecycle prefers ("hold the work item the question gates"),
was closed by bin/fm-teardown.sh's automatic backlog transition. The
completion gate passed, cleanup ran, and the captain's question moved to
Done with no recorded answer: the one thing the policy says must never
happen. `tasks-axi done` closes a held row silently, and nothing in
teardown asked whether the row was the captain's own call.
bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when
the task is still an open captain call, 1 when it is not, 2 when that
cannot be established. It reads the row through the transition library's
backend-aware probe, so it addresses the same backlog teardown does; the
script's other commands now address the configured data directory the
same way instead of FM_HOME, which also fixes captain holds in a home
with a relocated data directory.
Teardown asks `open` before any destructive step and refuses on 2. On 0
only the close changes: after cleanup and still under the task's own
lock, the row gets one "Deliverable of the finished work" line at the end
of its body and returns to Queued through `tasks-axi reopen`, keeping its
hold, so it lands in Captain's Call instead of reading as work under way.
--force does not lift this: it authorizes discarding unlanded work, never
the captain's question. The deliverable goes into the body because
`tasks-axi update --report` rewrites the title of a row that is not Done.
The crash window reuses the pending-close record teardown already stages:
a `mode=retain` line makes the existing replay record the deliverable and
reopen instead of closing, with the same validator, stale-generation
check, cleanup-incomplete marking, and non-blocking bootstrap lock as an
ordinary close. A retained row the captain answered first simply retires
the record. No parallel record type, recovery command, or second bootstrap
loop is introduced.
Regressions run the real executables: the captain-held scout survives
cleanup queued, held, with its deliverable and on the board, only
`answer` closes it, --force keeps it open, and an ordinary scout still
closes with its report; an interrupted cleanup leaves the row untouched
and the next session start retains it; a relocated backlog keeps the
retention in its one configured file; and a ship row whose hold cannot be
read refuses cleanup before anything destructive.
Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np
* no-mistakes(review): Serialize captain holds and fix backend-aware listing
* no-mistakes(document): Update captain-call retention documentation
* no-mistakes(document): Fix relocated captain-hold backlog diagnostics
AgardnerAU
added a commit
to AgardnerAU/firstmate
that referenced
this pull request
Sep 4, 2026
* refactor(quota): extract mid-task polling and candidate selection into dedicated scripts (#3221)
* Add quota exhaustion detection and safe fallback helpers
- bin/fm-procevent-quota.sh: generic procevent adapter that arms a
recurring quota-axi --json poll and wakes firstmate when a tracked
provider's effectivePercentRemaining drops below a threshold or its
runway.status becomes exhausted_now.
- bin/fm-quota-choose.sh: worker-side helper that picks the first ranked
harness:model candidate with positive effectivePercentRemaining.
- AGENTS.md and .agents/skills/quota-array-dispatch/SKILL.md: document
the new helpers and the mid-task quota-exhaustion wake path.
- tests/fm-quota-choose.test.sh: unit tests with a mocked quota-axi JSON
source.
* no-mistakes(review): Fix quota polling and scope bounds
* no-mistakes(review): Enforce safe default quota selection
* no-mistakes(review): Handle decimal quota values safely
* no-mistakes(review): Fail closed on invalid quota inputs
* no-mistakes(review): Reject empty quota candidate segments
* no-mistakes(review): Harden quota parsing and timeout ownership
* no-mistakes(review): Reuse captured quota snapshots consistently
* no-mistakes(review): Match quota using explicit candidate providers
* no-mistakes(review): Centralize fail-closed quota schema validation
* no-mistakes(review): Reject out-of-range quota percentages
* no-mistakes(review): Validate quota runway status enum
* no-mistakes(review): Tighten quota scope and status contracts
* no-mistakes(review): Preserve unknown quota and exact product bounds
* no-mistakes(review): Preserve provider-level unknown quota
* no-mistakes(review): Reuse canonical verified harness validation
* no-mistakes(document): Document mid-task quota handling
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* fix(docs): restore default routing contract, keep quota helper optional
Restore the AGENTS.md section 4 always-loaded routing paragraph the PR
had deleted, so the standing TOON-first intake, spendPriority ranker,
every-candidate accounting, and load-trigger contract stay exactly as
before this PR. The mid-task quota wake is optional and must not alter
default routing.
Restore the quota-array-dispatch skill ownership line to section 4 as
the always-loaded intake boundary owner; keep the worker-side helper
section as an addition only, without rewiring ownership or load
triggers to section 13.
* fix(bin): use harness-keyed quota matching in optional helper
Revert fm-quota-choose.sh from harness:provider:model tuples back to
harness:model candidates with harness-keyed provider matching, per the
resolved ask-user finding. The helper is optional; authoritative
multi-provider routing (provider discovery from the harness catalog and
quota matching by that explicit provider) stays owned by AGENTS.md
section 4 and the quota-array-dispatch skill intake procedure, not the
helper.
Document the multi-provider limitation in the helper header and the
quota-array-dispatch skill: the helper maps each harness to one primary
provider family only, so a candidate whose established provider differs
from that primary family is checked against the wrong quota row. Use it
only when the brief fixed the candidate order and every candidate's
provider is the harness's primary family.
The helper still consumes one already-captured default-TOON or JSON
snapshot via stdin or --snapshot and never calls quota-axi itself, so
it selects from the same quota state as the intake.
* no-mistakes(review): Fix Muse quota mapping and helper contract docs
* no-mistakes(review): Reject known-empty quotas and map quota tests explicitly
* no-mistakes(review): Preserve unmeasured candidates and enforce snapshot reuse
* no-mistakes(review): Fix quota retirement and dependent regression coverage
* no-mistakes(review): Accept zero-row quota TOON snapshots
* no-mistakes(review): Enforce quota semantics status consistency
* no-mistakes(review): Veto dispatch on any exhausted applicable scope
* no-mistakes(review): Record exhausted quota scope in wake details
* no-mistakes(review): Fix quota help and control dependency coverage
* no-mistakes(review): Decode quoted TOON fields and document quota wakes
* no-mistakes(review): Validate zero-row TOON and map timeout coverage
* no-mistakes(review): Reject multi-value JSON and malformed TOON envelopes
* no-mistakes(review): Validate complete nonzero TOON envelopes
* no-mistakes(review): Accept producer-shaped quota TOON envelopes
* no-mistakes(review): Support empty quota arrays and validate counted rows
* no-mistakes(review): Harden TOON completion, scopes, and quoted fields
* no-mistakes(review): Preserve unknown-headroom exhaustion and reject trailing fields
* no-mistakes(review): Allow unknown headroom under known semantics
* no-mistakes(review): Reject noncanonical quota identities
* no-mistakes(review): Preserve empty quota polling and validate attention identities
* no-mistakes(review): Reject noncanonical provider watches
* no-mistakes(review): Validate all candidates before quota selection
* no-mistakes(document): Correct quota helper safety documentation
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* fix: surface comments on Lavish annotations (#3371)
* fix(bin): keep typed Lavish comments when an element is also annotated
read preferred element text over prompt, so an annotate-and-comment
item dropped the captain's words. Surface prompt as its own field.
Co-authored-by: Cursor <cursoragent@cursor.com>
* no-mistakes(review): Filter non-comment prompts from Lavish reader output
* no-mistakes(document): Clarify Lavish comment presentation contract
* no-mistakes(ci): Fixed Lavish reader comment provenance: non-choice prompts are now emitted even when identical to element text. Added observable regression coverage for identical selector+comment input while retaining pure annotation/message coverage. Reader cases, bash syntax, and diff checks pass. Full fm-procevent suite stops earlier at unrelated “reconcile never claimed” setup failure
* no-mistakes(ci): Fixed duplicate pure-annotation prompts by emitting `prompt:` only when it differs from captured element text. Updated behavioral coverage for selector+comment, pure annotation, and pure message cases. Focused reader regressions, syntax checks, and diff checks pass. Full suite remains blocked by the pre-existing “reconcile never claimed the registered source” failure
* fix(bin): always emit Lavish comments and use real annotation fixtures
Stop inferring comment provenance from prompt==text. Real pure
annotations have no prompt, so always-emit does not duplicate.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: support first public-followup registration on Bash 3.2 (#3420)
* Fix public-followup register crashing on empty lock arrays under bash 3.2.
bash 3.2 with set -u treats "${arr[@]}" on an empty array as unbound, so the first register in a fresh home aborted before taking the registry lock.
The empty-lock regression also runs under the existing stock macOS Bash CI lane so pre-fix code would fail there.
* no-mistakes(document): Document stock Bash registration coverage
* no-mistakes(ci): Pinned the stock macOS Bash CI lane to tasks-axi@0.2.5, eliminating dependency drift. Verified workflow YAML parsing, git diff checks, and the focused regression under /bin/bash 3.2.57 with tasks-axi 0.2.5
* no-mistakes(ci): Fixed the flaky portable CI test: it treated exited zombie processes as live because `kill -0` succeeds for zombies. The watcher and descendant assertions now check process state and regard zombies as exited. Verified `tests/fm-pr-check-security.test.sh`, ShellCheck, `git diff --check`, and the focused Bash public-followup regression
* fix(bin): isolate new Herdr server environments (#2792)
* fix(herdr): isolate server launch environment
* no-mistakes(review): Clear inherited supervision model from Herdr launches
* no-mistakes(document): Document Herdr server launch environment isolation
* fix: surface inbound Relay media to responding agents (#3442)
* fix: surface inbound Relay attachments to the responding agent
A Discord support thread's screenshots were never seen by the agent
handling the mention. The relay delivered them and the poll stashed
them: the reporter's images arrived on the `thread_starter` entry of
`in_reply_to_chain` while the mention's own media list was empty. The
gap was in the responder's playbook, which enumerated a fixed field
list (`request_id`, `text`, `in_reply_to`, `in_reply_to_chain`) and so
made every other field, attachments included, invisible.
Fix it where the gap is, in prose:
- Read the complete payload object rather than a fixed field list, so
media and later relay fields are never skipped again.
- Fetch and view attached media with the agent's own tools, on the
mention and on every chain entry, and call out the common shape where
only the thread starter carries the screenshots.
- Restrict those fetches to known-good platform media hosts over https
(Discord: cdn.discordapp.com, media.discordapp.net,
images-ext-1.discordapp.net, images-ext-2.discordapp.net; X:
pbs.twimg.com, video.twimg.com), report a blocked host instead of
working around it, and treat everything fetched as untrusted public
input on the same terms as the surrounding thread text.
The poll stays out of it and downloads nothing, so no third-party bytes
are pulled on the polling path.
The new test pins the contract the playbook depends on: a mention in the
incident's shape, with an empty top-level media list and screenshots on
the thread starter, must reach the inbox with the payload intact and its
media URLs unfetched.
* no-mistakes(review): Preserve media authority and enforce poll-only fetching
* no-mistakes(document): Clarify Relay attachment safety prose
* fix(bin): defer inactive reconciliation during startup (#3480)
* Defer inactive startup reconciliation
* no-mistakes(review): Queue deferred inactive reconciliation diagnostics durably
* no-mistakes(review): Require worker phases to cover startup requests
* no-mistakes(review): Make diagnostic wakes safely acknowledgeable
* no-mistakes(document): Document deferred startup phase coverage
* fix(bin): bound wake drain presentation lock waits (#3475)
* fix: bound status presentation lock waits
* no-mistakes(review): Distinguish malformed presentation locks from live contention
* no-mistakes(review): Bound no-ack drain queue lock acquisition
* no-mistakes(document): Document bounded presentation-lock drain behavior
* no-mistakes(lint): Annotate bounded lock output global
* no-mistakes(ci): Added deterministic regression coverage for successful bounded-lock acquisition after live contention, verifying helper-to-caller PID ownership handoff and caller release. Verified with bash syntax checks, git diff checks, and the full fm-wake-queue test suite
* fix(bin): retire public follow-ups in remote homes (#3479)
* fix(relay): close a public loop whose work lives in a remote secondmate home
A public-followup loop bound to a REMOTE secondmate could never be closed.
`clear_public_followup_link` (bin/fm-public-followup.sh:701) required an
absolute recorded `work_home_path` for a `secondmate:*` work home, but a remote
route has no local path on this machine, so registration records that field
empty (bin/fm-public-followup.sh:291). Every close ran that clear first, so
`retire` died with "could not clear the legacy X link ... retained for
reconciliation" forever, and `deliver` posted the public reply and then stranded
the loop at `posted`. `--force` never covered that step.
The clear now goes to the remote home over that route's SSH transport, running
`fm-x-followup.sh --clear <work-id>` through `bin/fm-on.sh`. The route is decided
from `data/secondmates.md` before any local path is consulted, so a same-named
local directory can never stand in for a remote home, and registrations already
on disk retire without needing a new field. `fm-on.sh` passes ssh's status
through, so 255 stays the established "delivered but completion unknown" result
this codebase already reconciles: the close is refused, the registration and the
remote link are left exactly as they were, and the message names the unknown
completion instead of claiming a definite failure.
Local secondmate and `main` work homes are untouched, and `--force` still
governs only the unresolved-obligation refusal.
Three regression cases drive a remote route end to end, faking only the ssh
binary at the FM_SSH_BIN seam and then running the real remote entrypoint
against a local checkout, so the clear that must reach the remote home actually
happens there.
* no-mistakes(review): Guard remote link clears by request identity
* no-mistakes(review): Fail guarded clears on unreadable remote state
* no-mistakes(review): Reject guarded clears on non-writable remote state
* no-mistakes(review): Allow no-link retirement in non-writable remote state
* no-mistakes(document): Correct public-followup verification guarantee count
* no-mistakes(ci): Fixed the guarded link-clear race by ensuring absence is decided under the metadata lock whenever publication is possible. Added a behavioral concurrency regression test. Verified with fm-x-mode and fm-public-followup suites, Bash syntax checks, diff checks, and bin/fm-lint.sh
* no-mistakes(ci): Fixed the guarded link-clear race by refusing an unlocked absence decision when a publisher already owns the metadata lock in a non-writable directory. Added a behavioral concurrency regression test. Verified with fm-x-mode, fm-public-followup, syntax/diff checks, and fm-lint
* no-mistakes(ci): Fixed the guarded-clear race by refusing all guarded clears when the metadata parent is non-writable, including apparent link absence. Added a behavioral regression with a publisher waiting to create the lock, updated remote-retirement expectations and verification docs. Passed fm-x-mode, fm-public-followup, fm-lint, documentation audience, Bash syntax, and diff checks
* fix(relay): bound the guarded remote link clear so it refuses instead of hanging
The guarded clear checks that the remote state directory is writable before
taking the metadata lock, but that check cannot close the window: the parent can
turn non-writable between the check and lock creation, and a lock held by a live
holder is indistinguishable from that at the acquire. `fm_lock_acquire_wait` is
an unbounded `while ! try; do sleep 0.1; done`, so either case retried forever
and `deliver` or `retire` wedged with nothing reported, instead of returning the
retained-for-reconciliation refusal the guard exists to produce. This path runs
unattended over the secondmate transport, where a wedge is worse than either
outcome the guard defines.
The guarded clear now acquires through `fm_lock_acquire_wait_bounded`
(FMX_LINK_CLEAR_LOCK_TIMEOUT, default 10 seconds) and refuses on timeout through
the existing failure path. Unguarded local callers keep the ordinary unbounded
wait, so local behavior is unchanged.
The bounded primitive's header no longer claims presentation-only scope, since
this is a second authorized caller; nothing else in the shared lock
infrastructure changed.
The regression holds the metadata lock with a genuinely live process while
leaving the state directory writable, so the refusal can only come from the
bound and never from the writability precondition. Against the unbounded wait it
does not terminate at all; with the bound it refuses, retains the registration,
writes no receipt, and leaves the remote link untouched.
* no-mistakes(review): Harden lock-timeout regression with independent deadline
* no-mistakes(review): Restore no-op guarded clears on read-only state
* no-mistakes(document): Clarify remote public-followup cleanup contract
* fix(bin): support process events under symlinked homes (#3484)
* fix(bin): resolve process-event state roots before validating them
The process-event module validated the caller's spelling of a home's state
root instead of the directory it operates on: it required the supplied path
to equal its own lexical normalization, which rejects any path reached
through a symlinked ancestor. On macOS both /tmp and $TMPDIR are symlinks,
so an operator home under either could never claim a source. Reconcile still
reported the runner started, while the detached runner died writing "cannot
claim source" to the discarded stderr, and the source silently never fired.
Resolve the state root to its physical directory once, then apply the
existing private-directory validation to that resolved directory and derive
every path, recorded claim identity, and later confinement check from it.
This keeps the confinement contract for the directory actually operated on
rather than only for callers that already spelled it physically, and removes
the window where an ancestor symlink could be repointed between check and
use. Homes already spelled physically behave identically.
This was the single cause of both deterministic macOS failures in
tests/fm-procevent.test.sh ("reconcile never claimed the registered source")
and tests/fm-procevent-when.test.sh ("the winning concurrent arm did not
produce an outcome"). The new case pins the behavior with an explicit
symlinked-ancestor home, so it fails without the fix on any platform rather
than only where the temp root happens to be a symlink.
* fix(bin): pin the external capture staging boundary to its physical path
The extension capture path pinned its registry staging boundary by comparing
`pwd -P` against the caller-spelled registry directory, so a home reached
through a symlinked ancestor still refused to start an extension-backed
source after the state root itself resolved correctly. That left such a home
half working: built-in sources ran while external ones failed.
The staging preparer now prints the physical registry directory it validated,
matching the inbox and reservation preparers beside it, and the start path
pins on that returned path. The new end-to-end case drives the shipped
file-signal package from a symlinked home spelling.
* no-mistakes(review): Propagate canonical process-event state roots
* no-mistakes(review): Propagate canonical state to process-event adapters
* no-mistakes(document): Document physical process-event state roots
* fix(pi): deliver captain outcomes as deterministic transcript entries (#3312)
* fix(pi): persist captain outcomes visibly
* no-mistakes(review): Recover captain outcomes after cold-start lock acquisition
* no-mistakes(document): Document cold-start captain-outcome recovery
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes(review): Prove immediate Pi captain-outcome transcript delivery
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* fix(pi): process captain outcomes through a sequence-keyed turn
PR #3312 made every captain-facing supervision outcome a durable, exact-once
visible transcript entry with the read cursor advancing only after that entry
exists. That is the display half of the delivery contract. Left alone it turns
a probabilistic silent loss into a deterministic one: the captain sees an
anchor line, and firstmate never acts, because nothing opens a turn and
nothing records whether main ever processed the outcome.
The 2026-08-31 timeline showed the two shapes this must survive on the
previous hidden-turn path: seven delivered decision outcomes each answered by
an empty assistant message (cursor advanced, no retry, unanswered for close
to three hours), and two answered by an unrelated prior reply. Both happened
because delivery advanced the cursor at enqueue and accepted whatever the
next assistant message was.
Add the processing half on top of the persistence half:
- bin/fm-branch-outcome.sh keeps a processed marker separate from the read
cursor (`unprocessed`, `mark-processed --through`, `processed-init`). It
only advances through an explicit sequence-bound acknowledgement, never
past the read cursor and never backwards; an absent marker reads as zero
and `processed-init` migrates delivered history once so an upgraded home
is not re-presented its past.
- After the visible entry for a captain outcome exists, the extension hands
every still-unprocessed captain row to main as one hidden, typed
`fm-branch-process` request listing each `[seq N] task: summary`, opening
exactly one main turn. Main closes it only by calling the new
`fm_branch_processed` tool with the highest sequence listed. An unrelated,
empty, or paraphrased answer leaves the sequence open, and the same request
is presented again at the end of the next main run and at session start.
The first two presentations of a sequence set open a turn of their own;
after that the request rides the captain's next prompt so an ignored
request cannot loop, and a session replacement resets that budget.
Routine outcomes stay turn-free.
- The regressions cover exactly those incident shapes against the real store
scripts: an empty answer and an unrelated prior answer neither advance the
marker nor stop re-presentation, the acknowledgement is refused beyond the
read cursor and outside lock ownership, a partial acknowledgement keeps the
newer sequence open, and #3312's own assertions now forbid an unkeyed turn
rather than any turn. The store suite pins the marker's bounds and the
migration; the real-SDK guard for appendEntry persistence and model
exclusion is unchanged.
Docs move the protocol from "no model turn" to "one sequence-keyed processing
turn closed only by its acknowledgement", and the verification record carries
the dated run against Pi 0.84.4.
* no-mistakes(review): Harden outcome listing and sequence-bound acknowledgements
* no-mistakes(review): Harden outcome state validation and request pacing
* no-mistakes(review): Reject unsafe sidecars and unterminated outcome stores
* no-mistakes(review): Validate canonical mark-read cursor state
* no-mistakes(review): Guard cursor advancement against corrupt processed state
* no-mistakes(review): Bind acknowledgements to active processing requests
* no-mistakes(review): Reset pacing when processing sequence membership changes
* no-mistakes(review): Enforce silent outcome invariants at storage boundary
* no-mistakes(document): Document hardened captain outcome processing contracts
---------
Co-authored-by: kunchenguid <kun@kunchenguid.com>
* feat: add bounded concurrent Bearings ledger collection (#3481)
* feat: bound Bearings remote ledger collection
* no-mistakes(review): Clarify default remote-ledger collection behavior
* no-mistakes(review): Detach reconcile delivery from watcher loop
* no-mistakes(review): Enforce bounded snapshot and request captures
* no-mistakes(review): Bound legacy summary capture before parsing
* no-mistakes(review): Bound primary remote ledger captures
* no-mistakes(document): Correct snapshot and reconcile documentation
* no-mistakes(lint): Fix ShellCheck quoting in bounded collector
* no-mistakes(ci): Fixed all three CI failures: updated the macOS Bearings assertion to 44 tests, made the home-summary test deterministic and aligned with default ledger consumption, and increased the asynchronous reconcile retirement wait for loaded CI. Verified both focused suites, all 44 Bearings tests, ShellCheck, actionlint, Bash parsing, and git diff checks
* test: await reconcile request retirement
* no-mistakes(review): Avoid empty reconcile queue process churn
* no-mistakes(review): Read ledger summaries from immutable snapshots
* no-mistakes(review): Reject multi-document home ledger streams
* no-mistakes(review): Coalesce durable reconcile requests per target
* no-mistakes(review): Unify reconcile keys and reject snapshot streams
* no-mistakes(review): Key reconcile requests by stable target ID
* no-mistakes(document): Document per-target reconcile request coalescing
* no-mistakes(lint): Remove unused snapshot summary file variable
* no-mistakes(ci): Adjusted the concurrent collector regression’s end-to-end timing ceiling to account for stock macOS process/jq overhead outside the three-second remote collection budget, while remaining below the 15-second serial-read floor. Verified with stock /bin/bash 3.2: all 44 Bearings tests pass; bash syntax and git diff checks pass
* no-mistakes(ci): Fixed legacy summary validation to require exactly one top-level JSON document and added behavioral regression coverage. Stabilized CI by conditionally waiting longer for durable reconcile delivery and synchronously stopping the fm-on worker tree before fixture cleanup. Removed a redundant flaky healthy-path timing assertion; the wedged-reader test still proves concurrent bounded collection. Verified fm-bearings-snapshot, fm-secondmate-reconcile, and fm-on tests, plus project ShellCheck, bash syntax, and git diff checks
* ci: rebalance portable serial test shards (#3489)
* fix(ci): rebalance the portable serial shards on measured durations
The "Behavior portable serial 3" shard ran 17-20 minutes against its
20-minute job cap and intermittently timed out seconds after a passing
test, on branches and on main alike.
Shards are packed longest-processing-time from per-script duration hints,
and those hints were last measured on 2026-08-21 at 116 scripts. The lane
has since grown to 139 scripts and from ~42 to ~63 minutes: 17 scripts had
no hint at all and fell back to the 20 s default, and several existing
hints were low by 2-5x (fm-watch-triage 142 s hinted vs 263 s measured,
fm-public-followup 36 s vs 197 s). The partition therefore looked
perfectly balanced in hint space, 734.6 s per shard, while really running
11.5, 13.6, 18.8 and 16.5 minutes. Script-count balance, which is what the
tests asserted, stayed normal throughout and hid it.
Refresh the hints from the timing artifacts of three green runs, taking
the slowest measurement of each script so the balance holds on a slow
runner, and split the lane across five shards instead of four. Replayed
against those runs' real per-script durations the worst shard is now
12.54 minutes, 63% of the unchanged 20-minute cap, and the serial lane's
wall clock drops from ~20 to ~12.5 minutes.
Bound the drift that caused this rather than relying on the hints being
refreshed by hand: the coverage guard now reports the unmeasured share as
serial_unhinted= and refuses past PORTABLE_SERIAL_MAX_UNHINTED_PERCENT,
which leaves room for newly added tests while making a stale table fail
the guard instead of silently pushing one shard into its cap.
No test changes what it asserts and no test stops running; only the
partition across shards changes.
* no-mistakes(document): Clarify conservative shard timing aggregate
* fix(pi): fall back on incomplete supervision branch prompts (#3491)
* fix(pi): fall back after settled branch errors
* no-mistakes(review): Detect provider errors across prompt compaction
* no-mistakes(review): Preserve in-flight branch state across selection changes
* fix(pi): re-probe supervision branch after cooldown (#3497)
* fix(pi): recover supervision branch after cooldown
* no-mistakes(review): Defer branch recovery until prompt settlement
* no-mistakes(document): Clarify supervision cooldown recovery contract
* fix(bin): remove legacy remote snapshot reads (#3501)
* refactor: remove legacy remote summary reads
* no-mistakes(document): Document ledger-only snapshot reads
* no-mistakes(ci): Fixed the snapshot test fixture so ledger refreshes use the same fake executable PATH as the snapshot consumer. This preserves observable endpoint freshness after removing legacy summary computation. Verified stock Bash parsing and all 44 Bearings tests pass under /bin/bash; git diff checks pass
* no-mistakes(ci): Fixed the CI-only snapshot fixture failure by ensuring the bounded-ledger refresh uses its fake tmux backend. This removes host tmux availability as a source of nondeterminism. Verified all 44 Bearings tests pass, Bash syntax passes, and git diff checks are clean
* no-mistakes(ci): Fixed CI nondeterminism in the Bearings fixture: all local ledger refreshes now use the fixture’s fake tmux backend when available, instead of depending on host tmux state. Verified stock /bin/bash syntax, git diff checks, and all 44 Bearings tests with a deliberately failing host tmux
* fix(pi): preserve watcher continuity across session replacement (#3498)
* fix(pi): rearm watcher after session replacement
* no-mistakes(review): Queue actionable closes across Pi session replacement
* no-mistakes(review): Stop replacement arm when handoff persistence fails
* no-mistakes(review): Preserve actionable wakes through branch and late child races
* no-mistakes(review): Surface late handoff failures without crashing Pi
* no-mistakes(review): Coordinate replacement delivery settlement and unique handoff tokens
* no-mistakes(review): Retry stale deliveries and release settled claims
* no-mistakes(review): Distinguish branch settlement and retry handoff cleanup
* no-mistakes(review): Deduplicate persistent handoff cleanup alerts
* no-mistakes(review): Acknowledge watcher follow-ups only when consumed
* no-mistakes(review): Persist idle follow-ups until agent consumption
* no-mistakes(review): Preserve pending outcomes when handoff persistence fails
* no-mistakes(review): Arm replacement before awaiting prior delivery settlement
* no-mistakes(review): Adopt pending handoffs after lock reclamation
* no-mistakes(review): Prevent stale generations from adopting replacement handoffs
* no-mistakes(review): Scope replacement handoffs by watcher state
* no-mistakes(document): Clarify replacement handoff documentation
* no-mistakes(ci): Fixed the failing branch-extension tests to model the new settlement-promise contract. Failure cases now assert that delivery ownership returns to the watcher instead of expecting direct extension fallback. Verified the updated branch suite, Pi watcher suite, shell syntax, and diff checks
* no-mistakes(review): Update branch settlement tests and preserve chunked outcomes
* no-mistakes(document): Document watcher-owned replacement handoffs
* no-mistakes(document): Verify replacement handoff documentation
* test(pi): cover watcher-owned branch fallback
* no-mistakes(document): Refresh watcher-owned fallback documentation
* fix(bin): resurface task statuses missed by wake handling (#3495)
* fix(bin): resurface terminal statuses lost after branch handling
* test(watch): canonicalize process-event fixture homes
* no-mistakes(review): Index branch outcomes by causal status position
* no-mistakes(review): Recover outcome indexes and deduplicate resurfaced statuses
* no-mistakes(review): Handle legacy ambiguity and oversized status diagnostics
* no-mistakes(review): Keep unclassifiable oversized statuses silent
* no-mistakes(document): Document lost-wake outcome backstop
* no-mistakes(document): Update outcome backstop documentation
* no-mistakes(ci): Fixed CI regressions in wake-drain: parseable reserved-key decisions can no longer bypass the durable decision-fold guard, and status output is prepared and receipt-committed before presentation to prevent repeated one-shot outcomes after later failures. Added a behavioral regression for receipt commit failure and retry. Targeted backstop, correlation-token, decision-cursor, open-decision, unread-status, syntax, and diff checks pass locally. Shard-4 failures appeared unrelated/flaky; the network-parallel test passed locally
* no-mistakes(ci): Fixed the Greptile P1 data-loss issue by committing presentation receipts only after prepared output reaches stdout. Added behavioral coverage proving output failure leaves the backstop retryable and receipt failure may duplicate but never lose a presentation. Relevant wake-drain suites and syntax/diff checks pass. The shard-4 Pi extension failure is unrelated to this PR and did not warrant changes
* no-mistakes(ci): Stabilized tests/fm-bootstrap-network-parallel.test.sh by replacing scheduler-sensitive equal-sleep timing with bounded synchronization between mocked fetch and remote probes. This preserves detection of real serialization while avoiding false failures under CI load. Verified with five consecutive test runs, bash syntax validation, ShellCheck, and git diff checks. The separate Pi stock-rendering failure reproduces locally but is unrelated environment/version drift
* no-mistakes(ci): Fixed Behavior portable serial 4 by adding fm-classify-lib.sh and fm-timeout-lib.sh to the broken-root Pi test fixture; fm-branch-outcome.sh now depends on them. Verified the full Pi branch-extension suite with real-Pi checks skipped, the wake-drain outcome-backstop suite, Bash syntax, and git diff checks. Greptile findings are already addressed at HEAD; the no-mistakes attestation failure is external head-SHA state
* fix(bin): collect follow-up results from remote work homes (#3503)
* fix(bin): deliver typed terminal results from remote work homes
A public commitment whose work is bound to a REMOTE secondmate home could
never receive its typed terminal result. `fm-public-followup.sh brief`
printed an emit command carrying this home's own absolute path and this
checkout's own script path, neither of which exists on the machine the
worker runs on, so the worker had nothing it could write to that the
owning home would ever read - and `consume` kept finding nothing while
the promise stayed open.
The brief is now route-aware: for a remote work home it prints that
route's own code root and home with `--stage-in`, so the typed event is
staged in the home where the work actually runs, and the closing
paragraph names the owning home as the one on the other machine instead
of pointing at the path above it. The owning home collects those staged
results over the same SSH route it reaches that secondmate on, because
the transport only runs outbound: `consume` pulls them into its own
inbox and reconciles them exactly as it reconciles a local report.
Collection is non-destructive until the result is durably held, so a
dropped connection cannot lose a terminal result, and a route that could
not be reached is named in `consume`'s output with the promise left open
rather than reported as an empty inbox.
A local work home is untouched: the brief still prints `--home` with this
home and this checkout's script, and the event still lands directly in
this home's typed terminal-result inbox.
This is the emit-side counterpart of the retire/clear fix in #3479 and
reuses the remote-route resolution that landed with it. Reconciling a
loop bound to a remote route now reaches that route, so the existing
remote cases drive `consume` through the same faked transport their
other steps already use.
* no-mistakes(review): Fail loudly on unresolved routes and invalid staging homes
* no-mistakes(review): Fail collection when remote outbox is unreadable
* no-mistakes(review): Surface reassigned remote routes during empty collection
* no-mistakes(review): Fail remote collection on invalid registrations
* no-mistakes(review): Reject unsafe registration entries during remote collection
* no-mistakes(review): Restore healthy empty remote collection behavior
* no-mistakes(review): Skip remote collection for delivered registrations
* no-mistakes(review): Skip delivered registrations before route validation
* no-mistakes(document): Document remote follow-up collection semantics
* fix(bin): exclude secondmates from home-summary validity (#3504)
* fix(bin): exclude secondmates from home-summary child inventory
kind=secondmate meta records never have backlog rows, so counting them in unowned_children or terminal_in_flight made a clean main home look invalid once earlier ledger checks passed.
* no-mistakes(review): Cover terminal secondmate in-flight exclusion
* no-mistakes(ci): Updated the stock macOS Bash CI snapshot expectation from 15 to 16 tests. Verified all 16 snapshot/fleet-view tests pass under Bash 3.2.57 and `git diff --check` succeeds
* fix(bin): self-heal outcome indexes on first drain (#3509)
* fix(bin): self-heal status-outcome indexes on every drain
Missing ready markers were skipping the lost-wake backstop on non-Pi homes because only the Pi branch ran processed-init. Drain now rebuilds those indexes under the outcome lock and fails closed only on a real store fault.
* no-mistakes(review): Guard held-lock initialization and fail marker writes
* no-mistakes(document): Document cross-harness outcome-index self-healing
* fix(bearings): keep active children underway during captain holds (#3505)
* fix(bearings): keep active children underway beside a captain hold
Project each readable home's active children into Underway independently of the home-level captain-decision classification so a hold no longer hides live work.
* no-mistakes(review): Preserve Underway repos and disclose child truncation
* no-mistakes(review): Fall back to task project for Underway repos
* no-mistakes(ci): Updated the stock macOS Bash CI assertion from 44 to 45 Bearings tests, matching the newly added behavioral regression. Verified all 45 tests pass under /bin/bash, Bash syntax checks pass, and git diff validation is clean
* fix(pi): settle watcher delivery on Pi accepting the follow-up (#3513)
* fix(pi): settle watcher delivery on Pi accepting the follow-up
A follow-up queued while main is streaming joins the running run without
ever raising before_agent_start, so waiting on that event before clearing
the successor pipeline (#3498) stalled every later actionable close: no
successor started, no wake was delivered or offered to the branch, and the
turn-end guard woke main to re-arm by hand after every close.
The pipeline now settles once Pi accepts the follow-up. Consumption is
observed at before_agent_start for an idle main and at the user
message_start for a streaming main, and decides only what a replacement
session (/new, /resume, /fork, reload) replays. An exhausted restoration
delivers its typed failure without launching an arm past the retry bound,
which the stall had hidden. The replacement-coordinator map is typed so the
strict no-emit typecheck passes again.
Tests: the doubles no longer raise before_agent_start for a streaming send,
a portable regression drives two actionable closes while main streams and
proves the successor chain plus consumption-scoped replay, and a
credential-free real-SDK probe pins Pi's event contract for both the
streaming and the idle follow-up.
Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a
* fix(pi): retry a verified successor that fails during wake delivery
A verified successor can exit while the wake it was started for is still
being delivered, most plausibly during a branch turn that holds the
settlement for minutes. Its failure close arrived while the pipeline's
single-flight guard was set, so the close handler skipped the retry, and
the pipeline's end no longer launched an arm, which left the live
generation with no watcher and no retry timer.
The close handler now records that failure when the child had reported
readiness and was not retired by the restoration itself, and the pipeline
runs the ordinary bounded, lock-checked retry for it once the delivery
settles. A restoration started for a later pending supersedes it, and an
exhausted restoration still hands repair to main without a further arm.
The regression holds a branch settlement open while the verified
successor exits with a failure and proves one retry watcher starts after
the settlement releases, none while it is held.
Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a
* fix(bin): bound repeat stale wakes for parked workers (#3532)
* fix(bin): bound repeat stale wakes for a parked but live worker
A worker parked on a declared wait - `paused:` for an external or pipeline
wait, or a verified `captain-held` transfer - kept waking firstmate far inside
FM_PAUSE_RESURFACE_SECS. Observed as five consecutive alarms on one
captain-held worker and dozens across a day on a pipeline wait, and reported
upstream as four wakes in 75 minutes against a 3600s window.
pause_state_class deliberately answers `none` for a still-live agent even under
a declared wait, so a worker genuinely waiting on a decision is never silenced.
That classification is correct and is left alone; it routes every parked but
live worker through surface_nonterminal_stale on first sight of each distinct
stale hash, and an idle parked pane still churns its hash on a clock or a token
counter without changing what is being waited on.
Two places let that churn re-alarm:
- surface_nonterminal_stale queued the wake BEFORE consulting whether a wait was
declared, then wrote `.paused-resurfaced-<key>` - the very throttle that should
have suppressed it. The throttle was never read on this path and was advanced
by the wake it should have prevented.
- The hash-change path cleared that throttle through clear_pause_tracking
whenever the classification came back `none`, so each tick also bought the same
declared wait a fresh window. Fixing only the first site changes nothing.
Read the throttle before anything is queued and advance it only on a wake that
really fires, and on the hash-change path reset only the per-hash bookkeeping
while the declaration still stands, via a clear_stale_hash_tracking split so
neither half of clear_pause_tracking is duplicated. The throttle is keyed to the
declaration, not to the pane.
First sight still wakes, so an inconclusive state is still inspected, and the
window's end still re-surfaces once, so a forgotten wait cannot rot invisibly -
noise traded for a bounded cadence, never for silence. The wake identity stays
the plain `stale: <win>` the away-mode handoff depends on.
Tests cover both observed forms and were confirmed to fail against three
deliberate breaks: each site reverted on its own, and a re-surface that never
fires again.
* fix(document): Clarify declared-wait wake cadence documentation
* fix(ci): Captain, fixed the stale-throttle inheritance: cadence markers now bind to the current wait declaration, so replacement paused and captain-held waits each emit their first plain `stale:` wake. Added behavioral coverage for both forms. Bite proof failed as expected when identity matching was removed, then passed after restoration. Full watcher triage suite, `bin/fm-lint.sh`, syntax checks, and diff checks pass. Changes remain uncommitted for the outer executor
* fix(ci): Captain, fixed the confirmed Greptile finding. `resurface_absorbed` now applies a throttle only when its stored declaration scope matches the current wait, so replacement `paused:` and `captain-held` waits surface immediately without changing classification. Added executable coverage for both absorbed forms. Bite proof failed before the fix at the intended assertion; afterward the full watcher triage suite, `bin/fm-lint.sh`, shell syntax checks, and `git diff --check` passed
* fix(bin): accept the away-mode daemon as the turn-end supervision owner (#3567)
* fix(turnend): accept the away-mode daemon as the supervision owner
While state/.afk exists the away-mode daemon owns supervision and runs
bin/fm-watch.sh one-shot: the watcher exits on every wake and the daemon
starts its replacement. The turn-end guard tested for a live watcher
process holding the watch lock at that instant, so a turn boundary that
landed in the hand-off blocked with "TURN WOULD END BLIND" while
supervision was completely healthy, costing a full handling turn each
time.
Reproduced with the real daemon wrapping the real watcher and the real
guard sampling the same home: 6 of 40 samples blocked, every one of them
with the daemon alive and the beacon 2-3 seconds old, and a new watcher
pid on each cycle. After the fix the same reproduction blocks 0 of 40,
and killing the daemon and its watcher (away mode still on, beacon still
fresh) blocks again.
The guard now accepts a live, identity-matched daemon holding this home
as proof of supervision while away mode is active. The identity match is
the same discipline the watcher lock uses, so a recycled pid or a lock
left by a killed daemon proves nothing. The fresh-beacon half of the
predicate is unchanged: a daemon that stops restarting its watcher still
blocks once the beacon passes grace, a home with no supervisor blocks
exactly as before, and with away mode off the strict watcher predicate is
untouched.
The predicate reads only durable state, so it behaves identically for
every primary harness and runtime backend.
* no-mistakes(document): clarify away-mode daemon supervision proof and test coverage
* no-mistakes(document): generalize stale turn-end predicate summary in architecture.md
* fix(backlog): omit --file from row probes for non-markdown backends (#3582)
* fix(backlog): omit markdown file for beads probes
* no-mistakes(document): Narrow backlog addressing doc to mutations for backend-aware probes
* no-mistakes(ci): Fixed the Greptile P2 review comment (the only failing check) on tests/fm-backlog-atomicity.test.sh. The comment correctly noted that an exported TASKS_AXI_BACKEND environment variable would inherit into the spawned scripts and, because fm_tasks_axi_backend gives it top precedence, override each test case's .tasks.toml backend fixture — making the backend-specific argv assertions fail for environmental reasons. Fix: unset TASKS_AXI_BACKEND in the test harness right after sourcing tests/lib.sh, with a comment explaining why, so every case deterministically exercises its declared backend (4 lines added; no production code touched). Verified: reproduced the leak before the fix (TASKS_AXI_BACKEND=beads made the markdown dispatch case fail with 'beads show failed', exactly the reported failure mode); after the fix the full suite passes (0 failures, exit 0) both with and without TASKS_AXI_BACKEND=beads exported. The added lines are shellcheck-clean (the only shellcheck note, SC1091 on the lib.sh source line, pre-exists this change)
* fix(bin): classify progress updates on requested work as routine (#3589)
The supervision branch's verdict rule escalated every outcome that
answered a captain request, so "the work started" and "still working"
notes reached the captain with nothing to look at. The rule now keeps a
finished result of requested work captain-facing, even when healthy, and
treats start or still-working updates that bring no new artifact,
finding, or decision as routine. The captain list for review-ready PRs,
ask-user findings, exhausted blockers, credentials, and destructive or
security-sensitive cases is unchanged, as are the unsolicited-routine,
silent-fleet-review, and doubt-chooses-captain rules.
The fm_branch_report tool description and the two docs that restated the
old unconditional rule now point at the prompt's "Verdict: routine or
captain" section as the one owner instead of carrying a second copy.
* fix(bin): preserve captain calls during teardown (#3595)
* fix(bin): never close a captain call during cleanup
A scout that held its own work item for the captain, which is what
captain-hold-lifecycle prefers ("hold the work item the question gates"),
was closed by bin/fm-teardown.sh's automatic backlog transition. The
completion gate passed, cleanup ran, and the captain's question moved to
Done with no recorded answer: the one thing the policy says must never
happen. `tasks-axi done` closes a held row silently, and nothing in
teardown asked whether the row was the captain's own call.
bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when
the task is still an open captain call, 1 when it is not, 2 when that
cannot be established. It reads the row through the transition library's
backend-aware probe, so it addresses the same backlog teardown does; the
script's other commands now address the configured data directory the
same way instead of FM_HOME, which also fixes captain holds in a home
with a relocated data directory.
Teardown asks `open` before any destructive step and refuses on 2. On 0
only the close changes: after cleanup and still under the task's own
lock, the row gets one "Deliverable of the finished work" line at the end
of its body and returns to Queued through `tasks-axi reopen`, keeping its
hold, so it lands in Captain's Call instead of reading as work under way.
--force does not lift this: it authorizes discarding unlanded work, never
the captain's question. The deliverable goes into the body because
`tasks-axi update --report` rewrites the title of a row that is not Done.
The crash window reuses the pending-close record teardown already stages:
a `mode=retain` line makes the existing replay record the deliverable and
reopen instead of closing, with the same validator, stale-generation
check, cleanup-incomplete marking, and non-blocking bootstrap lock as an
ordinary close. A retained row the captain answered first simply retires
the record. No parallel record type, recovery command, or second bootstrap
loop is introduced.
Regressions run the real executables: the captain-held scout survives
cleanup queued, held, with its deliverable and on the board, only
`answer` closes it, --force keeps it open, and an ordinary scout still
closes with its report; an interrupted cleanup leaves the row untouched
and the next session start retains it; a relocated backlog keeps the
retention in its one configured file; and a ship row whose hold cannot be
read refuses cleanup before anything destructive.
Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np
* no-mistakes(review): Serialize captain holds and fix backend-aware listing
* no-mistakes(document): Update captain-call retention documentation
* no-mistakes(document): Fix relocated captain-hold backlog diagnostics
* fix(bin): deliver secondmate outcomes to the parent channel (#3592)
* fix(bin): deliver every secondmate outcome on the parent channel from the recording scripts
A secondmate's captain-facing outcomes could miss: the mate model addressed
the captain in its own unread chat instead of appending to the parent
channel, and a PR-ready report, a finding, a decision, a blocker, and a
failure all depended on that one remembered append. Make delivery
structural, so the parent channel never depends on the model:
- bin/fm-parent-channel-lib.sh is the one owner of channel resolution and
exact-line append-once; the merge outcome path and the inactive-outcome
scan now publish through it instead of two private copies.
- bin/fm-inactive-reconcile.sh gains a ledger-first path that runs on every
watcher poll in a secondmate home: a direct child's whole terminal done or
failed line is delivered at once with its note, recorded PR, mode, merge
posture, and scout report pointer, keyed and receipted so it is delivered
once, and the inactive path yields to it. `report <task-id>` runs the same
delivery for a caller holding the child's meta lock.
- bin/fm-pr-check.sh publishes the PR-ready line with the canonical URL at
registration.
- bin/fm-captain-hold.sh publishes a hold and its answer, keyed by task id
and resolution-record count, with no new persisted state.
- bin/fm-teardown.sh delivers the child's final line before removing its
record and refuses, retaining every record, while the channel cannot be
written.
- The charter opens with the parent-channel rule and confines the mate's own
appends to judgement; AGENTS.md carries the carve-out at the persona
address rule and the escalation list.
docs/secondmate-parent-channel.md records the design and its coverage, and
docs/verification/secondmate-parent-channel.md records the live run with real
tmux panes and both real watchers delivering every line with no model.
Supersedes #3569.
* no-mistakes(review): Fix parent outcome retries and reconciliation locking
* no-mistakes(review): Prevent busy children from starving ledger delivery
* no-mistakes(review): Correct ledger metadata and hold occurrence handling
* no-mistakes(review): Disambiguate ledger outcomes and normalize hold reasons
* no-mistakes(review): Close ledger races and preserve teardown records
* no-mistakes(document): Correct parent-channel receipt and scanner documentation
* no-mistakes(lint): Quote done arguments for ShellCheck compliance
* no-mistakes(ci): Fixed both CI failures. Updated GOTMP teardown fixtures for the new final-outcome reporter and isolated them from host tmux state. Updated the PR security assertion to distinguish the accepted PR-ready line from duplicate merge outcomes. Verified with both failing test suites, bash syntax checks, and git diff checks
* no-mistakes(ci): Fixed Greptile’s duplicate-delivery race in bin/fm-inactive-reconcile.sh. Ledger events now claim matching already-delivered inactive receipts using the prior status fingerprint, preventing duplicate parent reports while preserving later same-state completions. Added behavioral regression coverage. Verified inactive-reconcile tests, project lint, documentation audience checks, syntax, and diff checks. Teardown tests passed relevant cases before the documented pre-existing herdr-preflight-missing-adapter failure
* fix(bin): sync remote second mates to primary commit (#3599)
* fix(bin): sync remote second-mate homes to the parent primary commit
Session start and remote launch pointed a remote second-mate home at whatever
Firstmate copy its own host kept, so a home that had already advanced past that
copy refused as a non-fast-forward and every other home stopped at the host's
older commit while the primary ran ahead.
The parent now resolves ITS primary default-branch commit with the existing
helper and hands that commit to the host on both paths. Because a remote home
is a standalone clone, the host imports that one commit before advancing -
already present, else from that host's Firstmate copy without moving it, else
from the home's own origin - and then runs the SAME ff_target guards a local
home gets, so dirty, diverged, feature-branch, and unresolvable targets skip
untouched and the ancestry rules keep one owner. An unimportable target now
names /updatefirstmate instead of failing opaquely, and a host still running an
older Firstmate copy is reported the same way rather than echoing a bare
refusal.
The host-local launch leg no longer re-runs its own secondmate sync, so the
spawn it drives cannot re-target that host's copy after the parent has already
converged the home.
/updatefirstmate is unchanged: it still refreshes the remote code root from that
host's origin and then syncs the home to that refreshed copy, which is what the
sync call with no target commit means.
* no-mistakes(document): Document primary-targeted remote secondmate synchronization
* fix(bin): separate captain intent from firstmate specs (#3597)
* fix(bin): split brief task into captain intent and firstmate spec
Keep no-mistakes --intent as the captain's ask plus later captain words, not the build spec or worker tradeoffs.
* fix(bin): stop task-subsection copies at the next heading
Promotion was swallowing the scout Setup contract into Firstmate spec, and pre-subsection briefs lost their # Task body.
* no-mistakes(review): Validate brief content and preserve nested specifications
* no-mistakes(review): Scope placeholder validation to scaffold-only subsection bodies
* no-mistakes(review): Ignore fenced subsection headings during brief validation
* no-mistakes(review): Preserve captain intent across scout promotion
* no-mistakes(review): Enforce safe intent boundaries for legacy promotions
* no-mistakes(review): Allow marked legacy intent and reject empty promotions
* no-mistakes(review): Scope task parsing and overlay legacy intent contracts
* no-mistakes(review): Overlay current intent contract for all no-mistakes spawns
* no-mistakes(review): Preserve later captain clarifications in intent overlays
* no-mistakes(document): Document brief intent enforcement and ownership
* no-mistakes(ci): Updated spawn-related test fixtures to use valid Captain intent and Firstmate spec subsections, corrected launch-path expectations to launch-brief.md, and resolved ShellCheck quoting findings. Verified with fm-lint.sh and 15 affected behavior tests, including real Herdr tests; all passed
* no-mistakes(ci): Updated stale spawn/promotion fixtures in the Muse, Orca, secondmate-harness, and public-followup suites to provide valid Captain's intent and Firstmate spec subsections. Verified full Orca and secondmate-harness suites, targeted public-followup promotion behavior, Bash syntax, diff checks, and fm-lint
* fix: start a fresh supervision branch for every main session (#3600)
* fix(pi): start a new supervision branch conversation per main session
The supervision branch reopened one recorded conversation forever, so
every main session start reloaded the current generated prompt and then
weeks of accumulated thread, where a superseded rule could still outweigh
today's.
The branch conversation is now scoped to one main session: the session
generation owns the recorded conversation, so a cold start, /new,
/resume, /fork, or a reload always builds a new one, while a rebuild
inside one session (a model or effort change) still continues that
session's own conversation.
The dialog mirror re-anchors with it. Its durable cursor records what the
previous branch conversation received, so a /resume or reload - which
keeps main's own session file - would otherwise leave the new branch
blind to dialog main itself still has. The reset is bounded by the
current main session, and the cursor keeps advancing incrementally within
it. The durable outcome store and its processed marker are untouched, so
unacknowledged captain-facing outcomes still re-present on the new main
session.
* no-mistakes(document): Document fresh Pi supervision conversations
* no-mistakes(ci): Fixed the flaky concurrent inbox failure. Lock acquisition now retries when a competing lock disappears between a failed claim and inspection. Added a behavioral regression covering that race. Verified the full inbox test four times, project lint, and git diff checks
* feat: restart second mates after instruction updates (#3614)
* feat(update): restart second mates whose instructions changed
/updatefirstmate pulled new bytes onto disk and then asked each advanced
second mate to re-read them. A running agent holds AGENTS.md and every
loaded skill frozen from launch and no verified harness offers a reload,
so that steer could not reach a loaded skill at all and left the mate
holding two contradictory copies of its own job description.
An eligible mate is now restarted instead, in the same home and endpoint,
through the existing transactional relaunch. The restart is gated on the
mate first writing down the open work it holds only in conversation - the
open-record half of /stow, never its memory sweeps - so an unregistered
captain call is flushed before the conversation is spent. Anything that
leaves the reload unprovable falls back to the old re-read message and is
reported as exactly that, never as a clean reload.
Remote mates take the same path: fm-remote-secondmate-control.sh gains a
relaunch verb whose host-local leg runs that same control plane, since the
mate is an ordinary local secondmate from its host's point of view. The
primary resolves the profile and passes it explicitly, because
config/secondmate-harness is not inherited and the file on that host
belongs to a different home.
fm-update.sh now splits its advanced live mates into a restart set and a
nudge residual, and both sets require a changed instruction surface, which
also closes the over-nudge against the session-start sweep. Restart is
stricter still: a bin/-only advance reloads itself on the next call, so it
never costs a conversation.
Colocated tests cover the gating, the persist-then-restart order, the
task-subset persist request, each unsafe fallback, the remote hop, and the
remote sync's new instruction-surface report.
* no-mistakes(review): Fix restart correlation, concurrent waits, and lifecycle reporting
* no-mistakes(review): Parallelize relaunches and classify replacement incarnations
* no-mistakes(review): Gate restart actions on live agent state
* no-mistakes(review): Handle failed restart workers without hanging
* no-mistakes(review): Nudge legacy remotes and preserve persist recovery
* no-mistakes(review): Document one-time secondmate restart rollout
* no-mistakes(review): Honor arrived replies and refresh remote profiles
* no-mistakes(review): Revert remote parent profile reconciliation
* no-mistakes(review): Reset remote profile defaults and honor published results
* no-mistakes(review): Preserve fallback nudges for unverifiable secondmates
* no-mistakes(document): Document second-mate restart update flow
* no-mistakes(lint): Fix ShellCheck warnings in restart scripts
* perf: accelerate local validation with bounded concurrency (#3644)
* perf(tests): route gate verification through the bounded concurrent runner
Local validation was the pipeline's dominant cost: across 67 recorded
no-mistakes agent sessions on this repo, 99.3% of command execution was
`bash tests/*.test.sh`, run strictly one script at a time, and 2% of those
calls were killed by an agent-guessed timeout and paid for twice.
Three changes, each measured:
- `.no-mistakes.yaml` pins `commands.test` to
`bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner
already owns changed-file selection, bounded concurrency, the refusal of
unproven scripts, and a generous automatic per-script bound, so the gate's
baseline is neither a serial chain nor a guessed timeout. It stays
intent-targeted - the Test step still runs its evidence agent on top - and
excludes the live-Herdr family the required Herdr lane owns.
- `bin/fm-test-run.sh` gives a plain list of script paths the same bounded
automatic scheduler and automatic bound that `--changed` gets. Naming several
subjects is how a verification round asks for exactly those scripts. The
curated selections are untouched: `--lane` still composes CI shards whose
serial lane must stay serial, `--family` is what the required Herdr lane runs,
and `--all` stays a deliberate complete regression.
- `pr-forge` is admitted to the concurrent-safe family registry on two
consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those,
and records `secondmate` and `session-bootstrap` as refused with the exact
script and reason each failed on, so the refusals are actionable rather than
silent.
Measured on this host, 0 failures on both sides:
verification round, 4 scripts 448s chained -> 231s through the runner (-48%)
pr-forge family 409.2s at 1 worker -> 237.9s at 4 (1.72x)
watcher-wake-lock family 1311.1s at 1 worker -> 539.3s at 4 (2.43x)
A fourth lever was implemented and then removed because the measurement
refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made
`fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s
unchanged, back to back. Those sleeps are not overhead added to the clock -
they are how a test waits for a subject moving on fm-watch.sh's own one-second
cadence - so sampling less often only delays detection. It also broke
`fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a
settled condition. CONTRIBUTING.md records that result so the experiment is not
repeated.
* no-mistakes(review): Separate concurrent runs by isolation proof family
* no-mistakes(review): Limit automatic timeouts to changed-file validation
* no-mistakes(document): Clarify validation concurrency documentation
* fix: copy PR URLs from durable records (#3648)
* fix: copy PR URLs from records or abstain, never assemble them
Supervision reported a plausible but dead PR link three times because its
prompt demanded a full https:// URL at a moment when only a PR number was
observable, so the model assembled an owner/repository from memory, and the PR
check then accepted that URL and wrote it into the task record, after which the
model kept defending its own tool-endorsed guess over the worker's real link.
Three changes close that chain without any live forge lookup, so private
forges are treated exactly like public ones:
- bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy
or abstain" section requires a URL to be copied verbatim from a durable
record (the done: PR <url> status line, pr= metadata, or the backlog note),
forbids assembling owner, repository, host, or number from memory, and has
the branch report only the identifier it actually holds when no record names
the URL yet, leaving the PR check unarmed until the worker's ready line
arrives. AGENTS.md section 7 and 9 carry the same copy-or-abstain rule for
main in place of the bare full-URL mandate.
- Worker briefs (bin/fm-brief.sh, ship and scout rules) require …
lytv
pushed a commit
to lytv/mymate
that referenced
this pull request
Sep 8, 2026
* fix(bin): never close a captain call during cleanup
A scout that held its own work item for the captain, which is what
captain-hold-lifecycle prefers ("hold the work item the question gates"),
was closed by bin/fm-teardown.sh's automatic backlog transition. The
completion gate passed, cleanup ran, and the captain's question moved to
Done with no recorded answer: the one thing the policy says must never
happen. `tasks-axi done` closes a held row silently, and nothing in
teardown asked whether the row was the captain's own call.
bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when
the task is still an open captain call, 1 when it is not, 2 when that
cannot be established. It reads the row through the transition library's
backend-aware probe, so it addresses the same backlog teardown does; the
script's other commands now address the configured data directory the
same way instead of FM_HOME, which also fixes captain holds in a home
with a relocated data directory.
Teardown asks `open` before any destructive step and refuses on 2. On 0
only the close changes: after cleanup and still under the task's own
lock, the row gets one "Deliverable of the finished work" line at the end
of its body and returns to Queued through `tasks-axi reopen`, keeping its
hold, so it lands in Captain's Call instead of reading as work under way.
--force does not lift this: it authorizes discarding unlanded work, never
the captain's question. The deliverable goes into the body because
`tasks-axi update --report` rewrites the title of a row that is not Done.
The crash window reuses the pending-close record teardown already stages:
a `mode=retain` line makes the existing replay record the deliverable and
reopen instead of closing, with the same validator, stale-generation
check, cleanup-incomplete marking, and non-blocking bootstrap lock as an
ordinary close. A retained row the captain answered first simply retires
the record. No parallel record type, recovery command, or second bootstrap
loop is introduced.
Regressions run the real executables: the captain-held scout survives
cleanup queued, held, with its deliverable and on the board, only
`answer` closes it, --force keeps it open, and an ordinary scout still
closes with its report; an interrupted cleanup leaves the row untouched
and the next session start retains it; a relocated backlog keeps the
retention in its one configured file; and a ship row whose hold cannot be
read refuses cleanup before anything destructive.
Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np
* no-mistakes(review): Serialize captain holds and fix backend-aware listing
* no-mistakes(document): Update captain-call retention documentation
* no-mistakes(document): Fix relocated captain-hold backlog diagnostics
kunchenguid
added a commit
that referenced
this pull request
Sep 9, 2026
…#3417) * fix(backlog): honor configured task adapters * no-mistakes(review): Harden backend purity lint against prefixed Beads calls * no-mistakes(document): Document configured backend lifecycle transitions * fix(backlog): preserve markdown exemptions * no-mistakes(review): Enforce backend purity for explicit lint paths * no-mistakes(document): Update lifecycle backend documentation * no-mistakes(lint): Remove redundant backend lint pattern * fix(backlog): close adapter routing gaps * no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint * no-mistakes(document): Document environment-selected backlog adapters * no-mistakes(lint): Fix empty local variable assignment * fix(backlog): close quoted path gaps * no-mistakes(review): Reject partially quoted direct Beads commands * no-mistakes(document): Align lifecycle documentation with configured adapters * test(backlog): keep structural cases markdown-only * fix(backlog): honor configured task adapters * no-mistakes(review): Harden backend purity lint against prefixed Beads calls * no-mistakes(document): Document configured backend lifecycle transitions * fix(backlog): preserve markdown exemptions * no-mistakes(review): Enforce backend purity for explicit lint paths * no-mistakes(document): Update lifecycle backend documentation * no-mistakes(lint): Remove redundant backend lint pattern * fix(backlog): close adapter routing gaps * no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint * no-mistakes(document): Document environment-selected backlog adapters * no-mistakes(lint): Fix empty local variable assignment * fix(backlog): close quoted path gaps * no-mistakes(review): Reject partially quoted direct Beads commands * no-mistakes(document): Align lifecycle documentation with configured adapters * test(backlog): keep structural cases markdown-only * no-mistakes(review): Harden markdown lifecycle routing and close recovery * fix(lint): catch dollar-quoted beads commands * no-mistakes(review): Drop unused authorized-data-dir param; canonicalize fm-lint ROOT with pwd -P * no-mistakes(document): Document fm-lint backend-purity check in header and CONTRIBUTING * no-mistakes(ci): Two failing checks, one code-caused and one attestation-only. 1. Greptile Review (P1: ANSI-C quoting bypasses lint) — REAL DEFECT, FIXED. The backend-purity normalizer in bin/fm-lint.sh (invokes_bd awk function) stripped $'...' quote delimiters without decoding ANSI-C escapes, so a core script containing $'\x62\x64' close fm-example executes as `bd close` while the lint accepted it. Fix: the normalizer now tracks whether a single-quote context came from $' (ansi flag) and decodes ANSI-C escapes while building the command word: \xHH (1-2 hex), \uHHHH / \UHHHHHHHH (4/8 hex), \NNN (1-3 octal), \cX control characters, simple escapes (a/b/e/E/f/n/r/t/v decode to a placeholder that can never spell bd), NUL (value 0) truncates the word per bash C-string semantics, and unknown escapes drop the backslash per bash. Values outside printable ASCII decode to a placeholder so they can neither falsely match nor collide into `bd`. Regression coverage added to the existing lint-interface test test_rejects_direct_beads_cli_invocations in tests/fm-lint.test.sh: $'\x62\x64', $'\142\144' (octal), and b$'\x64' (split word), each written into a fixture and asserted rejected through the real fm-lint.sh executable. Verified regression property: with the fix stashed, the new hex case passes lint (reproduces the reported bypass); with the fix, it is rejected. Verification: all 29 tests in tests/fm-lint.test.sh pass, and the full CI-parity lint (CI=true bin/fm-lint.sh: ShellCheck 0.11.0 full analysis over the canonical set, backend-purity check, and actionlint workflow validation) exits 0. One iteration was needed because an awk comment containing a literal $'\x62\x64' sample broke shell-level quoting (bash -n / SC1001/SC2026); the comment was reworded without quotes. 2. PR must be raised via no-mistakes — NOT code-caused. The check failed with 'Pipeline attestation head_sha does not match the current PR head': the PR body attestation binds to 82b41c7 while the PR head is 1995cdb because a later pipeline push moved the head. This is exactly the stale-attestation condition the user intent describes; it clears when the outer pipeline re-runs 'git push no-mistakes' and re-binds the attestation to the new head (which now includes this Greptile fix). No code change can or should address it. Files changed: bin/fm-lint.sh (ANSI-C escape decoding in the backend-purity normalizer), tests/fm-lint.test.sh (three encoded-bd rejection cases) * no-mistakes(review): fix tasks.toml hang, root authorization, lint quoting * no-mistakes(review): validate tasks config before exemption; fix lint quote gap * fix(backlog): address the markdown backlog as <data>/backlog.md Resolving the markdown backlog through a configured `[markdown] path` was scope this task never asked for. It is absent from main, which addresses `<data>/backlog.md` everywhere, and it came from an earlier review round rather than the task brief. Making it effective on the transition path alone put that path at odds with every other consumer of the same backlog - fm-captain-hold.sh, fm-session-start.sh, fm-fleet-snapshot.sh, fm-inbox.sh, fm-backlog-handoff.sh - which all still address `<data>/backlog.md`. In fm-captain-hold.sh the split was live: its reads had already moved to the shared gate while its writes had not, so the two could address different files. Address `<data>/backlog.md` from the shared gate, delete the unused resolver, and drop the two tests that pinned the withdrawn behaviour. What this task actually changes is unaffected: a configured non-markdown adapter is still addressed by its own root, without `--file`. * fix(backlog): honor configured task adapters * no-mistakes(review): Harden backend purity lint against prefixed Beads calls * no-mistakes(document): Document configured backend lifecycle transitions * fix(backlog): preserve markdown exemptions * no-mistakes(review): Enforce backend purity for explicit lint paths * no-mistakes(document): Update lifecycle backend documentation * no-mistakes(lint): Remove redundant backend lint pattern * fix(backlog): close adapter routing gaps * no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint * no-mistakes(document): Document environment-selected backlog adapters * no-mistakes(lint): Fix empty local variable assignment * fix(backlog): close quoted path gaps * no-mistakes(review): Reject partially quoted direct Beads commands * no-mistakes(document): Align lifecycle documentation with configured adapters * test(backlog): keep structural cases markdown-only * fix(lint): catch dollar-quoted beads commands * no-mistakes(review): Drop unused authorized-data-dir param; canonicalize fm-lint ROOT with pwd -P * no-mistakes(document): Document fm-lint backend-purity check in header and CONTRIBUTING * no-mistakes(ci): Two failing checks, one code-caused and one attestation-only. 1. Greptile Review (P1: ANSI-C quoting bypasses lint) — REAL DEFECT, FIXED. The backend-purity normalizer in bin/fm-lint.sh (invokes_bd awk function) stripped $'...' quote delimiters without decoding ANSI-C escapes, so a core script containing $'\x62\x64' close fm-example executes as `bd close` while the lint accepted it. Fix: the normalizer now tracks whether a single-quote context came from $' (ansi flag) and decodes ANSI-C escapes while building the command word: \xHH (1-2 hex), \uHHHH / \UHHHHHHHH (4/8 hex), \NNN (1-3 octal), \cX control characters, simple escapes (a/b/e/E/f/n/r/t/v decode to a placeholder that can never spell bd), NUL (value 0) truncates the word per bash C-string semantics, and unknown escapes drop the backslash per bash. Values outside printable ASCII decode to a placeholder so they can neither falsely match nor collide into `bd`. Regression coverage added to the existing lint-interface test test_rejects_direct_beads_cli_invocations in tests/fm-lint.test.sh: $'\x62\x64', $'\142\144' (octal), and b$'\x64' (split word), each written into a fixture and asserted rejected through the real fm-lint.sh executable. Verified regression property: with the fix stashed, the new hex case passes lint (reproduces the reported bypass); with the fix, it is rejected. Verification: all 29 tests in tests/fm-lint.test.sh pass, and the full CI-parity lint (CI=true bin/fm-lint.sh: ShellCheck 0.11.0 full analysis over the canonical set, backend-purity check, and actionlint workflow validation) exits 0. One iteration was needed because an awk comment containing a literal $'\x62\x64' sample broke shell-level quoting (bash -n / SC1001/SC2026); the comment was reworded without quotes. 2. PR must be raised via no-mistakes — NOT code-caused. The check failed with 'Pipeline attestation head_sha does not match the current PR head': the PR body attestation binds to 82b41c7 while the PR head is 1995cdb because a later pipeline push moved the head. This is exactly the stale-attestation condition the user intent describes; it clears when the outer pipeline re-runs 'git push no-mistakes' and re-binds the attestation to the new head (which now includes this Greptile fix). No code change can or should address it. Files changed: bin/fm-lint.sh (ANSI-C escape decoding in the backend-purity normalizer), tests/fm-lint.test.sh (three encoded-bd rejection cases) * fix(bin): preserve captain calls during teardown (#3595) * fix(bin): never close a captain call during cleanup A scout that held its own work item for the captain, which is what captain-hold-lifecycle prefers ("hold the work item the question gates"), was closed by bin/fm-teardown.sh's automatic backlog transition. The completion gate passed, cleanup ran, and the captain's question moved to Done with no recorded answer: the one thing the policy says must never happen. `tasks-axi done` closes a held row silently, and nothing in teardown asked whether the row was the captain's own call. bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when the task is still an open captain call, 1 when it is not, 2 when that cannot be established. It reads the row through the transition library's backend-aware probe, so it addresses the same backlog teardown does; the script's other commands now address the configured data directory the same way instead of FM_HOME, which also fixes captain holds in a home with a relocated data directory. Teardown asks `open` before any destructive step and refuses on 2. On 0 only the close changes: after cleanup and still under the task's own lock, the row gets one "Deliverable of the finished work" line at the end of its body and returns to Queued through `tasks-axi reopen`, keeping its hold, so it lands in Captain's Call instead of reading as work under way. --force does not lift this: it authorizes discarding unlanded work, never the captain's question. The deliverable goes into the body because `tasks-axi update --report` rewrites the title of a row that is not Done. The crash window reuses the pending-close record teardown already stages: a `mode=retain` line makes the existing replay record the deliverable and reopen instead of closing, with the same validator, stale-generation check, cleanup-incomplete marking, and non-blocking bootstrap lock as an ordinary close. A retained row the captain answered first simply retires the record. No parallel record type, recovery command, or second bootstrap loop is introduced. Regressions run the real executables: the captain-held scout survives cleanup queued, held, with its deliverable and on the board, only `answer` closes it, --force keeps it open, and an ordinary scout still closes with its report; an interrupted cleanup leaves the row untouched and the next session start retains it; a relocated backlog keeps the retention in its one configured file; and a ship row whose hold cannot be read refuses cleanup before anything destructive. Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np * no-mistakes(review): Serialize captain holds and fix backend-aware listing * no-mistakes(document): Update captain-call retention documentation * no-mistakes(document): Fix relocated captain-hold backlog diagnostics * fix(backlog): honor configured task adapters * no-mistakes(document): Update lifecycle backend documentation * fix(backlog): close adapter routing gaps * no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint * no-mistakes(lint): Fix empty local variable assignment * fix(backlog): close quoted path gaps * test(backlog): keep structural cases markdown-only * no-mistakes(review): Harden markdown lifecycle routing and close recovery * no-mistakes(review): fix tasks.toml hang, root authorization, lint quoting * no-mistakes(review): validate tasks config before exemption; fix lint quote gap * fix(backlog): address the markdown backlog as <data>/backlog.md Resolving the markdown backlog through a configured `[markdown] path` was scope this task never asked for. It is absent from main, which addresses `<data>/backlog.md` everywhere, and it came from an earlier review round rather than the task brief. Making it effective on the transition path alone put that path at odds with every other consumer of the same backlog - fm-captain-hold.sh, fm-session-start.sh, fm-fleet-snapshot.sh, fm-inbox.sh, fm-backlog-handoff.sh - which all still address `<data>/backlog.md`. In fm-captain-hold.sh the split was live: its reads had already moved to the shared gate while its writes had not, so the two could address different files. Address `<data>/backlog.md` from the shared gate, delete the unused resolver, and drop the two tests that pinned the withdrawn behaviour. What this task actually changes is unaffected: a configured non-markdown adapter is still addressed by its own root, without `--file`. * no-mistakes(review): restore home boundary guard and tighten purity lint * no-mistakes(review): authorize home boundary for every backlog adapter * no-mistakes(test): complete tasks-axi stubs in fm-gotmp teardown fixtures * no-mistakes(document): align backlog transition docs with adapter-neutral addressing * no-mistakes(review): label adapter data-dir authorization, drop dead row_probe local * no-mistakes(review): pin markdown backend at relocated-data addressing roots * no-mistakes(document): point lint-definition mention at fm-lint.sh header * no-mistakes(document): point mutate comment at adapter addressing owner * no-mistakes(review): Fix leftover-symlink refusal on non-markdown homes; hoist config check and lint/dedup cleanups * no-mistakes(document): Align fm-lint purity scope header with bin/backends --------- Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
BenWilcox8
pushed a commit
to BenWilcox8/firstmate
that referenced
this pull request
Sep 12, 2026
* fix(bin): never close a captain call during cleanup
A scout that held its own work item for the captain, which is what
captain-hold-lifecycle prefers ("hold the work item the question gates"),
was closed by bin/fm-teardown.sh's automatic backlog transition. The
completion gate passed, cleanup ran, and the captain's question moved to
Done with no recorded answer: the one thing the policy says must never
happen. `tasks-axi done` closes a held row silently, and nothing in
teardown asked whether the row was the captain's own call.
bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when
the task is still an open captain call, 1 when it is not, 2 when that
cannot be established. It reads the row through the transition library's
backend-aware probe, so it addresses the same backlog teardown does; the
script's other commands now address the configured data directory the
same way instead of FM_HOME, which also fixes captain holds in a home
with a relocated data directory.
Teardown asks `open` before any destructive step and refuses on 2. On 0
only the close changes: after cleanup and still under the task's own
lock, the row gets one "Deliverable of the finished work" line at the end
of its body and returns to Queued through `tasks-axi reopen`, keeping its
hold, so it lands in Captain's Call instead of reading as work under way.
--force does not lift this: it authorizes discarding unlanded work, never
the captain's question. The deliverable goes into the body because
`tasks-axi update --report` rewrites the title of a row that is not Done.
The crash window reuses the pending-close record teardown already stages:
a `mode=retain` line makes the existing replay record the deliverable and
reopen instead of closing, with the same validator, stale-generation
check, cleanup-incomplete marking, and non-blocking bootstrap lock as an
ordinary close. A retained row the captain answered first simply retires
the record. No parallel record type, recovery command, or second bootstrap
loop is introduced.
Regressions run the real executables: the captain-held scout survives
cleanup queued, held, with its deliverable and on the board, only
`answer` closes it, --force keeps it open, and an ordinary scout still
closes with its report; an interrupted cleanup leaves the row untouched
and the next session start retains it; a relocated backlog keeps the
retention in its one configured file; and a ship row whose hold cannot be
read refuses cleanup before anything destructive.
Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np
* no-mistakes(review): Serialize captain holds and fix backend-aware listing
* no-mistakes(document): Update captain-call retention documentation
* no-mistakes(document): Fix relocated captain-hold backlog diagnostics
adibirzu
added a commit
to adibirzu/firstmate
that referenced
this pull request
Sep 15, 2026
* test: make timestamp fixtures portable across macOS and Linux (#4037)
* test(lib): set fixture mtimes through one portable epoch helper
On macOS the visible symptom was ONE red case in the turn-end guard suite. The
actual damage was TWO cases that had quietly stopped testing their subject. The
red one was the harmless half - people read one red case as one broken thing,
and here that intuition is wrong.
`touch -d @<epoch>` is a GNU extension; BSD touch rejects it outright and leaves
the file at its current mtime. So on macOS the three away-mode beacon cases
never aged their beacon at all.
The 400s case exists to pin that 400s is stale under the flat 300s default but
fresh under the poll-derived grace (660s at FM_POLL=600). Deleting the grace it
guards (FM_POLL=60, so max(300,120)=300) and re-running proves what it was
worth on this platform:
pre-fix input (beacon left at now): ok - passes with the feature DELETED
post-fix input (beacon 400s old): not ok - expected exit 0, got 2
It was green while measuring nothing, and could not have caught a regression in
the grace it names. Only the 700s case broke loudly.
`touch -t [[CC]YY]MMDDhhmm[.SS]` is POSIX and both platforms accept it, so the
only host-specific step left is formatting the epoch into that stamp, which
date(1) spells two incompatible ways. fm_touch_epoch in tests/lib.sh owns that
probe once and fails loudly rather than leaving an unset timestamp behind - the
failure mode that caused this. Verified on BSD touch/date here and on GNU
coreutils 9.7 in a container.
Three real sites, and one consistency change - not four fixes. The stale
destination lock in tests/fm-remote-backlog-handoff.test.sh was never a defect:
its `uname = Darwin` branch made the `touch -d` line unreachable on macOS, and
BSD touch accepts that space-separated form anyway. `touch -t` takes the date
directly on both platforms, so the branch goes rather than standing as a second
copy of the same platform assumption.
Known limit: the third case (away mode off) is only HALF recovered here. It now
receives the input its name claims, but it is still insensitive after this fix -
its verdict is identical with a 0s and a 400s beacon, because the fixture
records a daemon lock and no watcher lock, and with away mode off the daemon
lock proves nothing. Not fixed here; tracked separately, with the requirement
that any fix be shown to FAIL when the protection is removed.
FULL SUITE ON macOS: 189 scripts, four red, none of them this change. The
turn-end guard and remote-handoff suites are clean. Attribution was established
by running the four failures at the base commit and at this head on an idle
machine, because base-idle against head-under-load moves two variables at once:
script head/loaded base/idle head/idle verdict
fm-calm-pi-extension red red red pre-existing
fm-backlog-atomicity red red red pre-existing
fm-procevent red red red pre-existing
fm-startup-network red green green cause unestablished,
load-sensitive under
a full run
Reported, not fixed. fm-calm-pi-extension deserves its own note: it FAILS
because Chrome is absent instead of declaring the capability it needs and
standing aside, so its verdict is about the machine rather than its subject -
the same family as the defect above, with the red at least announcing itself.
Neighbouring class, reported not changed: `file_mode()` - a verbatim
`uname = Darwin ? stat -f %Lp : stat -c %a` - is copy-pasted across at least
five test scripts plus a `reread_mode` variant, and epoch-mtime reads are
open-coded as `date -r … || stat -c %Y` in three more; same one-owner shape as
the defect above. `git init` without `-b main` depends on the host's
init.defaultBranch in several scripts (branch-name case, tracked elsewhere).
timeout, sha256sum and sed -i uses are all correctly guarded where checked.
Observed while building the check rather than the fix: the first watcher I
wrote to wait for the suite matched its own command line, so it was waiting on
its own existence and could never fire. Same shape as the cases above -
machinery answering confidently about something other than its subject, by
including itself in the evidence it was meant to judge. The file sentinel it
was replaced with cannot be produced by the observer that reads it.
* fix(review): Pin fixture timestamps to UTC across DST transitions
* fix(document): Clarify shared fixture suite coverage
* feat(pi): accept native Codex ultra effort with progress-aware supervision (#4038)
* fix(pi): preserve native Codex effort and guarded supervision
* test(pi): identify native compatibility guard versions
* no-mistakes(review): share native-main follow rule between build and picker
* no-mistakes(document): document native progress marker and ultra effort owners
* no-mistakes(ci): Failing check "Behavior portable serial 1" was caused by this PR. The shard ran the default-on live guard tests/fm-pi-branch-responsiveness-live-e2e.test.sh, whose idle arm loads .pi/extensions/fm-branch-supervision.ts into a scratch project with a fixed list of copied libs. This PR added `import { registerFirstmateTool } from "./lib/fm-native-contract.ts"` to that extension and updated every other loading fixture's copy list, but missed this guard. Pi 0.85.1 therefore refused to load the extension ("Cannot find module './lib/fm-native-contract.ts'"), never drew its TUI, the test failed with "Pi 0.85.1 never drew its TUI in the idle arm", and the job hit its 20-minute cap. Fix (one line): added fm-native-contract to the lib copy loop in tests/fm-pi-branch-responsiveness-live-e2e.test.sh. Swept all other suites referencing fm-branch-supervision.ts / fm-primary-pi-watch.ts; the remaining ones without the new lib only hash, path-reference, or string-match the files and do not load them into Pi, so no further fixture changes are needed. Verification: reproduced the mechanism against the installed Pi 0.85.1 by building the lab copy with the old lib list (load error as above) and the fixed list (loads cleanly). tmux is not installed on this machine, so the live guard itself gate-skips locally ("skip: live: tmux absent") and could not be run end to end here; CI (which has tmux and Pi) will exercise it. shellcheck is clean on the edited file
---------
Co-authored-by: Talon Stark <talonstark@gmail.com>
* fix(herdr): bypass stale clients rejected by running servers (#4041)
* fix(herdr): step around a stale client the running server refuses
A remote host can carry a self-updated herdr in ~/.local/bin beside a
package-managed one, and the fixed remote-job PATH resolves ~/.local/bin
first. After the server upgraded to 0.9.0 (protocol 22) the stale 0.8.2
client (protocol 20) was answered with protocol_mismatch on every command,
which the read classifiers folded into `unreadable`: the live remote
secondmate read unknown, every doorbell into it failed, and both the spawn
and relaunch recovery paths refused, so the defect trapped itself.
The adapter's session-scoped CLI wrapper now recognizes that refusal, reads
status per session from each distinct herdr on PATH, adopts the first one
the running server reports compatible, retries once, and keeps it for the
process. The happy path makes no extra call and no other failure reselects.
An endpoint that still reads unreadable names the refused client, both
protocols, and the fix on stderr; the remote state read, fm-crew-state, and
the launch refusal carry that reason, and fm-remote-doctor reports the
selected client and rebinds the launch agent to it.
Regression coverage: fake two-client hosts in the herdr unit suite, the
doctor suite, the crew-state remote arm, and the real host-local control
script in the remote lifecycle e2e; the real-herdr smoke refreshes the
status shape the selection reads.
* no-mistakes(review): Reselect Herdr client after every protocol mismatch
* no-mistakes(review): Remove unrequired Herdr diagnostics and launch-agent rebinding
* no-mistakes(review): Scope cached Herdr clients to their selected session
* no-mistakes(review): Restrict herdr client selection to reactive CLI calls
* no-mistakes(document): Document session-scoped Herdr client reselection
* no-mistakes(document): Clarify Herdr client selection documentation
* no-mistakes(ci): Updated the trusted fm-remote-doctor.sh SHA-256 in bin/fm-remote-entrypoint.sh after the PR changed the doctor, restoring git-unavailable bootstrap authentication. Verified tests/fm-on.test.sh, tests/fm-backend-herdr.test.sh, bin/fm-lint.sh, and git diff --check all pass
* feat: add durable AFK posture lifecycle (#4048)
* feat(afk): record the away posture and its lifecycle (phase 1)
Away mode becomes a posture of the one supervision session, recorded in
state/.afk-contract by the new bin/fm-afk-contract.sh: the one owner of the
record schema, the mandate-clause grammar and compiler, refusal naming the
missing part, the read-back rendering, the entry announcement (hold-for-return
only, no phone channel), and the archive at return. This release records
clauses and does not execute them; the announcement and return brief say so.
bin/fm-afk-launch.sh gains propose and confirm, confirms the record before any
daemon launch, refuses to launch the daemon on Pi and pi-signed, and archives
the record last on stop. bin/fm-afk-return.sh snapshots supervisor health
before shutdown, renders the return brief (health, mandate, waiting on the
captain, could not fix, handled, cost) from the archived record, the outcome
store, the held set, and the status logs, and shrinks the blocker gate to what
the away session could not fix.
While the record exists the watcher and the daemon never recheck an item held
for the captain. Declared external waits get a four-hour default cadence and
honor `until <UTC ISO 8601>` on the paused line, in both postures, bounded by
FM_PAUSE_UNTIL_MAX_SECS.
The /afk skill, AGENTS.md's layout and away-mode stub, the session-start
digest, and the architecture, Pi branch, configuration, and scripts docs
describe the record. The Pi/Herdr e2e now proves the no-daemon posture on a
real Pi primary; its verification record carries the 2026-09-08 run.
* no-mistakes(review): Fix AFK confirmation, grammar, waits, and return gating
* no-mistakes(review): Harden AFK authority and posture lifecycle
* no-mistakes(review): Preserve AFK history and tighten authority grammar
* refactor(afk): record clause fields with no natural-language parser
By the captain's mandate the away-posture record keeps no static parser
that tries to understand natural language. A mandate clause is now given
as explicit fields (--action, --object, --when, optional --stop) that
bin/fm-afk-contract.sh records verbatim. The structural check asserts
only that the action, object, and precondition fields are present and
that the action is a listed verb; whether a precondition holds is the
supervision session's judgment at execution time in a later phase.
The never-set stays as a forbidden-concept safety scan: fields mentioning
credentials, passwords, logins, legal or financial acceptance, payments,
invoices, one-time codes, or an attended prompt are refused, matched at
token prefixes after punctuation normalization so compound and plural
spellings are caught. The red-check grammar, class-word rejection,
unconditional-word detection, clause-reference resolution, and condition
aliases are removed. --words-file keeps the captain's words verbatim,
trailing newline included.
The skill, docs, launcher help, and tests describe the field form.
* no-mistakes(review): Preserve AFK words and tighten safety refusals
* no-mistakes(review): Preserve clause bytes and honor declared waits
* no-mistakes(review): Harden deny-list and gate unreadable outcomes
* no-mistakes(review): Demote never-set scan and clarify authority
* no-mistakes(review): Gate return on unreadable held and status data
* no-mistakes(review): Validate posture archives and enforce Pi detection
* fix(afk): make the never-set a non-refusing flag and keep return fail-safe
Per the captain's decision the never-set scan is a coarse best-effort
flag, never a refusal and never the gate: a clause naming a listed
concept is still recorded with a flag the read-back, announcement, and
return brief show, and the scan matches listed terms exactly or with a
plain inflection at punctuation-delimited token boundaries, so unrelated
names such as ping-service or tokenize-worker are never flagged and
joined compounds remain a documented miss. Authoritative never-set and
forbidden-action enforcement is the supervision session's judgment at
execution time in phase 4.
A replacement copies the superseded record through a temporary name and
renames it atomically so a failed copy leaves no partial archive, the
record owner gains validate and flags subcommands, and the return keeps
catch-up gated when a superseded archive cannot be read.
* no-mistakes(review): Harden AFK record validation and return reconciliation
* no-mistakes(review): Harden AFK record validation and simplify commands
* no-mistakes(review): Harden mandate validation and retain missing records
* no-mistakes(review): Refuse blank explicit mandate stops
* no-mistakes(review): Recover restored posture epoch before return
* no-mistakes(review): Prevent return brief status symlink reads
* no-mistakes(document): Refresh AFK posture documentation
* no-mistakes(ci): Fixed both CI failures: updated lint telemetry for the new fourth source directive, quoted the hyphenated fixture value, and removed unreachable test cleanup. Verified with tests/fm-lint.test.sh, targeted CI-mode ShellCheck, bin/fm-lint.sh, bash syntax checks, and git diff checks
* no-mistakes(ci): Bound structured pause deadlines by FM_PAUSE_RESURFACE_SECS in watcher and daemon housekeeping, added distinct bounded-horizon reasons, regression coverage for near, passed, and wrong-year deadlines, and updated documentation. Verified targeted behavior tests, full daemon tests, ShellCheck source-following lint, syntax, and diff checks
* feat(bin): add IMAP/SMTP mail plane with standing poll (#3765)
Opt-in IMAP/SMTP mail plane (fm-mail.sh / fm-mail-check.sh). Absent FM_MAIL_* stays off.
Speaking as Kun's firstmate: this is merged. Thank you @feilipu — really appreciate you taking the time on this.
* fix: launch remote Herdr through the user login shell (#4061)
* fix(remote): start the fm-remote Herdr agent through a login shell
Launchd was exec-ing herdr directly, so the Aqua agent inherited a background session without login-keychain access. Start it via /bin/zsh -lc exec so panes keep login env and can refresh OAuth tokens after reboot.
* fix(remote): start fm-remote Herdr via the account login shell
Resolve UserShell from Directory Services and invoke it with separate -l and -c so bash, fish, and zsh all get login-keychain access. Fall back to SHELL, then /bin/zsh, then /bin/sh without failing the render.
* no-mistakes(review): Fix launch-agent shell fallback resolution
* no-mistakes(review): Preserve and escape Directory Services shell paths
* no-mistakes(document): Document login-shell LaunchAgent behavior
* no-mistakes(ci): Updated the trusted fm-remote-doctor SHA-256 identity in bin/fm-remote-entrypoint.sh. Verified with tests/fm-on.test.sh, tests/fm-remote-doctor.test.sh, bash syntax checks, and git diff --check
* no-mistakes(ci): Resolved the login shell exactly once per doctor invocation and threaded it through plist rendering, installed/loaded contract validation, repair reporting, and post-repair checks. Added a regression test proving repeated repair remains healthy and performs no reload when a hypothetical second Directory Services lookup would differ. Updated the trusted doctor hash. Verified doctor, fm-on, remote-entrypoint, lint tests, ShellCheck, syntax, and diff checks
* no-mistakes(ci): Made Darwin shell resolution hermetic with executable injection and a 2-second Directory Services timeout. Updated tests to inject shells by default, isolate dscl-specific cases, parse plists semantically, and verify stalled dscl fallback. Updated the trusted doctor hash. Doctor, fm-on, entrypoint, syntax, hash, and diff checks pass
* no-mistakes(ci): Raised portable serial CI timeout from 20 to 30 minutes, refreshed the specified timing hints, added missing hints, and recomputed shard documentation. Verified coverage, runner behavior tests, workflow lint tests, shell syntax, requested timing maxima, and diff checks
* fix: distinguish landed deliveries from resolved captain calls (#3710)
* fix(bearings): keep captain-approved deliveries in Recently Landed
A closed task is never held: tasks-axi clears the held flag when a task
closes and keeps hold-kind and the hold reason as the record of the call
that was made. Recently Landed excluded every Done row whose hold-kind was
captain, so the marker it treated as "closed while still waiting on the
captain" was in fact the proof that the captain had approved the work. Every
merge routed through a captain decision disappeared from the list of what
shipped, including under --all-landed.
The selector now asks whether the closed row delivered something. Recently
Landed is merged PRs, completed scouts, and finished local-only merges, so a
row carrying one of those artifacts belongs there whoever approved it. A
captain question closes with an answer and no artifact of its own, and that
is what still stays out, so an answered question is never rendered as
shipped work.
The same rule was written twice - the bearings projection selects this
home's Done rows and the fleet snapshot selects each secondmate home's Done
rows into the roll-up the same section merges in - which is why one defect
hid deliveries in every home. Both now share bin/fm-landed-lib.sh.
* fix(review): Normalize landed evidence and exclude answered captain questions
* fix(review): Normalize captain delivery evidence across relocated data
* fix(review): Record authoritative delivery provenance with legacy fallback
* fix(review): Harden delivery provenance across forced and pruned completions
* fix(review): Replace premature merge closure with existing release contract
* fix(review): Document provenance-based Recently Landed selection
* fix(document): Align documentation with completion provenance
* fix(lint): Fix targeted ShellCheck warnings
* fix(ci): order the pinned tasks-axi install before its stock-Bash consumers
In `.github/workflows/ci.yml` the pinned tasks-axi install now precedes both
stock-Bash consumers, and the Bearings expectation is updated from 49 to 50
tests.
Verified with macOS Bash 3.2: snapshot 16/16, Bearings 50/50, public-followup
1/1. Full repository lint and all three workflow validations pass, and
`git diff --check` is clean.
* fix(review): Make completion provenance unambiguous
* fix(review): Make completion verdict authoritative over quoted provenance
* fix(review): Preserve retained artifacts through resumed captain closes
* fix(review): Unify completion provenance ordering across writer and reader
* fix(review): Preserve artifacts across failed captain closes
* fix(review): Refresh v1 assertions; provenance authority remains unresolved
* fix(review): Remove unreliable provenance while preserving landed deliveries
* fix(review): Reject stale home summaries visibly
* fix(review): Restore retained deliverable recording
* fix(review): Match landed artifacts and restore retention documentation
* fix(review): Disambiguate captain calls and restore landed artifact matching
* fix(review): Persist retained report and PR artifacts
* fix(review): Preserve staged artifacts before captain answers
* fix(review): Avoid wedging answers on unsupported report paths
* fix(review): Exclude unreleased captain-held pull requests
* fix(review): Exclude held local-only answers from landed
* fix(review): Preserve retained scout reports across snapshot rendering
* fix(document): Align landed lifecycle documentation with release semantics
* fix(review): Enforce landed artifact-kind ownership
* fix(review): Infer canonical task kinds in snapshots
* fix(review): Require captain-hold release before merges
* fix(review): Qualify merge lifecycle regression evidence
* fix(review): Serialize captain holds with merge operations
* fix(review): Document merge cleanup residuals honestly
* fix(test): Replace vacuous Bearings regression with behavioral cases
* fix(document): Align Bearings verification and merge lifecycle documentation
* fix(review): Serialize merges and exclude captain calls from landed
* fix(review): Harden merge identity and landed selection
* fix(document): Clarify landed selector compatibility filtering
* fix(bin): keep merge entrypoints usable on records without an incarnation
The merge identity guard refused any task record with no spawn_gen field.
That field identifies one exact incarnation, so comparing it across the wait
for the merge lock is what catches a task relaunched while the merge was
queued. Requiring it to be present is a different rule, and it refused every
record written before the field existed: a legacy task could no longer be
merged at all, and five behaviour suites refused before reaching the check
they were written to exercise.
The comparison only needs to notice a change. An absent field is now read as
an empty incarnation and compared like any other value, so a record that
gains, loses, or alters one is still refused, while a record that simply
predates the field merges. An ambiguous or unreadable field stays an error,
because a record that cannot name one incarnation cannot be compared. The
missing-record message each entrypoint had before the guard is restored, so
a genuinely absent record still says so in its own words.
The role partition now precedes reading the record. Refusing the supervision
branch is a statement about the actor, not about the task, so it cannot
depend on a record the wrong actor may not have.
A backlog file that does not exist meant "no longer an open captain call".
For a caller that asked to tell absence apart it now means absent, so a board
card whose home carries no backlog stays visible instead of being dropped as
resolved.
Fixture repositories pin their initial branch instead of inheriting
init.defaultBranch, which resolved to main on a developer machine and master
on a runner, so a fixture naming main failed only in CI.
* fix(review): read local-only note from body; surface pending-close failures
* fix(review): keep kindless local-only landings in Recently Landed
* fix(review): bind local-only note scan to the tasks-axi note line
* fix(review): Guard unavailable captain-hold authority records
* fix(document): Document unreadable authority predicate outcome
* fix(bin): read an absent backlog as absence, not an unreadable record
The merge gate refused every task whose home carries no backlog file. A
backlog that does not exist holds no captain call, so nothing can be held and
the merge is safe; only a backlog that exists and cannot be read may hide a
live hold. Those two states were collapsed into one refusal, which stopped
merges in any home that keeps no backlog.
The predicate now reports a missing backlog file as absence, alongside a row
the backlog does not carry. A record that exists but cannot be read still
leaves by the existing cannot-tell path, which both merge entrypoints already
refuse, so the restrictive direction is unchanged.
That leaves no way to reach the separate unavailable-record result, so the
result and the two branches that handled it are removed rather than left
describing an outcome that can no longer occur. The lifecycle documentation
loses the same claim.
Regressions cover both directions in each entrypoint: a home with a task
record and no backlog merges, and a backlog present but unreadable refuses
without reaching the forge.
* fix(review): Fail closed unreadable backend configuration
* fix(tests): pin the bare origin's initial branch in the remote seed fixture
The fixture created its bare origin with no initial branch, so that
repository's HEAD followed init.defaultBranch while the source repository
pushed the branch fm_git_init_commit pins. On a host that still defaults to
master the two disagreed: the bare origin's HEAD named a branch the push never
created, cloning it warned that the remote HEAD referred to a nonexistent ref
and checked out nothing, and the seed assertion for the cloned README failed.
A machine whose default is already main paired the two by accident and hid it,
which is why the fixture passed locally and failed on the runner.
Pinning the bare origin to the same branch removes the dependency on the
ambient default from both sides. Verified under both conditions: with
init.defaultBranch set to master, and set to main, the suite passes 26 of 26.
* fix(review): Fail closed unreadable user backend configuration
* fix(bin): republish the home summary as v1 and record two load-bearing rules
The published home-summary schema had moved to v3, which routed every
secondmate home still emitting the earlier version to the stale branch: their
landed rows, open decisions and holds all came back empty and their state read
as unknown until each home was updated. The payload never justified that. Its
field set, field order, truncations and the landed array construction are
byte-identical to v1, so only which rows the selector places in landed
differs, and a v1 consumer reads that the same way.
Republishing as v1 removes the rollout regression and, with it, the tolerance
machinery that existed only to soften the bump: the stale-schema predicate,
its two collection branches, the flag and its provenance branch, the omitted
surface that can no longer be reached, and the fixtures and assertions that
covered them.
Two rules that a scope review proposed removing are kept, each now carrying
the reason it exists, because both were measured to be load-bearing:
The artifact-kind ownership clause is what keeps an explicit scout that
recorded no report out of Recently Landed. Without it such a row has none of
the three artifacts, satisfies the compatibility fallback and renders as
shipped work with an empty artifact.
The kind fallback is needed because tasks-axi omits the kind metadata
entirely when a title begins with a canonical keyword. Without it a scout
titled "SCOUT ..." reports no kind, its recorded report stops counting as a
delivery, and it drops out of the section this selector exists to repair.
* fix(bin): move the scout guard note onto the rule and drop two dead pieces
The LOAD-BEARING note sat on an unreachable branch. Measured in both
directions: removing that branch together with the kind-is-not-scout guards
lets an explicit reportless scout into Recently Landed and fails
tests/fm-captain-hold-lifecycle.test.sh, while removing the branch alone
leaves that suite passing at 49 assertions. The guards carry the rule, so the
note now sits on them and the unreachable branch is gone. A note pointing a
later reader at the wrong line is the hazard this change corrects elsewhere.
summary_file_has_schema lost its only caller when the stale-schema machinery
was removed, so it goes with it.
* fix(review): Fix legacy report artifacts and canonical keyword boundaries
* fix(review): Update pinned Bearings test count to 56
* fix(document): Clarify landed summary compatibility documentation
* fix(review): Preserve unreadable backend configuration errors
* fix(review): Honor backend resolution errors at existing call sites
* fix(test): Stabilize remote collector tests under host load
* fix(document): Document backend resolution failure contracts
* fix(lint): Suppress intentional deferred probe expansion warnings
* fix(ci): Captain, quoted the two literal test IDs in tests/fm-backlog-atomicity.test.sh to fix SC2100 without changing behavior. Both warnings reproduced before the fix; the targeted fm-lint.sh run now passes with ShellCheck 0.11.0. Bash syntax and git diff --check also pass
* fix(ci): Fixed the resolver’s two configuration-parent checks to return 2 for inaccessible directories while preserving genuine absence. Added two behavioral tests; RED/GREEN and both requested mutation proofs confirmed. All 10 focused checks, targeted lint, syntax, and whitespace checks passed. Broader merge suite stopped after 10 passing cases under host load. Declined portable checks and merge-authority code remain unchanged
* fix(remote): keep fm-remote Herdr servers in the Aqua session (#4090)
* fix(remote): let the Aqua launch agent own the fm-remote Herdr session
A herdr server keeps the macOS audit session of whatever started it, and
only the Aqua login session (gui/<uid>) can read the login keychain
without a prompt. Herdr's SSH remote attach starts the fm-remote server
as its own child when it finds none, wins the socket at boot because sshd
accepts connections before the login session exists, and every claude
pane under that server then gets `security` exit 36, falls back to a stale
plaintext credentials file, and reports "Login expired". launchd's own job
lost the socket on every KeepAlive retry and the doctor still reported the
session ready because it only asked whether any server answered.
- Add bin/fm-remote-herdr-guard.sh, the launch agent's exec target: start
the server in the foreground when nothing owns the socket, exit 0 when an
Aqua-born server does, and otherwise stop the foreign server, wait for the
socket, and exec the server at once.
- Add bin/fm-remote-herdr-owner-lib.sh, the single owner of socket-owner
discovery (lsof; pgrep cannot see herdr's argv on macOS) and the birth
markers (SSH_*, XPC_SERVICE_NAME, FM_REMOTE_JOB_ACTIVE, sshd or
remote-client-bridge ancestry matched on argv[0] and whole arguments).
- Render the agent as the login shell exec'ing the guard with
KeepAlive={SuccessfulExit=false} and ThrottleInterval=10, check the loaded
job's successful-exit semaphore, and report a session served outside the
Aqua login session as fixable so --fix retakes it through launchd; the
reload waits for an Aqua-born owner rather than any running server.
- Correct the doctor and docs: the launch shell provides environment parity,
the launchd domain provides keychain access.
- Pin the guard's decision table and the doctor's verdicts against real
marker-carrying processes, and record the dated audit-session evidence.
* no-mistakes(review): Verify Aqua ownership through launchd domains
* no-mistakes(document): Document macOS lsof ownership requirement
* fix(bin): prefer a live no-mistakes run over a terminal one (#2881)
* fix(bin): prefer a live no-mistakes run over a terminal one
A worktree can bind to more than one recorded no-mistakes run at once.
The branch-and-code-identity rule in bin/fm-nm-run-lib.sh accepts both an
exact-equal commit and a worktree-is-an-ancestor match, but never stated
which wins when both bind, so the tie fell to whichever candidate the
caller reached first.
Observed on a live fleet: a crashed validation daemon left a FAILED run
at the worktree's own commit while the live run that replaced it
validated a descendant commit on the same branch. Bare `axi status`
answers with the most-recently-touched run - the corpse - and it bound by
the equal-commit rule, so every recomputation reported `failed` for a
task whose real run was healthy. The same label had also read `failed`
earlier while the work was genuinely stalled, so the signal was wrong in
both directions.
State the live-over-terminal policy in the matching rule's own contract,
where the equal-commit and ancestor rules already live, and add
fm_nm_run_status_class as the one classifier that decides liveness from a
recorded status word. fm-crew-state.sh applies it on both selection
paths: the runs listing now scans past a terminal row for a live one, and
a terminal `axi status` answer is provisional until the listing has been
asked whether this worktree also has a live run.
Same-liveness-class candidates keep the listing's newest-first
precedence, and a status word the classifier cannot place keeps the
caller's own ordering rather than displacing a known result, so a
single-run task and a task whose runs are all terminal are unchanged.
Regression coverage reproduces the proven case (terminal run at the
worktree's exact commit plus a live run descending from it) and its
runs-list twin; both fail under the old tie-break. Two companion cases
pin the no-widening half - two terminal rows still resolve newest-first,
and a terminal run with no live sibling keeps its full run-step detail -
and both pass before and after the change.
* no-mistakes(review): accept unfetched live sibling anchored at exact worktree head
* docs(bin): name both ledger reads behind the runs-limit setting
The FM_CREW_STATE_RUNS_LIMIT comment in bin/fm-crew-state.sh still described
the runs ledger as scanned only by the cross-branch fallback, but the
live-over-terminal fix also consults it as the live-sibling probe behind a
terminal axi status answer. Point the comment at docs/configuration.md as the
setting's owner instead of restating a second copy.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* fix(bin): repair process-event shutdown and tighten guard timing (#4009)
* fix(procevent): make the ordinary stop signal actually stop a runner
The owner guard that shipped in #3904 half-reaps. Against a poll child that
handles the ordinary stop signal and keeps waiting, the guard signals the group,
loses the runner leader to its own signal, then reads that success as a
leaderless group and exits without escalating. It destroys the only proof of
ownership that would have authorised the forced signal, so the survivor becomes
unreachable by retire, reconcile, sweep-home and the guard alike. A guard that
turns a leaking-but-identifiable generation into a permanently unreachable one
is worse than no guard at all.
Two defects, and they hid each other:
- The escalation re-derived ownership from the leader. `runner_group_signal`
now takes a `proved` mode, passed only by the escalation inside the stop that
already proved and signalled that exact generation moments earlier. A leader
dying to our own signal is the ordinary outcome, not fresh ambiguity.
- Every stop held the per-source lock across its wait while the runner's own
exit cleanup waited unboundedly for that same lock. That circular wait was
broken only by the forced signal, so the forced signal silently became the
normal path - and, by keeping the leader alive through the whole window, it
masked the escalation defect above. The runner's exit cleanup now refuses that
lock instead of waiting for it, which is what its existing `return 0` already
said it did.
Fixing the lock alone would have turned every stop of a signal-proof child into
a refusal that leaves it running, so both land together and the tests pin that.
Measured on macOS with a stand-in poll child that traps TERM, INT and HUP:
the guard left it running past 70s and now clears the group within the lease
plus one check; retiring a healthy runner fell from ~2.8s with a forced group
signal every time to ~0.6s on the ordinary signal alone.
Unchanged and stated deliberately: a leader lost to anything other than the
stop's own signal still leaves a group that retire, reconcile, sweep-home and
the guard all refuse, permanently - and that source stops listening without
saying so. Whether such a group may ever be signalled is an open decision and
is not answered here.
* fix(review): Fix proved escalation race and stop regression assertions
* fix(review): Preserve proved escalation through transient identity failures
* fix(review): Simplify proved escalation and correct guard timing documentation
* fix(document): Clarify process-event stop ownership and cleanup limits
* fix(document): Clarify process-event stop ownership and fixture comments
* revert(skills): restore the leaderless-ambiguity limit to the loaded skill
An automatic documentation step in this branch's validation edited
.agents/skills/process-event-sources/SKILL.md, which no instruction in this
change asked it to touch. That file is not documentation about the code: it is
the agent-loaded instruction surface, what an agent reads to know what it is
permitted to do.
The step deleted this line:
- leaderless PID/PGID-reuse ambiguity preserves the claim without signalling
or replacement, as owned by the operating contract in
[`docs/configuration.md`](../../../docs/configuration.md#process-to-event-sources-stateprocevent);
and folded it, with its neighbour, into a generic "registration and ownership
transitions, stop authority, and claim reclamation follow the operating
contract".
That deleted line states a PROHIBITION - that such a group is preserved WITHOUT
SIGNALLING - and it is the exact limit an open captain decision currently rests
on. Folded into a pointer, an agent reading the skill to learn what it may do
would have to chase a second document to discover it may not signal. A
prohibition that requires a second lookup is not a prohibition. The effect was
to weaken, in the instructions themselves, the boundary that keeps one home from
signalling another's process group - while the question of whether that boundary
should move at all is still open.
This is a deliberate revert, not an oversight, and it restores the file exactly
to its pre-branch state. The full statement also survives in
docs/configuration.md; that does not rescue it, because the agent handling a
process-event wake loads the skill and not the documentation.
* revert(procevent): restore the open-question marking beside the escalation
The same automatic documentation step that edited the loaded skill also removed
this from the comment above runner_group_signal:
A leaderless group nobody in this call ever proved remains refused too, for
every caller. That untouched refusal is what makes a crashed leader's group
permanent, and relaxing it is a separate open question, not something this
path assumes.
and replaced it with a pointer to docs/configuration.md.
This one fails differently from the skill deletion, which is why it is restored
separately. There, a prohibition was moved out of the reader's path, and a
missing prohibition gets violated. Here the prohibition survives in code - the
unproved path still refuses - and what was removed is the fact that the limit is
UNDECIDED. A prohibition that has quietly lost its "this is still open" reads as
settled design, and settled design gets relied on, extended, and eventually
relaxed by someone confident they understand why it is there. That question is
open right now.
The rule this branch's four instances produce, stated once here because this is
the point of decision: an unresolved question must be marked unresolved AT THE
POINT OF DECISION, not only where the contract is documented. A reader who does
not know something is open will treat it as closed, and that default is stronger
than any pointer overcomes.
The pointer added by that step is kept alongside; this restores what it replaced
rather than reverting it.
* docs(verification): restore the measured guard bound and its reason
The document step's rewrite of this record dropped the concrete figure while
keeping the surrounding measurements. What went missing was the bound itself -
lease plus two consecutive failed checks plus the stop's grace, roughly 630
seconds at the shipped 600-second lease and 15-second check - together with the
reason there are two checks rather than one: a single unreadable read must not
be enough to kill a live runner.
The mechanism survived elsewhere and the reason survived in
docs/configuration.md, so nothing was lost from the repository. The concreteness
was, and that is what this restores. A number recorded without why it is that
number is the one a later reader shortens; the reason is the whole safety
argument for the debounce, and the debounce is what stops the reaper killing a
live runner on one bad read.
* fix(document): Replace stale stop-authority summaries with owner pointers
* test(procevent): make the guard-bound case able to fail for its own reason
An automated reviewer observed that this case allowed sixty seconds for a bound
of roughly eight, so it could not go red for the reason it names: it would have
passed a guard that took fifty-five seconds. That is correct, and it is the same
family as the defect the case exists to defend against - a check that is green
because it cannot fail, rather than because the thing it guards is working.
The deadline is now derived from the bound itself - the lease, plus the two
consecutive failed checks the guard debounces on, plus the stop's own ordinary
and forced signal windows - rather than from a flat wall-clock number, and the
shortened lease and check the fixtures run under have a single definition so a
derived deadline cannot silently diverge from the settings the guard is given.
The doubling that remains is a load allowance and is documented as one; widening
it to make a slow guard pass would convert the assertion back into decoration.
Proven by mutation rather than by argument. Against the repaired case:
correct code ok
guard debounces on 20 misses instead of 2 not ok - "still holding
the group after 16s,
against a documented
bound of 8s"
proved escalation removed (the original defect) not ok - same
code restored ok
The previous sixty-second version passes every one of those mutations.
The reviewer's other claim, that the guard can survive past the announced bound
when an owner disappears immediately after a check, was measured and does not
hold against what this branch announces. Sweeping the phase deliberately at
0.0, 0.2, 0.4, 0.6 and 0.8 of a check interval gave 7.21s, 7.31s, 6.75s, 6.49s
and 6.31s, worst 7.31s, against the announced lease plus two consecutive failed
checks plus stop grace, which is up to 8s at those settings. The mechanism the
reviewer describes is real and is the announced mechanism; the bound it was
measured against is a phrasing this branch no longer carries.
* revert(scope): return the instruction surfaces to their base state
This delivery is being split. It carries the two proven process fixes alone; the
instruction text travels separately, through a run that removes the
documentation step rather than refusing it at its gate.
Two surfaces are therefore returned to exactly what the base branch has, so this
delivery neither adds to them nor removes from them:
.agents/skills/process-event-sources/SKILL.md - identical to base again. Three
bullets an automatic documentation step had folded into a pointer, including
that leaderless PID/PGID-reuse ambiguity preserves the claim WITHOUT
SIGNALLING and that there is one identity-matched owner per canonical source
across homes sharing one store.
The header comment block of bin/fm-procevent.sh, which is what the script
prints as its own help. Seven lines were removed from it: that a live owner is
never displaced, that only a claim whose stale owner and independently absent
process group prove its whole generation gone is reclaimed, that a crashed
leader or reused pid whose process group still has members cannot relax
ownership cleanup, and that reconcile signals only a live identity-matched
runner group and otherwise keeps the claim without starting a replacement.
The help output is now byte-identical to base.
Neither removal was requested by any instruction in this change, and both were
made to text that predates it. Returning them is scoping, not a third
restoration: nothing is being added to those files here.
* fix(ci): Captain, live CI revealed a fixture deadlock: it suspended the runner before startup released its lock. Added a public-list synchronization barrier in tests/fm-procevent.test.sh. Forced-delay reproduction detected the deadlock before the fix; all four cases passed afterward. Targeted lint, Bash syntax, and whitespace checks passed. Greptile’s watchdog requirement conflicts with the recorded R2 decision; runtime behavior and documentation remain unchanged. Full CI rerun belongs to the outer executor
* test(procevent): make the post-TERM cases report what they saw when they fail
On the failure path only, these cases now print what they actually saw: the
identity recorded at claim time, the identity readable at that moment, the size
of the signals file, the leader's state and wchan, every live member of the
runner's process group with its own state and wchan, the elapsed time since the
stop began, and what retire said. None of it runs when a case passes.
WHY THIS IS KEPT, stated accurately rather than by its original reason. It was
written to make an unexplained CI failure verifiable. That failure is now
explained - it was a fixture deadlock, diagnosed and repaired in the preceding
commit - so that justification has expired and is not the reason given here.
The reason it stays is smaller and independent of that failure: it is already
written, it is small, it sits in the file whose assertion this change reworked,
and an assertion that could not say why it failed cost most of a morning to
diagnose from the outside. The next failure will not be this one.
WHAT A PASSING RUN WOULD NOT MEAN: a pass is a sample of behaviour already
observed many times, not proof that anything is fixed. Only a failure carrying
the evidence above establishes a cause.
* fix(document): Clarify process-event fixture diagnostic rationale
* fix(ci): Captain, fixed two cleanup races in tests/fm-procevent.test.sh: removed premature child completion and waited for runner exit before retiring the restart fixture. Controlled Linux reproductions demonstrated failure before and success after. The full Linux process-event suite, six focused macOS checks, targeted ShellCheck, Bash syntax, and whitespace checks passed. Runtime behavior, guard debounce, and documentation remain unchanged. CI rerun belongs to the outer executor
* fix(procevent): bound owner-guard cleanup at one check interval, not two
A THIRD WAY, not a capitulation to the reviewer and not a refusal of it.
The automated reviewer's grievance was the LOOSENESS OF THE BOUND, not the
number of observations the guard makes before it acts. It asked for a single
read because that was the only route it could see to an acceptable bound. There
was another route, and this change takes it: the bound is reached and both reads
are kept.
TIGHTENED - the SPACING of the guard's two reads, not their number. The owner
watchdog now sleeps half the configured check interval and still requires two
consecutive failing reads, so the pair completes inside one check interval
instead of costing two. Worst-case detection falls from the lease term plus TWO
check intervals to the lease term plus ONE. At the shipped 600s lease and 15s
interval the stated bound falls from ~635s to ~620s.
PRESERVED - the second read. bin/fm-procevent.sh's two-consecutive-miss rule is
untouched. WHY IT PROTECTS: the guard's inputs are a lease read and a state-root
identity read, and either can fail transiently on a live, healthy home. Acting
on the first failure would let one isolated unreadable read kill a live service.
Requiring a second, independent read is what makes that impossible, and it is a
protection rather than padding. Nothing was traded away to reach the bound.
Both properties are now guarded by their own case, and each was proven by
MUTATION rather than asserted:
- putting a full interval back between the two reads fails the bound case:
"still running 17.0s after the last owner activity, against a documented
bound of 15s";
- acting on one failed read fails the new debounce case: "one unreadable lease
read ended a runner whose home was still alive" - while the bound case then
passes FASTER, 9.9s against 13.1s. The unsafe variant being the quicker one
is exactly why these are two cases: one elapsed-time case would have
registered the removal of the protection as an improvement.
MEASURED, sampling the phase between the guard's check clock and the lease clock
across eight runs per variant, on macOS (Darwin 25.5.0). Reaping an orphaned
listener whose home stopped refreshing its lease:
lease 2s / interval 1s: 4.41-5.29s before, 3.48-4.65s after
lease 2s / interval 4s: 7.69-8.12s before, 5.94-6.13s after
The 4s configuration is the informative one: the gap is about one check
interval, which is precisely the term that was removed.
A previously unstated term of the bound surfaced while measuring: the lease age
is compared in whole seconds, so a configured lease of N is honoured until that
age reads N+1. It is now part of the documented bound and of the regression's
derivation instead of being absorbed into a fudge factor.
The bound regression derives its deadline from the documented bound instead of a
flat number, and PINS the phase between the guard's check clock and the lease
clock rather than sampling it, because with a sampled phase a guard spending two
intervals passes about half the time on a lucky alignment. Its load slack is
additive and stays under half a check interval, so an extra whole interval
cannot hide inside it. The two flat deadlines that were there before (40s and
20s) and the doubling allowance on the derived one are gone; that looseness was
the reviewer's third complaint.
The stop's own grace is untouched: 2s for the ordinary signal, then 2s for the
forced one. It is a ceiling paid only by a group that outlives the signal it was
sent, not a delay every stop pays - a healthy runner's whole retire measures
0.40-0.66s on this host. The reviewer's literal "lease plus one tick" is
unreachable by any implementation, since signalling a process and giving it any
chance to exit takes non-zero time; detection now meets it and the stop runs
inside its own ceiling, and the contract says so rather than glossing it.
NECESSARY BUT NOT SUFFICIENT, and written BEFORE this head's integration runs
start rather than after they report. On the previous head, "Behavior portable
serial 1" and "Behavior portable serial 4" were both CANCELLED at the job
ceiling, independently of this finding. A new head triggers fresh runs, so those
two lanes MAY complete this time. IF THEY DO, THAT IS NOT EVIDENCE THE CEILING
DEFECT IS FIXED. It is one more sample of a lane that has been cut repeatedly
and sometimes is not; the shard-packing repair for it is open separately. Do not
reread a lucky pass here as a resolution.
Relatedly, and deliberately: the per-script duration hint in bin/fm-test-run.sh
was NOT updated even though the two new cases add ~19s of wall clock.
docs/fm-test-portable-shards.md says those hints are replaced wholesale from CI
timing artifacts of green runs, and that repair is the open request doing it; a
hand-edited estimate here would collide with it and silently repack the shards.
This suite runs in portable serial shard 3, which was green in the last run.
Verification: tests/fm-procevent.test.sh green, plus
tests/fm-captain-hold-lifecycle.test.sh, the test-coverage guard, and
bin/fm-lint.sh. The unrelated "reconcile stops a runner whose registration was
removed" case flaked in 4 of 7 local full runs; an isolated 20-trial
reproduction measured it at 13/20 unclean before this change and 11/20 after, so
it is issue 4080 and is not aggravated here.
* fix(procevent): repair our decimal-interval regression and enforce the timing phase
REPAIRED BEFORE PUBLICATION, AND IT WAS OURS. The half-interval arithmetic added
by the previous commit read a zero-prefixed interval as octal: 010 halved to 4
instead of 5, and 08 was not a number at all, so the owner guard died before
reporting ready and the runner failed closed and never listened. The validator
accepts those values and `[` compares them as decimal, so this broke a
configuration that worked before. Introduced by this delivery, found in review,
repaired here. Forcing base ten before the arithmetic is the whole runtime fix.
Proven by driving it rather than by reading the source: a new case starts a real
listener at 08 and at 010 and observes the guard's actual sleep argument - 4s and
5s. Removing the normalisation turns that case red with "a zero-prefixed decimal
interval (08) prevented the listener from starting".
THE TIMING PHASE IS NOW OBSERVED AND ENFORCED, NOT ASSUMED. The bound case
pinned its phase by CONSTRUCTION, from an assumed startup time, and enforced
nothing. Review was right that this is not enough: once startup reaches about two
seconds the expiry lands in a different part of the interval and the case
silently stops rejecting a two-interval guard while still reporting success. A
bound that cannot fail for the reason it names is the defect this whole delivery
exists to correct, so it must not ship inside the fix for it.
Now the lease is synchronised to the guard's own FIRST observed lease read,
every later real read is recorded, and the case REFUSES unless one recorded read
proves the required phase: it read the synchronised reference, it was still
fresh, and it began late enough that two further full intervals could not finish
before the deadline. An unestablished precondition refuses; it does not proceed
on trust. The derived deadline, the two-read debounce and the additive slack are
unchanged, and the slack invariant is now asserted rather than left to a comment.
Review also found the deadline was only ever checked while the group was still
alive, so a sampler descheduled past it would see the group gone and certify
success. The observed completion time is now checked too.
PROVEN BY MUTATION, each one run against this code:
- remove the decimal normalisation -> the interval case fails on 08;
- a full interval between the two reads -> "the guard exceeded its bound:
group still running 17.1s ... against a documented bound of 15s";
- a full interval WITH startup forced to ~2.5s, which is exactly the condition
the old construction pin could not survive -> still red, same message;
- the same ~2.5s startup with the correct guard -> still passes, 13.0s against
the 15s bound, so the delay alone does not break the case;
- phase evidence made unavailable -> "could not establish the required
pre-expiry guard-read phase", a refusal rather than a pass, even though the
group stopped quickly;
- act on one failed read -> the debounce case fails and the bound case passes
FASTER, 9.5s against 12.7s, which is why these remain separate cases.
Verification: full tests/fm-procevent.test.sh green, and bin/fm-lint.sh clean.
* fix(document): Correct process-event timing and debounce comments
* fix(backlog): route lifecycle transitions through configured adapters (#3417)
* fix(backlog): honor configured task adapters
* no-mistakes(review): Harden backend purity lint against prefixed Beads calls
* no-mistakes(document): Document configured backend lifecycle transitions
* fix(backlog): preserve markdown exemptions
* no-mistakes(review): Enforce backend purity for explicit lint paths
* no-mistakes(document): Update lifecycle backend documentation
* no-mistakes(lint): Remove redundant backend lint pattern
* fix(backlog): close adapter routing gaps
* no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint
* no-mistakes(document): Document environment-selected backlog adapters
* no-mistakes(lint): Fix empty local variable assignment
* fix(backlog): close quoted path gaps
* no-mistakes(review): Reject partially quoted direct Beads commands
* no-mistakes(document): Align lifecycle documentation with configured adapters
* test(backlog): keep structural cases markdown-only
* fix(backlog): honor configured task adapters
* no-mistakes(review): Harden backend purity lint against prefixed Beads calls
* no-mistakes(document): Document configured backend lifecycle transitions
* fix(backlog): preserve markdown exemptions
* no-mistakes(review): Enforce backend purity for explicit lint paths
* no-mistakes(document): Update lifecycle backend documentation
* no-mistakes(lint): Remove redundant backend lint pattern
* fix(backlog): close adapter routing gaps
* no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint
* no-mistakes(document): Document environment-selected backlog adapters
* no-mistakes(lint): Fix empty local variable assignment
* fix(backlog): close quoted path gaps
* no-mistakes(review): Reject partially quoted direct Beads commands
* no-mistakes(document): Align lifecycle documentation with configured adapters
* test(backlog): keep structural cases markdown-only
* no-mistakes(review): Harden markdown lifecycle routing and close recovery
* fix(lint): catch dollar-quoted beads commands
* no-mistakes(review): Drop unused authorized-data-dir param; canonicalize fm-lint ROOT with pwd -P
* no-mistakes(document): Document fm-lint backend-purity check in header and CONTRIBUTING
* no-mistakes(ci): Two failing checks, one code-caused and one attestation-only. 1. Greptile Review (P1: ANSI-C quoting bypasses lint) — REAL DEFECT, FIXED. The backend-purity normalizer in bin/fm-lint.sh (invokes_bd awk function) stripped $'...' quote delimiters without decoding ANSI-C escapes, so a core script containing $'\x62\x64' close fm-example executes as `bd close` while the lint accepted it. Fix: the normalizer now tracks whether a single-quote context came from $' (ansi flag) and decodes ANSI-C escapes while building the command word: \xHH (1-2 hex), \uHHHH / \UHHHHHHHH (4/8 hex), \NNN (1-3 octal), \cX control characters, simple escapes (a/b/e/E/f/n/r/t/v decode to a placeholder that can never spell bd), NUL (value 0) truncates the word per bash C-string semantics, and unknown escapes drop the backslash per bash. Values outside printable ASCII decode to a placeholder so they can neither falsely match nor collide into `bd`. Regression coverage added to the existing lint-interface test test_rejects_direct_beads_cli_invocations in tests/fm-lint.test.sh: $'\x62\x64', $'\142\144' (octal), and b$'\x64' (split word), each written into a fixture and asserted rejected through the real fm-lint.sh executable. Verified regression property: with the fix stashed, the new hex case passes lint (reproduces the reported bypass); with the fix, it is rejected. Verification: all 29 tests in tests/fm-lint.test.sh pass, and the full CI-parity lint (CI=true bin/fm-lint.sh: ShellCheck 0.11.0 full analysis over the canonical set, backend-purity check, and actionlint workflow validation) exits 0. One iteration was needed because an awk comment containing a literal $'\x62\x64' sample broke shell-level quoting (bash -n / SC1001/SC2026); the comment was reworded without quotes. 2. PR must be raised via no-mistakes — NOT code-caused. The check failed with 'Pipeline attestation head_sha does not match the current PR head': the PR body attestation binds to 82b41c7 while the PR head is 1995cdb because a later pipeline push moved the head. This is exactly the stale-attestation condition the user intent describes; it clears when the outer pipeline re-runs 'git push no-mistakes' and re-binds the attestation to the new head (which now includes this Greptile fix). No code change can or should address it. Files changed: bin/fm-lint.sh (ANSI-C escape decoding in the backend-purity normalizer), tests/fm-lint.test.sh (three encoded-bd rejection cases)
* no-mistakes(review): fix tasks.toml hang, root authorization, lint quoting
* no-mistakes(review): validate tasks config before exemption; fix lint quote gap
* fix(backlog): address the markdown backlog as <data>/backlog.md
Resolving the markdown backlog through a configured `[markdown] path` was
scope this task never asked for. It is absent from main, which addresses
`<data>/backlog.md` everywhere, and it came from an earlier review round
rather than the task brief.
Making it effective on the transition path alone put that path at odds
with every other consumer of the same backlog - fm-captain-hold.sh,
fm-session-start.sh, fm-fleet-snapshot.sh, fm-inbox.sh,
fm-backlog-handoff.sh - which all still address `<data>/backlog.md`. In
fm-captain-hold.sh the split was live: its reads had already moved to the
shared gate while its writes had not, so the two could address different
files.
Address `<data>/backlog.md` from the shared gate, delete the unused
resolver, and drop the two tests that pinned the withdrawn behaviour.
What this task actually changes is unaffected: a configured non-markdown
adapter is still addressed by its own root, without `--file`.
* fix(backlog): honor configured task adapters
* no-mistakes(review): Harden backend purity lint against prefixed Beads calls
* no-mistakes(document): Document configured backend lifecycle transitions
* fix(backlog): preserve markdown exemptions
* no-mistakes(review): Enforce backend purity for explicit lint paths
* no-mistakes(document): Update lifecycle backend documentation
* no-mistakes(lint): Remove redundant backend lint pattern
* fix(backlog): close adapter routing gaps
* no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint
* no-mistakes(document): Document environment-selected backlog adapters
* no-mistakes(lint): Fix empty local variable assignment
* fix(backlog): close quoted path gaps
* no-mistakes(review): Reject partially quoted direct Beads commands
* no-mistakes(document): Align lifecycle documentation with configured adapters
* test(backlog): keep structural cases markdown-only
* fix(lint): catch dollar-quoted beads commands
* no-mistakes(review): Drop unused authorized-data-dir param; canonicalize fm-lint ROOT with pwd -P
* no-mistakes(document): Document fm-lint backend-purity check in header and CONTRIBUTING
* no-mistakes(ci): Two failing checks, one code-caused and one attestation-only. 1. Greptile Review (P1: ANSI-C quoting bypasses lint) — REAL DEFECT, FIXED. The backend-purity normalizer in bin/fm-lint.sh (invokes_bd awk function) stripped $'...' quote delimiters without decoding ANSI-C escapes, so a core script containing $'\x62\x64' close fm-example executes as `bd close` while the lint accepted it. Fix: the normalizer now tracks whether a single-quote context came from $' (ansi flag) and decodes ANSI-C escapes while building the command word: \xHH (1-2 hex), \uHHHH / \UHHHHHHHH (4/8 hex), \NNN (1-3 octal), \cX control characters, simple escapes (a/b/e/E/f/n/r/t/v decode to a placeholder that can never spell bd), NUL (value 0) truncates the word per bash C-string semantics, and unknown escapes drop the backslash per bash. Values outside printable ASCII decode to a placeholder so they can neither falsely match nor collide into `bd`. Regression coverage added to the existing lint-interface test test_rejects_direct_beads_cli_invocations in tests/fm-lint.test.sh: $'\x62\x64', $'\142\144' (octal), and b$'\x64' (split word), each written into a fixture and asserted rejected through the real fm-lint.sh executable. Verified regression property: with the fix stashed, the new hex case passes lint (reproduces the reported bypass); with the fix, it is rejected. Verification: all 29 tests in tests/fm-lint.test.sh pass, and the full CI-parity lint (CI=true bin/fm-lint.sh: ShellCheck 0.11.0 full analysis over the canonical set, backend-purity check, and actionlint workflow validation) exits 0. One iteration was needed because an awk comment containing a literal $'\x62\x64' sample broke shell-level quoting (bash -n / SC1001/SC2026); the comment was reworded without quotes. 2. PR must be raised via no-mistakes — NOT code-caused. The check failed with 'Pipeline attestation head_sha does not match the current PR head': the PR body attestation binds to 82b41c7 while the PR head is 1995cdb because a later pipeline push moved the head. This is exactly the stale-attestation condition the user intent describes; it clears when the outer pipeline re-runs 'git push no-mistakes' and re-binds the attestation to the new head (which now includes this Greptile fix). No code change can or should address it. Files changed: bin/fm-lint.sh (ANSI-C escape decoding in the backend-purity normalizer), tests/fm-lint.test.sh (three encoded-bd rejection cases)
* fix(bin): preserve captain calls during teardown (#3595)
* fix(bin): never close a captain call during cleanup
A scout that held its own work item for the captain, which is what
captain-hold-lifecycle prefers ("hold the work item the question gates"),
was closed by bin/fm-teardown.sh's automatic backlog transition. The
completion gate passed, cleanup ran, and the captain's question moved to
Done with no recorded answer: the one thing the policy says must never
happen. `tasks-axi done` clos…
jorguez96
added a commit
to jorguez96/firstmate
that referenced
this pull request
Sep 17, 2026
…s resolved (#14) * fix(bin): defer inactive reconciliation during startup (#3480) * Defer inactive startup reconciliation * no-mistakes(review): Queue deferred inactive reconciliation diagnostics durably * no-mistakes(review): Require worker phases to cover startup requests * no-mistakes(review): Make diagnostic wakes safely acknowledgeable * no-mistakes(document): Document deferred startup phase coverage * fix(bin): bound wake drain presentation lock waits (#3475) * fix: bound status presentation lock waits * no-mistakes(review): Distinguish malformed presentation locks from live contention * no-mistakes(review): Bound no-ack drain queue lock acquisition * no-mistakes(document): Document bounded presentation-lock drain behavior * no-mistakes(lint): Annotate bounded lock output global * no-mistakes(ci): Added deterministic regression coverage for successful bounded-lock acquisition after live contention, verifying helper-to-caller PID ownership handoff and caller release. Verified with bash syntax checks, git diff checks, and the full fm-wake-queue test suite * fix(bin): retire public follow-ups in remote homes (#3479) * fix(relay): close a public loop whose work lives in a remote secondmate home A public-followup loop bound to a REMOTE secondmate could never be closed. `clear_public_followup_link` (bin/fm-public-followup.sh:701) required an absolute recorded `work_home_path` for a `secondmate:*` work home, but a remote route has no local path on this machine, so registration records that field empty (bin/fm-public-followup.sh:291). Every close ran that clear first, so `retire` died with "could not clear the legacy X link ... retained for reconciliation" forever, and `deliver` posted the public reply and then stranded the loop at `posted`. `--force` never covered that step. The clear now goes to the remote home over that route's SSH transport, running `fm-x-followup.sh --clear <work-id>` through `bin/fm-on.sh`. The route is decided from `data/secondmates.md` before any local path is consulted, so a same-named local directory can never stand in for a remote home, and registrations already on disk retire without needing a new field. `fm-on.sh` passes ssh's status through, so 255 stays the established "delivered but completion unknown" result this codebase already reconciles: the close is refused, the registration and the remote link are left exactly as they were, and the message names the unknown completion instead of claiming a definite failure. Local secondmate and `main` work homes are untouched, and `--force` still governs only the unresolved-obligation refusal. Three regression cases drive a remote route end to end, faking only the ssh binary at the FM_SSH_BIN seam and then running the real remote entrypoint against a local checkout, so the clear that must reach the remote home actually happens there. * no-mistakes(review): Guard remote link clears by request identity * no-mistakes(review): Fail guarded clears on unreadable remote state * no-mistakes(review): Reject guarded clears on non-writable remote state * no-mistakes(review): Allow no-link retirement in non-writable remote state * no-mistakes(document): Correct public-followup verification guarantee count * no-mistakes(ci): Fixed the guarded link-clear race by ensuring absence is decided under the metadata lock whenever publication is possible. Added a behavioral concurrency regression test. Verified with fm-x-mode and fm-public-followup suites, Bash syntax checks, diff checks, and bin/fm-lint.sh * no-mistakes(ci): Fixed the guarded link-clear race by refusing an unlocked absence decision when a publisher already owns the metadata lock in a non-writable directory. Added a behavioral concurrency regression test. Verified with fm-x-mode, fm-public-followup, syntax/diff checks, and fm-lint * no-mistakes(ci): Fixed the guarded-clear race by refusing all guarded clears when the metadata parent is non-writable, including apparent link absence. Added a behavioral regression with a publisher waiting to create the lock, updated remote-retirement expectations and verification docs. Passed fm-x-mode, fm-public-followup, fm-lint, documentation audience, Bash syntax, and diff checks * fix(relay): bound the guarded remote link clear so it refuses instead of hanging The guarded clear checks that the remote state directory is writable before taking the metadata lock, but that check cannot close the window: the parent can turn non-writable between the check and lock creation, and a lock held by a live holder is indistinguishable from that at the acquire. `fm_lock_acquire_wait` is an unbounded `while ! try; do sleep 0.1; done`, so either case retried forever and `deliver` or `retire` wedged with nothing reported, instead of returning the retained-for-reconciliation refusal the guard exists to produce. This path runs unattended over the secondmate transport, where a wedge is worse than either outcome the guard defines. The guarded clear now acquires through `fm_lock_acquire_wait_bounded` (FMX_LINK_CLEAR_LOCK_TIMEOUT, default 10 seconds) and refuses on timeout through the existing failure path. Unguarded local callers keep the ordinary unbounded wait, so local behavior is unchanged. The bounded primitive's header no longer claims presentation-only scope, since this is a second authorized caller; nothing else in the shared lock infrastructure changed. The regression holds the metadata lock with a genuinely live process while leaving the state directory writable, so the refusal can only come from the bound and never from the writability precondition. Against the unbounded wait it does not terminate at all; with the bound it refuses, retains the registration, writes no receipt, and leaves the remote link untouched. * no-mistakes(review): Harden lock-timeout regression with independent deadline * no-mistakes(review): Restore no-op guarded clears on read-only state * no-mistakes(document): Clarify remote public-followup cleanup contract * fix(bin): support process events under symlinked homes (#3484) * fix(bin): resolve process-event state roots before validating them The process-event module validated the caller's spelling of a home's state root instead of the directory it operates on: it required the supplied path to equal its own lexical normalization, which rejects any path reached through a symlinked ancestor. On macOS both /tmp and $TMPDIR are symlinks, so an operator home under either could never claim a source. Reconcile still reported the runner started, while the detached runner died writing "cannot claim source" to the discarded stderr, and the source silently never fired. Resolve the state root to its physical directory once, then apply the existing private-directory validation to that resolved directory and derive every path, recorded claim identity, and later confinement check from it. This keeps the confinement contract for the directory actually operated on rather than only for callers that already spelled it physically, and removes the window where an ancestor symlink could be repointed between check and use. Homes already spelled physically behave identically. This was the single cause of both deterministic macOS failures in tests/fm-procevent.test.sh ("reconcile never claimed the registered source") and tests/fm-procevent-when.test.sh ("the winning concurrent arm did not produce an outcome"). The new case pins the behavior with an explicit symlinked-ancestor home, so it fails without the fix on any platform rather than only where the temp root happens to be a symlink. * fix(bin): pin the external capture staging boundary to its physical path The extension capture path pinned its registry staging boundary by comparing `pwd -P` against the caller-spelled registry directory, so a home reached through a symlinked ancestor still refused to start an extension-backed source after the state root itself resolved correctly. That left such a home half working: built-in sources ran while external ones failed. The staging preparer now prints the physical registry directory it validated, matching the inbox and reservation preparers beside it, and the start path pins on that returned path. The new end-to-end case drives the shipped file-signal package from a symlinked home spelling. * no-mistakes(review): Propagate canonical process-event state roots * no-mistakes(review): Propagate canonical state to process-event adapters * no-mistakes(document): Document physical process-event state roots * fix(pi): deliver captain outcomes as deterministic transcript entries (#3312) * fix(pi): persist captain outcomes visibly * no-mistakes(review): Recover captain outcomes after cold-start lock acquisition * no-mistakes(document): Document cold-start captain-outcome recovery * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes(review): Prove immediate Pi captain-outcome transcript delivery * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * fix(pi): process captain outcomes through a sequence-keyed turn PR #3312 made every captain-facing supervision outcome a durable, exact-once visible transcript entry with the read cursor advancing only after that entry exists. That is the display half of the delivery contract. Left alone it turns a probabilistic silent loss into a deterministic one: the captain sees an anchor line, and firstmate never acts, because nothing opens a turn and nothing records whether main ever processed the outcome. The 2026-08-31 timeline showed the two shapes this must survive on the previous hidden-turn path: seven delivered decision outcomes each answered by an empty assistant message (cursor advanced, no retry, unanswered for close to three hours), and two answered by an unrelated prior reply. Both happened because delivery advanced the cursor at enqueue and accepted whatever the next assistant message was. Add the processing half on top of the persistence half: - bin/fm-branch-outcome.sh keeps a processed marker separate from the read cursor (`unprocessed`, `mark-processed --through`, `processed-init`). It only advances through an explicit sequence-bound acknowledgement, never past the read cursor and never backwards; an absent marker reads as zero and `processed-init` migrates delivered history once so an upgraded home is not re-presented its past. - After the visible entry for a captain outcome exists, the extension hands every still-unprocessed captain row to main as one hidden, typed `fm-branch-process` request listing each `[seq N] task: summary`, opening exactly one main turn. Main closes it only by calling the new `fm_branch_processed` tool with the highest sequence listed. An unrelated, empty, or paraphrased answer leaves the sequence open, and the same request is presented again at the end of the next main run and at session start. The first two presentations of a sequence set open a turn of their own; after that the request rides the captain's next prompt so an ignored request cannot loop, and a session replacement resets that budget. Routine outcomes stay turn-free. - The regressions cover exactly those incident shapes against the real store scripts: an empty answer and an unrelated prior answer neither advance the marker nor stop re-presentation, the acknowledgement is refused beyond the read cursor and outside lock ownership, a partial acknowledgement keeps the newer sequence open, and #3312's own assertions now forbid an unkeyed turn rather than any turn. The store suite pins the marker's bounds and the migration; the real-SDK guard for appendEntry persistence and model exclusion is unchanged. Docs move the protocol from "no model turn" to "one sequence-keyed processing turn closed only by its acknowledgement", and the verification record carries the dated run against Pi 0.84.4. * no-mistakes(review): Harden outcome listing and sequence-bound acknowledgements * no-mistakes(review): Harden outcome state validation and request pacing * no-mistakes(review): Reject unsafe sidecars and unterminated outcome stores * no-mistakes(review): Validate canonical mark-read cursor state * no-mistakes(review): Guard cursor advancement against corrupt processed state * no-mistakes(review): Bind acknowledgements to active processing requests * no-mistakes(review): Reset pacing when processing sequence membership changes * no-mistakes(review): Enforce silent outcome invariants at storage boundary * no-mistakes(document): Document hardened captain outcome processing contracts --------- Co-authored-by: kunchenguid <kun@kunchenguid.com> * feat: add bounded concurrent Bearings ledger collection (#3481) * feat: bound Bearings remote ledger collection * no-mistakes(review): Clarify default remote-ledger collection behavior * no-mistakes(review): Detach reconcile delivery from watcher loop * no-mistakes(review): Enforce bounded snapshot and request captures * no-mistakes(review): Bound legacy summary capture before parsing * no-mistakes(review): Bound primary remote ledger captures * no-mistakes(document): Correct snapshot and reconcile documentation * no-mistakes(lint): Fix ShellCheck quoting in bounded collector * no-mistakes(ci): Fixed all three CI failures: updated the macOS Bearings assertion to 44 tests, made the home-summary test deterministic and aligned with default ledger consumption, and increased the asynchronous reconcile retirement wait for loaded CI. Verified both focused suites, all 44 Bearings tests, ShellCheck, actionlint, Bash parsing, and git diff checks * test: await reconcile request retirement * no-mistakes(review): Avoid empty reconcile queue process churn * no-mistakes(review): Read ledger summaries from immutable snapshots * no-mistakes(review): Reject multi-document home ledger streams * no-mistakes(review): Coalesce durable reconcile requests per target * no-mistakes(review): Unify reconcile keys and reject snapshot streams * no-mistakes(review): Key reconcile requests by stable target ID * no-mistakes(document): Document per-target reconcile request coalescing * no-mistakes(lint): Remove unused snapshot summary file variable * no-mistakes(ci): Adjusted the concurrent collector regression’s end-to-end timing ceiling to account for stock macOS process/jq overhead outside the three-second remote collection budget, while remaining below the 15-second serial-read floor. Verified with stock /bin/bash 3.2: all 44 Bearings tests pass; bash syntax and git diff checks pass * no-mistakes(ci): Fixed legacy summary validation to require exactly one top-level JSON document and added behavioral regression coverage. Stabilized CI by conditionally waiting longer for durable reconcile delivery and synchronously stopping the fm-on worker tree before fixture cleanup. Removed a redundant flaky healthy-path timing assertion; the wedged-reader test still proves concurrent bounded collection. Verified fm-bearings-snapshot, fm-secondmate-reconcile, and fm-on tests, plus project ShellCheck, bash syntax, and git diff checks * ci: rebalance portable serial test shards (#3489) * fix(ci): rebalance the portable serial shards on measured durations The "Behavior portable serial 3" shard ran 17-20 minutes against its 20-minute job cap and intermittently timed out seconds after a passing test, on branches and on main alike. Shards are packed longest-processing-time from per-script duration hints, and those hints were last measured on 2026-08-21 at 116 scripts. The lane has since grown to 139 scripts and from ~42 to ~63 minutes: 17 scripts had no hint at all and fell back to the 20 s default, and several existing hints were low by 2-5x (fm-watch-triage 142 s hinted vs 263 s measured, fm-public-followup 36 s vs 197 s). The partition therefore looked perfectly balanced in hint space, 734.6 s per shard, while really running 11.5, 13.6, 18.8 and 16.5 minutes. Script-count balance, which is what the tests asserted, stayed normal throughout and hid it. Refresh the hints from the timing artifacts of three green runs, taking the slowest measurement of each script so the balance holds on a slow runner, and split the lane across five shards instead of four. Replayed against those runs' real per-script durations the worst shard is now 12.54 minutes, 63% of the unchanged 20-minute cap, and the serial lane's wall clock drops from ~20 to ~12.5 minutes. Bound the drift that caused this rather than relying on the hints being refreshed by hand: the coverage guard now reports the unmeasured share as serial_unhinted= and refuses past PORTABLE_SERIAL_MAX_UNHINTED_PERCENT, which leaves room for newly added tests while making a stale table fail the guard instead of silently pushing one shard into its cap. No test changes what it asserts and no test stops running; only the partition across shards changes. * no-mistakes(document): Clarify conservative shard timing aggregate * fix(pi): fall back on incomplete supervision branch prompts (#3491) * fix(pi): fall back after settled branch errors * no-mistakes(review): Detect provider errors across prompt compaction * no-mistakes(review): Preserve in-flight branch state across selection changes * fix(pi): re-probe supervision branch after cooldown (#3497) * fix(pi): recover supervision branch after cooldown * no-mistakes(review): Defer branch recovery until prompt settlement * no-mistakes(document): Clarify supervision cooldown recovery contract * fix(bin): remove legacy remote snapshot reads (#3501) * refactor: remove legacy remote summary reads * no-mistakes(document): Document ledger-only snapshot reads * no-mistakes(ci): Fixed the snapshot test fixture so ledger refreshes use the same fake executable PATH as the snapshot consumer. This preserves observable endpoint freshness after removing legacy summary computation. Verified stock Bash parsing and all 44 Bearings tests pass under /bin/bash; git diff checks pass * no-mistakes(ci): Fixed the CI-only snapshot fixture failure by ensuring the bounded-ledger refresh uses its fake tmux backend. This removes host tmux availability as a source of nondeterminism. Verified all 44 Bearings tests pass, Bash syntax passes, and git diff checks are clean * no-mistakes(ci): Fixed CI nondeterminism in the Bearings fixture: all local ledger refreshes now use the fixture’s fake tmux backend when available, instead of depending on host tmux state. Verified stock /bin/bash syntax, git diff checks, and all 44 Bearings tests with a deliberately failing host tmux * fix(pi): preserve watcher continuity across session replacement (#3498) * fix(pi): rearm watcher after session replacement * no-mistakes(review): Queue actionable closes across Pi session replacement * no-mistakes(review): Stop replacement arm when handoff persistence fails * no-mistakes(review): Preserve actionable wakes through branch and late child races * no-mistakes(review): Surface late handoff failures without crashing Pi * no-mistakes(review): Coordinate replacement delivery settlement and unique handoff tokens * no-mistakes(review): Retry stale deliveries and release settled claims * no-mistakes(review): Distinguish branch settlement and retry handoff cleanup * no-mistakes(review): Deduplicate persistent handoff cleanup alerts * no-mistakes(review): Acknowledge watcher follow-ups only when consumed * no-mistakes(review): Persist idle follow-ups until agent consumption * no-mistakes(review): Preserve pending outcomes when handoff persistence fails * no-mistakes(review): Arm replacement before awaiting prior delivery settlement * no-mistakes(review): Adopt pending handoffs after lock reclamation * no-mistakes(review): Prevent stale generations from adopting replacement handoffs * no-mistakes(review): Scope replacement handoffs by watcher state * no-mistakes(document): Clarify replacement handoff documentation * no-mistakes(ci): Fixed the failing branch-extension tests to model the new settlement-promise contract. Failure cases now assert that delivery ownership returns to the watcher instead of expecting direct extension fallback. Verified the updated branch suite, Pi watcher suite, shell syntax, and diff checks * no-mistakes(review): Update branch settlement tests and preserve chunked outcomes * no-mistakes(document): Document watcher-owned replacement handoffs * no-mistakes(document): Verify replacement handoff documentation * test(pi): cover watcher-owned branch fallback * no-mistakes(document): Refresh watcher-owned fallback documentation * fix(bin): resurface task statuses missed by wake handling (#3495) * fix(bin): resurface terminal statuses lost after branch handling * test(watch): canonicalize process-event fixture homes * no-mistakes(review): Index branch outcomes by causal status position * no-mistakes(review): Recover outcome indexes and deduplicate resurfaced statuses * no-mistakes(review): Handle legacy ambiguity and oversized status diagnostics * no-mistakes(review): Keep unclassifiable oversized statuses silent * no-mistakes(document): Document lost-wake outcome backstop * no-mistakes(document): Update outcome backstop documentation * no-mistakes(ci): Fixed CI regressions in wake-drain: parseable reserved-key decisions can no longer bypass the durable decision-fold guard, and status output is prepared and receipt-committed before presentation to prevent repeated one-shot outcomes after later failures. Added a behavioral regression for receipt commit failure and retry. Targeted backstop, correlation-token, decision-cursor, open-decision, unread-status, syntax, and diff checks pass locally. Shard-4 failures appeared unrelated/flaky; the network-parallel test passed locally * no-mistakes(ci): Fixed the Greptile P1 data-loss issue by committing presentation receipts only after prepared output reaches stdout. Added behavioral coverage proving output failure leaves the backstop retryable and receipt failure may duplicate but never lose a presentation. Relevant wake-drain suites and syntax/diff checks pass. The shard-4 Pi extension failure is unrelated to this PR and did not warrant changes * no-mistakes(ci): Stabilized tests/fm-bootstrap-network-parallel.test.sh by replacing scheduler-sensitive equal-sleep timing with bounded synchronization between mocked fetch and remote probes. This preserves detection of real serialization while avoiding false failures under CI load. Verified with five consecutive test runs, bash syntax validation, ShellCheck, and git diff checks. The separate Pi stock-rendering failure reproduces locally but is unrelated environment/version drift * no-mistakes(ci): Fixed Behavior portable serial 4 by adding fm-classify-lib.sh and fm-timeout-lib.sh to the broken-root Pi test fixture; fm-branch-outcome.sh now depends on them. Verified the full Pi branch-extension suite with real-Pi checks skipped, the wake-drain outcome-backstop suite, Bash syntax, and git diff checks. Greptile findings are already addressed at HEAD; the no-mistakes attestation failure is external head-SHA state * fix(bin): collect follow-up results from remote work homes (#3503) * fix(bin): deliver typed terminal results from remote work homes A public commitment whose work is bound to a REMOTE secondmate home could never receive its typed terminal result. `fm-public-followup.sh brief` printed an emit command carrying this home's own absolute path and this checkout's own script path, neither of which exists on the machine the worker runs on, so the worker had nothing it could write to that the owning home would ever read - and `consume` kept finding nothing while the promise stayed open. The brief is now route-aware: for a remote work home it prints that route's own code root and home with `--stage-in`, so the typed event is staged in the home where the work actually runs, and the closing paragraph names the owning home as the one on the other machine instead of pointing at the path above it. The owning home collects those staged results over the same SSH route it reaches that secondmate on, because the transport only runs outbound: `consume` pulls them into its own inbox and reconciles them exactly as it reconciles a local report. Collection is non-destructive until the result is durably held, so a dropped connection cannot lose a terminal result, and a route that could not be reached is named in `consume`'s output with the promise left open rather than reported as an empty inbox. A local work home is untouched: the brief still prints `--home` with this home and this checkout's script, and the event still lands directly in this home's typed terminal-result inbox. This is the emit-side counterpart of the retire/clear fix in #3479 and reuses the remote-route resolution that landed with it. Reconciling a loop bound to a remote route now reaches that route, so the existing remote cases drive `consume` through the same faked transport their other steps already use. * no-mistakes(review): Fail loudly on unresolved routes and invalid staging homes * no-mistakes(review): Fail collection when remote outbox is unreadable * no-mistakes(review): Surface reassigned remote routes during empty collection * no-mistakes(review): Fail remote collection on invalid registrations * no-mistakes(review): Reject unsafe registration entries during remote collection * no-mistakes(review): Restore healthy empty remote collection behavior * no-mistakes(review): Skip remote collection for delivered registrations * no-mistakes(review): Skip delivered registrations before route validation * no-mistakes(document): Document remote follow-up collection semantics * fix(bin): exclude secondmates from home-summary validity (#3504) * fix(bin): exclude secondmates from home-summary child inventory kind=secondmate meta records never have backlog rows, so counting them in unowned_children or terminal_in_flight made a clean main home look invalid once earlier ledger checks passed. * no-mistakes(review): Cover terminal secondmate in-flight exclusion * no-mistakes(ci): Updated the stock macOS Bash CI snapshot expectation from 15 to 16 tests. Verified all 16 snapshot/fleet-view tests pass under Bash 3.2.57 and `git diff --check` succeeds * fix(bin): self-heal outcome indexes on first drain (#3509) * fix(bin): self-heal status-outcome indexes on every drain Missing ready markers were skipping the lost-wake backstop on non-Pi homes because only the Pi branch ran processed-init. Drain now rebuilds those indexes under the outcome lock and fails closed only on a real store fault. * no-mistakes(review): Guard held-lock initialization and fail marker writes * no-mistakes(document): Document cross-harness outcome-index self-healing * fix(bearings): keep active children underway during captain holds (#3505) * fix(bearings): keep active children underway beside a captain hold Project each readable home's active children into Underway independently of the home-level captain-decision classification so a hold no longer hides live work. * no-mistakes(review): Preserve Underway repos and disclose child truncation * no-mistakes(review): Fall back to task project for Underway repos * no-mistakes(ci): Updated the stock macOS Bash CI assertion from 44 to 45 Bearings tests, matching the newly added behavioral regression. Verified all 45 tests pass under /bin/bash, Bash syntax checks pass, and git diff validation is clean * fix(pi): settle watcher delivery on Pi accepting the follow-up (#3513) * fix(pi): settle watcher delivery on Pi accepting the follow-up A follow-up queued while main is streaming joins the running run without ever raising before_agent_start, so waiting on that event before clearing the successor pipeline (#3498) stalled every later actionable close: no successor started, no wake was delivered or offered to the branch, and the turn-end guard woke main to re-arm by hand after every close. The pipeline now settles once Pi accepts the follow-up. Consumption is observed at before_agent_start for an idle main and at the user message_start for a streaming main, and decides only what a replacement session (/new, /resume, /fork, reload) replays. An exhausted restoration delivers its typed failure without launching an arm past the retry bound, which the stall had hidden. The replacement-coordinator map is typed so the strict no-emit typecheck passes again. Tests: the doubles no longer raise before_agent_start for a streaming send, a portable regression drives two actionable closes while main streams and proves the successor chain plus consumption-scoped replay, and a credential-free real-SDK probe pins Pi's event contract for both the streaming and the idle follow-up. Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a * fix(pi): retry a verified successor that fails during wake delivery A verified successor can exit while the wake it was started for is still being delivered, most plausibly during a branch turn that holds the settlement for minutes. Its failure close arrived while the pipeline's single-flight guard was set, so the close handler skipped the retry, and the pipeline's end no longer launched an arm, which left the live generation with no watcher and no retry timer. The close handler now records that failure when the child had reported readiness and was not retired by the restoration itself, and the pipeline runs the ordinary bounded, lock-checked retry for it once the delivery settles. A restoration started for a later pending supersedes it, and an exhausted restoration still hands repair to main without a further arm. The regression holds a branch settlement open while the verified successor exits with a failure and proves one retry watcher starts after the settlement releases, none while it is held. Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a * fix(bin): bound repeat stale wakes for parked workers (#3532) * fix(bin): bound repeat stale wakes for a parked but live worker A worker parked on a declared wait - `paused:` for an external or pipeline wait, or a verified `captain-held` transfer - kept waking firstmate far inside FM_PAUSE_RESURFACE_SECS. Observed as five consecutive alarms on one captain-held worker and dozens across a day on a pipeline wait, and reported upstream as four wakes in 75 minutes against a 3600s window. pause_state_class deliberately answers `none` for a still-live agent even under a declared wait, so a worker genuinely waiting on a decision is never silenced. That classification is correct and is left alone; it routes every parked but live worker through surface_nonterminal_stale on first sight of each distinct stale hash, and an idle parked pane still churns its hash on a clock or a token counter without changing what is being waited on. Two places let that churn re-alarm: - surface_nonterminal_stale queued the wake BEFORE consulting whether a wait was declared, then wrote `.paused-resurfaced-<key>` - the very throttle that should have suppressed it. The throttle was never read on this path and was advanced by the wake it should have prevented. - The hash-change path cleared that throttle through clear_pause_tracking whenever the classification came back `none`, so each tick also bought the same declared wait a fresh window. Fixing only the first site changes nothing. Read the throttle before anything is queued and advance it only on a wake that really fires, and on the hash-change path reset only the per-hash bookkeeping while the declaration still stands, via a clear_stale_hash_tracking split so neither half of clear_pause_tracking is duplicated. The throttle is keyed to the declaration, not to the pane. First sight still wakes, so an inconclusive state is still inspected, and the window's end still re-surfaces once, so a forgotten wait cannot rot invisibly - noise traded for a bounded cadence, never for silence. The wake identity stays the plain `stale: <win>` the away-mode handoff depends on. Tests cover both observed forms and were confirmed to fail against three deliberate breaks: each site reverted on its own, and a re-surface that never fires again. * fix(document): Clarify declared-wait wake cadence documentation * fix(ci): Captain, fixed the stale-throttle inheritance: cadence markers now bind to the current wait declaration, so replacement paused and captain-held waits each emit their first plain `stale:` wake. Added behavioral coverage for both forms. Bite proof failed as expected when identity matching was removed, then passed after restoration. Full watcher triage suite, `bin/fm-lint.sh`, syntax checks, and diff checks pass. Changes remain uncommitted for the outer executor * fix(ci): Captain, fixed the confirmed Greptile finding. `resurface_absorbed` now applies a throttle only when its stored declaration scope matches the current wait, so replacement `paused:` and `captain-held` waits surface immediately without changing classification. Added executable coverage for both absorbed forms. Bite proof failed before the fix at the intended assertion; afterward the full watcher triage suite, `bin/fm-lint.sh`, shell syntax checks, and `git diff --check` passed * fix(bin): accept the away-mode daemon as the turn-end supervision owner (#3567) * fix(turnend): accept the away-mode daemon as the supervision owner While state/.afk exists the away-mode daemon owns supervision and runs bin/fm-watch.sh one-shot: the watcher exits on every wake and the daemon starts its replacement. The turn-end guard tested for a live watcher process holding the watch lock at that instant, so a turn boundary that landed in the hand-off blocked with "TURN WOULD END BLIND" while supervision was completely healthy, costing a full handling turn each time. Reproduced with the real daemon wrapping the real watcher and the real guard sampling the same home: 6 of 40 samples blocked, every one of them with the daemon alive and the beacon 2-3 seconds old, and a new watcher pid on each cycle. After the fix the same reproduction blocks 0 of 40, and killing the daemon and its watcher (away mode still on, beacon still fresh) blocks again. The guard now accepts a live, identity-matched daemon holding this home as proof of supervision while away mode is active. The identity match is the same discipline the watcher lock uses, so a recycled pid or a lock left by a killed daemon proves nothing. The fresh-beacon half of the predicate is unchanged: a daemon that stops restarting its watcher still blocks once the beacon passes grace, a home with no supervisor blocks exactly as before, and with away mode off the strict watcher predicate is untouched. The predicate reads only durable state, so it behaves identically for every primary harness and runtime backend. * no-mistakes(document): clarify away-mode daemon supervision proof and test coverage * no-mistakes(document): generalize stale turn-end predicate summary in architecture.md * fix(backlog): omit --file from row probes for non-markdown backends (#3582) * fix(backlog): omit markdown file for beads probes * no-mistakes(document): Narrow backlog addressing doc to mutations for backend-aware probes * no-mistakes(ci): Fixed the Greptile P2 review comment (the only failing check) on tests/fm-backlog-atomicity.test.sh. The comment correctly noted that an exported TASKS_AXI_BACKEND environment variable would inherit into the spawned scripts and, because fm_tasks_axi_backend gives it top precedence, override each test case's .tasks.toml backend fixture — making the backend-specific argv assertions fail for environmental reasons. Fix: unset TASKS_AXI_BACKEND in the test harness right after sourcing tests/lib.sh, with a comment explaining why, so every case deterministically exercises its declared backend (4 lines added; no production code touched). Verified: reproduced the leak before the fix (TASKS_AXI_BACKEND=beads made the markdown dispatch case fail with 'beads show failed', exactly the reported failure mode); after the fix the full suite passes (0 failures, exit 0) both with and without TASKS_AXI_BACKEND=beads exported. The added lines are shellcheck-clean (the only shellcheck note, SC1091 on the lib.sh source line, pre-exists this change) * fix(bin): classify progress updates on requested work as routine (#3589) The supervision branch's verdict rule escalated every outcome that answered a captain request, so "the work started" and "still working" notes reached the captain with nothing to look at. The rule now keeps a finished result of requested work captain-facing, even when healthy, and treats start or still-working updates that bring no new artifact, finding, or decision as routine. The captain list for review-ready PRs, ask-user findings, exhausted blockers, credentials, and destructive or security-sensitive cases is unchanged, as are the unsolicited-routine, silent-fleet-review, and doubt-chooses-captain rules. The fm_branch_report tool description and the two docs that restated the old unconditional rule now point at the prompt's "Verdict: routine or captain" section as the one owner instead of carrying a second copy. * fix(bin): preserve captain calls during teardown (#3595) * fix(bin): never close a captain call during cleanup A scout that held its own work item for the captain, which is what captain-hold-lifecycle prefers ("hold the work item the question gates"), was closed by bin/fm-teardown.sh's automatic backlog transition. The completion gate passed, cleanup ran, and the captain's question moved to Done with no recorded answer: the one thing the policy says must never happen. `tasks-axi done` closes a held row silently, and nothing in teardown asked whether the row was the captain's own call. bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when the task is still an open captain call, 1 when it is not, 2 when that cannot be established. It reads the row through the transition library's backend-aware probe, so it addresses the same backlog teardown does; the script's other commands now address the configured data directory the same way instead of FM_HOME, which also fixes captain holds in a home with a relocated data directory. Teardown asks `open` before any destructive step and refuses on 2. On 0 only the close changes: after cleanup and still under the task's own lock, the row gets one "Deliverable of the finished work" line at the end of its body and returns to Queued through `tasks-axi reopen`, keeping its hold, so it lands in Captain's Call instead of reading as work under way. --force does not lift this: it authorizes discarding unlanded work, never the captain's question. The deliverable goes into the body because `tasks-axi update --report` rewrites the title of a row that is not Done. The crash window reuses the pending-close record teardown already stages: a `mode=retain` line makes the existing replay record the deliverable and reopen instead of closing, with the same validator, stale-generation check, cleanup-incomplete marking, and non-blocking bootstrap lock as an ordinary close. A retained row the captain answered first simply retires the record. No parallel record type, recovery command, or second bootstrap loop is introduced. Regressions run the real executables: the captain-held scout survives cleanup queued, held, with its deliverable and on the board, only `answer` closes it, --force keeps it open, and an ordinary scout still closes with its report; an interrupted cleanup leaves the row untouched and the next session start retains it; a relocated backlog keeps the retention in its one configured file; and a ship row whose hold cannot be read refuses cleanup before anything destructive. Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np * no-mistakes(review): Serialize captain holds and fix backend-aware listing * no-mistakes(document): Update captain-call retention documentation * no-mistakes(document): Fix relocated captain-hold backlog diagnostics * fix(bin): deliver secondmate outcomes to the parent channel (#3592) * fix(bin): deliver every secondmate outcome on the parent channel from the recording scripts A secondmate's captain-facing outcomes could miss: the mate model addressed the captain in its own unread chat instead of appending to the parent channel, and a PR-ready report, a finding, a decision, a blocker, and a failure all depended on that one remembered append. Make delivery structural, so the parent channel never depends on the model: - bin/fm-parent-channel-lib.sh is the one owner of channel resolution and exact-line append-once; the merge outcome path and the inactive-outcome scan now publish through it instead of two private copies. - bin/fm-inactive-reconcile.sh gains a ledger-first path that runs on every watcher poll in a secondmate home: a direct child's whole terminal done or failed line is delivered at once with its note, recorded PR, mode, merge posture, and scout report pointer, keyed and receipted so it is delivered once, and the inactive path yields to it. `report <task-id>` runs the same delivery for a caller holding the child's meta lock. - bin/fm-pr-check.sh publishes the PR-ready line with the canonical URL at registration. - bin/fm-captain-hold.sh publishes a hold and its answer, keyed by task id and resolution-record count, with no new persisted state. - bin/fm-teardown.sh delivers the child's final line before removing its record and refuses, retaining every record, while the channel cannot be written. - The charter opens with the parent-channel rule and confines the mate's own appends to judgement; AGENTS.md carries the carve-out at the persona address rule and the escalation list. docs/secondmate-parent-channel.md records the design and its coverage, and docs/verification/secondmate-parent-channel.md records the live run with real tmux panes and both real watchers delivering every line with no model. Supersedes #3569. * no-mistakes(review): Fix parent outcome retries and reconciliation locking * no-mistakes(review): Prevent busy children from starving ledger delivery * no-mistakes(review): Correct ledger metadata and hold occurrence handling * no-mistakes(review): Disambiguate ledger outcomes and normalize hold reasons * no-mistakes(review): Close ledger races and preserve teardown records * no-mistakes(document): Correct parent-channel receipt and scanner documentation * no-mistakes(lint): Quote done arguments for ShellCheck compliance * no-mistakes(ci): Fixed both CI failures. Updated GOTMP teardown fixtures for the new final-outcome reporter and isolated them from host tmux state. Updated the PR security assertion to distinguish the accepted PR-ready line from duplicate merge outcomes. Verified with both failing test suites, bash syntax checks, and git diff checks * no-mistakes(ci): Fixed Greptile’s duplicate-delivery race in bin/fm-inactive-reconcile.sh. Ledger events now claim matching already-delivered inactive receipts using the prior status fingerprint, preventing duplicate parent reports while preserving later same-state completions. Added behavioral regression coverage. Verified inactive-reconcile tests, project lint, documentation audience checks, syntax, and diff checks. Teardown tests passed relevant cases before the documented pre-existing herdr-preflight-missing-adapter failure * fix(bin): sync remote second mates to primary commit (#3599) * fix(bin): sync remote second-mate homes to the parent primary commit Session start and remote launch pointed a remote second-mate home at whatever Firstmate copy its own host kept, so a home that had already advanced past that copy refused as a non-fast-forward and every other home stopped at the host's older commit while the primary ran ahead. The parent now resolves ITS primary default-branch commit with the existing helper and hands that commit to the host on both paths. Because a remote home is a standalone clone, the host imports that one commit before advancing - already present, else from that host's Firstmate copy without moving it, else from the home's own origin - and then runs the SAME ff_target guards a local home gets, so dirty, diverged, feature-branch, and unresolvable targets skip untouched and the ancestry rules keep one owner. An unimportable target now names /updatefirstmate instead of failing opaquely, and a host still running an older Firstmate copy is reported the same way rather than echoing a bare refusal. The host-local launch leg no longer re-runs its own secondmate sync, so the spawn it drives cannot re-target that host's copy after the parent has already converged the home. /updatefirstmate is unchanged: it still refreshes the remote code root from that host's origin and then syncs the home to that refreshed copy, which is what the sync call with no target commit means. * no-mistakes(document): Document primary-targeted remote secondmate synchronization * fix(bin): separate captain intent from firstmate specs (#3597) * fix(bin): split brief task into captain intent and firstmate spec Keep no-mistakes --intent as the captain's ask plus later captain words, not the build spec or worker tradeoffs. * fix(bin): stop task-subsection copies at the next heading Promotion was swallowing the scout Setup contract into Firstmate spec, and pre-subsection briefs lost their # Task body. * no-mistakes(review): Validate brief content and preserve nested specifications * no-mistakes(review): Scope placeholder validation to scaffold-only subsection bodies * no-mistakes(review): Ignore fenced subsection headings during brief validation * no-mistakes(review): Preserve captain intent across scout promotion * no-mistakes(review): Enforce safe intent boundaries for legacy promotions * no-mistakes(review): Allow marked legacy intent and reject empty promotions * no-mistakes(review): Scope task parsing and overlay legacy intent contracts * no-mistakes(review): Overlay current intent contract for all no-mistakes spawns * no-mistakes(review): Preserve later captain clarifications in intent overlays * no-mistakes(document): Document brief intent enforcement and ownership * no-mistakes(ci): Updated spawn-related test fixtures to use valid Captain intent and Firstmate spec subsections, corrected launch-path expectations to launch-brief.md, and resolved ShellCheck quoting findings. Verified with fm-lint.sh and 15 affected behavior tests, including real Herdr tests; all passed * no-mistakes(ci): Updated stale spawn/promotion fixtures in the Muse, Orca, secondmate-harness, and public-followup suites to provide valid Captain's intent and Firstmate spec subsections. Verified full Orca and secondmate-harness suites, targeted public-followup promotion behavior, Bash syntax, diff checks, and fm-lint * fix: start a fresh supervision branch for every main session (#3600) * fix(pi): start a new supervision branch conversation per main session The supervision branch reopened one recorded conversation forever, so every main session start reloaded the current generated prompt and then weeks of accumulated thread, where a superseded rule could still outweigh today's. The branch conversation is now scoped to one main session: the session generation owns the recorded conversation, so a cold start, /new, /resume, /fork, or a reload always builds a new one, while a rebuild inside one session (a model or effort change) still continues that session's own conversation. The dialog mirror re-anchors with it. Its durable cursor records what the previous branch conversation received, so a /resume or reload - which keeps main's own session file - would otherwise leave the new branch blind to dialog main itself still has. The reset is bounded by the current main session, and the cursor keeps advancing incrementally within it. The durable outcome store and its processed marker are untouched, so unacknowledged captain-facing outcomes still re-present on the new main session. * no-mistakes(document): Document fresh Pi supervision conversations * no-mistakes(ci): Fixed the flaky concurrent inbox failure. Lock acquisition now retries when a competing lock disappears between a failed claim and inspection. Added a behavioral regression covering that race. Verified the full inbox test four times, project lint, and git diff checks * feat: restart second mates after instruction updates (#3614) * feat(update): restart second mates whose instructions changed /updatefirstmate pulled new bytes onto disk and then asked each advanced second mate to re-read them. A running agent holds AGENTS.md and every loaded skill frozen from launch and no verified harness offers a reload, so that steer could not reach a loaded skill at all and left the mate holding two contradictory copies of its own job description. An eligible mate is now restarted instead, in the same home and endpoint, through the existing transactional relaunch. The restart is gated on the mate first writing down the open work it holds only in conversation - the open-record half of /stow, never its memory sweeps - so an unregistered captain call is flushed before the conversation is spent. Anything that leaves the reload unprovable falls back to the old re-read message and is reported as exactly that, never as a clean reload. Remote mates take the same path: fm-remote-secondmate-control.sh gains a relaunch verb whose host-local leg runs that same control plane, since the mate is an ordinary local secondmate from its host's point of view. The primary resolves the profile and passes it explicitly, because config/secondmate-harness is not inherited and the file on that host belongs to a different home. fm-update.sh now splits its advanced live mates into a restart set and a nudge residual, and both sets require a changed instruction surface, which also closes the over-nudge against the session-start sweep. Restart is stricter still: a bin/-only advance reloads itself on the next call, so it never costs a conversation. Colocated tests cover the gating, the persist-then-restart order, the task-subset persist request, each unsafe fallback, the remote hop, and the remote sync's new instruction-surface report. * no-mistakes(review): Fix restart correlation, concurrent waits, and lifecycle reporting * no-mistakes(review): Parallelize relaunches and classify replacement incarnations * no-mistakes(review): Gate restart actions on live agent state * no-mistakes(review): Handle failed restart workers without hanging * no-mistakes(review): Nudge legacy remotes and preserve persist recovery * no-mistakes(review): Document one-time secondmate restart rollout * no-mistakes(review): Honor arrived replies and refresh remote profiles * no-mistakes(review): Revert remote parent profile reconciliation * no-mistakes(review): Reset remote profile defaults and honor published results * no-mistakes(review): Preserve fallback nudges for unverifiable secondmates * no-mistakes(document): Document second-mate restart update flow * no-mistakes(lint): Fix ShellCheck warnings in restart scripts * perf: accelerate local validation with bounded concurrency (#3644) * perf(tests): route gate verification through the bounded concurrent runner Local validation was the pipeline's dominant cost: across 67 recorded no-mistakes agent sessions on this repo, 99.3% of command execution was `bash tests/*.test.sh`, run strictly one script at a time, and 2% of those calls were killed by an agent-guessed timeout and paid for twice. Three changes, each measured: - `.no-mistakes.yaml` pins `commands.test` to `bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner already owns changed-file selection, bounded concurrency, the refusal of unproven scripts, and a generous automatic per-script bound, so the gate's baseline is neither a serial chain nor a guessed timeout. It stays intent-targeted - the Test step still runs its evidence agent on top - and excludes the live-Herdr family the required Herdr lane owns. - `bin/fm-test-run.sh` gives a plain list of script paths the same bounded automatic scheduler and automatic bound that `--changed` gets. Naming several subjects is how a verification round asks for exactly those scripts. The curated selections are untouched: `--lane` still composes CI shards whose serial lane must stay serial, `--family` is what the required Herdr lane runs, and `--all` stays a deliberate complete regression. - `pr-forge` is admitted to the concurrent-safe family registry on two consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those, and records `secondmate` and `session-bootstrap` as refused with the exact script and reason each failed on, so the refusals are actionable rather than silent. Measured on this host, 0 failures on both sides: verification round, 4 scripts 448s chained -> 231s through the runner (-48%) pr-forge family 409.2s at 1 worker -> 237.9s at 4 (1.72x) watcher-wake-lock family 1311.1s at 1 worker -> 539.3s at 4 (2.43x) A fourth lever was implemented and then removed because the measurement refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made `fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s unchanged, back to back. Those sleeps are not overhead added to the clock - they are how a test waits for a subject moving on fm-watch.sh's own one-second cadence - so sampling less often only delays detection. It also broke `fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a settled condition. CONTRIBUTING.md records that result so the experiment is not repeated. * no-mistakes(review): Separate concurrent runs by isolation proof family * no-mistakes(review): Limit automatic timeouts to changed-file validation * no-mistakes(document): Clarify validation concurrency documentation * fix: copy PR URLs from durable records (#3648) * fix: copy PR URLs from records or abstain, never assemble them Supervision reported a plausible but dead PR link three times because its prompt demanded a full https:// URL at a moment when only a PR number was observable, so the model assembled an owner/repository from memory, and the PR check then accepted that URL and wrote it into the task record, after which the model kept defending its own tool-endorsed guess over the worker's real link. Three changes close that chain without any live forge lookup, so private forges are treated exactly like public ones: - bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy or abstain" section requires a URL to be copied verbatim from a durable record (the done: PR <url> status line, pr= metadata, or the backlog note), forbids assembling owner, repository, host, or number from memory, and has the branch report only the identifier it actually holds when no record names the URL yet, leaving the PR check unarmed until the worker's ready line arrives. AGENTS.md section 7 and 9 carry the same copy-or-abstain rule for main in place of the bare full-URL mandate. - Worker briefs (bin/fm-brief.sh, ship and scout rules) require the full https:// URL wherever a PR is mentioned - status line, terminal, or summary - never a bare "PR 108", so the link is in view as early as the number is. - bin/fm-pr-check.sh refuses, offline and before any side effect, a URL that the task's own done lines contradict, printing both spellings; a log naming no URL still records the argument as before. fm_pr_status_ready_urls in bin/fm-pr-lib.sh owns reading those lines. The refusal also reaches bin/fm-pr-merge.sh, so nothing merges under a contradicted URL. Tests cover the offline refusal with zero side effects, the recorded spelling being accepted, markdown-wrapped and punctuated URLs, working lines not counting, the merge wrapper propagation, a self-hosted merge request with no forge call, the prompt carrying the rule, and the brief carrying the worker rule. * no-mistakes(review): Remove stale PR URL enforcement * no-mistakes(ci): Removed backlog notes as an accepted PR identity source. PR URLs may now be copied only from the task’s `done: PR <url>` status or canonical `pr=` metadata; otherwise supervision reports only the known identifier and leaves PR checking unarmed. Updated related guidance/docs and verified with branch-supervision tests, brief tests, ShellCheck, and `git diff --check` * fix(bin): disable Claude feedback drafts for fleet launches (#3661) * fix(bin): disable Claude's feedback-draft flow for fleet-launched agents Scope --settings '{"feedbackDrafts":"off"}' to every Firstmate-launched Claude crewmate and secondmate, so /bug and /feedback never queue or submit a bug report on the captain's behalf. feedbackDrafts is the documented settings key (Claude Code changelog 2.1.247); the per-launch CLI flag never touches the captain's global settings.json. Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3 * no-mistakes(review): Prevent managed settings from re-enabling Claude feedback drafts * no-mistakes(document): Fix Claude feedback documentation formatting * fix(bin): layer both feedback-draft controls for defense in depth The prior --settings-only fix can be overridden by a managed Claude settings policy (feedbackDrafts precedence). Keep CLAUDE_CODE_SEND_FEEDBACK=0 alongside --settings '{"feedbackDrafts":"off"}': either control alone disables the SendFeedback tool, so a managed override of one still leaves the other in force. Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3 * no-mistakes(document): Document Claude feedback-draft suppression ownership * feat(tests): run three more validation families concurrently (#3662) * perf(tests): admit three more families to concurrent validation The three families that `docs/fm-test-isolation-proof.md` recorded as refused were not refused for concurrency. Each blocker was a test that decided a property by wall clock, or a script filed where it cannot run. Fixing those three things admits all three families and recovers 28.6 minutes of local validation with no assertion removed or weakened. - `tests/fm-backlog-handoff.test.sh` injected its pre-move crash by killing the handoff, sleeping a fixed second, then delegating the move to the real binary. Nothing ever killed the fake, so on a host slow enough for the case's next assertions to take longer than a second, the orphan woke and completed the very move the case requires left undone, and recovery then failed with `Task "pre-move-crash" not found in this backlog`. Watching the two backlogs during the injected crash showed exactly that, the item moving one second after the crash. All four crash injections in the file now go through a new `fm_fake_crash_injector` shim that signals the target and returns only once it is observably gone, and the pre-move fake never delegates the move at all. - `tests/fm-session-start.test.sh` proved the startup digest does not block on a slow current-state read by timing the whole digest against a fixed eight-second sleep, which a loaded host exceeds without the property being violated. It now holds that read open until the case releases it and asserts, the moment the digest returns, that the read has not finished. A digest that waited would wait indefinitely rather than for an interval a slow host can out-run, so the assertion is stronger than the bound it replaces. Its scan budget moves to the maximum, because the old value left two seconds of margin over the fixed sleep and measured the host rather than the deadline that `tests/fm-inactive-reconcile.test.sh` owns. - `fm-backend-herdr-focus-flash-e2e` was filed in the family map's catch-all, which put it in the portable serial lane, where Linux CI gate-skips it: that real-Herdr regression was running nowhere. It moves to `real-herdr-gated` and the required Herdr lane. `fm-claude-stop-autoarm-live-e2e` gate-skips on its opt-in variable and moves to `live-harness-optin`. The 28 remaining ungrouped scripts become an enumerated `standalone` family instead of admitting `unclassified` itself. `unclassified` is the family map's `*)` arm, so admitting it would silently grant concurrency to every test added afterwards, which is exactly the population with no proof. A new test still lands in `unclassified` and stays serial, and `tests/fm-test-run.test.sh` covers that split behaviorally. Each family passes two consecutive four-worker proofs with zero failures. On the production runner, `secondmate` goes 1233.1s to 453.4s, `session-bootstrap` 756.4s to 286.4s, and `standalone` 724.6s to 261.1s: 2.71x overall and 1713.2s recovered. The whole suite runs 177 scripts in 52.6 minutes of wall clock against 121 minutes of summed script time. * no-mistakes(document): Refresh concurrent validation and shard documentation * no-mistakes(ci): Fixed the real-Herdr focus-flash E2E race exposed by reclassification. Part C now starts its persistent child atomically via `pane run` and verifies stable child identity through Herdr’s public `process-info` interface, avoiding the racy send-text/send-keys sequence and platform-specific `ps` matching. Verified with bash syntax checking, ShellCheck, git diff checks, and the complete E2E test on Herdr 0.8.2 * feat: structure no-mistakes ask-user escalations (#3670) * feat(brief): structure no-mistakes ask-user escalation as event + snapshot file Crewmates escalating a no-mistakes ask-user gate now report one status event naming every finding id plus a snapshot file holding the gate's axi finding records verbatim (id, severity, file, line, description, authority), using the same shape even for a single finding. The status line never paraphrases. The format is defined once in fm-dod-lib.sh and rendered into both the scout and ship rule 6 in fm-brief.sh, so a promoted scout - whose rule 6 fm-promote.sh preserves unchanged - gets the identical contract as a freshly-spawned no-mistakes ship worker. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PpiWaDerbYavTLPPtEjQei * no-mistakes(review): Preserve ask-user escalation output contract * no-mistakes(review): Align escalation format test expectation * no-mistakes(review): Scope ask-user escalation instructions correctly * no-mistakes(review): Remove ask-user from generic decision rules --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> * fix(bin): require self-sufficient no-mistakes intent (#3671) * fix(bin): require a self-sufficient no-mistakes intent A no-mistakes worker's --intent is only as useful as the string it passes. PR #3604 shipped with an intent that was only "do 1, 2, 3, 7 from the report": the real contract lived in a private scout report and never reached --intent, so nobody holding that string plus the codebase could have derived the specification. This is pure instruction at the contract's one owner; no spawn-side or promotion-side check is added. - bin/fm-dod-lib.sh: the generated no-mistakes Definition of done now states that the --intent string must be self-sufficient (the string plus the codebase reconstructs roughly the same specification) and tells the worker to write the substance of any report, decision, or PR the captain's intent refers to into --intent rather than the pointer, while Firstmate build instructions and the worker's own decisions still stay out. The spawn-time overlay points back at that rule so its "supersedes" wording cannot cancel it, and the header's owner statement carries the rule. - AGENTS.md section 11 and bin/fm-brief.sh's header ask Firstmate to include the substance of referenced material when filling ## Captain's intent, and section 11 points at the owner of the rule. - tests/fm-brief.test.sh and tests/fm-task-delivery.test.sh assert the rendered brief and launch contract carry the rule. Claude-Session: https://claude.ai/code/session_01YMhEe42q7BAAoN6RxNuzim * no-mistakes(document): Replace incident-specific intent test commentary * fix: accelerate local Bearings snapshot composition (#3499) * Speed local fleet snapshot composition * no-mistakes(review): Stabilize task inventory during concurrent snapshot composition * no-mistakes(document): Document local snapshot observation concurrency * no-mistakes(ci): Fixed CI failures by making empty task manifests compatible with stock macOS Bash 3.2, snapshotting task metadata before concurrent observations to prevent generation drift, strengthening the behavioral race regression, and updating the stock-Bash Bearings test count to 45. Verified fleet snapshot tests (15), Bearings tests (45), workflow lint tests, project lint, Bash 3.2 parsing, and diff checks * no-mistakes(ci): Fixed the Linux CI failure caused by passing large backlog/task JSON through jq command-line arguments, which exceeded the per-argument size limit. Both inventory projections now stream large JSON inputs through stdin. Verified with fm-bearings-snapshot.test.sh (45 tests), fm-fleet-snapshot-view.test.sh (15 tests), Bash s…
sctru
added a commit
to sctru/firstmate
that referenced
this pull request
Sep 19, 2026
* fix(bin): safely unregister custom checks (#3369)
* fix(bin): add a safe owner for custom-check retirement
Agents were improvising rm of check files with unset STATE/ID, which wedges
headless panes. Unregister validates the id and state directory first.
Co-authored-by: Cursor <cursoragent@cursor.com>
* no-mistakes(review): Refuse explicitly empty custom-check state overrides
* no-mistakes(document): Document custom-check retirement safety contract
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* refactor(quota): extract mid-task polling and candidate selection into dedicated scripts (#3221)
* Add quota exhaustion detection and safe fallback helpers
- bin/fm-procevent-quota.sh: generic procevent adapter that arms a
recurring quota-axi --json poll and wakes firstmate when a tracked
provider's effectivePercentRemaining drops below a threshold or its
runway.status becomes exhausted_now.
- bin/fm-quota-choose.sh: worker-side helper that picks the first ranked
harness:model candidate with positive effectivePercentRemaining.
- AGENTS.md and .agents/skills/quota-array-dispatch/SKILL.md: document
the new helpers and the mid-task quota-exhaustion wake path.
- tests/fm-quota-choose.test.sh: unit tests with a mocked quota-axi JSON
source.
* no-mistakes(review): Fix quota polling and scope bounds
* no-mistakes(review): Enforce safe default quota selection
* no-mistakes(review): Handle decimal quota values safely
* no-mistakes(review): Fail closed on invalid quota inputs
* no-mistakes(review): Reject empty quota candidate segments
* no-mistakes(review): Harden quota parsing and timeout ownership
* no-mistakes(review): Reuse captured quota snapshots consistently
* no-mistakes(review): Match quota using explicit candidate providers
* no-mistakes(review): Centralize fail-closed quota schema validation
* no-mistakes(review): Reject out-of-range quota percentages
* no-mistakes(review): Validate quota runway status enum
* no-mistakes(review): Tighten quota scope and status contracts
* no-mistakes(review): Preserve unknown quota and exact product bounds
* no-mistakes(review): Preserve provider-level unknown quota
* no-mistakes(review): Reuse canonical verified harness validation
* no-mistakes(document): Document mid-task quota handling
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* fix(docs): restore default routing contract, keep quota helper optional
Restore the AGENTS.md section 4 always-loaded routing paragraph the PR
had deleted, so the standing TOON-first intake, spendPriority ranker,
every-candidate accounting, and load-trigger contract stay exactly as
before this PR. The mid-task quota wake is optional and must not alter
default routing.
Restore the quota-array-dispatch skill ownership line to section 4 as
the always-loaded intake boundary owner; keep the worker-side helper
section as an addition only, without rewiring ownership or load
triggers to section 13.
* fix(bin): use harness-keyed quota matching in optional helper
Revert fm-quota-choose.sh from harness:provider:model tuples back to
harness:model candidates with harness-keyed provider matching, per the
resolved ask-user finding. The helper is optional; authoritative
multi-provider routing (provider discovery from the harness catalog and
quota matching by that explicit provider) stays owned by AGENTS.md
section 4 and the quota-array-dispatch skill intake procedure, not the
helper.
Document the multi-provider limitation in the helper header and the
quota-array-dispatch skill: the helper maps each harness to one primary
provider family only, so a candidate whose established provider differs
from that primary family is checked against the wrong quota row. Use it
only when the brief fixed the candidate order and every candidate's
provider is the harness's primary family.
The helper still consumes one already-captured default-TOON or JSON
snapshot via stdin or --snapshot and never calls quota-axi itself, so
it selects from the same quota state as the intake.
* no-mistakes(review): Fix Muse quota mapping and helper contract docs
* no-mistakes(review): Reject known-empty quotas and map quota tests explicitly
* no-mistakes(review): Preserve unmeasured candidates and enforce snapshot reuse
* no-mistakes(review): Fix quota retirement and dependent regression coverage
* no-mistakes(review): Accept zero-row quota TOON snapshots
* no-mistakes(review): Enforce quota semantics status consistency
* no-mistakes(review): Veto dispatch on any exhausted applicable scope
* no-mistakes(review): Record exhausted quota scope in wake details
* no-mistakes(review): Fix quota help and control dependency coverage
* no-mistakes(review): Decode quoted TOON fields and document quota wakes
* no-mistakes(review): Validate zero-row TOON and map timeout coverage
* no-mistakes(review): Reject multi-value JSON and malformed TOON envelopes
* no-mistakes(review): Validate complete nonzero TOON envelopes
* no-mistakes(review): Accept producer-shaped quota TOON envelopes
* no-mistakes(review): Support empty quota arrays and validate counted rows
* no-mistakes(review): Harden TOON completion, scopes, and quoted fields
* no-mistakes(review): Preserve unknown-headroom exhaustion and reject trailing fields
* no-mistakes(review): Allow unknown headroom under known semantics
* no-mistakes(review): Reject noncanonical quota identities
* no-mistakes(review): Preserve empty quota polling and validate attention identities
* no-mistakes(review): Reject noncanonical provider watches
* no-mistakes(review): Validate all candidates before quota selection
* no-mistakes(document): Correct quota helper safety documentation
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* fix: surface comments on Lavish annotations (#3371)
* fix(bin): keep typed Lavish comments when an element is also annotated
read preferred element text over prompt, so an annotate-and-comment
item dropped the captain's words. Surface prompt as its own field.
Co-authored-by: Cursor <cursoragent@cursor.com>
* no-mistakes(review): Filter non-comment prompts from Lavish reader output
* no-mistakes(document): Clarify Lavish comment presentation contract
* no-mistakes(ci): Fixed Lavish reader comment provenance: non-choice prompts are now emitted even when identical to element text. Added observable regression coverage for identical selector+comment input while retaining pure annotation/message coverage. Reader cases, bash syntax, and diff checks pass. Full fm-procevent suite stops earlier at unrelated “reconcile never claimed” setup failure
* no-mistakes(ci): Fixed duplicate pure-annotation prompts by emitting `prompt:` only when it differs from captured element text. Updated behavioral coverage for selector+comment, pure annotation, and pure message cases. Focused reader regressions, syntax checks, and diff checks pass. Full suite remains blocked by the pre-existing “reconcile never claimed the registered source” failure
* fix(bin): always emit Lavish comments and use real annotation fixtures
Stop inferring comment provenance from prompt==text. Real pure
annotations have no prompt, so always-emit does not duplicate.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: support first public-followup registration on Bash 3.2 (#3420)
* Fix public-followup register crashing on empty lock arrays under bash 3.2.
bash 3.2 with set -u treats "${arr[@]}" on an empty array as unbound, so the first register in a fresh home aborted before taking the registry lock.
The empty-lock regression also runs under the existing stock macOS Bash CI lane so pre-fix code would fail there.
* no-mistakes(document): Document stock Bash registration coverage
* no-mistakes(ci): Pinned the stock macOS Bash CI lane to tasks-axi@0.2.5, eliminating dependency drift. Verified workflow YAML parsing, git diff checks, and the focused regression under /bin/bash 3.2.57 with tasks-axi 0.2.5
* no-mistakes(ci): Fixed the flaky portable CI test: it treated exited zombie processes as live because `kill -0` succeeds for zombies. The watcher and descendant assertions now check process state and regard zombies as exited. Verified `tests/fm-pr-check-security.test.sh`, ShellCheck, `git diff --check`, and the focused Bash public-followup regression
* fix(bin): isolate new Herdr server environments (#2792)
* fix(herdr): isolate server launch environment
* no-mistakes(review): Clear inherited supervision model from Herdr launches
* no-mistakes(document): Document Herdr server launch environment isolation
* fix: surface inbound Relay media to responding agents (#3442)
* fix: surface inbound Relay attachments to the responding agent
A Discord support thread's screenshots were never seen by the agent
handling the mention. The relay delivered them and the poll stashed
them: the reporter's images arrived on the `thread_starter` entry of
`in_reply_to_chain` while the mention's own media list was empty. The
gap was in the responder's playbook, which enumerated a fixed field
list (`request_id`, `text`, `in_reply_to`, `in_reply_to_chain`) and so
made every other field, attachments included, invisible.
Fix it where the gap is, in prose:
- Read the complete payload object rather than a fixed field list, so
media and later relay fields are never skipped again.
- Fetch and view attached media with the agent's own tools, on the
mention and on every chain entry, and call out the common shape where
only the thread starter carries the screenshots.
- Restrict those fetches to known-good platform media hosts over https
(Discord: cdn.discordapp.com, media.discordapp.net,
images-ext-1.discordapp.net, images-ext-2.discordapp.net; X:
pbs.twimg.com, video.twimg.com), report a blocked host instead of
working around it, and treat everything fetched as untrusted public
input on the same terms as the surrounding thread text.
The poll stays out of it and downloads nothing, so no third-party bytes
are pulled on the polling path.
The new test pins the contract the playbook depends on: a mention in the
incident's shape, with an empty top-level media list and screenshots on
the thread starter, must reach the inbox with the payload intact and its
media URLs unfetched.
* no-mistakes(review): Preserve media authority and enforce poll-only fetching
* no-mistakes(document): Clarify Relay attachment safety prose
* fix(bin): defer inactive reconciliation during startup (#3480)
* Defer inactive startup reconciliation
* no-mistakes(review): Queue deferred inactive reconciliation diagnostics durably
* no-mistakes(review): Require worker phases to cover startup requests
* no-mistakes(review): Make diagnostic wakes safely acknowledgeable
* no-mistakes(document): Document deferred startup phase coverage
* fix(bin): bound wake drain presentation lock waits (#3475)
* fix: bound status presentation lock waits
* no-mistakes(review): Distinguish malformed presentation locks from live contention
* no-mistakes(review): Bound no-ack drain queue lock acquisition
* no-mistakes(document): Document bounded presentation-lock drain behavior
* no-mistakes(lint): Annotate bounded lock output global
* no-mistakes(ci): Added deterministic regression coverage for successful bounded-lock acquisition after live contention, verifying helper-to-caller PID ownership handoff and caller release. Verified with bash syntax checks, git diff checks, and the full fm-wake-queue test suite
* fix(bin): retire public follow-ups in remote homes (#3479)
* fix(relay): close a public loop whose work lives in a remote secondmate home
A public-followup loop bound to a REMOTE secondmate could never be closed.
`clear_public_followup_link` (bin/fm-public-followup.sh:701) required an
absolute recorded `work_home_path` for a `secondmate:*` work home, but a remote
route has no local path on this machine, so registration records that field
empty (bin/fm-public-followup.sh:291). Every close ran that clear first, so
`retire` died with "could not clear the legacy X link ... retained for
reconciliation" forever, and `deliver` posted the public reply and then stranded
the loop at `posted`. `--force` never covered that step.
The clear now goes to the remote home over that route's SSH transport, running
`fm-x-followup.sh --clear <work-id>` through `bin/fm-on.sh`. The route is decided
from `data/secondmates.md` before any local path is consulted, so a same-named
local directory can never stand in for a remote home, and registrations already
on disk retire without needing a new field. `fm-on.sh` passes ssh's status
through, so 255 stays the established "delivered but completion unknown" result
this codebase already reconciles: the close is refused, the registration and the
remote link are left exactly as they were, and the message names the unknown
completion instead of claiming a definite failure.
Local secondmate and `main` work homes are untouched, and `--force` still
governs only the unresolved-obligation refusal.
Three regression cases drive a remote route end to end, faking only the ssh
binary at the FM_SSH_BIN seam and then running the real remote entrypoint
against a local checkout, so the clear that must reach the remote home actually
happens there.
* no-mistakes(review): Guard remote link clears by request identity
* no-mistakes(review): Fail guarded clears on unreadable remote state
* no-mistakes(review): Reject guarded clears on non-writable remote state
* no-mistakes(review): Allow no-link retirement in non-writable remote state
* no-mistakes(document): Correct public-followup verification guarantee count
* no-mistakes(ci): Fixed the guarded link-clear race by ensuring absence is decided under the metadata lock whenever publication is possible. Added a behavioral concurrency regression test. Verified with fm-x-mode and fm-public-followup suites, Bash syntax checks, diff checks, and bin/fm-lint.sh
* no-mistakes(ci): Fixed the guarded link-clear race by refusing an unlocked absence decision when a publisher already owns the metadata lock in a non-writable directory. Added a behavioral concurrency regression test. Verified with fm-x-mode, fm-public-followup, syntax/diff checks, and fm-lint
* no-mistakes(ci): Fixed the guarded-clear race by refusing all guarded clears when the metadata parent is non-writable, including apparent link absence. Added a behavioral regression with a publisher waiting to create the lock, updated remote-retirement expectations and verification docs. Passed fm-x-mode, fm-public-followup, fm-lint, documentation audience, Bash syntax, and diff checks
* fix(relay): bound the guarded remote link clear so it refuses instead of hanging
The guarded clear checks that the remote state directory is writable before
taking the metadata lock, but that check cannot close the window: the parent can
turn non-writable between the check and lock creation, and a lock held by a live
holder is indistinguishable from that at the acquire. `fm_lock_acquire_wait` is
an unbounded `while ! try; do sleep 0.1; done`, so either case retried forever
and `deliver` or `retire` wedged with nothing reported, instead of returning the
retained-for-reconciliation refusal the guard exists to produce. This path runs
unattended over the secondmate transport, where a wedge is worse than either
outcome the guard defines.
The guarded clear now acquires through `fm_lock_acquire_wait_bounded`
(FMX_LINK_CLEAR_LOCK_TIMEOUT, default 10 seconds) and refuses on timeout through
the existing failure path. Unguarded local callers keep the ordinary unbounded
wait, so local behavior is unchanged.
The bounded primitive's header no longer claims presentation-only scope, since
this is a second authorized caller; nothing else in the shared lock
infrastructure changed.
The regression holds the metadata lock with a genuinely live process while
leaving the state directory writable, so the refusal can only come from the
bound and never from the writability precondition. Against the unbounded wait it
does not terminate at all; with the bound it refuses, retains the registration,
writes no receipt, and leaves the remote link untouched.
* no-mistakes(review): Harden lock-timeout regression with independent deadline
* no-mistakes(review): Restore no-op guarded clears on read-only state
* no-mistakes(document): Clarify remote public-followup cleanup contract
* fix(bin): support process events under symlinked homes (#3484)
* fix(bin): resolve process-event state roots before validating them
The process-event module validated the caller's spelling of a home's state
root instead of the directory it operates on: it required the supplied path
to equal its own lexical normalization, which rejects any path reached
through a symlinked ancestor. On macOS both /tmp and $TMPDIR are symlinks,
so an operator home under either could never claim a source. Reconcile still
reported the runner started, while the detached runner died writing "cannot
claim source" to the discarded stderr, and the source silently never fired.
Resolve the state root to its physical directory once, then apply the
existing private-directory validation to that resolved directory and derive
every path, recorded claim identity, and later confinement check from it.
This keeps the confinement contract for the directory actually operated on
rather than only for callers that already spelled it physically, and removes
the window where an ancestor symlink could be repointed between check and
use. Homes already spelled physically behave identically.
This was the single cause of both deterministic macOS failures in
tests/fm-procevent.test.sh ("reconcile never claimed the registered source")
and tests/fm-procevent-when.test.sh ("the winning concurrent arm did not
produce an outcome"). The new case pins the behavior with an explicit
symlinked-ancestor home, so it fails without the fix on any platform rather
than only where the temp root happens to be a symlink.
* fix(bin): pin the external capture staging boundary to its physical path
The extension capture path pinned its registry staging boundary by comparing
`pwd -P` against the caller-spelled registry directory, so a home reached
through a symlinked ancestor still refused to start an extension-backed
source after the state root itself resolved correctly. That left such a home
half working: built-in sources ran while external ones failed.
The staging preparer now prints the physical registry directory it validated,
matching the inbox and reservation preparers beside it, and the start path
pins on that returned path. The new end-to-end case drives the shipped
file-signal package from a symlinked home spelling.
* no-mistakes(review): Propagate canonical process-event state roots
* no-mistakes(review): Propagate canonical state to process-event adapters
* no-mistakes(document): Document physical process-event state roots
* fix(pi): deliver captain outcomes as deterministic transcript entries (#3312)
* fix(pi): persist captain outcomes visibly
* no-mistakes(review): Recover captain outcomes after cold-start lock acquisition
* no-mistakes(document): Document cold-start captain-outcome recovery
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes(review): Prove immediate Pi captain-outcome transcript delivery
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* fix(pi): process captain outcomes through a sequence-keyed turn
PR #3312 made every captain-facing supervision outcome a durable, exact-once
visible transcript entry with the read cursor advancing only after that entry
exists. That is the display half of the delivery contract. Left alone it turns
a probabilistic silent loss into a deterministic one: the captain sees an
anchor line, and firstmate never acts, because nothing opens a turn and
nothing records whether main ever processed the outcome.
The 2026-08-31 timeline showed the two shapes this must survive on the
previous hidden-turn path: seven delivered decision outcomes each answered by
an empty assistant message (cursor advanced, no retry, unanswered for close
to three hours), and two answered by an unrelated prior reply. Both happened
because delivery advanced the cursor at enqueue and accepted whatever the
next assistant message was.
Add the processing half on top of the persistence half:
- bin/fm-branch-outcome.sh keeps a processed marker separate from the read
cursor (`unprocessed`, `mark-processed --through`, `processed-init`). It
only advances through an explicit sequence-bound acknowledgement, never
past the read cursor and never backwards; an absent marker reads as zero
and `processed-init` migrates delivered history once so an upgraded home
is not re-presented its past.
- After the visible entry for a captain outcome exists, the extension hands
every still-unprocessed captain row to main as one hidden, typed
`fm-branch-process` request listing each `[seq N] task: summary`, opening
exactly one main turn. Main closes it only by calling the new
`fm_branch_processed` tool with the highest sequence listed. An unrelated,
empty, or paraphrased answer leaves the sequence open, and the same request
is presented again at the end of the next main run and at session start.
The first two presentations of a sequence set open a turn of their own;
after that the request rides the captain's next prompt so an ignored
request cannot loop, and a session replacement resets that budget.
Routine outcomes stay turn-free.
- The regressions cover exactly those incident shapes against the real store
scripts: an empty answer and an unrelated prior answer neither advance the
marker nor stop re-presentation, the acknowledgement is refused beyond the
read cursor and outside lock ownership, a partial acknowledgement keeps the
newer sequence open, and #3312's own assertions now forbid an unkeyed turn
rather than any turn. The store suite pins the marker's bounds and the
migration; the real-SDK guard for appendEntry persistence and model
exclusion is unchanged.
Docs move the protocol from "no model turn" to "one sequence-keyed processing
turn closed only by its acknowledgement", and the verification record carries
the dated run against Pi 0.84.4.
* no-mistakes(review): Harden outcome listing and sequence-bound acknowledgements
* no-mistakes(review): Harden outcome state validation and request pacing
* no-mistakes(review): Reject unsafe sidecars and unterminated outcome stores
* no-mistakes(review): Validate canonical mark-read cursor state
* no-mistakes(review): Guard cursor advancement against corrupt processed state
* no-mistakes(review): Bind acknowledgements to active processing requests
* no-mistakes(review): Reset pacing when processing sequence membership changes
* no-mistakes(review): Enforce silent outcome invariants at storage boundary
* no-mistakes(document): Document hardened captain outcome processing contracts
---------
Co-authored-by: kunchenguid <kun@kunchenguid.com>
* feat: add bounded concurrent Bearings ledger collection (#3481)
* feat: bound Bearings remote ledger collection
* no-mistakes(review): Clarify default remote-ledger collection behavior
* no-mistakes(review): Detach reconcile delivery from watcher loop
* no-mistakes(review): Enforce bounded snapshot and request captures
* no-mistakes(review): Bound legacy summary capture before parsing
* no-mistakes(review): Bound primary remote ledger captures
* no-mistakes(document): Correct snapshot and reconcile documentation
* no-mistakes(lint): Fix ShellCheck quoting in bounded collector
* no-mistakes(ci): Fixed all three CI failures: updated the macOS Bearings assertion to 44 tests, made the home-summary test deterministic and aligned with default ledger consumption, and increased the asynchronous reconcile retirement wait for loaded CI. Verified both focused suites, all 44 Bearings tests, ShellCheck, actionlint, Bash parsing, and git diff checks
* test: await reconcile request retirement
* no-mistakes(review): Avoid empty reconcile queue process churn
* no-mistakes(review): Read ledger summaries from immutable snapshots
* no-mistakes(review): Reject multi-document home ledger streams
* no-mistakes(review): Coalesce durable reconcile requests per target
* no-mistakes(review): Unify reconcile keys and reject snapshot streams
* no-mistakes(review): Key reconcile requests by stable target ID
* no-mistakes(document): Document per-target reconcile request coalescing
* no-mistakes(lint): Remove unused snapshot summary file variable
* no-mistakes(ci): Adjusted the concurrent collector regression’s end-to-end timing ceiling to account for stock macOS process/jq overhead outside the three-second remote collection budget, while remaining below the 15-second serial-read floor. Verified with stock /bin/bash 3.2: all 44 Bearings tests pass; bash syntax and git diff checks pass
* no-mistakes(ci): Fixed legacy summary validation to require exactly one top-level JSON document and added behavioral regression coverage. Stabilized CI by conditionally waiting longer for durable reconcile delivery and synchronously stopping the fm-on worker tree before fixture cleanup. Removed a redundant flaky healthy-path timing assertion; the wedged-reader test still proves concurrent bounded collection. Verified fm-bearings-snapshot, fm-secondmate-reconcile, and fm-on tests, plus project ShellCheck, bash syntax, and git diff checks
* ci: rebalance portable serial test shards (#3489)
* fix(ci): rebalance the portable serial shards on measured durations
The "Behavior portable serial 3" shard ran 17-20 minutes against its
20-minute job cap and intermittently timed out seconds after a passing
test, on branches and on main alike.
Shards are packed longest-processing-time from per-script duration hints,
and those hints were last measured on 2026-08-21 at 116 scripts. The lane
has since grown to 139 scripts and from ~42 to ~63 minutes: 17 scripts had
no hint at all and fell back to the 20 s default, and several existing
hints were low by 2-5x (fm-watch-triage 142 s hinted vs 263 s measured,
fm-public-followup 36 s vs 197 s). The partition therefore looked
perfectly balanced in hint space, 734.6 s per shard, while really running
11.5, 13.6, 18.8 and 16.5 minutes. Script-count balance, which is what the
tests asserted, stayed normal throughout and hid it.
Refresh the hints from the timing artifacts of three green runs, taking
the slowest measurement of each script so the balance holds on a slow
runner, and split the lane across five shards instead of four. Replayed
against those runs' real per-script durations the worst shard is now
12.54 minutes, 63% of the unchanged 20-minute cap, and the serial lane's
wall clock drops from ~20 to ~12.5 minutes.
Bound the drift that caused this rather than relying on the hints being
refreshed by hand: the coverage guard now reports the unmeasured share as
serial_unhinted= and refuses past PORTABLE_SERIAL_MAX_UNHINTED_PERCENT,
which leaves room for newly added tests while making a stale table fail
the guard instead of silently pushing one shard into its cap.
No test changes what it asserts and no test stops running; only the
partition across shards changes.
* no-mistakes(document): Clarify conservative shard timing aggregate
* fix(pi): fall back on incomplete supervision branch prompts (#3491)
* fix(pi): fall back after settled branch errors
* no-mistakes(review): Detect provider errors across prompt compaction
* no-mistakes(review): Preserve in-flight branch state across selection changes
* fix(pi): re-probe supervision branch after cooldown (#3497)
* fix(pi): recover supervision branch after cooldown
* no-mistakes(review): Defer branch recovery until prompt settlement
* no-mistakes(document): Clarify supervision cooldown recovery contract
* fix(bin): remove legacy remote snapshot reads (#3501)
* refactor: remove legacy remote summary reads
* no-mistakes(document): Document ledger-only snapshot reads
* no-mistakes(ci): Fixed the snapshot test fixture so ledger refreshes use the same fake executable PATH as the snapshot consumer. This preserves observable endpoint freshness after removing legacy summary computation. Verified stock Bash parsing and all 44 Bearings tests pass under /bin/bash; git diff checks pass
* no-mistakes(ci): Fixed the CI-only snapshot fixture failure by ensuring the bounded-ledger refresh uses its fake tmux backend. This removes host tmux availability as a source of nondeterminism. Verified all 44 Bearings tests pass, Bash syntax passes, and git diff checks are clean
* no-mistakes(ci): Fixed CI nondeterminism in the Bearings fixture: all local ledger refreshes now use the fixture’s fake tmux backend when available, instead of depending on host tmux state. Verified stock /bin/bash syntax, git diff checks, and all 44 Bearings tests with a deliberately failing host tmux
* fix(pi): preserve watcher continuity across session replacement (#3498)
* fix(pi): rearm watcher after session replacement
* no-mistakes(review): Queue actionable closes across Pi session replacement
* no-mistakes(review): Stop replacement arm when handoff persistence fails
* no-mistakes(review): Preserve actionable wakes through branch and late child races
* no-mistakes(review): Surface late handoff failures without crashing Pi
* no-mistakes(review): Coordinate replacement delivery settlement and unique handoff tokens
* no-mistakes(review): Retry stale deliveries and release settled claims
* no-mistakes(review): Distinguish branch settlement and retry handoff cleanup
* no-mistakes(review): Deduplicate persistent handoff cleanup alerts
* no-mistakes(review): Acknowledge watcher follow-ups only when consumed
* no-mistakes(review): Persist idle follow-ups until agent consumption
* no-mistakes(review): Preserve pending outcomes when handoff persistence fails
* no-mistakes(review): Arm replacement before awaiting prior delivery settlement
* no-mistakes(review): Adopt pending handoffs after lock reclamation
* no-mistakes(review): Prevent stale generations from adopting replacement handoffs
* no-mistakes(review): Scope replacement handoffs by watcher state
* no-mistakes(document): Clarify replacement handoff documentation
* no-mistakes(ci): Fixed the failing branch-extension tests to model the new settlement-promise contract. Failure cases now assert that delivery ownership returns to the watcher instead of expecting direct extension fallback. Verified the updated branch suite, Pi watcher suite, shell syntax, and diff checks
* no-mistakes(review): Update branch settlement tests and preserve chunked outcomes
* no-mistakes(document): Document watcher-owned replacement handoffs
* no-mistakes(document): Verify replacement handoff documentation
* test(pi): cover watcher-owned branch fallback
* no-mistakes(document): Refresh watcher-owned fallback documentation
* fix(bin): resurface task statuses missed by wake handling (#3495)
* fix(bin): resurface terminal statuses lost after branch handling
* test(watch): canonicalize process-event fixture homes
* no-mistakes(review): Index branch outcomes by causal status position
* no-mistakes(review): Recover outcome indexes and deduplicate resurfaced statuses
* no-mistakes(review): Handle legacy ambiguity and oversized status diagnostics
* no-mistakes(review): Keep unclassifiable oversized statuses silent
* no-mistakes(document): Document lost-wake outcome backstop
* no-mistakes(document): Update outcome backstop documentation
* no-mistakes(ci): Fixed CI regressions in wake-drain: parseable reserved-key decisions can no longer bypass the durable decision-fold guard, and status output is prepared and receipt-committed before presentation to prevent repeated one-shot outcomes after later failures. Added a behavioral regression for receipt commit failure and retry. Targeted backstop, correlation-token, decision-cursor, open-decision, unread-status, syntax, and diff checks pass locally. Shard-4 failures appeared unrelated/flaky; the network-parallel test passed locally
* no-mistakes(ci): Fixed the Greptile P1 data-loss issue by committing presentation receipts only after prepared output reaches stdout. Added behavioral coverage proving output failure leaves the backstop retryable and receipt failure may duplicate but never lose a presentation. Relevant wake-drain suites and syntax/diff checks pass. The shard-4 Pi extension failure is unrelated to this PR and did not warrant changes
* no-mistakes(ci): Stabilized tests/fm-bootstrap-network-parallel.test.sh by replacing scheduler-sensitive equal-sleep timing with bounded synchronization between mocked fetch and remote probes. This preserves detection of real serialization while avoiding false failures under CI load. Verified with five consecutive test runs, bash syntax validation, ShellCheck, and git diff checks. The separate Pi stock-rendering failure reproduces locally but is unrelated environment/version drift
* no-mistakes(ci): Fixed Behavior portable serial 4 by adding fm-classify-lib.sh and fm-timeout-lib.sh to the broken-root Pi test fixture; fm-branch-outcome.sh now depends on them. Verified the full Pi branch-extension suite with real-Pi checks skipped, the wake-drain outcome-backstop suite, Bash syntax, and git diff checks. Greptile findings are already addressed at HEAD; the no-mistakes attestation failure is external head-SHA state
* fix(bin): collect follow-up results from remote work homes (#3503)
* fix(bin): deliver typed terminal results from remote work homes
A public commitment whose work is bound to a REMOTE secondmate home could
never receive its typed terminal result. `fm-public-followup.sh brief`
printed an emit command carrying this home's own absolute path and this
checkout's own script path, neither of which exists on the machine the
worker runs on, so the worker had nothing it could write to that the
owning home would ever read - and `consume` kept finding nothing while
the promise stayed open.
The brief is now route-aware: for a remote work home it prints that
route's own code root and home with `--stage-in`, so the typed event is
staged in the home where the work actually runs, and the closing
paragraph names the owning home as the one on the other machine instead
of pointing at the path above it. The owning home collects those staged
results over the same SSH route it reaches that secondmate on, because
the transport only runs outbound: `consume` pulls them into its own
inbox and reconciles them exactly as it reconciles a local report.
Collection is non-destructive until the result is durably held, so a
dropped connection cannot lose a terminal result, and a route that could
not be reached is named in `consume`'s output with the promise left open
rather than reported as an empty inbox.
A local work home is untouched: the brief still prints `--home` with this
home and this checkout's script, and the event still lands directly in
this home's typed terminal-result inbox.
This is the emit-side counterpart of the retire/clear fix in #3479 and
reuses the remote-route resolution that landed with it. Reconciling a
loop bound to a remote route now reaches that route, so the existing
remote cases drive `consume` through the same faked transport their
other steps already use.
* no-mistakes(review): Fail loudly on unresolved routes and invalid staging homes
* no-mistakes(review): Fail collection when remote outbox is unreadable
* no-mistakes(review): Surface reassigned remote routes during empty collection
* no-mistakes(review): Fail remote collection on invalid registrations
* no-mistakes(review): Reject unsafe registration entries during remote collection
* no-mistakes(review): Restore healthy empty remote collection behavior
* no-mistakes(review): Skip remote collection for delivered registrations
* no-mistakes(review): Skip delivered registrations before route validation
* no-mistakes(document): Document remote follow-up collection semantics
* fix(bin): exclude secondmates from home-summary validity (#3504)
* fix(bin): exclude secondmates from home-summary child inventory
kind=secondmate meta records never have backlog rows, so counting them in unowned_children or terminal_in_flight made a clean main home look invalid once earlier ledger checks passed.
* no-mistakes(review): Cover terminal secondmate in-flight exclusion
* no-mistakes(ci): Updated the stock macOS Bash CI snapshot expectation from 15 to 16 tests. Verified all 16 snapshot/fleet-view tests pass under Bash 3.2.57 and `git diff --check` succeeds
* fix(bin): self-heal outcome indexes on first drain (#3509)
* fix(bin): self-heal status-outcome indexes on every drain
Missing ready markers were skipping the lost-wake backstop on non-Pi homes because only the Pi branch ran processed-init. Drain now rebuilds those indexes under the outcome lock and fails closed only on a real store fault.
* no-mistakes(review): Guard held-lock initialization and fail marker writes
* no-mistakes(document): Document cross-harness outcome-index self-healing
* fix(bearings): keep active children underway during captain holds (#3505)
* fix(bearings): keep active children underway beside a captain hold
Project each readable home's active children into Underway independently of the home-level captain-decision classification so a hold no longer hides live work.
* no-mistakes(review): Preserve Underway repos and disclose child truncation
* no-mistakes(review): Fall back to task project for Underway repos
* no-mistakes(ci): Updated the stock macOS Bash CI assertion from 44 to 45 Bearings tests, matching the newly added behavioral regression. Verified all 45 tests pass under /bin/bash, Bash syntax checks pass, and git diff validation is clean
* fix(pi): settle watcher delivery on Pi accepting the follow-up (#3513)
* fix(pi): settle watcher delivery on Pi accepting the follow-up
A follow-up queued while main is streaming joins the running run without
ever raising before_agent_start, so waiting on that event before clearing
the successor pipeline (#3498) stalled every later actionable close: no
successor started, no wake was delivered or offered to the branch, and the
turn-end guard woke main to re-arm by hand after every close.
The pipeline now settles once Pi accepts the follow-up. Consumption is
observed at before_agent_start for an idle main and at the user
message_start for a streaming main, and decides only what a replacement
session (/new, /resume, /fork, reload) replays. An exhausted restoration
delivers its typed failure without launching an arm past the retry bound,
which the stall had hidden. The replacement-coordinator map is typed so the
strict no-emit typecheck passes again.
Tests: the doubles no longer raise before_agent_start for a streaming send,
a portable regression drives two actionable closes while main streams and
proves the successor chain plus consumption-scoped replay, and a
credential-free real-SDK probe pins Pi's event contract for both the
streaming and the idle follow-up.
Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a
* fix(pi): retry a verified successor that fails during wake delivery
A verified successor can exit while the wake it was started for is still
being delivered, most plausibly during a branch turn that holds the
settlement for minutes. Its failure close arrived while the pipeline's
single-flight guard was set, so the close handler skipped the retry, and
the pipeline's end no longer launched an arm, which left the live
generation with no watcher and no retry timer.
The close handler now records that failure when the child had reported
readiness and was not retired by the restoration itself, and the pipeline
runs the ordinary bounded, lock-checked retry for it once the delivery
settles. A restoration started for a later pending supersedes it, and an
exhausted restoration still hands repair to main without a further arm.
The regression holds a branch settlement open while the verified
successor exits with a failure and proves one retry watcher starts after
the settlement releases, none while it is held.
Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a
* fix(bin): bound repeat stale wakes for parked workers (#3532)
* fix(bin): bound repeat stale wakes for a parked but live worker
A worker parked on a declared wait - `paused:` for an external or pipeline
wait, or a verified `captain-held` transfer - kept waking firstmate far inside
FM_PAUSE_RESURFACE_SECS. Observed as five consecutive alarms on one
captain-held worker and dozens across a day on a pipeline wait, and reported
upstream as four wakes in 75 minutes against a 3600s window.
pause_state_class deliberately answers `none` for a still-live agent even under
a declared wait, so a worker genuinely waiting on a decision is never silenced.
That classification is correct and is left alone; it routes every parked but
live worker through surface_nonterminal_stale on first sight of each distinct
stale hash, and an idle parked pane still churns its hash on a clock or a token
counter without changing what is being waited on.
Two places let that churn re-alarm:
- surface_nonterminal_stale queued the wake BEFORE consulting whether a wait was
declared, then wrote `.paused-resurfaced-<key>` - the very throttle that should
have suppressed it. The throttle was never read on this path and was advanced
by the wake it should have prevented.
- The hash-change path cleared that throttle through clear_pause_tracking
whenever the classification came back `none`, so each tick also bought the same
declared wait a fresh window. Fixing only the first site changes nothing.
Read the throttle before anything is queued and advance it only on a wake that
really fires, and on the hash-change path reset only the per-hash bookkeeping
while the declaration still stands, via a clear_stale_hash_tracking split so
neither half of clear_pause_tracking is duplicated. The throttle is keyed to the
declaration, not to the pane.
First sight still wakes, so an inconclusive state is still inspected, and the
window's end still re-surfaces once, so a forgotten wait cannot rot invisibly -
noise traded for a bounded cadence, never for silence. The wake identity stays
the plain `stale: <win>` the away-mode handoff depends on.
Tests cover both observed forms and were confirmed to fail against three
deliberate breaks: each site reverted on its own, and a re-surface that never
fires again.
* fix(document): Clarify declared-wait wake cadence documentation
* fix(ci): Captain, fixed the stale-throttle inheritance: cadence markers now bind to the current wait declaration, so replacement paused and captain-held waits each emit their first plain `stale:` wake. Added behavioral coverage for both forms. Bite proof failed as expected when identity matching was removed, then passed after restoration. Full watcher triage suite, `bin/fm-lint.sh`, syntax checks, and diff checks pass. Changes remain uncommitted for the outer executor
* fix(ci): Captain, fixed the confirmed Greptile finding. `resurface_absorbed` now applies a throttle only when its stored declaration scope matches the current wait, so replacement `paused:` and `captain-held` waits surface immediately without changing classification. Added executable coverage for both absorbed forms. Bite proof failed before the fix at the intended assertion; afterward the full watcher triage suite, `bin/fm-lint.sh`, shell syntax checks, and `git diff --check` passed
* fix(bin): accept the away-mode daemon as the turn-end supervision owner (#3567)
* fix(turnend): accept the away-mode daemon as the supervision owner
While state/.afk exists the away-mode daemon owns supervision and runs
bin/fm-watch.sh one-shot: the watcher exits on every wake and the daemon
starts its replacement. The turn-end guard tested for a live watcher
process holding the watch lock at that instant, so a turn boundary that
landed in the hand-off blocked with "TURN WOULD END BLIND" while
supervision was completely healthy, costing a full handling turn each
time.
Reproduced with the real daemon wrapping the real watcher and the real
guard sampling the same home: 6 of 40 samples blocked, every one of them
with the daemon alive and the beacon 2-3 seconds old, and a new watcher
pid on each cycle. After the fix the same reproduction blocks 0 of 40,
and killing the daemon and its watcher (away mode still on, beacon still
fresh) blocks again.
The guard now accepts a live, identity-matched daemon holding this home
as proof of supervision while away mode is active. The identity match is
the same discipline the watcher lock uses, so a recycled pid or a lock
left by a killed daemon proves nothing. The fresh-beacon half of the
predicate is unchanged: a daemon that stops restarting its watcher still
blocks once the beacon passes grace, a home with no supervisor blocks
exactly as before, and with away mode off the strict watcher predicate is
untouched.
The predicate reads only durable state, so it behaves identically for
every primary harness and runtime backend.
* no-mistakes(document): clarify away-mode daemon supervision proof and test coverage
* no-mistakes(document): generalize stale turn-end predicate summary in architecture.md
* fix(backlog): omit --file from row probes for non-markdown backends (#3582)
* fix(backlog): omit markdown file for beads probes
* no-mistakes(document): Narrow backlog addressing doc to mutations for backend-aware probes
* no-mistakes(ci): Fixed the Greptile P2 review comment (the only failing check) on tests/fm-backlog-atomicity.test.sh. The comment correctly noted that an exported TASKS_AXI_BACKEND environment variable would inherit into the spawned scripts and, because fm_tasks_axi_backend gives it top precedence, override each test case's .tasks.toml backend fixture — making the backend-specific argv assertions fail for environmental reasons. Fix: unset TASKS_AXI_BACKEND in the test harness right after sourcing tests/lib.sh, with a comment explaining why, so every case deterministically exercises its declared backend (4 lines added; no production code touched). Verified: reproduced the leak before the fix (TASKS_AXI_BACKEND=beads made the markdown dispatch case fail with 'beads show failed', exactly the reported failure mode); after the fix the full suite passes (0 failures, exit 0) both with and without TASKS_AXI_BACKEND=beads exported. The added lines are shellcheck-clean (the only shellcheck note, SC1091 on the lib.sh source line, pre-exists this change)
* fix(bin): classify progress updates on requested work as routine (#3589)
The supervision branch's verdict rule escalated every outcome that
answered a captain request, so "the work started" and "still working"
notes reached the captain with nothing to look at. The rule now keeps a
finished result of requested work captain-facing, even when healthy, and
treats start or still-working updates that bring no new artifact,
finding, or decision as routine. The captain list for review-ready PRs,
ask-user findings, exhausted blockers, credentials, and destructive or
security-sensitive cases is unchanged, as are the unsolicited-routine,
silent-fleet-review, and doubt-chooses-captain rules.
The fm_branch_report tool description and the two docs that restated the
old unconditional rule now point at the prompt's "Verdict: routine or
captain" section as the one owner instead of carrying a second copy.
* fix(bin): preserve captain calls during teardown (#3595)
* fix(bin): never close a captain call during cleanup
A scout that held its own work item for the captain, which is what
captain-hold-lifecycle prefers ("hold the work item the question gates"),
was closed by bin/fm-teardown.sh's automatic backlog transition. The
completion gate passed, cleanup ran, and the captain's question moved to
Done with no recorded answer: the one thing the policy says must never
happen. `tasks-axi done` closes a held row silently, and nothing in
teardown asked whether the row was the captain's own call.
bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when
the task is still an open captain call, 1 when it is not, 2 when that
cannot be established. It reads the row through the transition library's
backend-aware probe, so it addresses the same backlog teardown does; the
script's other commands now address the configured data directory the
same way instead of FM_HOME, which also fixes captain holds in a home
with a relocated data directory.
Teardown asks `open` before any destructive step and refuses on 2. On 0
only the close changes: after cleanup and still under the task's own
lock, the row gets one "Deliverable of the finished work" line at the end
of its body and returns to Queued through `tasks-axi reopen`, keeping its
hold, so it lands in Captain's Call instead of reading as work under way.
--force does not lift this: it authorizes discarding unlanded work, never
the captain's question. The deliverable goes into the body because
`tasks-axi update --report` rewrites the title of a row that is not Done.
The crash window reuses the pending-close record teardown already stages:
a `mode=retain` line makes the existing replay record the deliverable and
reopen instead of closing, with the same validator, stale-generation
check, cleanup-incomplete marking, and non-blocking bootstrap lock as an
ordinary close. A retained row the captain answered first simply retires
the record. No parallel record type, recovery command, or second bootstrap
loop is introduced.
Regressions run the real executables: the captain-held scout survives
cleanup queued, held, with its deliverable and on the board, only
`answer` closes it, --force keeps it open, and an ordinary scout still
closes with its report; an interrupted cleanup leaves the row untouched
and the next session start retains it; a relocated backlog keeps the
retention in its one configured file; and a ship row whose hold cannot be
read refuses cleanup before anything destructive.
Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np
* no-mistakes(review): Serialize captain holds and fix backend-aware listing
* no-mistakes(document): Update captain-call retention documentation
* no-mistakes(document): Fix relocated captain-hold backlog diagnostics
* fix(bin): deliver secondmate outcomes to the parent channel (#3592)
* fix(bin): deliver every secondmate outcome on the parent channel from the recording scripts
A secondmate's captain-facing outcomes could miss: the mate model addressed
the captain in its own unread chat instead of appending to the parent
channel, and a PR-ready report, a finding, a decision, a blocker, and a
failure all depended on that one remembered append. Make delivery
structural, so the parent channel never depends on the model:
- bin/fm-parent-channel-lib.sh is the one owner of channel resolution and
exact-line append-once; the merge outcome path and the inactive-outcome
scan now publish through it instead of two private copies.
- bin/fm-inactive-reconcile.sh gains a ledger-first path that runs on every
watcher poll in a secondmate home: a direct child's whole terminal done or
failed line is delivered at once with its note, recorded PR, mode, merge
posture, and scout report pointer, keyed and receipted so it is delivered
once, and the inactive path yields to it. `report <task-id>` runs the same
delivery for a caller holding the child's meta lock.
- bin/fm-pr-check.sh publishes the PR-ready line with the canonical URL at
registration.
- bin/fm-captain-hold.sh publishes a hold and its answer, keyed by task id
and resolution-record count, with no new persisted state.
- bin/fm-teardown.sh delivers the child's final line before removing its
record and refuses, retaining every record, while the channel cannot be
written.
- The charter opens with the parent-channel rule and confines the mate's own
appends to judgement; AGENTS.md carries the carve-out at the persona
address rule and the escalation list.
docs/secondmate-parent-channel.md records the design and its coverage, and
docs/verification/secondmate-parent-channel.md records the live run with real
tmux panes and both real watchers delivering every line with no model.
Supersedes #3569.
* no-mistakes(review): Fix parent outcome retries and reconciliation locking
* no-mistakes(review): Prevent busy children from starving ledger delivery
* no-mistakes(review): Correct ledger metadata and hold occurrence handling
* no-mistakes(review): Disambiguate ledger outcomes and normalize hold reasons
* no-mistakes(review): Close ledger races and preserve teardown records
* no-mistakes(document): Correct parent-channel receipt and scanner documentation
* no-mistakes(lint): Quote done arguments for ShellCheck compliance
* no-mistakes(ci): Fixed both CI failures. Updated GOTMP teardown fixtures for the new final-outcome reporter and isolated them from host tmux state. Updated the PR security assertion to distinguish the accepted PR-ready line from duplicate merge outcomes. Verified with both failing test suites, bash syntax checks, and git diff checks
* no-mistakes(ci): Fixed Greptile’s duplicate-delivery race in bin/fm-inactive-reconcile.sh. Ledger events now claim matching already-delivered inactive receipts using the prior status fingerprint, preventing duplicate parent reports while preserving later same-state completions. Added behavioral regression coverage. Verified inactive-reconcile tests, project lint, documentation audience checks, syntax, and diff checks. Teardown tests passed relevant cases before the documented pre-existing herdr-preflight-missing-adapter failure
* fix(bin): sync remote second mates to primary commit (#3599)
* fix(bin): sync remote second-mate homes to the parent primary commit
Session start and remote launch pointed a remote second-mate home at whatever
Firstmate copy its own host kept, so a home that had already advanced past that
copy refused as a non-fast-forward and every other home stopped at the host's
older commit while the primary ran ahead.
The parent now resolves ITS primary default-branch commit with the existing
helper and hands that commit to the host on both paths. Because a remote home
is a standalone clone, the host imports that one commit before advancing -
already present, else from that host's Firstmate copy without moving it, else
from the home's own origin - and then runs the SAME ff_target guards a local
home gets, so dirty, diverged, feature-branch, and unresolvable targets skip
untouched and the ancestry rules keep one owner. An unimportable target now
names /updatefirstmate instead of failing opaquely, and a host still running an
older Firstmate copy is reported the same way rather than echoing a bare
refusal.
The host-local launch leg no longer re-runs its own secondmate sync, so the
spawn it drives cannot re-target that host's copy after the parent has already
converged the home.
/updatefirstmate is unchanged: it still refreshes the remote code root from that
host's origin and then syncs the home to that refreshed copy, which is what the
sync call with no target commit means.
* no-mistakes(document): Document primary-targeted remote secondmate synchronization
* fix(bin): separate captain intent from firstmate specs (#3597)
* fix(bin): split brief task into captain intent and firstmate spec
Keep no-mistakes --intent as the captain's ask plus later captain words, not the build spec or worker tradeoffs.
* fix(bin): stop task-subsection copies at the next heading
Promotion was swallowing the scout Setup contract into Firstmate spec, and pre-subsection briefs lost their # Task body.
* no-mistakes(review): Validate brief content and preserve nested specifications
* no-mistakes(review): Scope placeholder validation to scaffold-only subsection bodies
* no-mistakes(review): Ignore fenced subsection headings during brief validation
* no-mistakes(review): Preserve captain intent across scout promotion
* no-mistakes(review): Enforce safe intent boundaries for legacy promotions
* no-mistakes(review): Allow marked legacy intent and reject empty promotions
* no-mistakes(review): Scope task parsing and overlay legacy intent contracts
* no-mistakes(review): Overlay current intent contract for all no-mistakes spawns
* no-mistakes(review): Preserve later captain clarifications in intent overlays
* no-mistakes(document): Document brief intent enforcement and ownership
* no-mistakes(ci): Updated spawn-related test fixtures to use valid Captain intent and Firstmate spec subsections, corrected launch-path expectations to launch-brief.md, and resolved ShellCheck quoting findings. Verified with fm-lint.sh and 15 affected behavior tests, including real Herdr tests; all passed
* no-mistakes(ci): Updated stale spawn/promotion fixtures in the Muse, Orca, secondmate-harness, and public-followup suites to provide valid Captain's intent and Firstmate spec subsections. Verified full Orca and secondmate-harness suites, targeted public-followup promotion behavior, Bash syntax, diff checks, and fm-lint
* fix: start a fresh supervision branch for every main session (#3600)
* fix(pi): start a new supervision branch conversation per main session
The supervision branch reopened one recorded conversation forever, so
every main session start reloaded the current generated prompt and then
weeks of accumulated thread, where a superseded rule could still outweigh
today's.
The branch conversation is now scoped to one main session: the session
generation owns the recorded conversation, so a cold start, /new,
/resume, /fork, or a reload always builds a new one, while a rebuild
inside one session (a model or effort change) still continues that
session's own conversation.
The dialog mirror re-anchors with it. Its durable cursor records what the
previous branch conversation received, so a /resume or reload - which
keeps main's own session file - would otherwise leave the new branch
blind to dialog main itself still has. The reset is bounded by the
current main session, and the cursor keeps advancing incrementally within
it. The durable outcome store and its processed marker are untouched, so
unacknowledged captain-facing outcomes still re-present on the new main
session.
* no-mistakes(document): Document fresh Pi supervision conversations
* no-mistakes(ci): Fixed the flaky concurrent inbox failure. Lock acquisition now retries when a competing lock disappears between a failed claim and inspection. Added a behavioral regression covering that race. Verified the full inbox test four times, project lint, and git diff checks
* feat: restart second mates after instruction updates (#3614)
* feat(update): restart second mates whose instructions changed
/updatefirstmate pulled new bytes onto disk and then asked each advanced
second mate to re-read them. A running agent holds AGENTS.md and every
loaded skill frozen from launch and no verified harness offers a reload,
so that steer could not reach a loaded skill at all and left the mate
holding two contradictory copies of its own job description.
An eligible mate is now restarted instead, in the same home and endpoint,
through the existing transactional relaunch. The restart is gated on the
mate first writing down the open work it holds only in conversation - the
open-record half of /stow, never its memory sweeps - so an unregistered
captain call is flushed before the conversation is spent. Anything that
leaves the reload unprovable falls back to the old re-read message and is
reported as exactly that, never as a clean reload.
Remote mates take the same path: fm-remote-secondmate-control.sh gains a
relaunch verb whose host-local leg runs that same control plane, since the
mate is an ordinary local secondmate from its host's point of view. The
primary resolves the profile and passes it explicitly, because
config/secondmate-harness is not inherited and the file on that host
belongs to a different home.
fm-update.sh now splits its advanced live mates into a restart set and a
nudge residual, and both sets require a changed instruction surface, which
also closes the over-nudge against the session-start sweep. Restart is
stricter still: a bin/-only advance reloads itself on the next call, so it
never costs a conversation.
Colocated tests cover the gating, the persist-then-restart order, the
task-subset persist request, each unsafe fallback, the remote hop, and the
remote sync's new instruction-surface report.
* no-mistakes(review): Fix restart correlation, concurrent waits, and lifecycle reporting
* no-mistakes(review): Parallelize relaunches and classify replacement incarnations
* no-mistakes(review): Gate restart actions on live agent state
* no-mistakes(review): Handle failed restart workers without hanging
* no-mistakes(review): Nudge legacy remotes and preserve persist recovery
* no-mistakes(review): Document one-time secondmate restart rollout
* no-mistakes(review): Honor arrived replies and refresh remote profiles
* no-mistakes(review): Revert remote parent profile reconciliation
* no-mistakes(review): Reset remote profile defaults and honor published results
* no-mistakes(review): Preserve fallback nudges for unverifiable secondmates
* no-mistakes(document): Document second-mate restart update flow
* no-mistakes(lint): Fix ShellCheck warnings in restart scripts
* perf: accelerate local validation with bounded concurrency (#3644)
* perf(tests): route gate verification through the bounded concurrent runner
Local validation was the pipeline's dominant cost: across 67 recorded
no-mistakes agent sessions on this repo, 99.3% of command execution was
`bash tests/*.test.sh`, run strictly one script at a time, and 2% of those
calls were killed by an agent-guessed timeout and paid for twice.
Three changes, each measured:
- `.no-mistakes.yaml` pins `commands.test` to
`bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner
already owns changed-file selection, bounded concurrency, the refusal of
unproven scripts, and a generous automatic per-script bound, so the gate's
baseline is neither a serial chain nor a guessed timeout. It stays
intent-targeted - the Test step still runs its evidence agent on top - and
excludes the live-Herdr family the required Herdr lane owns.
- `bin/fm-test-run.sh` gives a plain list of script paths the same bounded
automatic scheduler and automatic bound that `--changed` gets. Naming several
subjects is how a verification round asks for exactly those scripts. The
curated selections are untouched: `--lane` still composes CI shards whose
serial lane must stay serial, `--family` is what the required Herdr lane runs,
and `--all` stays a deliberate complete regression.
- `pr-forge` is admitted to the concurrent-safe family registry on two
consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those,
and records `secondmate` and `session-bootstrap` as refused with the exact
script and reason each failed on, so the refusals are actionable rather than
silent.
Measured on this host, 0 failures on both sides:
verification round, 4 scripts 448s chained -> 231s through the runner (-48%)
pr-forge family 409.2s at 1 worker -> 237.9s at 4 (1.72x)
watcher-wake-lock family 1311.1s at 1 worker -> 539.3s at 4 (2.43x)
A fourth lever was implemented and then removed because the measurement
refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made
`fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s
unchanged, back to back. Those sleeps are not overhead added to the clock -
they are how a test waits for a subject moving on fm-watch.sh's own one-second
cadence - so sampling less often only delays detection. It also broke
`fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a
settled condition. CONTRIBUTING.md records that result so the experiment is not
repeated.
* no-mistakes(review): Separate concurrent runs by isolation proof family
* no-mistakes(review): Limit automatic timeouts to changed-file validation
* no-mistakes(document): Clarify validation concurrency documentation
* fix: copy PR URLs from durable records (#3648)
* fix: copy PR URLs from records or abstain, never assemble them
Supervision reported a plausible but dead PR link three times because its
prompt demanded a full https:// URL at a moment when only a PR number was
observable, so the model assembled an owner/repository from memory, and the PR
check then accepted that URL and wrote it into the task record, after which the
model kept defending its own tool-endorsed guess over the worker's real link.
Three changes close that chain without any live forge lookup, so private
forges are treated exactly like public ones:
- bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy
or abstain" section requires a URL to be copied verbatim from a dura…
peterOC26
added a commit
to peterOC26/firstmate
that referenced
this pull request
Sep 20, 2026
* fix: surface inbound Relay media to responding agents (#3442)
* fix: surface inbound Relay attachments to the responding agent
A Discord support thread's screenshots were never seen by the agent
handling the mention. The relay delivered them and the poll stashed
them: the reporter's images arrived on the `thread_starter` entry of
`in_reply_to_chain` while the mention's own media list was empty. The
gap was in the responder's playbook, which enumerated a fixed field
list (`request_id`, `text`, `in_reply_to`, `in_reply_to_chain`) and so
made every other field, attachments included, invisible.
Fix it where the gap is, in prose:
- Read the complete payload object rather than a fixed field list, so
media and later relay fields are never skipped again.
- Fetch and view attached media with the agent's own tools, on the
mention and on every chain entry, and call out the common shape where
only the thread starter carries the screenshots.
- Restrict those fetches to known-good platform media hosts over https
(Discord: cdn.discordapp.com, media.discordapp.net,
images-ext-1.discordapp.net, images-ext-2.discordapp.net; X:
pbs.twimg.com, video.twimg.com), report a blocked host instead of
working around it, and treat everything fetched as untrusted public
input on the same terms as the surrounding thread text.
The poll stays out of it and downloads nothing, so no third-party bytes
are pulled on the polling path.
The new test pins the contract the playbook depends on: a mention in the
incident's shape, with an empty top-level media list and screenshots on
the thread starter, must reach the inbox with the payload intact and its
media URLs unfetched.
* no-mistakes(review): Preserve media authority and enforce poll-only fetching
* no-mistakes(document): Clarify Relay attachment safety prose
* fix(bin): defer inactive reconciliation during startup (#3480)
* Defer inactive startup reconciliation
* no-mistakes(review): Queue deferred inactive reconciliation diagnostics durably
* no-mistakes(review): Require worker phases to cover startup requests
* no-mistakes(review): Make diagnostic wakes safely acknowledgeable
* no-mistakes(document): Document deferred startup phase coverage
* fix(bin): bound wake drain presentation lock waits (#3475)
* fix: bound status presentation lock waits
* no-mistakes(review): Distinguish malformed presentation locks from live contention
* no-mistakes(review): Bound no-ack drain queue lock acquisition
* no-mistakes(document): Document bounded presentation-lock drain behavior
* no-mistakes(lint): Annotate bounded lock output global
* no-mistakes(ci): Added deterministic regression coverage for successful bounded-lock acquisition after live contention, verifying helper-to-caller PID ownership handoff and caller release. Verified with bash syntax checks, git diff checks, and the full fm-wake-queue test suite
* fix(bin): retire public follow-ups in remote homes (#3479)
* fix(relay): close a public loop whose work lives in a remote secondmate home
A public-followup loop bound to a REMOTE secondmate could never be closed.
`clear_public_followup_link` (bin/fm-public-followup.sh:701) required an
absolute recorded `work_home_path` for a `secondmate:*` work home, but a remote
route has no local path on this machine, so registration records that field
empty (bin/fm-public-followup.sh:291). Every close ran that clear first, so
`retire` died with "could not clear the legacy X link ... retained for
reconciliation" forever, and `deliver` posted the public reply and then stranded
the loop at `posted`. `--force` never covered that step.
The clear now goes to the remote home over that route's SSH transport, running
`fm-x-followup.sh --clear <work-id>` through `bin/fm-on.sh`. The route is decided
from `data/secondmates.md` before any local path is consulted, so a same-named
local directory can never stand in for a remote home, and registrations already
on disk retire without needing a new field. `fm-on.sh` passes ssh's status
through, so 255 stays the established "delivered but completion unknown" result
this codebase already reconciles: the close is refused, the registration and the
remote link are left exactly as they were, and the message names the unknown
completion instead of claiming a definite failure.
Local secondmate and `main` work homes are untouched, and `--force` still
governs only the unresolved-obligation refusal.
Three regression cases drive a remote route end to end, faking only the ssh
binary at the FM_SSH_BIN seam and then running the real remote entrypoint
against a local checkout, so the clear that must reach the remote home actually
happens there.
* no-mistakes(review): Guard remote link clears by request identity
* no-mistakes(review): Fail guarded clears on unreadable remote state
* no-mistakes(review): Reject guarded clears on non-writable remote state
* no-mistakes(review): Allow no-link retirement in non-writable remote state
* no-mistakes(document): Correct public-followup verification guarantee count
* no-mistakes(ci): Fixed the guarded link-clear race by ensuring absence is decided under the metadata lock whenever publication is possible. Added a behavioral concurrency regression test. Verified with fm-x-mode and fm-public-followup suites, Bash syntax checks, diff checks, and bin/fm-lint.sh
* no-mistakes(ci): Fixed the guarded link-clear race by refusing an unlocked absence decision when a publisher already owns the metadata lock in a non-writable directory. Added a behavioral concurrency regression test. Verified with fm-x-mode, fm-public-followup, syntax/diff checks, and fm-lint
* no-mistakes(ci): Fixed the guarded-clear race by refusing all guarded clears when the metadata parent is non-writable, including apparent link absence. Added a behavioral regression with a publisher waiting to create the lock, updated remote-retirement expectations and verification docs. Passed fm-x-mode, fm-public-followup, fm-lint, documentation audience, Bash syntax, and diff checks
* fix(relay): bound the guarded remote link clear so it refuses instead of hanging
The guarded clear checks that the remote state directory is writable before
taking the metadata lock, but that check cannot close the window: the parent can
turn non-writable between the check and lock creation, and a lock held by a live
holder is indistinguishable from that at the acquire. `fm_lock_acquire_wait` is
an unbounded `while ! try; do sleep 0.1; done`, so either case retried forever
and `deliver` or `retire` wedged with nothing reported, instead of returning the
retained-for-reconciliation refusal the guard exists to produce. This path runs
unattended over the secondmate transport, where a wedge is worse than either
outcome the guard defines.
The guarded clear now acquires through `fm_lock_acquire_wait_bounded`
(FMX_LINK_CLEAR_LOCK_TIMEOUT, default 10 seconds) and refuses on timeout through
the existing failure path. Unguarded local callers keep the ordinary unbounded
wait, so local behavior is unchanged.
The bounded primitive's header no longer claims presentation-only scope, since
this is a second authorized caller; nothing else in the shared lock
infrastructure changed.
The regression holds the metadata lock with a genuinely live process while
leaving the state directory writable, so the refusal can only come from the
bound and never from the writability precondition. Against the unbounded wait it
does not terminate at all; with the bound it refuses, retains the registration,
writes no receipt, and leaves the remote link untouched.
* no-mistakes(review): Harden lock-timeout regression with independent deadline
* no-mistakes(review): Restore no-op guarded clears on read-only state
* no-mistakes(document): Clarify remote public-followup cleanup contract
* fix(bin): support process events under symlinked homes (#3484)
* fix(bin): resolve process-event state roots before validating them
The process-event module validated the caller's spelling of a home's state
root instead of the directory it operates on: it required the supplied path
to equal its own lexical normalization, which rejects any path reached
through a symlinked ancestor. On macOS both /tmp and $TMPDIR are symlinks,
so an operator home under either could never claim a source. Reconcile still
reported the runner started, while the detached runner died writing "cannot
claim source" to the discarded stderr, and the source silently never fired.
Resolve the state root to its physical directory once, then apply the
existing private-directory validation to that resolved directory and derive
every path, recorded claim identity, and later confinement check from it.
This keeps the confinement contract for the directory actually operated on
rather than only for callers that already spelled it physically, and removes
the window where an ancestor symlink could be repointed between check and
use. Homes already spelled physically behave identically.
This was the single cause of both deterministic macOS failures in
tests/fm-procevent.test.sh ("reconcile never claimed the registered source")
and tests/fm-procevent-when.test.sh ("the winning concurrent arm did not
produce an outcome"). The new case pins the behavior with an explicit
symlinked-ancestor home, so it fails without the fix on any platform rather
than only where the temp root happens to be a symlink.
* fix(bin): pin the external capture staging boundary to its physical path
The extension capture path pinned its registry staging boundary by comparing
`pwd -P` against the caller-spelled registry directory, so a home reached
through a symlinked ancestor still refused to start an extension-backed
source after the state root itself resolved correctly. That left such a home
half working: built-in sources ran while external ones failed.
The staging preparer now prints the physical registry directory it validated,
matching the inbox and reservation preparers beside it, and the start path
pins on that returned path. The new end-to-end case drives the shipped
file-signal package from a symlinked home spelling.
* no-mistakes(review): Propagate canonical process-event state roots
* no-mistakes(review): Propagate canonical state to process-event adapters
* no-mistakes(document): Document physical process-event state roots
* fix(pi): deliver captain outcomes as deterministic transcript entries (#3312)
* fix(pi): persist captain outcomes visibly
* no-mistakes(review): Recover captain outcomes after cold-start lock acquisition
* no-mistakes(document): Document cold-start captain-outcome recovery
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes(review): Prove immediate Pi captain-outcome transcript delivery
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* fix(pi): process captain outcomes through a sequence-keyed turn
PR #3312 made every captain-facing supervision outcome a durable, exact-once
visible transcript entry with the read cursor advancing only after that entry
exists. That is the display half of the delivery contract. Left alone it turns
a probabilistic silent loss into a deterministic one: the captain sees an
anchor line, and firstmate never acts, because nothing opens a turn and
nothing records whether main ever processed the outcome.
The 2026-08-31 timeline showed the two shapes this must survive on the
previous hidden-turn path: seven delivered decision outcomes each answered by
an empty assistant message (cursor advanced, no retry, unanswered for close
to three hours), and two answered by an unrelated prior reply. Both happened
because delivery advanced the cursor at enqueue and accepted whatever the
next assistant message was.
Add the processing half on top of the persistence half:
- bin/fm-branch-outcome.sh keeps a processed marker separate from the read
cursor (`unprocessed`, `mark-processed --through`, `processed-init`). It
only advances through an explicit sequence-bound acknowledgement, never
past the read cursor and never backwards; an absent marker reads as zero
and `processed-init` migrates delivered history once so an upgraded home
is not re-presented its past.
- After the visible entry for a captain outcome exists, the extension hands
every still-unprocessed captain row to main as one hidden, typed
`fm-branch-process` request listing each `[seq N] task: summary`, opening
exactly one main turn. Main closes it only by calling the new
`fm_branch_processed` tool with the highest sequence listed. An unrelated,
empty, or paraphrased answer leaves the sequence open, and the same request
is presented again at the end of the next main run and at session start.
The first two presentations of a sequence set open a turn of their own;
after that the request rides the captain's next prompt so an ignored
request cannot loop, and a session replacement resets that budget.
Routine outcomes stay turn-free.
- The regressions cover exactly those incident shapes against the real store
scripts: an empty answer and an unrelated prior answer neither advance the
marker nor stop re-presentation, the acknowledgement is refused beyond the
read cursor and outside lock ownership, a partial acknowledgement keeps the
newer sequence open, and #3312's own assertions now forbid an unkeyed turn
rather than any turn. The store suite pins the marker's bounds and the
migration; the real-SDK guard for appendEntry persistence and model
exclusion is unchanged.
Docs move the protocol from "no model turn" to "one sequence-keyed processing
turn closed only by its acknowledgement", and the verification record carries
the dated run against Pi 0.84.4.
* no-mistakes(review): Harden outcome listing and sequence-bound acknowledgements
* no-mistakes(review): Harden outcome state validation and request pacing
* no-mistakes(review): Reject unsafe sidecars and unterminated outcome stores
* no-mistakes(review): Validate canonical mark-read cursor state
* no-mistakes(review): Guard cursor advancement against corrupt processed state
* no-mistakes(review): Bind acknowledgements to active processing requests
* no-mistakes(review): Reset pacing when processing sequence membership changes
* no-mistakes(review): Enforce silent outcome invariants at storage boundary
* no-mistakes(document): Document hardened captain outcome processing contracts
---------
Co-authored-by: kunchenguid <kun@kunchenguid.com>
* feat: add bounded concurrent Bearings ledger collection (#3481)
* feat: bound Bearings remote ledger collection
* no-mistakes(review): Clarify default remote-ledger collection behavior
* no-mistakes(review): Detach reconcile delivery from watcher loop
* no-mistakes(review): Enforce bounded snapshot and request captures
* no-mistakes(review): Bound legacy summary capture before parsing
* no-mistakes(review): Bound primary remote ledger captures
* no-mistakes(document): Correct snapshot and reconcile documentation
* no-mistakes(lint): Fix ShellCheck quoting in bounded collector
* no-mistakes(ci): Fixed all three CI failures: updated the macOS Bearings assertion to 44 tests, made the home-summary test deterministic and aligned with default ledger consumption, and increased the asynchronous reconcile retirement wait for loaded CI. Verified both focused suites, all 44 Bearings tests, ShellCheck, actionlint, Bash parsing, and git diff checks
* test: await reconcile request retirement
* no-mistakes(review): Avoid empty reconcile queue process churn
* no-mistakes(review): Read ledger summaries from immutable snapshots
* no-mistakes(review): Reject multi-document home ledger streams
* no-mistakes(review): Coalesce durable reconcile requests per target
* no-mistakes(review): Unify reconcile keys and reject snapshot streams
* no-mistakes(review): Key reconcile requests by stable target ID
* no-mistakes(document): Document per-target reconcile request coalescing
* no-mistakes(lint): Remove unused snapshot summary file variable
* no-mistakes(ci): Adjusted the concurrent collector regression’s end-to-end timing ceiling to account for stock macOS process/jq overhead outside the three-second remote collection budget, while remaining below the 15-second serial-read floor. Verified with stock /bin/bash 3.2: all 44 Bearings tests pass; bash syntax and git diff checks pass
* no-mistakes(ci): Fixed legacy summary validation to require exactly one top-level JSON document and added behavioral regression coverage. Stabilized CI by conditionally waiting longer for durable reconcile delivery and synchronously stopping the fm-on worker tree before fixture cleanup. Removed a redundant flaky healthy-path timing assertion; the wedged-reader test still proves concurrent bounded collection. Verified fm-bearings-snapshot, fm-secondmate-reconcile, and fm-on tests, plus project ShellCheck, bash syntax, and git diff checks
* ci: rebalance portable serial test shards (#3489)
* fix(ci): rebalance the portable serial shards on measured durations
The "Behavior portable serial 3" shard ran 17-20 minutes against its
20-minute job cap and intermittently timed out seconds after a passing
test, on branches and on main alike.
Shards are packed longest-processing-time from per-script duration hints,
and those hints were last measured on 2026-08-21 at 116 scripts. The lane
has since grown to 139 scripts and from ~42 to ~63 minutes: 17 scripts had
no hint at all and fell back to the 20 s default, and several existing
hints were low by 2-5x (fm-watch-triage 142 s hinted vs 263 s measured,
fm-public-followup 36 s vs 197 s). The partition therefore looked
perfectly balanced in hint space, 734.6 s per shard, while really running
11.5, 13.6, 18.8 and 16.5 minutes. Script-count balance, which is what the
tests asserted, stayed normal throughout and hid it.
Refresh the hints from the timing artifacts of three green runs, taking
the slowest measurement of each script so the balance holds on a slow
runner, and split the lane across five shards instead of four. Replayed
against those runs' real per-script durations the worst shard is now
12.54 minutes, 63% of the unchanged 20-minute cap, and the serial lane's
wall clock drops from ~20 to ~12.5 minutes.
Bound the drift that caused this rather than relying on the hints being
refreshed by hand: the coverage guard now reports the unmeasured share as
serial_unhinted= and refuses past PORTABLE_SERIAL_MAX_UNHINTED_PERCENT,
which leaves room for newly added tests while making a stale table fail
the guard instead of silently pushing one shard into its cap.
No test changes what it asserts and no test stops running; only the
partition across shards changes.
* no-mistakes(document): Clarify conservative shard timing aggregate
* fix(pi): fall back on incomplete supervision branch prompts (#3491)
* fix(pi): fall back after settled branch errors
* no-mistakes(review): Detect provider errors across prompt compaction
* no-mistakes(review): Preserve in-flight branch state across selection changes
* fix(pi): re-probe supervision branch after cooldown (#3497)
* fix(pi): recover supervision branch after cooldown
* no-mistakes(review): Defer branch recovery until prompt settlement
* no-mistakes(document): Clarify supervision cooldown recovery contract
* fix(bin): remove legacy remote snapshot reads (#3501)
* refactor: remove legacy remote summary reads
* no-mistakes(document): Document ledger-only snapshot reads
* no-mistakes(ci): Fixed the snapshot test fixture so ledger refreshes use the same fake executable PATH as the snapshot consumer. This preserves observable endpoint freshness after removing legacy summary computation. Verified stock Bash parsing and all 44 Bearings tests pass under /bin/bash; git diff checks pass
* no-mistakes(ci): Fixed the CI-only snapshot fixture failure by ensuring the bounded-ledger refresh uses its fake tmux backend. This removes host tmux availability as a source of nondeterminism. Verified all 44 Bearings tests pass, Bash syntax passes, and git diff checks are clean
* no-mistakes(ci): Fixed CI nondeterminism in the Bearings fixture: all local ledger refreshes now use the fixture’s fake tmux backend when available, instead of depending on host tmux state. Verified stock /bin/bash syntax, git diff checks, and all 44 Bearings tests with a deliberately failing host tmux
* fix(pi): preserve watcher continuity across session replacement (#3498)
* fix(pi): rearm watcher after session replacement
* no-mistakes(review): Queue actionable closes across Pi session replacement
* no-mistakes(review): Stop replacement arm when handoff persistence fails
* no-mistakes(review): Preserve actionable wakes through branch and late child races
* no-mistakes(review): Surface late handoff failures without crashing Pi
* no-mistakes(review): Coordinate replacement delivery settlement and unique handoff tokens
* no-mistakes(review): Retry stale deliveries and release settled claims
* no-mistakes(review): Distinguish branch settlement and retry handoff cleanup
* no-mistakes(review): Deduplicate persistent handoff cleanup alerts
* no-mistakes(review): Acknowledge watcher follow-ups only when consumed
* no-mistakes(review): Persist idle follow-ups until agent consumption
* no-mistakes(review): Preserve pending outcomes when handoff persistence fails
* no-mistakes(review): Arm replacement before awaiting prior delivery settlement
* no-mistakes(review): Adopt pending handoffs after lock reclamation
* no-mistakes(review): Prevent stale generations from adopting replacement handoffs
* no-mistakes(review): Scope replacement handoffs by watcher state
* no-mistakes(document): Clarify replacement handoff documentation
* no-mistakes(ci): Fixed the failing branch-extension tests to model the new settlement-promise contract. Failure cases now assert that delivery ownership returns to the watcher instead of expecting direct extension fallback. Verified the updated branch suite, Pi watcher suite, shell syntax, and diff checks
* no-mistakes(review): Update branch settlement tests and preserve chunked outcomes
* no-mistakes(document): Document watcher-owned replacement handoffs
* no-mistakes(document): Verify replacement handoff documentation
* test(pi): cover watcher-owned branch fallback
* no-mistakes(document): Refresh watcher-owned fallback documentation
* fix(bin): resurface task statuses missed by wake handling (#3495)
* fix(bin): resurface terminal statuses lost after branch handling
* test(watch): canonicalize process-event fixture homes
* no-mistakes(review): Index branch outcomes by causal status position
* no-mistakes(review): Recover outcome indexes and deduplicate resurfaced statuses
* no-mistakes(review): Handle legacy ambiguity and oversized status diagnostics
* no-mistakes(review): Keep unclassifiable oversized statuses silent
* no-mistakes(document): Document lost-wake outcome backstop
* no-mistakes(document): Update outcome backstop documentation
* no-mistakes(ci): Fixed CI regressions in wake-drain: parseable reserved-key decisions can no longer bypass the durable decision-fold guard, and status output is prepared and receipt-committed before presentation to prevent repeated one-shot outcomes after later failures. Added a behavioral regression for receipt commit failure and retry. Targeted backstop, correlation-token, decision-cursor, open-decision, unread-status, syntax, and diff checks pass locally. Shard-4 failures appeared unrelated/flaky; the network-parallel test passed locally
* no-mistakes(ci): Fixed the Greptile P1 data-loss issue by committing presentation receipts only after prepared output reaches stdout. Added behavioral coverage proving output failure leaves the backstop retryable and receipt failure may duplicate but never lose a presentation. Relevant wake-drain suites and syntax/diff checks pass. The shard-4 Pi extension failure is unrelated to this PR and did not warrant changes
* no-mistakes(ci): Stabilized tests/fm-bootstrap-network-parallel.test.sh by replacing scheduler-sensitive equal-sleep timing with bounded synchronization between mocked fetch and remote probes. This preserves detection of real serialization while avoiding false failures under CI load. Verified with five consecutive test runs, bash syntax validation, ShellCheck, and git diff checks. The separate Pi stock-rendering failure reproduces locally but is unrelated environment/version drift
* no-mistakes(ci): Fixed Behavior portable serial 4 by adding fm-classify-lib.sh and fm-timeout-lib.sh to the broken-root Pi test fixture; fm-branch-outcome.sh now depends on them. Verified the full Pi branch-extension suite with real-Pi checks skipped, the wake-drain outcome-backstop suite, Bash syntax, and git diff checks. Greptile findings are already addressed at HEAD; the no-mistakes attestation failure is external head-SHA state
* fix(bin): collect follow-up results from remote work homes (#3503)
* fix(bin): deliver typed terminal results from remote work homes
A public commitment whose work is bound to a REMOTE secondmate home could
never receive its typed terminal result. `fm-public-followup.sh brief`
printed an emit command carrying this home's own absolute path and this
checkout's own script path, neither of which exists on the machine the
worker runs on, so the worker had nothing it could write to that the
owning home would ever read - and `consume` kept finding nothing while
the promise stayed open.
The brief is now route-aware: for a remote work home it prints that
route's own code root and home with `--stage-in`, so the typed event is
staged in the home where the work actually runs, and the closing
paragraph names the owning home as the one on the other machine instead
of pointing at the path above it. The owning home collects those staged
results over the same SSH route it reaches that secondmate on, because
the transport only runs outbound: `consume` pulls them into its own
inbox and reconciles them exactly as it reconciles a local report.
Collection is non-destructive until the result is durably held, so a
dropped connection cannot lose a terminal result, and a route that could
not be reached is named in `consume`'s output with the promise left open
rather than reported as an empty inbox.
A local work home is untouched: the brief still prints `--home` with this
home and this checkout's script, and the event still lands directly in
this home's typed terminal-result inbox.
This is the emit-side counterpart of the retire/clear fix in #3479 and
reuses the remote-route resolution that landed with it. Reconciling a
loop bound to a remote route now reaches that route, so the existing
remote cases drive `consume` through the same faked transport their
other steps already use.
* no-mistakes(review): Fail loudly on unresolved routes and invalid staging homes
* no-mistakes(review): Fail collection when remote outbox is unreadable
* no-mistakes(review): Surface reassigned remote routes during empty collection
* no-mistakes(review): Fail remote collection on invalid registrations
* no-mistakes(review): Reject unsafe registration entries during remote collection
* no-mistakes(review): Restore healthy empty remote collection behavior
* no-mistakes(review): Skip remote collection for delivered registrations
* no-mistakes(review): Skip delivered registrations before route validation
* no-mistakes(document): Document remote follow-up collection semantics
* fix(bin): exclude secondmates from home-summary validity (#3504)
* fix(bin): exclude secondmates from home-summary child inventory
kind=secondmate meta records never have backlog rows, so counting them in unowned_children or terminal_in_flight made a clean main home look invalid once earlier ledger checks passed.
* no-mistakes(review): Cover terminal secondmate in-flight exclusion
* no-mistakes(ci): Updated the stock macOS Bash CI snapshot expectation from 15 to 16 tests. Verified all 16 snapshot/fleet-view tests pass under Bash 3.2.57 and `git diff --check` succeeds
* fix(bin): self-heal outcome indexes on first drain (#3509)
* fix(bin): self-heal status-outcome indexes on every drain
Missing ready markers were skipping the lost-wake backstop on non-Pi homes because only the Pi branch ran processed-init. Drain now rebuilds those indexes under the outcome lock and fails closed only on a real store fault.
* no-mistakes(review): Guard held-lock initialization and fail marker writes
* no-mistakes(document): Document cross-harness outcome-index self-healing
* fix(bearings): keep active children underway during captain holds (#3505)
* fix(bearings): keep active children underway beside a captain hold
Project each readable home's active children into Underway independently of the home-level captain-decision classification so a hold no longer hides live work.
* no-mistakes(review): Preserve Underway repos and disclose child truncation
* no-mistakes(review): Fall back to task project for Underway repos
* no-mistakes(ci): Updated the stock macOS Bash CI assertion from 44 to 45 Bearings tests, matching the newly added behavioral regression. Verified all 45 tests pass under /bin/bash, Bash syntax checks pass, and git diff validation is clean
* fix(pi): settle watcher delivery on Pi accepting the follow-up (#3513)
* fix(pi): settle watcher delivery on Pi accepting the follow-up
A follow-up queued while main is streaming joins the running run without
ever raising before_agent_start, so waiting on that event before clearing
the successor pipeline (#3498) stalled every later actionable close: no
successor started, no wake was delivered or offered to the branch, and the
turn-end guard woke main to re-arm by hand after every close.
The pipeline now settles once Pi accepts the follow-up. Consumption is
observed at before_agent_start for an idle main and at the user
message_start for a streaming main, and decides only what a replacement
session (/new, /resume, /fork, reload) replays. An exhausted restoration
delivers its typed failure without launching an arm past the retry bound,
which the stall had hidden. The replacement-coordinator map is typed so the
strict no-emit typecheck passes again.
Tests: the doubles no longer raise before_agent_start for a streaming send,
a portable regression drives two actionable closes while main streams and
proves the successor chain plus consumption-scoped replay, and a
credential-free real-SDK probe pins Pi's event contract for both the
streaming and the idle follow-up.
Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a
* fix(pi): retry a verified successor that fails during wake delivery
A verified successor can exit while the wake it was started for is still
being delivered, most plausibly during a branch turn that holds the
settlement for minutes. Its failure close arrived while the pipeline's
single-flight guard was set, so the close handler skipped the retry, and
the pipeline's end no longer launched an arm, which left the live
generation with no watcher and no retry timer.
The close handler now records that failure when the child had reported
readiness and was not retired by the restoration itself, and the pipeline
runs the ordinary bounded, lock-checked retry for it once the delivery
settles. A restoration started for a later pending supersedes it, and an
exhausted restoration still hands repair to main without a further arm.
The regression holds a branch settlement open while the verified
successor exits with a failure and proves one retry watcher starts after
the settlement releases, none while it is held.
Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a
* fix(bin): bound repeat stale wakes for parked workers (#3532)
* fix(bin): bound repeat stale wakes for a parked but live worker
A worker parked on a declared wait - `paused:` for an external or pipeline
wait, or a verified `captain-held` transfer - kept waking firstmate far inside
FM_PAUSE_RESURFACE_SECS. Observed as five consecutive alarms on one
captain-held worker and dozens across a day on a pipeline wait, and reported
upstream as four wakes in 75 minutes against a 3600s window.
pause_state_class deliberately answers `none` for a still-live agent even under
a declared wait, so a worker genuinely waiting on a decision is never silenced.
That classification is correct and is left alone; it routes every parked but
live worker through surface_nonterminal_stale on first sight of each distinct
stale hash, and an idle parked pane still churns its hash on a clock or a token
counter without changing what is being waited on.
Two places let that churn re-alarm:
- surface_nonterminal_stale queued the wake BEFORE consulting whether a wait was
declared, then wrote `.paused-resurfaced-<key>` - the very throttle that should
have suppressed it. The throttle was never read on this path and was advanced
by the wake it should have prevented.
- The hash-change path cleared that throttle through clear_pause_tracking
whenever the classification came back `none`, so each tick also bought the same
declared wait a fresh window. Fixing only the first site changes nothing.
Read the throttle before anything is queued and advance it only on a wake that
really fires, and on the hash-change path reset only the per-hash bookkeeping
while the declaration still stands, via a clear_stale_hash_tracking split so
neither half of clear_pause_tracking is duplicated. The throttle is keyed to the
declaration, not to the pane.
First sight still wakes, so an inconclusive state is still inspected, and the
window's end still re-surfaces once, so a forgotten wait cannot rot invisibly -
noise traded for a bounded cadence, never for silence. The wake identity stays
the plain `stale: <win>` the away-mode handoff depends on.
Tests cover both observed forms and were confirmed to fail against three
deliberate breaks: each site reverted on its own, and a re-surface that never
fires again.
* fix(document): Clarify declared-wait wake cadence documentation
* fix(ci): Captain, fixed the stale-throttle inheritance: cadence markers now bind to the current wait declaration, so replacement paused and captain-held waits each emit their first plain `stale:` wake. Added behavioral coverage for both forms. Bite proof failed as expected when identity matching was removed, then passed after restoration. Full watcher triage suite, `bin/fm-lint.sh`, syntax checks, and diff checks pass. Changes remain uncommitted for the outer executor
* fix(ci): Captain, fixed the confirmed Greptile finding. `resurface_absorbed` now applies a throttle only when its stored declaration scope matches the current wait, so replacement `paused:` and `captain-held` waits surface immediately without changing classification. Added executable coverage for both absorbed forms. Bite proof failed before the fix at the intended assertion; afterward the full watcher triage suite, `bin/fm-lint.sh`, shell syntax checks, and `git diff --check` passed
* fix(bin): accept the away-mode daemon as the turn-end supervision owner (#3567)
* fix(turnend): accept the away-mode daemon as the supervision owner
While state/.afk exists the away-mode daemon owns supervision and runs
bin/fm-watch.sh one-shot: the watcher exits on every wake and the daemon
starts its replacement. The turn-end guard tested for a live watcher
process holding the watch lock at that instant, so a turn boundary that
landed in the hand-off blocked with "TURN WOULD END BLIND" while
supervision was completely healthy, costing a full handling turn each
time.
Reproduced with the real daemon wrapping the real watcher and the real
guard sampling the same home: 6 of 40 samples blocked, every one of them
with the daemon alive and the beacon 2-3 seconds old, and a new watcher
pid on each cycle. After the fix the same reproduction blocks 0 of 40,
and killing the daemon and its watcher (away mode still on, beacon still
fresh) blocks again.
The guard now accepts a live, identity-matched daemon holding this home
as proof of supervision while away mode is active. The identity match is
the same discipline the watcher lock uses, so a recycled pid or a lock
left by a killed daemon proves nothing. The fresh-beacon half of the
predicate is unchanged: a daemon that stops restarting its watcher still
blocks once the beacon passes grace, a home with no supervisor blocks
exactly as before, and with away mode off the strict watcher predicate is
untouched.
The predicate reads only durable state, so it behaves identically for
every primary harness and runtime backend.
* no-mistakes(document): clarify away-mode daemon supervision proof and test coverage
* no-mistakes(document): generalize stale turn-end predicate summary in architecture.md
* fix(backlog): omit --file from row probes for non-markdown backends (#3582)
* fix(backlog): omit markdown file for beads probes
* no-mistakes(document): Narrow backlog addressing doc to mutations for backend-aware probes
* no-mistakes(ci): Fixed the Greptile P2 review comment (the only failing check) on tests/fm-backlog-atomicity.test.sh. The comment correctly noted that an exported TASKS_AXI_BACKEND environment variable would inherit into the spawned scripts and, because fm_tasks_axi_backend gives it top precedence, override each test case's .tasks.toml backend fixture — making the backend-specific argv assertions fail for environmental reasons. Fix: unset TASKS_AXI_BACKEND in the test harness right after sourcing tests/lib.sh, with a comment explaining why, so every case deterministically exercises its declared backend (4 lines added; no production code touched). Verified: reproduced the leak before the fix (TASKS_AXI_BACKEND=beads made the markdown dispatch case fail with 'beads show failed', exactly the reported failure mode); after the fix the full suite passes (0 failures, exit 0) both with and without TASKS_AXI_BACKEND=beads exported. The added lines are shellcheck-clean (the only shellcheck note, SC1091 on the lib.sh source line, pre-exists this change)
* fix(bin): classify progress updates on requested work as routine (#3589)
The supervision branch's verdict rule escalated every outcome that
answered a captain request, so "the work started" and "still working"
notes reached the captain with nothing to look at. The rule now keeps a
finished result of requested work captain-facing, even when healthy, and
treats start or still-working updates that bring no new artifact,
finding, or decision as routine. The captain list for review-ready PRs,
ask-user findings, exhausted blockers, credentials, and destructive or
security-sensitive cases is unchanged, as are the unsolicited-routine,
silent-fleet-review, and doubt-chooses-captain rules.
The fm_branch_report tool description and the two docs that restated the
old unconditional rule now point at the prompt's "Verdict: routine or
captain" section as the one owner instead of carrying a second copy.
* fix(bin): preserve captain calls during teardown (#3595)
* fix(bin): never close a captain call during cleanup
A scout that held its own work item for the captain, which is what
captain-hold-lifecycle prefers ("hold the work item the question gates"),
was closed by bin/fm-teardown.sh's automatic backlog transition. The
completion gate passed, cleanup ran, and the captain's question moved to
Done with no recorded answer: the one thing the policy says must never
happen. `tasks-axi done` closes a held row silently, and nothing in
teardown asked whether the row was the captain's own call.
bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when
the task is still an open captain call, 1 when it is not, 2 when that
cannot be established. It reads the row through the transition library's
backend-aware probe, so it addresses the same backlog teardown does; the
script's other commands now address the configured data directory the
same way instead of FM_HOME, which also fixes captain holds in a home
with a relocated data directory.
Teardown asks `open` before any destructive step and refuses on 2. On 0
only the close changes: after cleanup and still under the task's own
lock, the row gets one "Deliverable of the finished work" line at the end
of its body and returns to Queued through `tasks-axi reopen`, keeping its
hold, so it lands in Captain's Call instead of reading as work under way.
--force does not lift this: it authorizes discarding unlanded work, never
the captain's question. The deliverable goes into the body because
`tasks-axi update --report` rewrites the title of a row that is not Done.
The crash window reuses the pending-close record teardown already stages:
a `mode=retain` line makes the existing replay record the deliverable and
reopen instead of closing, with the same validator, stale-generation
check, cleanup-incomplete marking, and non-blocking bootstrap lock as an
ordinary close. A retained row the captain answered first simply retires
the record. No parallel record type, recovery command, or second bootstrap
loop is introduced.
Regressions run the real executables: the captain-held scout survives
cleanup queued, held, with its deliverable and on the board, only
`answer` closes it, --force keeps it open, and an ordinary scout still
closes with its report; an interrupted cleanup leaves the row untouched
and the next session start retains it; a relocated backlog keeps the
retention in its one configured file; and a ship row whose hold cannot be
read refuses cleanup before anything destructive.
Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np
* no-mistakes(review): Serialize captain holds and fix backend-aware listing
* no-mistakes(document): Update captain-call retention documentation
* no-mistakes(document): Fix relocated captain-hold backlog diagnostics
* fix(bin): deliver secondmate outcomes to the parent channel (#3592)
* fix(bin): deliver every secondmate outcome on the parent channel from the recording scripts
A secondmate's captain-facing outcomes could miss: the mate model addressed
the captain in its own unread chat instead of appending to the parent
channel, and a PR-ready report, a finding, a decision, a blocker, and a
failure all depended on that one remembered append. Make delivery
structural, so the parent channel never depends on the model:
- bin/fm-parent-channel-lib.sh is the one owner of channel resolution and
exact-line append-once; the merge outcome path and the inactive-outcome
scan now publish through it instead of two private copies.
- bin/fm-inactive-reconcile.sh gains a ledger-first path that runs on every
watcher poll in a secondmate home: a direct child's whole terminal done or
failed line is delivered at once with its note, recorded PR, mode, merge
posture, and scout report pointer, keyed and receipted so it is delivered
once, and the inactive path yields to it. `report <task-id>` runs the same
delivery for a caller holding the child's meta lock.
- bin/fm-pr-check.sh publishes the PR-ready line with the canonical URL at
registration.
- bin/fm-captain-hold.sh publishes a hold and its answer, keyed by task id
and resolution-record count, with no new persisted state.
- bin/fm-teardown.sh delivers the child's final line before removing its
record and refuses, retaining every record, while the channel cannot be
written.
- The charter opens with the parent-channel rule and confines the mate's own
appends to judgement; AGENTS.md carries the carve-out at the persona
address rule and the escalation list.
docs/secondmate-parent-channel.md records the design and its coverage, and
docs/verification/secondmate-parent-channel.md records the live run with real
tmux panes and both real watchers delivering every line with no model.
Supersedes #3569.
* no-mistakes(review): Fix parent outcome retries and reconciliation locking
* no-mistakes(review): Prevent busy children from starving ledger delivery
* no-mistakes(review): Correct ledger metadata and hold occurrence handling
* no-mistakes(review): Disambiguate ledger outcomes and normalize hold reasons
* no-mistakes(review): Close ledger races and preserve teardown records
* no-mistakes(document): Correct parent-channel receipt and scanner documentation
* no-mistakes(lint): Quote done arguments for ShellCheck compliance
* no-mistakes(ci): Fixed both CI failures. Updated GOTMP teardown fixtures for the new final-outcome reporter and isolated them from host tmux state. Updated the PR security assertion to distinguish the accepted PR-ready line from duplicate merge outcomes. Verified with both failing test suites, bash syntax checks, and git diff checks
* no-mistakes(ci): Fixed Greptile’s duplicate-delivery race in bin/fm-inactive-reconcile.sh. Ledger events now claim matching already-delivered inactive receipts using the prior status fingerprint, preventing duplicate parent reports while preserving later same-state completions. Added behavioral regression coverage. Verified inactive-reconcile tests, project lint, documentation audience checks, syntax, and diff checks. Teardown tests passed relevant cases before the documented pre-existing herdr-preflight-missing-adapter failure
* fix(bin): sync remote second mates to primary commit (#3599)
* fix(bin): sync remote second-mate homes to the parent primary commit
Session start and remote launch pointed a remote second-mate home at whatever
Firstmate copy its own host kept, so a home that had already advanced past that
copy refused as a non-fast-forward and every other home stopped at the host's
older commit while the primary ran ahead.
The parent now resolves ITS primary default-branch commit with the existing
helper and hands that commit to the host on both paths. Because a remote home
is a standalone clone, the host imports that one commit before advancing -
already present, else from that host's Firstmate copy without moving it, else
from the home's own origin - and then runs the SAME ff_target guards a local
home gets, so dirty, diverged, feature-branch, and unresolvable targets skip
untouched and the ancestry rules keep one owner. An unimportable target now
names /updatefirstmate instead of failing opaquely, and a host still running an
older Firstmate copy is reported the same way rather than echoing a bare
refusal.
The host-local launch leg no longer re-runs its own secondmate sync, so the
spawn it drives cannot re-target that host's copy after the parent has already
converged the home.
/updatefirstmate is unchanged: it still refreshes the remote code root from that
host's origin and then syncs the home to that refreshed copy, which is what the
sync call with no target commit means.
* no-mistakes(document): Document primary-targeted remote secondmate synchronization
* fix(bin): separate captain intent from firstmate specs (#3597)
* fix(bin): split brief task into captain intent and firstmate spec
Keep no-mistakes --intent as the captain's ask plus later captain words, not the build spec or worker tradeoffs.
* fix(bin): stop task-subsection copies at the next heading
Promotion was swallowing the scout Setup contract into Firstmate spec, and pre-subsection briefs lost their # Task body.
* no-mistakes(review): Validate brief content and preserve nested specifications
* no-mistakes(review): Scope placeholder validation to scaffold-only subsection bodies
* no-mistakes(review): Ignore fenced subsection headings during brief validation
* no-mistakes(review): Preserve captain intent across scout promotion
* no-mistakes(review): Enforce safe intent boundaries for legacy promotions
* no-mistakes(review): Allow marked legacy intent and reject empty promotions
* no-mistakes(review): Scope task parsing and overlay legacy intent contracts
* no-mistakes(review): Overlay current intent contract for all no-mistakes spawns
* no-mistakes(review): Preserve later captain clarifications in intent overlays
* no-mistakes(document): Document brief intent enforcement and ownership
* no-mistakes(ci): Updated spawn-related test fixtures to use valid Captain intent and Firstmate spec subsections, corrected launch-path expectations to launch-brief.md, and resolved ShellCheck quoting findings. Verified with fm-lint.sh and 15 affected behavior tests, including real Herdr tests; all passed
* no-mistakes(ci): Updated stale spawn/promotion fixtures in the Muse, Orca, secondmate-harness, and public-followup suites to provide valid Captain's intent and Firstmate spec subsections. Verified full Orca and secondmate-harness suites, targeted public-followup promotion behavior, Bash syntax, diff checks, and fm-lint
* fix: start a fresh supervision branch for every main session (#3600)
* fix(pi): start a new supervision branch conversation per main session
The supervision branch reopened one recorded conversation forever, so
every main session start reloaded the current generated prompt and then
weeks of accumulated thread, where a superseded rule could still outweigh
today's.
The branch conversation is now scoped to one main session: the session
generation owns the recorded conversation, so a cold start, /new,
/resume, /fork, or a reload always builds a new one, while a rebuild
inside one session (a model or effort change) still continues that
session's own conversation.
The dialog mirror re-anchors with it. Its durable cursor records what the
previous branch conversation received, so a /resume or reload - which
keeps main's own session file - would otherwise leave the new branch
blind to dialog main itself still has. The reset is bounded by the
current main session, and the cursor keeps advancing incrementally within
it. The durable outcome store and its processed marker are untouched, so
unacknowledged captain-facing outcomes still re-present on the new main
session.
* no-mistakes(document): Document fresh Pi supervision conversations
* no-mistakes(ci): Fixed the flaky concurrent inbox failure. Lock acquisition now retries when a competing lock disappears between a failed claim and inspection. Added a behavioral regression covering that race. Verified the full inbox test four times, project lint, and git diff checks
* feat: restart second mates after instruction updates (#3614)
* feat(update): restart second mates whose instructions changed
/updatefirstmate pulled new bytes onto disk and then asked each advanced
second mate to re-read them. A running agent holds AGENTS.md and every
loaded skill frozen from launch and no verified harness offers a reload,
so that steer could not reach a loaded skill at all and left the mate
holding two contradictory copies of its own job description.
An eligible mate is now restarted instead, in the same home and endpoint,
through the existing transactional relaunch. The restart is gated on the
mate first writing down the open work it holds only in conversation - the
open-record half of /stow, never its memory sweeps - so an unregistered
captain call is flushed before the conversation is spent. Anything that
leaves the reload unprovable falls back to the old re-read message and is
reported as exactly that, never as a clean reload.
Remote mates take the same path: fm-remote-secondmate-control.sh gains a
relaunch verb whose host-local leg runs that same control plane, since the
mate is an ordinary local secondmate from its host's point of view. The
primary resolves the profile and passes it explicitly, because
config/secondmate-harness is not inherited and the file on that host
belongs to a different home.
fm-update.sh now splits its advanced live mates into a restart set and a
nudge residual, and both sets require a changed instruction surface, which
also closes the over-nudge against the session-start sweep. Restart is
stricter still: a bin/-only advance reloads itself on the next call, so it
never costs a conversation.
Colocated tests cover the gating, the persist-then-restart order, the
task-subset persist request, each unsafe fallback, the remote hop, and the
remote sync's new instruction-surface report.
* no-mistakes(review): Fix restart correlation, concurrent waits, and lifecycle reporting
* no-mistakes(review): Parallelize relaunches and classify replacement incarnations
* no-mistakes(review): Gate restart actions on live agent state
* no-mistakes(review): Handle failed restart workers without hanging
* no-mistakes(review): Nudge legacy remotes and preserve persist recovery
* no-mistakes(review): Document one-time secondmate restart rollout
* no-mistakes(review): Honor arrived replies and refresh remote profiles
* no-mistakes(review): Revert remote parent profile reconciliation
* no-mistakes(review): Reset remote profile defaults and honor published results
* no-mistakes(review): Preserve fallback nudges for unverifiable secondmates
* no-mistakes(document): Document second-mate restart update flow
* no-mistakes(lint): Fix ShellCheck warnings in restart scripts
* perf: accelerate local validation with bounded concurrency (#3644)
* perf(tests): route gate verification through the bounded concurrent runner
Local validation was the pipeline's dominant cost: across 67 recorded
no-mistakes agent sessions on this repo, 99.3% of command execution was
`bash tests/*.test.sh`, run strictly one script at a time, and 2% of those
calls were killed by an agent-guessed timeout and paid for twice.
Three changes, each measured:
- `.no-mistakes.yaml` pins `commands.test` to
`bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner
already owns changed-file selection, bounded concurrency, the refusal of
unproven scripts, and a generous automatic per-script bound, so the gate's
baseline is neither a serial chain nor a guessed timeout. It stays
intent-targeted - the Test step still runs its evidence agent on top - and
excludes the live-Herdr family the required Herdr lane owns.
- `bin/fm-test-run.sh` gives a plain list of script paths the same bounded
automatic scheduler and automatic bound that `--changed` gets. Naming several
subjects is how a verification round asks for exactly those scripts. The
curated selections are untouched: `--lane` still composes CI shards whose
serial lane must stay serial, `--family` is what the required Herdr lane runs,
and `--all` stays a deliberate complete regression.
- `pr-forge` is admitted to the concurrent-safe family registry on two
consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those,
and records `secondmate` and `session-bootstrap` as refused with the exact
script and reason each failed on, so the refusals are actionable rather than
silent.
Measured on this host, 0 failures on both sides:
verification round, 4 scripts 448s chained -> 231s through the runner (-48%)
pr-forge family 409.2s at 1 worker -> 237.9s at 4 (1.72x)
watcher-wake-lock family 1311.1s at 1 worker -> 539.3s at 4 (2.43x)
A fourth lever was implemented and then removed because the measurement
refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made
`fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s
unchanged, back to back. Those sleeps are not overhead added to the clock -
they are how a test waits for a subject moving on fm-watch.sh's own one-second
cadence - so sampling less often only delays detection. It also broke
`fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a
settled condition. CONTRIBUTING.md records that result so the experiment is not
repeated.
* no-mistakes(review): Separate concurrent runs by isolation proof family
* no-mistakes(review): Limit automatic timeouts to changed-file validation
* no-mistakes(document): Clarify validation concurrency documentation
* fix: copy PR URLs from durable records (#3648)
* fix: copy PR URLs from records or abstain, never assemble them
Supervision reported a plausible but dead PR link three times because its
prompt demanded a full https:// URL at a moment when only a PR number was
observable, so the model assembled an owner/repository from memory, and the PR
check then accepted that URL and wrote it into the task record, after which the
model kept defending its own tool-endorsed guess over the worker's real link.
Three changes close that chain without any live forge lookup, so private
forges are treated exactly like public ones:
- bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy
or abstain" section requires a URL to be copied verbatim from a durable
record (the done: PR <url> status line, pr= metadata, or the backlog note),
forbids assembling owner, repository, host, or number from memory, and has
the branch report only the identifier it actually holds when no record names
the URL yet, leaving the PR check unarmed until the worker's ready line
arrives. AGENTS.md section 7 and 9 carry the same copy-or-abstain rule for
main in place of the bare full-URL mandate.
- Worker briefs (bin/fm-brief.sh, ship and scout rules) require the full
https:// URL wherever a PR is mentioned - status line, terminal, or summary -
never a bare "PR 108", so the link is in view as early as the number is.
- bin/fm-pr-check.sh refuses, offline and before any side effect, a URL that
the task's own done lines contradict, printing both spellings; a log naming
no URL still records the argument as before. fm_pr_status_ready_urls in
bin/fm-pr-lib.sh owns reading those lines. The refusal also reaches
bin/fm-pr-merge.sh, so nothing merges under a contradicted URL.
Tests cover the offline refusal with zero side effects, the recorded spelling
being accepted, markdown-wrapped and punctuated URLs, working lines not
counting, the merge wrapper propagation, a self-hosted merge request with no
forge call, the prompt carrying the rule, and the brief carrying the worker
rule.
* no-mistakes(review): Remove stale PR URL enforcement
* no-mistakes(ci): Removed backlog notes as an accepted PR identity source. PR URLs may now be copied only from the task’s `done: PR <url>` status or canonical `pr=` metadata; otherwise supervision reports only the known identifier and leaves PR checking unarmed. Updated related guidance/docs and verified with branch-supervision tests, brief tests, ShellCheck, and `git diff --check`
* fix(bin): disable Claude feedback drafts for fleet launches (#3661)
* fix(bin): disable Claude's feedback-draft flow for fleet-launched agents
Scope --settings '{"feedbackDrafts":"off"}' to every Firstmate-launched
Claude crewmate and secondmate, so /bug and /feedback never queue or
submit a bug report on the captain's behalf. feedbackDrafts is the
documented settings key (Claude Code changelog 2.1.247); the
per-launch CLI flag never touches the captain's global settings.json.
Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3
* no-mistakes(review): Prevent managed settings from re-enabling Claude feedback drafts
* no-mistakes(document): Fix Claude feedback documentation formatting
* fix(bin): layer both feedback-draft controls for defense in depth
The prior --settings-only fix can be overridden by a managed Claude
settings policy (feedbackDrafts precedence). Keep CLAUDE_CODE_SEND_FEEDBACK=0
alongside --settings '{"feedbackDrafts":"off"}': either control alone
disables the SendFeedback tool, so a managed override of one still
leaves the other in force.
Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3
* no-mistakes(document): Document Claude feedback-draft suppression ownership
* feat(tests): run three more validation families concurrently (#3662)
* perf(tests): admit three more families to concurrent validation
The three families that `docs/fm-test-isolation-proof.md` recorded as refused
were not refused for concurrency. Each blocker was a test that decided a
property by wall clock, or a script filed where it cannot run. Fixing those
three things admits all three families and recovers 28.6 minutes of local
validation with no assertion removed or weakened.
- `tests/fm-backlog-handoff.test.sh` injected its pre-move crash by killing the
handoff, sleeping a fixed second, then delegating the move to the real
binary. Nothing ever killed the fake, so on a host slow enough for the case's
next assertions to take longer than a second, the orphan woke and completed
the very move the case requires left undone, and recovery then failed with
`Task "pre-move-crash" not found in this backlog`. Watching the two backlogs
during the injected crash showed exactly that, the item moving one second
after the crash. All four crash injections in the file now go through a new
`fm_fake_crash_injector` shim that signals the target and returns only once
it is observably gone, and the pre-move fake never delegates the move at all.
- `tests/fm-session-start.test.sh` proved the startup digest does not block on
a slow current-state read by timing the whole digest against a fixed
eight-second sleep, which a loaded host exceeds without the property being
violated. It now holds that read open until the case releases it and asserts,
the moment the digest returns, that the read has not finished. A digest that
waited would wait indefinitely rather than for an interval a slow host can
out-run, so the assertion is stronger than the bound it replaces. Its scan
budget moves to the maximum, because the old value left two seconds of margin
over the fixed sleep and measured the host rather than the deadline that
`tests/fm-inactive-reconcile.test.sh` owns.
- `fm-backend-herdr-focus-flash-e2e` was filed in the family map's catch-all,
which put it in the portable serial lane, where Linux CI gate-skips it: that
real-Herdr regression was running nowhere. It moves to `real-herdr-gated` and
the required Herdr lane. `fm-claude-stop-autoarm-live-e2e` gate-skips on its
opt-in variable and moves to `live-harness-optin`.
The 28 remaining ungrouped scripts become an enumerated `standalone` family
instead of admitting `unclassified` itself. `unclassified` is the family map's
`*)` arm, so admitting it would silently grant concurrency to every test added
afterwards, which is exactly the population with no proof. A new test still
lands in `unclassified` and stays serial, and `tests/fm-test-run.test.sh`
covers that split behaviorally.
Each family passes two consecutive four-worker proofs with zero failures. On
the production runner, `secondmate` goes 1233.1s to 453.4s, `session-bootstrap`
756.4s to 286.4s, and `standalone` 724.6s to 261.1s: 2.71x overall and 1713.2s
recovered. The whole suite runs 177 scripts in 52.6 minutes of wall clock
against 121 minutes of summed script time.
* no-mistakes(document): Refresh concurrent validation and shard documentation
* no-mistakes(ci): Fixed the real-Herdr focus-flash E2E race exposed by reclassification. Part C now starts its persistent child atomically via `pane run` and verifies stable child identity through Herdr’s public `process-info` interface, avoiding the racy send-text/send-keys sequence and platform-specific `ps` matching. Verified with bash syntax checking, ShellCheck, git diff checks, and the complete E2E test on Herdr 0.8.2
* feat: structure no-mistakes ask-user escalations (#3670)
* feat(brief): structure no-mistakes ask-user escalation as event + snapshot file
Crewmates escalating a no-mistakes ask-user gate now report one status
event naming every finding id plus a snapshot file holding the gate's
axi finding records verbatim (id, severity, file, line, description,
authority), using the same shape even for a single finding. The status
line never paraphrases. The format is defined once in fm-dod-lib.sh and
rendered into both the scout and ship rule 6 in fm-brief.sh, so a
promoted scout - whose rule 6 fm-promote.sh preserves unchanged - gets
the identical contract as a freshly-spawned no-mistakes ship worker.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PpiWaDerbYavTLPPtEjQei
* no-mistakes(review): Preserve ask-user escalation output contract
* no-mistakes(review): Align escalation format test expectation
* no-mistakes(review): Scope ask-user escalation instructions correctly
* no-mistakes(review): Remove ask-user from generic decision rules
---------
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* fix(bin): require self-sufficient no-mistakes intent (#3671)
* fix(bin): require a self-sufficient no-mistakes intent
A no-mistakes worker's --intent is only as useful as the string it
passes. PR #3604 shipped with an intent that was only "do 1, 2, 3, 7
from the report": the real contract lived in a private scout report and
never reached --intent, so nobody holding that string plus the codebase
could have derived the specification.
This is pure instruction at the contract's one owner; no spawn-side or
promotion-side check is added.
- bin/fm-dod-lib.sh: the generated no-mistakes Definition of done now
states that the --intent string must be self-sufficient (the string
plus the codebase reconstructs roughly the same specification) and
tells the worker to write the substance of any report, decision,…
sctru
added a commit
to sctru/firstmate
that referenced
this pull request
Sep 21, 2026
* fix(bin): safely unregister custom checks (#3369)
* fix(bin): add a safe owner for custom-check retirement
Agents were improvising rm of check files with unset STATE/ID, which wedges
headless panes. Unregister validates the id and state directory first.
Co-authored-by: Cursor <cursoragent@cursor.com>
* no-mistakes(review): Refuse explicitly empty custom-check state overrides
* no-mistakes(document): Document custom-check retirement safety contract
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* refactor(quota): extract mid-task polling and candidate selection into dedicated scripts (#3221)
* Add quota exhaustion detection and safe fallback helpers
- bin/fm-procevent-quota.sh: generic procevent adapter that arms a
recurring quota-axi --json poll and wakes firstmate when a tracked
provider's effectivePercentRemaining drops below a threshold or its
runway.status becomes exhausted_now.
- bin/fm-quota-choose.sh: worker-side helper that picks the first ranked
harness:model candidate with positive effectivePercentRemaining.
- AGENTS.md and .agents/skills/quota-array-dispatch/SKILL.md: document
the new helpers and the mid-task quota-exhaustion wake path.
- tests/fm-quota-choose.test.sh: unit tests with a mocked quota-axi JSON
source.
* no-mistakes(review): Fix quota polling and scope bounds
* no-mistakes(review): Enforce safe default quota selection
* no-mistakes(review): Handle decimal quota values safely
* no-mistakes(review): Fail closed on invalid quota inputs
* no-mistakes(review): Reject empty quota candidate segments
* no-mistakes(review): Harden quota parsing and timeout ownership
* no-mistakes(review): Reuse captured quota snapshots consistently
* no-mistakes(review): Match quota using explicit candidate providers
* no-mistakes(review): Centralize fail-closed quota schema validation
* no-mistakes(review): Reject out-of-range quota percentages
* no-mistakes(review): Validate quota runway status enum
* no-mistakes(review): Tighten quota scope and status contracts
* no-mistakes(review): Preserve unknown quota and exact product bounds
* no-mistakes(review): Preserve provider-level unknown quota
* no-mistakes(review): Reuse canonical verified harness validation
* no-mistakes(document): Document mid-task quota handling
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* fix(docs): restore default routing contract, keep quota helper optional
Restore the AGENTS.md section 4 always-loaded routing paragraph the PR
had deleted, so the standing TOON-first intake, spendPriority ranker,
every-candidate accounting, and load-trigger contract stay exactly as
before this PR. The mid-task quota wake is optional and must not alter
default routing.
Restore the quota-array-dispatch skill ownership line to section 4 as
the always-loaded intake boundary owner; keep the worker-side helper
section as an addition only, without rewiring ownership or load
triggers to section 13.
* fix(bin): use harness-keyed quota matching in optional helper
Revert fm-quota-choose.sh from harness:provider:model tuples back to
harness:model candidates with harness-keyed provider matching, per the
resolved ask-user finding. The helper is optional; authoritative
multi-provider routing (provider discovery from the harness catalog and
quota matching by that explicit provider) stays owned by AGENTS.md
section 4 and the quota-array-dispatch skill intake procedure, not the
helper.
Document the multi-provider limitation in the helper header and the
quota-array-dispatch skill: the helper maps each harness to one primary
provider family only, so a candidate whose established provider differs
from that primary family is checked against the wrong quota row. Use it
only when the brief fixed the candidate order and every candidate's
provider is the harness's primary family.
The helper still consumes one already-captured default-TOON or JSON
snapshot via stdin or --snapshot and never calls quota-axi itself, so
it selects from the same quota state as the intake.
* no-mistakes(review): Fix Muse quota mapping and helper contract docs
* no-mistakes(review): Reject known-empty quotas and map quota tests explicitly
* no-mistakes(review): Preserve unmeasured candidates and enforce snapshot reuse
* no-mistakes(review): Fix quota retirement and dependent regression coverage
* no-mistakes(review): Accept zero-row quota TOON snapshots
* no-mistakes(review): Enforce quota semantics status consistency
* no-mistakes(review): Veto dispatch on any exhausted applicable scope
* no-mistakes(review): Record exhausted quota scope in wake details
* no-mistakes(review): Fix quota help and control dependency coverage
* no-mistakes(review): Decode quoted TOON fields and document quota wakes
* no-mistakes(review): Validate zero-row TOON and map timeout coverage
* no-mistakes(review): Reject multi-value JSON and malformed TOON envelopes
* no-mistakes(review): Validate complete nonzero TOON envelopes
* no-mistakes(review): Accept producer-shaped quota TOON envelopes
* no-mistakes(review): Support empty quota arrays and validate counted rows
* no-mistakes(review): Harden TOON completion, scopes, and quoted fields
* no-mistakes(review): Preserve unknown-headroom exhaustion and reject trailing fields
* no-mistakes(review): Allow unknown headroom under known semantics
* no-mistakes(review): Reject noncanonical quota identities
* no-mistakes(review): Preserve empty quota polling and validate attention identities
* no-mistakes(review): Reject noncanonical provider watches
* no-mistakes(review): Validate all candidates before quota selection
* no-mistakes(document): Correct quota helper safety documentation
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* fix: surface comments on Lavish annotations (#3371)
* fix(bin): keep typed Lavish comments when an element is also annotated
read preferred element text over prompt, so an annotate-and-comment
item dropped the captain's words. Surface prompt as its own field.
Co-authored-by: Cursor <cursoragent@cursor.com>
* no-mistakes(review): Filter non-comment prompts from Lavish reader output
* no-mistakes(document): Clarify Lavish comment presentation contract
* no-mistakes(ci): Fixed Lavish reader comment provenance: non-choice prompts are now emitted even when identical to element text. Added observable regression coverage for identical selector+comment input while retaining pure annotation/message coverage. Reader cases, bash syntax, and diff checks pass. Full fm-procevent suite stops earlier at unrelated “reconcile never claimed” setup failure
* no-mistakes(ci): Fixed duplicate pure-annotation prompts by emitting `prompt:` only when it differs from captured element text. Updated behavioral coverage for selector+comment, pure annotation, and pure message cases. Focused reader regressions, syntax checks, and diff checks pass. Full suite remains blocked by the pre-existing “reconcile never claimed the registered source” failure
* fix(bin): always emit Lavish comments and use real annotation fixtures
Stop inferring comment provenance from prompt==text. Real pure
annotations have no prompt, so always-emit does not duplicate.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: support first public-followup registration on Bash 3.2 (#3420)
* Fix public-followup register crashing on empty lock arrays under bash 3.2.
bash 3.2 with set -u treats "${arr[@]}" on an empty array as unbound, so the first register in a fresh home aborted before taking the registry lock.
The empty-lock regression also runs under the existing stock macOS Bash CI lane so pre-fix code would fail there.
* no-mistakes(document): Document stock Bash registration coverage
* no-mistakes(ci): Pinned the stock macOS Bash CI lane to tasks-axi@0.2.5, eliminating dependency drift. Verified workflow YAML parsing, git diff checks, and the focused regression under /bin/bash 3.2.57 with tasks-axi 0.2.5
* no-mistakes(ci): Fixed the flaky portable CI test: it treated exited zombie processes as live because `kill -0` succeeds for zombies. The watcher and descendant assertions now check process state and regard zombies as exited. Verified `tests/fm-pr-check-security.test.sh`, ShellCheck, `git diff --check`, and the focused Bash public-followup regression
* fix(bin): isolate new Herdr server environments (#2792)
* fix(herdr): isolate server launch environment
* no-mistakes(review): Clear inherited supervision model from Herdr launches
* no-mistakes(document): Document Herdr server launch environment isolation
* fix: surface inbound Relay media to responding agents (#3442)
* fix: surface inbound Relay attachments to the responding agent
A Discord support thread's screenshots were never seen by the agent
handling the mention. The relay delivered them and the poll stashed
them: the reporter's images arrived on the `thread_starter` entry of
`in_reply_to_chain` while the mention's own media list was empty. The
gap was in the responder's playbook, which enumerated a fixed field
list (`request_id`, `text`, `in_reply_to`, `in_reply_to_chain`) and so
made every other field, attachments included, invisible.
Fix it where the gap is, in prose:
- Read the complete payload object rather than a fixed field list, so
media and later relay fields are never skipped again.
- Fetch and view attached media with the agent's own tools, on the
mention and on every chain entry, and call out the common shape where
only the thread starter carries the screenshots.
- Restrict those fetches to known-good platform media hosts over https
(Discord: cdn.discordapp.com, media.discordapp.net,
images-ext-1.discordapp.net, images-ext-2.discordapp.net; X:
pbs.twimg.com, video.twimg.com), report a blocked host instead of
working around it, and treat everything fetched as untrusted public
input on the same terms as the surrounding thread text.
The poll stays out of it and downloads nothing, so no third-party bytes
are pulled on the polling path.
The new test pins the contract the playbook depends on: a mention in the
incident's shape, with an empty top-level media list and screenshots on
the thread starter, must reach the inbox with the payload intact and its
media URLs unfetched.
* no-mistakes(review): Preserve media authority and enforce poll-only fetching
* no-mistakes(document): Clarify Relay attachment safety prose
* fix(bin): defer inactive reconciliation during startup (#3480)
* Defer inactive startup reconciliation
* no-mistakes(review): Queue deferred inactive reconciliation diagnostics durably
* no-mistakes(review): Require worker phases to cover startup requests
* no-mistakes(review): Make diagnostic wakes safely acknowledgeable
* no-mistakes(document): Document deferred startup phase coverage
* fix(bin): bound wake drain presentation lock waits (#3475)
* fix: bound status presentation lock waits
* no-mistakes(review): Distinguish malformed presentation locks from live contention
* no-mistakes(review): Bound no-ack drain queue lock acquisition
* no-mistakes(document): Document bounded presentation-lock drain behavior
* no-mistakes(lint): Annotate bounded lock output global
* no-mistakes(ci): Added deterministic regression coverage for successful bounded-lock acquisition after live contention, verifying helper-to-caller PID ownership handoff and caller release. Verified with bash syntax checks, git diff checks, and the full fm-wake-queue test suite
* fix(bin): retire public follow-ups in remote homes (#3479)
* fix(relay): close a public loop whose work lives in a remote secondmate home
A public-followup loop bound to a REMOTE secondmate could never be closed.
`clear_public_followup_link` (bin/fm-public-followup.sh:701) required an
absolute recorded `work_home_path` for a `secondmate:*` work home, but a remote
route has no local path on this machine, so registration records that field
empty (bin/fm-public-followup.sh:291). Every close ran that clear first, so
`retire` died with "could not clear the legacy X link ... retained for
reconciliation" forever, and `deliver` posted the public reply and then stranded
the loop at `posted`. `--force` never covered that step.
The clear now goes to the remote home over that route's SSH transport, running
`fm-x-followup.sh --clear <work-id>` through `bin/fm-on.sh`. The route is decided
from `data/secondmates.md` before any local path is consulted, so a same-named
local directory can never stand in for a remote home, and registrations already
on disk retire without needing a new field. `fm-on.sh` passes ssh's status
through, so 255 stays the established "delivered but completion unknown" result
this codebase already reconciles: the close is refused, the registration and the
remote link are left exactly as they were, and the message names the unknown
completion instead of claiming a definite failure.
Local secondmate and `main` work homes are untouched, and `--force` still
governs only the unresolved-obligation refusal.
Three regression cases drive a remote route end to end, faking only the ssh
binary at the FM_SSH_BIN seam and then running the real remote entrypoint
against a local checkout, so the clear that must reach the remote home actually
happens there.
* no-mistakes(review): Guard remote link clears by request identity
* no-mistakes(review): Fail guarded clears on unreadable remote state
* no-mistakes(review): Reject guarded clears on non-writable remote state
* no-mistakes(review): Allow no-link retirement in non-writable remote state
* no-mistakes(document): Correct public-followup verification guarantee count
* no-mistakes(ci): Fixed the guarded link-clear race by ensuring absence is decided under the metadata lock whenever publication is possible. Added a behavioral concurrency regression test. Verified with fm-x-mode and fm-public-followup suites, Bash syntax checks, diff checks, and bin/fm-lint.sh
* no-mistakes(ci): Fixed the guarded link-clear race by refusing an unlocked absence decision when a publisher already owns the metadata lock in a non-writable directory. Added a behavioral concurrency regression test. Verified with fm-x-mode, fm-public-followup, syntax/diff checks, and fm-lint
* no-mistakes(ci): Fixed the guarded-clear race by refusing all guarded clears when the metadata parent is non-writable, including apparent link absence. Added a behavioral regression with a publisher waiting to create the lock, updated remote-retirement expectations and verification docs. Passed fm-x-mode, fm-public-followup, fm-lint, documentation audience, Bash syntax, and diff checks
* fix(relay): bound the guarded remote link clear so it refuses instead of hanging
The guarded clear checks that the remote state directory is writable before
taking the metadata lock, but that check cannot close the window: the parent can
turn non-writable between the check and lock creation, and a lock held by a live
holder is indistinguishable from that at the acquire. `fm_lock_acquire_wait` is
an unbounded `while ! try; do sleep 0.1; done`, so either case retried forever
and `deliver` or `retire` wedged with nothing reported, instead of returning the
retained-for-reconciliation refusal the guard exists to produce. This path runs
unattended over the secondmate transport, where a wedge is worse than either
outcome the guard defines.
The guarded clear now acquires through `fm_lock_acquire_wait_bounded`
(FMX_LINK_CLEAR_LOCK_TIMEOUT, default 10 seconds) and refuses on timeout through
the existing failure path. Unguarded local callers keep the ordinary unbounded
wait, so local behavior is unchanged.
The bounded primitive's header no longer claims presentation-only scope, since
this is a second authorized caller; nothing else in the shared lock
infrastructure changed.
The regression holds the metadata lock with a genuinely live process while
leaving the state directory writable, so the refusal can only come from the
bound and never from the writability precondition. Against the unbounded wait it
does not terminate at all; with the bound it refuses, retains the registration,
writes no receipt, and leaves the remote link untouched.
* no-mistakes(review): Harden lock-timeout regression with independent deadline
* no-mistakes(review): Restore no-op guarded clears on read-only state
* no-mistakes(document): Clarify remote public-followup cleanup contract
* fix(bin): support process events under symlinked homes (#3484)
* fix(bin): resolve process-event state roots before validating them
The process-event module validated the caller's spelling of a home's state
root instead of the directory it operates on: it required the supplied path
to equal its own lexical normalization, which rejects any path reached
through a symlinked ancestor. On macOS both /tmp and $TMPDIR are symlinks,
so an operator home under either could never claim a source. Reconcile still
reported the runner started, while the detached runner died writing "cannot
claim source" to the discarded stderr, and the source silently never fired.
Resolve the state root to its physical directory once, then apply the
existing private-directory validation to that resolved directory and derive
every path, recorded claim identity, and later confinement check from it.
This keeps the confinement contract for the directory actually operated on
rather than only for callers that already spelled it physically, and removes
the window where an ancestor symlink could be repointed between check and
use. Homes already spelled physically behave identically.
This was the single cause of both deterministic macOS failures in
tests/fm-procevent.test.sh ("reconcile never claimed the registered source")
and tests/fm-procevent-when.test.sh ("the winning concurrent arm did not
produce an outcome"). The new case pins the behavior with an explicit
symlinked-ancestor home, so it fails without the fix on any platform rather
than only where the temp root happens to be a symlink.
* fix(bin): pin the external capture staging boundary to its physical path
The extension capture path pinned its registry staging boundary by comparing
`pwd -P` against the caller-spelled registry directory, so a home reached
through a symlinked ancestor still refused to start an extension-backed
source after the state root itself resolved correctly. That left such a home
half working: built-in sources ran while external ones failed.
The staging preparer now prints the physical registry directory it validated,
matching the inbox and reservation preparers beside it, and the start path
pins on that returned path. The new end-to-end case drives the shipped
file-signal package from a symlinked home spelling.
* no-mistakes(review): Propagate canonical process-event state roots
* no-mistakes(review): Propagate canonical state to process-event adapters
* no-mistakes(document): Document physical process-event state roots
* fix(pi): deliver captain outcomes as deterministic transcript entries (#3312)
* fix(pi): persist captain outcomes visibly
* no-mistakes(review): Recover captain outcomes after cold-start lock acquisition
* no-mistakes(document): Document cold-start captain-outcome recovery
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes(review): Prove immediate Pi captain-outcome transcript delivery
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* fix(pi): process captain outcomes through a sequence-keyed turn
PR #3312 made every captain-facing supervision outcome a durable, exact-once
visible transcript entry with the read cursor advancing only after that entry
exists. That is the display half of the delivery contract. Left alone it turns
a probabilistic silent loss into a deterministic one: the captain sees an
anchor line, and firstmate never acts, because nothing opens a turn and
nothing records whether main ever processed the outcome.
The 2026-08-31 timeline showed the two shapes this must survive on the
previous hidden-turn path: seven delivered decision outcomes each answered by
an empty assistant message (cursor advanced, no retry, unanswered for close
to three hours), and two answered by an unrelated prior reply. Both happened
because delivery advanced the cursor at enqueue and accepted whatever the
next assistant message was.
Add the processing half on top of the persistence half:
- bin/fm-branch-outcome.sh keeps a processed marker separate from the read
cursor (`unprocessed`, `mark-processed --through`, `processed-init`). It
only advances through an explicit sequence-bound acknowledgement, never
past the read cursor and never backwards; an absent marker reads as zero
and `processed-init` migrates delivered history once so an upgraded home
is not re-presented its past.
- After the visible entry for a captain outcome exists, the extension hands
every still-unprocessed captain row to main as one hidden, typed
`fm-branch-process` request listing each `[seq N] task: summary`, opening
exactly one main turn. Main closes it only by calling the new
`fm_branch_processed` tool with the highest sequence listed. An unrelated,
empty, or paraphrased answer leaves the sequence open, and the same request
is presented again at the end of the next main run and at session start.
The first two presentations of a sequence set open a turn of their own;
after that the request rides the captain's next prompt so an ignored
request cannot loop, and a session replacement resets that budget.
Routine outcomes stay turn-free.
- The regressions cover exactly those incident shapes against the real store
scripts: an empty answer and an unrelated prior answer neither advance the
marker nor stop re-presentation, the acknowledgement is refused beyond the
read cursor and outside lock ownership, a partial acknowledgement keeps the
newer sequence open, and #3312's own assertions now forbid an unkeyed turn
rather than any turn. The store suite pins the marker's bounds and the
migration; the real-SDK guard for appendEntry persistence and model
exclusion is unchanged.
Docs move the protocol from "no model turn" to "one sequence-keyed processing
turn closed only by its acknowledgement", and the verification record carries
the dated run against Pi 0.84.4.
* no-mistakes(review): Harden outcome listing and sequence-bound acknowledgements
* no-mistakes(review): Harden outcome state validation and request pacing
* no-mistakes(review): Reject unsafe sidecars and unterminated outcome stores
* no-mistakes(review): Validate canonical mark-read cursor state
* no-mistakes(review): Guard cursor advancement against corrupt processed state
* no-mistakes(review): Bind acknowledgements to active processing requests
* no-mistakes(review): Reset pacing when processing sequence membership changes
* no-mistakes(review): Enforce silent outcome invariants at storage boundary
* no-mistakes(document): Document hardened captain outcome processing contracts
---------
Co-authored-by: kunchenguid <kun@kunchenguid.com>
* feat: add bounded concurrent Bearings ledger collection (#3481)
* feat: bound Bearings remote ledger collection
* no-mistakes(review): Clarify default remote-ledger collection behavior
* no-mistakes(review): Detach reconcile delivery from watcher loop
* no-mistakes(review): Enforce bounded snapshot and request captures
* no-mistakes(review): Bound legacy summary capture before parsing
* no-mistakes(review): Bound primary remote ledger captures
* no-mistakes(document): Correct snapshot and reconcile documentation
* no-mistakes(lint): Fix ShellCheck quoting in bounded collector
* no-mistakes(ci): Fixed all three CI failures: updated the macOS Bearings assertion to 44 tests, made the home-summary test deterministic and aligned with default ledger consumption, and increased the asynchronous reconcile retirement wait for loaded CI. Verified both focused suites, all 44 Bearings tests, ShellCheck, actionlint, Bash parsing, and git diff checks
* test: await reconcile request retirement
* no-mistakes(review): Avoid empty reconcile queue process churn
* no-mistakes(review): Read ledger summaries from immutable snapshots
* no-mistakes(review): Reject multi-document home ledger streams
* no-mistakes(review): Coalesce durable reconcile requests per target
* no-mistakes(review): Unify reconcile keys and reject snapshot streams
* no-mistakes(review): Key reconcile requests by stable target ID
* no-mistakes(document): Document per-target reconcile request coalescing
* no-mistakes(lint): Remove unused snapshot summary file variable
* no-mistakes(ci): Adjusted the concurrent collector regression’s end-to-end timing ceiling to account for stock macOS process/jq overhead outside the three-second remote collection budget, while remaining below the 15-second serial-read floor. Verified with stock /bin/bash 3.2: all 44 Bearings tests pass; bash syntax and git diff checks pass
* no-mistakes(ci): Fixed legacy summary validation to require exactly one top-level JSON document and added behavioral regression coverage. Stabilized CI by conditionally waiting longer for durable reconcile delivery and synchronously stopping the fm-on worker tree before fixture cleanup. Removed a redundant flaky healthy-path timing assertion; the wedged-reader test still proves concurrent bounded collection. Verified fm-bearings-snapshot, fm-secondmate-reconcile, and fm-on tests, plus project ShellCheck, bash syntax, and git diff checks
* ci: rebalance portable serial test shards (#3489)
* fix(ci): rebalance the portable serial shards on measured durations
The "Behavior portable serial 3" shard ran 17-20 minutes against its
20-minute job cap and intermittently timed out seconds after a passing
test, on branches and on main alike.
Shards are packed longest-processing-time from per-script duration hints,
and those hints were last measured on 2026-08-21 at 116 scripts. The lane
has since grown to 139 scripts and from ~42 to ~63 minutes: 17 scripts had
no hint at all and fell back to the 20 s default, and several existing
hints were low by 2-5x (fm-watch-triage 142 s hinted vs 263 s measured,
fm-public-followup 36 s vs 197 s). The partition therefore looked
perfectly balanced in hint space, 734.6 s per shard, while really running
11.5, 13.6, 18.8 and 16.5 minutes. Script-count balance, which is what the
tests asserted, stayed normal throughout and hid it.
Refresh the hints from the timing artifacts of three green runs, taking
the slowest measurement of each script so the balance holds on a slow
runner, and split the lane across five shards instead of four. Replayed
against those runs' real per-script durations the worst shard is now
12.54 minutes, 63% of the unchanged 20-minute cap, and the serial lane's
wall clock drops from ~20 to ~12.5 minutes.
Bound the drift that caused this rather than relying on the hints being
refreshed by hand: the coverage guard now reports the unmeasured share as
serial_unhinted= and refuses past PORTABLE_SERIAL_MAX_UNHINTED_PERCENT,
which leaves room for newly added tests while making a stale table fail
the guard instead of silently pushing one shard into its cap.
No test changes what it asserts and no test stops running; only the
partition across shards changes.
* no-mistakes(document): Clarify conservative shard timing aggregate
* fix(pi): fall back on incomplete supervision branch prompts (#3491)
* fix(pi): fall back after settled branch errors
* no-mistakes(review): Detect provider errors across prompt compaction
* no-mistakes(review): Preserve in-flight branch state across selection changes
* fix(pi): re-probe supervision branch after cooldown (#3497)
* fix(pi): recover supervision branch after cooldown
* no-mistakes(review): Defer branch recovery until prompt settlement
* no-mistakes(document): Clarify supervision cooldown recovery contract
* fix(bin): remove legacy remote snapshot reads (#3501)
* refactor: remove legacy remote summary reads
* no-mistakes(document): Document ledger-only snapshot reads
* no-mistakes(ci): Fixed the snapshot test fixture so ledger refreshes use the same fake executable PATH as the snapshot consumer. This preserves observable endpoint freshness after removing legacy summary computation. Verified stock Bash parsing and all 44 Bearings tests pass under /bin/bash; git diff checks pass
* no-mistakes(ci): Fixed the CI-only snapshot fixture failure by ensuring the bounded-ledger refresh uses its fake tmux backend. This removes host tmux availability as a source of nondeterminism. Verified all 44 Bearings tests pass, Bash syntax passes, and git diff checks are clean
* no-mistakes(ci): Fixed CI nondeterminism in the Bearings fixture: all local ledger refreshes now use the fixture’s fake tmux backend when available, instead of depending on host tmux state. Verified stock /bin/bash syntax, git diff checks, and all 44 Bearings tests with a deliberately failing host tmux
* fix(pi): preserve watcher continuity across session replacement (#3498)
* fix(pi): rearm watcher after session replacement
* no-mistakes(review): Queue actionable closes across Pi session replacement
* no-mistakes(review): Stop replacement arm when handoff persistence fails
* no-mistakes(review): Preserve actionable wakes through branch and late child races
* no-mistakes(review): Surface late handoff failures without crashing Pi
* no-mistakes(review): Coordinate replacement delivery settlement and unique handoff tokens
* no-mistakes(review): Retry stale deliveries and release settled claims
* no-mistakes(review): Distinguish branch settlement and retry handoff cleanup
* no-mistakes(review): Deduplicate persistent handoff cleanup alerts
* no-mistakes(review): Acknowledge watcher follow-ups only when consumed
* no-mistakes(review): Persist idle follow-ups until agent consumption
* no-mistakes(review): Preserve pending outcomes when handoff persistence fails
* no-mistakes(review): Arm replacement before awaiting prior delivery settlement
* no-mistakes(review): Adopt pending handoffs after lock reclamation
* no-mistakes(review): Prevent stale generations from adopting replacement handoffs
* no-mistakes(review): Scope replacement handoffs by watcher state
* no-mistakes(document): Clarify replacement handoff documentation
* no-mistakes(ci): Fixed the failing branch-extension tests to model the new settlement-promise contract. Failure cases now assert that delivery ownership returns to the watcher instead of expecting direct extension fallback. Verified the updated branch suite, Pi watcher suite, shell syntax, and diff checks
* no-mistakes(review): Update branch settlement tests and preserve chunked outcomes
* no-mistakes(document): Document watcher-owned replacement handoffs
* no-mistakes(document): Verify replacement handoff documentation
* test(pi): cover watcher-owned branch fallback
* no-mistakes(document): Refresh watcher-owned fallback documentation
* fix(bin): resurface task statuses missed by wake handling (#3495)
* fix(bin): resurface terminal statuses lost after branch handling
* test(watch): canonicalize process-event fixture homes
* no-mistakes(review): Index branch outcomes by causal status position
* no-mistakes(review): Recover outcome indexes and deduplicate resurfaced statuses
* no-mistakes(review): Handle legacy ambiguity and oversized status diagnostics
* no-mistakes(review): Keep unclassifiable oversized statuses silent
* no-mistakes(document): Document lost-wake outcome backstop
* no-mistakes(document): Update outcome backstop documentation
* no-mistakes(ci): Fixed CI regressions in wake-drain: parseable reserved-key decisions can no longer bypass the durable decision-fold guard, and status output is prepared and receipt-committed before presentation to prevent repeated one-shot outcomes after later failures. Added a behavioral regression for receipt commit failure and retry. Targeted backstop, correlation-token, decision-cursor, open-decision, unread-status, syntax, and diff checks pass locally. Shard-4 failures appeared unrelated/flaky; the network-parallel test passed locally
* no-mistakes(ci): Fixed the Greptile P1 data-loss issue by committing presentation receipts only after prepared output reaches stdout. Added behavioral coverage proving output failure leaves the backstop retryable and receipt failure may duplicate but never lose a presentation. Relevant wake-drain suites and syntax/diff checks pass. The shard-4 Pi extension failure is unrelated to this PR and did not warrant changes
* no-mistakes(ci): Stabilized tests/fm-bootstrap-network-parallel.test.sh by replacing scheduler-sensitive equal-sleep timing with bounded synchronization between mocked fetch and remote probes. This preserves detection of real serialization while avoiding false failures under CI load. Verified with five consecutive test runs, bash syntax validation, ShellCheck, and git diff checks. The separate Pi stock-rendering failure reproduces locally but is unrelated environment/version drift
* no-mistakes(ci): Fixed Behavior portable serial 4 by adding fm-classify-lib.sh and fm-timeout-lib.sh to the broken-root Pi test fixture; fm-branch-outcome.sh now depends on them. Verified the full Pi branch-extension suite with real-Pi checks skipped, the wake-drain outcome-backstop suite, Bash syntax, and git diff checks. Greptile findings are already addressed at HEAD; the no-mistakes attestation failure is external head-SHA state
* fix(bin): collect follow-up results from remote work homes (#3503)
* fix(bin): deliver typed terminal results from remote work homes
A public commitment whose work is bound to a REMOTE secondmate home could
never receive its typed terminal result. `fm-public-followup.sh brief`
printed an emit command carrying this home's own absolute path and this
checkout's own script path, neither of which exists on the machine the
worker runs on, so the worker had nothing it could write to that the
owning home would ever read - and `consume` kept finding nothing while
the promise stayed open.
The brief is now route-aware: for a remote work home it prints that
route's own code root and home with `--stage-in`, so the typed event is
staged in the home where the work actually runs, and the closing
paragraph names the owning home as the one on the other machine instead
of pointing at the path above it. The owning home collects those staged
results over the same SSH route it reaches that secondmate on, because
the transport only runs outbound: `consume` pulls them into its own
inbox and reconciles them exactly as it reconciles a local report.
Collection is non-destructive until the result is durably held, so a
dropped connection cannot lose a terminal result, and a route that could
not be reached is named in `consume`'s output with the promise left open
rather than reported as an empty inbox.
A local work home is untouched: the brief still prints `--home` with this
home and this checkout's script, and the event still lands directly in
this home's typed terminal-result inbox.
This is the emit-side counterpart of the retire/clear fix in #3479 and
reuses the remote-route resolution that landed with it. Reconciling a
loop bound to a remote route now reaches that route, so the existing
remote cases drive `consume` through the same faked transport their
other steps already use.
* no-mistakes(review): Fail loudly on unresolved routes and invalid staging homes
* no-mistakes(review): Fail collection when remote outbox is unreadable
* no-mistakes(review): Surface reassigned remote routes during empty collection
* no-mistakes(review): Fail remote collection on invalid registrations
* no-mistakes(review): Reject unsafe registration entries during remote collection
* no-mistakes(review): Restore healthy empty remote collection behavior
* no-mistakes(review): Skip remote collection for delivered registrations
* no-mistakes(review): Skip delivered registrations before route validation
* no-mistakes(document): Document remote follow-up collection semantics
* fix(bin): exclude secondmates from home-summary validity (#3504)
* fix(bin): exclude secondmates from home-summary child inventory
kind=secondmate meta records never have backlog rows, so counting them in unowned_children or terminal_in_flight made a clean main home look invalid once earlier ledger checks passed.
* no-mistakes(review): Cover terminal secondmate in-flight exclusion
* no-mistakes(ci): Updated the stock macOS Bash CI snapshot expectation from 15 to 16 tests. Verified all 16 snapshot/fleet-view tests pass under Bash 3.2.57 and `git diff --check` succeeds
* fix(bin): self-heal outcome indexes on first drain (#3509)
* fix(bin): self-heal status-outcome indexes on every drain
Missing ready markers were skipping the lost-wake backstop on non-Pi homes because only the Pi branch ran processed-init. Drain now rebuilds those indexes under the outcome lock and fails closed only on a real store fault.
* no-mistakes(review): Guard held-lock initialization and fail marker writes
* no-mistakes(document): Document cross-harness outcome-index self-healing
* fix(bearings): keep active children underway during captain holds (#3505)
* fix(bearings): keep active children underway beside a captain hold
Project each readable home's active children into Underway independently of the home-level captain-decision classification so a hold no longer hides live work.
* no-mistakes(review): Preserve Underway repos and disclose child truncation
* no-mistakes(review): Fall back to task project for Underway repos
* no-mistakes(ci): Updated the stock macOS Bash CI assertion from 44 to 45 Bearings tests, matching the newly added behavioral regression. Verified all 45 tests pass under /bin/bash, Bash syntax checks pass, and git diff validation is clean
* fix(pi): settle watcher delivery on Pi accepting the follow-up (#3513)
* fix(pi): settle watcher delivery on Pi accepting the follow-up
A follow-up queued while main is streaming joins the running run without
ever raising before_agent_start, so waiting on that event before clearing
the successor pipeline (#3498) stalled every later actionable close: no
successor started, no wake was delivered or offered to the branch, and the
turn-end guard woke main to re-arm by hand after every close.
The pipeline now settles once Pi accepts the follow-up. Consumption is
observed at before_agent_start for an idle main and at the user
message_start for a streaming main, and decides only what a replacement
session (/new, /resume, /fork, reload) replays. An exhausted restoration
delivers its typed failure without launching an arm past the retry bound,
which the stall had hidden. The replacement-coordinator map is typed so the
strict no-emit typecheck passes again.
Tests: the doubles no longer raise before_agent_start for a streaming send,
a portable regression drives two actionable closes while main streams and
proves the successor chain plus consumption-scoped replay, and a
credential-free real-SDK probe pins Pi's event contract for both the
streaming and the idle follow-up.
Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a
* fix(pi): retry a verified successor that fails during wake delivery
A verified successor can exit while the wake it was started for is still
being delivered, most plausibly during a branch turn that holds the
settlement for minutes. Its failure close arrived while the pipeline's
single-flight guard was set, so the close handler skipped the retry, and
the pipeline's end no longer launched an arm, which left the live
generation with no watcher and no retry timer.
The close handler now records that failure when the child had reported
readiness and was not retired by the restoration itself, and the pipeline
runs the ordinary bounded, lock-checked retry for it once the delivery
settles. A restoration started for a later pending supersedes it, and an
exhausted restoration still hands repair to main without a further arm.
The regression holds a branch settlement open while the verified
successor exits with a failure and proves one retry watcher starts after
the settlement releases, none while it is held.
Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a
* fix(bin): bound repeat stale wakes for parked workers (#3532)
* fix(bin): bound repeat stale wakes for a parked but live worker
A worker parked on a declared wait - `paused:` for an external or pipeline
wait, or a verified `captain-held` transfer - kept waking firstmate far inside
FM_PAUSE_RESURFACE_SECS. Observed as five consecutive alarms on one
captain-held worker and dozens across a day on a pipeline wait, and reported
upstream as four wakes in 75 minutes against a 3600s window.
pause_state_class deliberately answers `none` for a still-live agent even under
a declared wait, so a worker genuinely waiting on a decision is never silenced.
That classification is correct and is left alone; it routes every parked but
live worker through surface_nonterminal_stale on first sight of each distinct
stale hash, and an idle parked pane still churns its hash on a clock or a token
counter without changing what is being waited on.
Two places let that churn re-alarm:
- surface_nonterminal_stale queued the wake BEFORE consulting whether a wait was
declared, then wrote `.paused-resurfaced-<key>` - the very throttle that should
have suppressed it. The throttle was never read on this path and was advanced
by the wake it should have prevented.
- The hash-change path cleared that throttle through clear_pause_tracking
whenever the classification came back `none`, so each tick also bought the same
declared wait a fresh window. Fixing only the first site changes nothing.
Read the throttle before anything is queued and advance it only on a wake that
really fires, and on the hash-change path reset only the per-hash bookkeeping
while the declaration still stands, via a clear_stale_hash_tracking split so
neither half of clear_pause_tracking is duplicated. The throttle is keyed to the
declaration, not to the pane.
First sight still wakes, so an inconclusive state is still inspected, and the
window's end still re-surfaces once, so a forgotten wait cannot rot invisibly -
noise traded for a bounded cadence, never for silence. The wake identity stays
the plain `stale: <win>` the away-mode handoff depends on.
Tests cover both observed forms and were confirmed to fail against three
deliberate breaks: each site reverted on its own, and a re-surface that never
fires again.
* fix(document): Clarify declared-wait wake cadence documentation
* fix(ci): Captain, fixed the stale-throttle inheritance: cadence markers now bind to the current wait declaration, so replacement paused and captain-held waits each emit their first plain `stale:` wake. Added behavioral coverage for both forms. Bite proof failed as expected when identity matching was removed, then passed after restoration. Full watcher triage suite, `bin/fm-lint.sh`, syntax checks, and diff checks pass. Changes remain uncommitted for the outer executor
* fix(ci): Captain, fixed the confirmed Greptile finding. `resurface_absorbed` now applies a throttle only when its stored declaration scope matches the current wait, so replacement `paused:` and `captain-held` waits surface immediately without changing classification. Added executable coverage for both absorbed forms. Bite proof failed before the fix at the intended assertion; afterward the full watcher triage suite, `bin/fm-lint.sh`, shell syntax checks, and `git diff --check` passed
* fix(bin): accept the away-mode daemon as the turn-end supervision owner (#3567)
* fix(turnend): accept the away-mode daemon as the supervision owner
While state/.afk exists the away-mode daemon owns supervision and runs
bin/fm-watch.sh one-shot: the watcher exits on every wake and the daemon
starts its replacement. The turn-end guard tested for a live watcher
process holding the watch lock at that instant, so a turn boundary that
landed in the hand-off blocked with "TURN WOULD END BLIND" while
supervision was completely healthy, costing a full handling turn each
time.
Reproduced with the real daemon wrapping the real watcher and the real
guard sampling the same home: 6 of 40 samples blocked, every one of them
with the daemon alive and the beacon 2-3 seconds old, and a new watcher
pid on each cycle. After the fix the same reproduction blocks 0 of 40,
and killing the daemon and its watcher (away mode still on, beacon still
fresh) blocks again.
The guard now accepts a live, identity-matched daemon holding this home
as proof of supervision while away mode is active. The identity match is
the same discipline the watcher lock uses, so a recycled pid or a lock
left by a killed daemon proves nothing. The fresh-beacon half of the
predicate is unchanged: a daemon that stops restarting its watcher still
blocks once the beacon passes grace, a home with no supervisor blocks
exactly as before, and with away mode off the strict watcher predicate is
untouched.
The predicate reads only durable state, so it behaves identically for
every primary harness and runtime backend.
* no-mistakes(document): clarify away-mode daemon supervision proof and test coverage
* no-mistakes(document): generalize stale turn-end predicate summary in architecture.md
* fix(backlog): omit --file from row probes for non-markdown backends (#3582)
* fix(backlog): omit markdown file for beads probes
* no-mistakes(document): Narrow backlog addressing doc to mutations for backend-aware probes
* no-mistakes(ci): Fixed the Greptile P2 review comment (the only failing check) on tests/fm-backlog-atomicity.test.sh. The comment correctly noted that an exported TASKS_AXI_BACKEND environment variable would inherit into the spawned scripts and, because fm_tasks_axi_backend gives it top precedence, override each test case's .tasks.toml backend fixture — making the backend-specific argv assertions fail for environmental reasons. Fix: unset TASKS_AXI_BACKEND in the test harness right after sourcing tests/lib.sh, with a comment explaining why, so every case deterministically exercises its declared backend (4 lines added; no production code touched). Verified: reproduced the leak before the fix (TASKS_AXI_BACKEND=beads made the markdown dispatch case fail with 'beads show failed', exactly the reported failure mode); after the fix the full suite passes (0 failures, exit 0) both with and without TASKS_AXI_BACKEND=beads exported. The added lines are shellcheck-clean (the only shellcheck note, SC1091 on the lib.sh source line, pre-exists this change)
* fix(bin): classify progress updates on requested work as routine (#3589)
The supervision branch's verdict rule escalated every outcome that
answered a captain request, so "the work started" and "still working"
notes reached the captain with nothing to look at. The rule now keeps a
finished result of requested work captain-facing, even when healthy, and
treats start or still-working updates that bring no new artifact,
finding, or decision as routine. The captain list for review-ready PRs,
ask-user findings, exhausted blockers, credentials, and destructive or
security-sensitive cases is unchanged, as are the unsolicited-routine,
silent-fleet-review, and doubt-chooses-captain rules.
The fm_branch_report tool description and the two docs that restated the
old unconditional rule now point at the prompt's "Verdict: routine or
captain" section as the one owner instead of carrying a second copy.
* fix(bin): preserve captain calls during teardown (#3595)
* fix(bin): never close a captain call during cleanup
A scout that held its own work item for the captain, which is what
captain-hold-lifecycle prefers ("hold the work item the question gates"),
was closed by bin/fm-teardown.sh's automatic backlog transition. The
completion gate passed, cleanup ran, and the captain's question moved to
Done with no recorded answer: the one thing the policy says must never
happen. `tasks-axi done` closes a held row silently, and nothing in
teardown asked whether the row was the captain's own call.
bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when
the task is still an open captain call, 1 when it is not, 2 when that
cannot be established. It reads the row through the transition library's
backend-aware probe, so it addresses the same backlog teardown does; the
script's other commands now address the configured data directory the
same way instead of FM_HOME, which also fixes captain holds in a home
with a relocated data directory.
Teardown asks `open` before any destructive step and refuses on 2. On 0
only the close changes: after cleanup and still under the task's own
lock, the row gets one "Deliverable of the finished work" line at the end
of its body and returns to Queued through `tasks-axi reopen`, keeping its
hold, so it lands in Captain's Call instead of reading as work under way.
--force does not lift this: it authorizes discarding unlanded work, never
the captain's question. The deliverable goes into the body because
`tasks-axi update --report` rewrites the title of a row that is not Done.
The crash window reuses the pending-close record teardown already stages:
a `mode=retain` line makes the existing replay record the deliverable and
reopen instead of closing, with the same validator, stale-generation
check, cleanup-incomplete marking, and non-blocking bootstrap lock as an
ordinary close. A retained row the captain answered first simply retires
the record. No parallel record type, recovery command, or second bootstrap
loop is introduced.
Regressions run the real executables: the captain-held scout survives
cleanup queued, held, with its deliverable and on the board, only
`answer` closes it, --force keeps it open, and an ordinary scout still
closes with its report; an interrupted cleanup leaves the row untouched
and the next session start retains it; a relocated backlog keeps the
retention in its one configured file; and a ship row whose hold cannot be
read refuses cleanup before anything destructive.
Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np
* no-mistakes(review): Serialize captain holds and fix backend-aware listing
* no-mistakes(document): Update captain-call retention documentation
* no-mistakes(document): Fix relocated captain-hold backlog diagnostics
* fix(bin): deliver secondmate outcomes to the parent channel (#3592)
* fix(bin): deliver every secondmate outcome on the parent channel from the recording scripts
A secondmate's captain-facing outcomes could miss: the mate model addressed
the captain in its own unread chat instead of appending to the parent
channel, and a PR-ready report, a finding, a decision, a blocker, and a
failure all depended on that one remembered append. Make delivery
structural, so the parent channel never depends on the model:
- bin/fm-parent-channel-lib.sh is the one owner of channel resolution and
exact-line append-once; the merge outcome path and the inactive-outcome
scan now publish through it instead of two private copies.
- bin/fm-inactive-reconcile.sh gains a ledger-first path that runs on every
watcher poll in a secondmate home: a direct child's whole terminal done or
failed line is delivered at once with its note, recorded PR, mode, merge
posture, and scout report pointer, keyed and receipted so it is delivered
once, and the inactive path yields to it. `report <task-id>` runs the same
delivery for a caller holding the child's meta lock.
- bin/fm-pr-check.sh publishes the PR-ready line with the canonical URL at
registration.
- bin/fm-captain-hold.sh publishes a hold and its answer, keyed by task id
and resolution-record count, with no new persisted state.
- bin/fm-teardown.sh delivers the child's final line before removing its
record and refuses, retaining every record, while the channel cannot be
written.
- The charter opens with the parent-channel rule and confines the mate's own
appends to judgement; AGENTS.md carries the carve-out at the persona
address rule and the escalation list.
docs/secondmate-parent-channel.md records the design and its coverage, and
docs/verification/secondmate-parent-channel.md records the live run with real
tmux panes and both real watchers delivering every line with no model.
Supersedes #3569.
* no-mistakes(review): Fix parent outcome retries and reconciliation locking
* no-mistakes(review): Prevent busy children from starving ledger delivery
* no-mistakes(review): Correct ledger metadata and hold occurrence handling
* no-mistakes(review): Disambiguate ledger outcomes and normalize hold reasons
* no-mistakes(review): Close ledger races and preserve teardown records
* no-mistakes(document): Correct parent-channel receipt and scanner documentation
* no-mistakes(lint): Quote done arguments for ShellCheck compliance
* no-mistakes(ci): Fixed both CI failures. Updated GOTMP teardown fixtures for the new final-outcome reporter and isolated them from host tmux state. Updated the PR security assertion to distinguish the accepted PR-ready line from duplicate merge outcomes. Verified with both failing test suites, bash syntax checks, and git diff checks
* no-mistakes(ci): Fixed Greptile’s duplicate-delivery race in bin/fm-inactive-reconcile.sh. Ledger events now claim matching already-delivered inactive receipts using the prior status fingerprint, preventing duplicate parent reports while preserving later same-state completions. Added behavioral regression coverage. Verified inactive-reconcile tests, project lint, documentation audience checks, syntax, and diff checks. Teardown tests passed relevant cases before the documented pre-existing herdr-preflight-missing-adapter failure
* fix(bin): sync remote second mates to primary commit (#3599)
* fix(bin): sync remote second-mate homes to the parent primary commit
Session start and remote launch pointed a remote second-mate home at whatever
Firstmate copy its own host kept, so a home that had already advanced past that
copy refused as a non-fast-forward and every other home stopped at the host's
older commit while the primary ran ahead.
The parent now resolves ITS primary default-branch commit with the existing
helper and hands that commit to the host on both paths. Because a remote home
is a standalone clone, the host imports that one commit before advancing -
already present, else from that host's Firstmate copy without moving it, else
from the home's own origin - and then runs the SAME ff_target guards a local
home gets, so dirty, diverged, feature-branch, and unresolvable targets skip
untouched and the ancestry rules keep one owner. An unimportable target now
names /updatefirstmate instead of failing opaquely, and a host still running an
older Firstmate copy is reported the same way rather than echoing a bare
refusal.
The host-local launch leg no longer re-runs its own secondmate sync, so the
spawn it drives cannot re-target that host's copy after the parent has already
converged the home.
/updatefirstmate is unchanged: it still refreshes the remote code root from that
host's origin and then syncs the home to that refreshed copy, which is what the
sync call with no target commit means.
* no-mistakes(document): Document primary-targeted remote secondmate synchronization
* fix(bin): separate captain intent from firstmate specs (#3597)
* fix(bin): split brief task into captain intent and firstmate spec
Keep no-mistakes --intent as the captain's ask plus later captain words, not the build spec or worker tradeoffs.
* fix(bin): stop task-subsection copies at the next heading
Promotion was swallowing the scout Setup contract into Firstmate spec, and pre-subsection briefs lost their # Task body.
* no-mistakes(review): Validate brief content and preserve nested specifications
* no-mistakes(review): Scope placeholder validation to scaffold-only subsection bodies
* no-mistakes(review): Ignore fenced subsection headings during brief validation
* no-mistakes(review): Preserve captain intent across scout promotion
* no-mistakes(review): Enforce safe intent boundaries for legacy promotions
* no-mistakes(review): Allow marked legacy intent and reject empty promotions
* no-mistakes(review): Scope task parsing and overlay legacy intent contracts
* no-mistakes(review): Overlay current intent contract for all no-mistakes spawns
* no-mistakes(review): Preserve later captain clarifications in intent overlays
* no-mistakes(document): Document brief intent enforcement and ownership
* no-mistakes(ci): Updated spawn-related test fixtures to use valid Captain intent and Firstmate spec subsections, corrected launch-path expectations to launch-brief.md, and resolved ShellCheck quoting findings. Verified with fm-lint.sh and 15 affected behavior tests, including real Herdr tests; all passed
* no-mistakes(ci): Updated stale spawn/promotion fixtures in the Muse, Orca, secondmate-harness, and public-followup suites to provide valid Captain's intent and Firstmate spec subsections. Verified full Orca and secondmate-harness suites, targeted public-followup promotion behavior, Bash syntax, diff checks, and fm-lint
* fix: start a fresh supervision branch for every main session (#3600)
* fix(pi): start a new supervision branch conversation per main session
The supervision branch reopened one recorded conversation forever, so
every main session start reloaded the current generated prompt and then
weeks of accumulated thread, where a superseded rule could still outweigh
today's.
The branch conversation is now scoped to one main session: the session
generation owns the recorded conversation, so a cold start, /new,
/resume, /fork, or a reload always builds a new one, while a rebuild
inside one session (a model or effort change) still continues that
session's own conversation.
The dialog mirror re-anchors with it. Its durable cursor records what the
previous branch conversation received, so a /resume or reload - which
keeps main's own session file - would otherwise leave the new branch
blind to dialog main itself still has. The reset is bounded by the
current main session, and the cursor keeps advancing incrementally within
it. The durable outcome store and its processed marker are untouched, so
unacknowledged captain-facing outcomes still re-present on the new main
session.
* no-mistakes(document): Document fresh Pi supervision conversations
* no-mistakes(ci): Fixed the flaky concurrent inbox failure. Lock acquisition now retries when a competing lock disappears between a failed claim and inspection. Added a behavioral regression covering that race. Verified the full inbox test four times, project lint, and git diff checks
* feat: restart second mates after instruction updates (#3614)
* feat(update): restart second mates whose instructions changed
/updatefirstmate pulled new bytes onto disk and then asked each advanced
second mate to re-read them. A running agent holds AGENTS.md and every
loaded skill frozen from launch and no verified harness offers a reload,
so that steer could not reach a loaded skill at all and left the mate
holding two contradictory copies of its own job description.
An eligible mate is now restarted instead, in the same home and endpoint,
through the existing transactional relaunch. The restart is gated on the
mate first writing down the open work it holds only in conversation - the
open-record half of /stow, never its memory sweeps - so an unregistered
captain call is flushed before the conversation is spent. Anything that
leaves the reload unprovable falls back to the old re-read message and is
reported as exactly that, never as a clean reload.
Remote mates take the same path: fm-remote-secondmate-control.sh gains a
relaunch verb whose host-local leg runs that same control plane, since the
mate is an ordinary local secondmate from its host's point of view. The
primary resolves the profile and passes it explicitly, because
config/secondmate-harness is not inherited and the file on that host
belongs to a different home.
fm-update.sh now splits its advanced live mates into a restart set and a
nudge residual, and both sets require a changed instruction surface, which
also closes the over-nudge against the session-start sweep. Restart is
stricter still: a bin/-only advance reloads itself on the next call, so it
never costs a conversation.
Colocated tests cover the gating, the persist-then-restart order, the
task-subset persist request, each unsafe fallback, the remote hop, and the
remote sync's new instruction-surface report.
* no-mistakes(review): Fix restart correlation, concurrent waits, and lifecycle reporting
* no-mistakes(review): Parallelize relaunches and classify replacement incarnations
* no-mistakes(review): Gate restart actions on live agent state
* no-mistakes(review): Handle failed restart workers without hanging
* no-mistakes(review): Nudge legacy remotes and preserve persist recovery
* no-mistakes(review): Document one-time secondmate restart rollout
* no-mistakes(review): Honor arrived replies and refresh remote profiles
* no-mistakes(review): Revert remote parent profile reconciliation
* no-mistakes(review): Reset remote profile defaults and honor published results
* no-mistakes(review): Preserve fallback nudges for unverifiable secondmates
* no-mistakes(document): Document second-mate restart update flow
* no-mistakes(lint): Fix ShellCheck warnings in restart scripts
* perf: accelerate local validation with bounded concurrency (#3644)
* perf(tests): route gate verification through the bounded concurrent runner
Local validation was the pipeline's dominant cost: across 67 recorded
no-mistakes agent sessions on this repo, 99.3% of command execution was
`bash tests/*.test.sh`, run strictly one script at a time, and 2% of those
calls were killed by an agent-guessed timeout and paid for twice.
Three changes, each measured:
- `.no-mistakes.yaml` pins `commands.test` to
`bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner
already owns changed-file selection, bounded concurrency, the refusal of
unproven scripts, and a generous automatic per-script bound, so the gate's
baseline is neither a serial chain nor a guessed timeout. It stays
intent-targeted - the Test step still runs its evidence agent on top - and
excludes the live-Herdr family the required Herdr lane owns.
- `bin/fm-test-run.sh` gives a plain list of script paths the same bounded
automatic scheduler and automatic bound that `--changed` gets. Naming several
subjects is how a verification round asks for exactly those scripts. The
curated selections are untouched: `--lane` still composes CI shards whose
serial lane must stay serial, `--family` is what the required Herdr lane runs,
and `--all` stays a deliberate complete regression.
- `pr-forge` is admitted to the concurrent-safe family registry on two
consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those,
and records `secondmate` and `session-bootstrap` as refused with the exact
script and reason each failed on, so the refusals are actionable rather than
silent.
Measured on this host, 0 failures on both sides:
verification round, 4 scripts 448s chained -> 231s through the runner (-48%)
pr-forge family 409.2s at 1 worker -> 237.9s at 4 (1.72x)
watcher-wake-lock family 1311.1s at 1 worker -> 539.3s at 4 (2.43x)
A fourth lever was implemented and then removed because the measurement
refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made
`fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s
unchanged, back to back. Those sleeps are not overhead added to the clock -
they are how a test waits for a subject moving on fm-watch.sh's own one-second
cadence - so sampling less often only delays detection. It also broke
`fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a
settled condition. CONTRIBUTING.md records that result so the experiment is not
repeated.
* no-mistakes(review): Separate concurrent runs by isolation proof family
* no-mistakes(review): Limit automatic timeouts to changed-file validation
* no-mistakes(document): Clarify validation concurrency documentation
* fix: copy PR URLs from durable records (#3648)
* fix: copy PR URLs from records or abstain, never assemble them
Supervision reported a plausible but dead PR link three times because its
prompt demanded a full https:// URL at a moment when only a PR number was
observable, so the model assembled an owner/repository from memory, and the PR
check then accepted that URL and wrote it into the task record, after which the
model kept defending its own tool-endorsed guess over the worker's real link.
Three changes close that chain without any live forge lookup, so private
forges are treated exactly like public ones:
- bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy
or abstain" section requires a URL to be copied verbatim from a …
friesentius
pushed a commit
to friesentius/firstmate
that referenced
this pull request
Sep 21, 2026
…kunchenguid#3417) * fix(backlog): honor configured task adapters * no-mistakes(review): Harden backend purity lint against prefixed Beads calls * no-mistakes(document): Document configured backend lifecycle transitions * fix(backlog): preserve markdown exemptions * no-mistakes(review): Enforce backend purity for explicit lint paths * no-mistakes(document): Update lifecycle backend documentation * no-mistakes(lint): Remove redundant backend lint pattern * fix(backlog): close adapter routing gaps * no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint * no-mistakes(document): Document environment-selected backlog adapters * no-mistakes(lint): Fix empty local variable assignment * fix(backlog): close quoted path gaps * no-mistakes(review): Reject partially quoted direct Beads commands * no-mistakes(document): Align lifecycle documentation with configured adapters * test(backlog): keep structural cases markdown-only * fix(backlog): honor configured task adapters * no-mistakes(review): Harden backend purity lint against prefixed Beads calls * no-mistakes(document): Document configured backend lifecycle transitions * fix(backlog): preserve markdown exemptions * no-mistakes(review): Enforce backend purity for explicit lint paths * no-mistakes(document): Update lifecycle backend documentation * no-mistakes(lint): Remove redundant backend lint pattern * fix(backlog): close adapter routing gaps * no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint * no-mistakes(document): Document environment-selected backlog adapters * no-mistakes(lint): Fix empty local variable assignment * fix(backlog): close quoted path gaps * no-mistakes(review): Reject partially quoted direct Beads commands * no-mistakes(document): Align lifecycle documentation with configured adapters * test(backlog): keep structural cases markdown-only * no-mistakes(review): Harden markdown lifecycle routing and close recovery * fix(lint): catch dollar-quoted beads commands * no-mistakes(review): Drop unused authorized-data-dir param; canonicalize fm-lint ROOT with pwd -P * no-mistakes(document): Document fm-lint backend-purity check in header and CONTRIBUTING * no-mistakes(ci): Two failing checks, one code-caused and one attestation-only. 1. Greptile Review (P1: ANSI-C quoting bypasses lint) — REAL DEFECT, FIXED. The backend-purity normalizer in bin/fm-lint.sh (invokes_bd awk function) stripped $'...' quote delimiters without decoding ANSI-C escapes, so a core script containing $'\x62\x64' close fm-example executes as `bd close` while the lint accepted it. Fix: the normalizer now tracks whether a single-quote context came from $' (ansi flag) and decodes ANSI-C escapes while building the command word: \xHH (1-2 hex), \uHHHH / \UHHHHHHHH (4/8 hex), \NNN (1-3 octal), \cX control characters, simple escapes (a/b/e/E/f/n/r/t/v decode to a placeholder that can never spell bd), NUL (value 0) truncates the word per bash C-string semantics, and unknown escapes drop the backslash per bash. Values outside printable ASCII decode to a placeholder so they can neither falsely match nor collide into `bd`. Regression coverage added to the existing lint-interface test test_rejects_direct_beads_cli_invocations in tests/fm-lint.test.sh: $'\x62\x64', $'\142\144' (octal), and b$'\x64' (split word), each written into a fixture and asserted rejected through the real fm-lint.sh executable. Verified regression property: with the fix stashed, the new hex case passes lint (reproduces the reported bypass); with the fix, it is rejected. Verification: all 29 tests in tests/fm-lint.test.sh pass, and the full CI-parity lint (CI=true bin/fm-lint.sh: ShellCheck 0.11.0 full analysis over the canonical set, backend-purity check, and actionlint workflow validation) exits 0. One iteration was needed because an awk comment containing a literal $'\x62\x64' sample broke shell-level quoting (bash -n / SC1001/SC2026); the comment was reworded without quotes. 2. PR must be raised via no-mistakes — NOT code-caused. The check failed with 'Pipeline attestation head_sha does not match the current PR head': the PR body attestation binds to 82b41c7 while the PR head is 1995cdb because a later pipeline push moved the head. This is exactly the stale-attestation condition the user intent describes; it clears when the outer pipeline re-runs 'git push no-mistakes' and re-binds the attestation to the new head (which now includes this Greptile fix). No code change can or should address it. Files changed: bin/fm-lint.sh (ANSI-C escape decoding in the backend-purity normalizer), tests/fm-lint.test.sh (three encoded-bd rejection cases) * no-mistakes(review): fix tasks.toml hang, root authorization, lint quoting * no-mistakes(review): validate tasks config before exemption; fix lint quote gap * fix(backlog): address the markdown backlog as <data>/backlog.md Resolving the markdown backlog through a configured `[markdown] path` was scope this task never asked for. It is absent from main, which addresses `<data>/backlog.md` everywhere, and it came from an earlier review round rather than the task brief. Making it effective on the transition path alone put that path at odds with every other consumer of the same backlog - fm-captain-hold.sh, fm-session-start.sh, fm-fleet-snapshot.sh, fm-inbox.sh, fm-backlog-handoff.sh - which all still address `<data>/backlog.md`. In fm-captain-hold.sh the split was live: its reads had already moved to the shared gate while its writes had not, so the two could address different files. Address `<data>/backlog.md` from the shared gate, delete the unused resolver, and drop the two tests that pinned the withdrawn behaviour. What this task actually changes is unaffected: a configured non-markdown adapter is still addressed by its own root, without `--file`. * fix(backlog): honor configured task adapters * no-mistakes(review): Harden backend purity lint against prefixed Beads calls * no-mistakes(document): Document configured backend lifecycle transitions * fix(backlog): preserve markdown exemptions * no-mistakes(review): Enforce backend purity for explicit lint paths * no-mistakes(document): Update lifecycle backend documentation * no-mistakes(lint): Remove redundant backend lint pattern * fix(backlog): close adapter routing gaps * no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint * no-mistakes(document): Document environment-selected backlog adapters * no-mistakes(lint): Fix empty local variable assignment * fix(backlog): close quoted path gaps * no-mistakes(review): Reject partially quoted direct Beads commands * no-mistakes(document): Align lifecycle documentation with configured adapters * test(backlog): keep structural cases markdown-only * fix(lint): catch dollar-quoted beads commands * no-mistakes(review): Drop unused authorized-data-dir param; canonicalize fm-lint ROOT with pwd -P * no-mistakes(document): Document fm-lint backend-purity check in header and CONTRIBUTING * no-mistakes(ci): Two failing checks, one code-caused and one attestation-only. 1. Greptile Review (P1: ANSI-C quoting bypasses lint) — REAL DEFECT, FIXED. The backend-purity normalizer in bin/fm-lint.sh (invokes_bd awk function) stripped $'...' quote delimiters without decoding ANSI-C escapes, so a core script containing $'\x62\x64' close fm-example executes as `bd close` while the lint accepted it. Fix: the normalizer now tracks whether a single-quote context came from $' (ansi flag) and decodes ANSI-C escapes while building the command word: \xHH (1-2 hex), \uHHHH / \UHHHHHHHH (4/8 hex), \NNN (1-3 octal), \cX control characters, simple escapes (a/b/e/E/f/n/r/t/v decode to a placeholder that can never spell bd), NUL (value 0) truncates the word per bash C-string semantics, and unknown escapes drop the backslash per bash. Values outside printable ASCII decode to a placeholder so they can neither falsely match nor collide into `bd`. Regression coverage added to the existing lint-interface test test_rejects_direct_beads_cli_invocations in tests/fm-lint.test.sh: $'\x62\x64', $'\142\144' (octal), and b$'\x64' (split word), each written into a fixture and asserted rejected through the real fm-lint.sh executable. Verified regression property: with the fix stashed, the new hex case passes lint (reproduces the reported bypass); with the fix, it is rejected. Verification: all 29 tests in tests/fm-lint.test.sh pass, and the full CI-parity lint (CI=true bin/fm-lint.sh: ShellCheck 0.11.0 full analysis over the canonical set, backend-purity check, and actionlint workflow validation) exits 0. One iteration was needed because an awk comment containing a literal $'\x62\x64' sample broke shell-level quoting (bash -n / SC1001/SC2026); the comment was reworded without quotes. 2. PR must be raised via no-mistakes — NOT code-caused. The check failed with 'Pipeline attestation head_sha does not match the current PR head': the PR body attestation binds to 82b41c7 while the PR head is 1995cdb because a later pipeline push moved the head. This is exactly the stale-attestation condition the user intent describes; it clears when the outer pipeline re-runs 'git push no-mistakes' and re-binds the attestation to the new head (which now includes this Greptile fix). No code change can or should address it. Files changed: bin/fm-lint.sh (ANSI-C escape decoding in the backend-purity normalizer), tests/fm-lint.test.sh (three encoded-bd rejection cases) * fix(bin): preserve captain calls during teardown (kunchenguid#3595) * fix(bin): never close a captain call during cleanup A scout that held its own work item for the captain, which is what captain-hold-lifecycle prefers ("hold the work item the question gates"), was closed by bin/fm-teardown.sh's automatic backlog transition. The completion gate passed, cleanup ran, and the captain's question moved to Done with no recorded answer: the one thing the policy says must never happen. `tasks-axi done` closes a held row silently, and nothing in teardown asked whether the row was the captain's own call. bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when the task is still an open captain call, 1 when it is not, 2 when that cannot be established. It reads the row through the transition library's backend-aware probe, so it addresses the same backlog teardown does; the script's other commands now address the configured data directory the same way instead of FM_HOME, which also fixes captain holds in a home with a relocated data directory. Teardown asks `open` before any destructive step and refuses on 2. On 0 only the close changes: after cleanup and still under the task's own lock, the row gets one "Deliverable of the finished work" line at the end of its body and returns to Queued through `tasks-axi reopen`, keeping its hold, so it lands in Captain's Call instead of reading as work under way. --force does not lift this: it authorizes discarding unlanded work, never the captain's question. The deliverable goes into the body because `tasks-axi update --report` rewrites the title of a row that is not Done. The crash window reuses the pending-close record teardown already stages: a `mode=retain` line makes the existing replay record the deliverable and reopen instead of closing, with the same validator, stale-generation check, cleanup-incomplete marking, and non-blocking bootstrap lock as an ordinary close. A retained row the captain answered first simply retires the record. No parallel record type, recovery command, or second bootstrap loop is introduced. Regressions run the real executables: the captain-held scout survives cleanup queued, held, with its deliverable and on the board, only `answer` closes it, --force keeps it open, and an ordinary scout still closes with its report; an interrupted cleanup leaves the row untouched and the next session start retains it; a relocated backlog keeps the retention in its one configured file; and a ship row whose hold cannot be read refuses cleanup before anything destructive. Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np * no-mistakes(review): Serialize captain holds and fix backend-aware listing * no-mistakes(document): Update captain-call retention documentation * no-mistakes(document): Fix relocated captain-hold backlog diagnostics * fix(backlog): honor configured task adapters * no-mistakes(document): Update lifecycle backend documentation * fix(backlog): close adapter routing gaps * no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint * no-mistakes(lint): Fix empty local variable assignment * fix(backlog): close quoted path gaps * test(backlog): keep structural cases markdown-only * no-mistakes(review): Harden markdown lifecycle routing and close recovery * no-mistakes(review): fix tasks.toml hang, root authorization, lint quoting * no-mistakes(review): validate tasks config before exemption; fix lint quote gap * fix(backlog): address the markdown backlog as <data>/backlog.md Resolving the markdown backlog through a configured `[markdown] path` was scope this task never asked for. It is absent from main, which addresses `<data>/backlog.md` everywhere, and it came from an earlier review round rather than the task brief. Making it effective on the transition path alone put that path at odds with every other consumer of the same backlog - fm-captain-hold.sh, fm-session-start.sh, fm-fleet-snapshot.sh, fm-inbox.sh, fm-backlog-handoff.sh - which all still address `<data>/backlog.md`. In fm-captain-hold.sh the split was live: its reads had already moved to the shared gate while its writes had not, so the two could address different files. Address `<data>/backlog.md` from the shared gate, delete the unused resolver, and drop the two tests that pinned the withdrawn behaviour. What this task actually changes is unaffected: a configured non-markdown adapter is still addressed by its own root, without `--file`. * no-mistakes(review): restore home boundary guard and tighten purity lint * no-mistakes(review): authorize home boundary for every backlog adapter * no-mistakes(test): complete tasks-axi stubs in fm-gotmp teardown fixtures * no-mistakes(document): align backlog transition docs with adapter-neutral addressing * no-mistakes(review): label adapter data-dir authorization, drop dead row_probe local * no-mistakes(review): pin markdown backend at relocated-data addressing roots * no-mistakes(document): point lint-definition mention at fm-lint.sh header * no-mistakes(document): point mutate comment at adapter addressing owner * no-mistakes(review): Fix leftover-symlink refusal on non-markdown homes; hoist config check and lint/dedup cleanups * no-mistakes(document): Align fm-lint purity scope header with bin/backends --------- Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
friesentius
pushed a commit
to friesentius/firstmate
that referenced
this pull request
Sep 21, 2026
…kunchenguid#3417) * fix(backlog): honor configured task adapters * no-mistakes(review): Harden backend purity lint against prefixed Beads calls * no-mistakes(document): Document configured backend lifecycle transitions * fix(backlog): preserve markdown exemptions * no-mistakes(review): Enforce backend purity for explicit lint paths * no-mistakes(document): Update lifecycle backend documentation * no-mistakes(lint): Remove redundant backend lint pattern * fix(backlog): close adapter routing gaps * no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint * no-mistakes(document): Document environment-selected backlog adapters * no-mistakes(lint): Fix empty local variable assignment * fix(backlog): close quoted path gaps * no-mistakes(review): Reject partially quoted direct Beads commands * no-mistakes(document): Align lifecycle documentation with configured adapters * test(backlog): keep structural cases markdown-only * fix(backlog): honor configured task adapters * no-mistakes(review): Harden backend purity lint against prefixed Beads calls * no-mistakes(document): Document configured backend lifecycle transitions * fix(backlog): preserve markdown exemptions * no-mistakes(review): Enforce backend purity for explicit lint paths * no-mistakes(document): Update lifecycle backend documentation * no-mistakes(lint): Remove redundant backend lint pattern * fix(backlog): close adapter routing gaps * no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint * no-mistakes(document): Document environment-selected backlog adapters * no-mistakes(lint): Fix empty local variable assignment * fix(backlog): close quoted path gaps * no-mistakes(review): Reject partially quoted direct Beads commands * no-mistakes(document): Align lifecycle documentation with configured adapters * test(backlog): keep structural cases markdown-only * no-mistakes(review): Harden markdown lifecycle routing and close recovery * fix(lint): catch dollar-quoted beads commands * no-mistakes(review): Drop unused authorized-data-dir param; canonicalize fm-lint ROOT with pwd -P * no-mistakes(document): Document fm-lint backend-purity check in header and CONTRIBUTING * no-mistakes(ci): Two failing checks, one code-caused and one attestation-only. 1. Greptile Review (P1: ANSI-C quoting bypasses lint) — REAL DEFECT, FIXED. The backend-purity normalizer in bin/fm-lint.sh (invokes_bd awk function) stripped $'...' quote delimiters without decoding ANSI-C escapes, so a core script containing $'\x62\x64' close fm-example executes as `bd close` while the lint accepted it. Fix: the normalizer now tracks whether a single-quote context came from $' (ansi flag) and decodes ANSI-C escapes while building the command word: \xHH (1-2 hex), \uHHHH / \UHHHHHHHH (4/8 hex), \NNN (1-3 octal), \cX control characters, simple escapes (a/b/e/E/f/n/r/t/v decode to a placeholder that can never spell bd), NUL (value 0) truncates the word per bash C-string semantics, and unknown escapes drop the backslash per bash. Values outside printable ASCII decode to a placeholder so they can neither falsely match nor collide into `bd`. Regression coverage added to the existing lint-interface test test_rejects_direct_beads_cli_invocations in tests/fm-lint.test.sh: $'\x62\x64', $'\142\144' (octal), and b$'\x64' (split word), each written into a fixture and asserted rejected through the real fm-lint.sh executable. Verified regression property: with the fix stashed, the new hex case passes lint (reproduces the reported bypass); with the fix, it is rejected. Verification: all 29 tests in tests/fm-lint.test.sh pass, and the full CI-parity lint (CI=true bin/fm-lint.sh: ShellCheck 0.11.0 full analysis over the canonical set, backend-purity check, and actionlint workflow validation) exits 0. One iteration was needed because an awk comment containing a literal $'\x62\x64' sample broke shell-level quoting (bash -n / SC1001/SC2026); the comment was reworded without quotes. 2. PR must be raised via no-mistakes — NOT code-caused. The check failed with 'Pipeline attestation head_sha does not match the current PR head': the PR body attestation binds to 82b41c7 while the PR head is 1995cdb because a later pipeline push moved the head. This is exactly the stale-attestation condition the user intent describes; it clears when the outer pipeline re-runs 'git push no-mistakes' and re-binds the attestation to the new head (which now includes this Greptile fix). No code change can or should address it. Files changed: bin/fm-lint.sh (ANSI-C escape decoding in the backend-purity normalizer), tests/fm-lint.test.sh (three encoded-bd rejection cases) * no-mistakes(review): fix tasks.toml hang, root authorization, lint quoting * no-mistakes(review): validate tasks config before exemption; fix lint quote gap * fix(backlog): address the markdown backlog as <data>/backlog.md Resolving the markdown backlog through a configured `[markdown] path` was scope this task never asked for. It is absent from main, which addresses `<data>/backlog.md` everywhere, and it came from an earlier review round rather than the task brief. Making it effective on the transition path alone put that path at odds with every other consumer of the same backlog - fm-captain-hold.sh, fm-session-start.sh, fm-fleet-snapshot.sh, fm-inbox.sh, fm-backlog-handoff.sh - which all still address `<data>/backlog.md`. In fm-captain-hold.sh the split was live: its reads had already moved to the shared gate while its writes had not, so the two could address different files. Address `<data>/backlog.md` from the shared gate, delete the unused resolver, and drop the two tests that pinned the withdrawn behaviour. What this task actually changes is unaffected: a configured non-markdown adapter is still addressed by its own root, without `--file`. * fix(backlog): honor configured task adapters * no-mistakes(review): Harden backend purity lint against prefixed Beads calls * no-mistakes(document): Document configured backend lifecycle transitions * fix(backlog): preserve markdown exemptions * no-mistakes(review): Enforce backend purity for explicit lint paths * no-mistakes(document): Update lifecycle backend documentation * no-mistakes(lint): Remove redundant backend lint pattern * fix(backlog): close adapter routing gaps * no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint * no-mistakes(document): Document environment-selected backlog adapters * no-mistakes(lint): Fix empty local variable assignment * fix(backlog): close quoted path gaps * no-mistakes(review): Reject partially quoted direct Beads commands * no-mistakes(document): Align lifecycle documentation with configured adapters * test(backlog): keep structural cases markdown-only * fix(lint): catch dollar-quoted beads commands * no-mistakes(review): Drop unused authorized-data-dir param; canonicalize fm-lint ROOT with pwd -P * no-mistakes(document): Document fm-lint backend-purity check in header and CONTRIBUTING * no-mistakes(ci): Two failing checks, one code-caused and one attestation-only. 1. Greptile Review (P1: ANSI-C quoting bypasses lint) — REAL DEFECT, FIXED. The backend-purity normalizer in bin/fm-lint.sh (invokes_bd awk function) stripped $'...' quote delimiters without decoding ANSI-C escapes, so a core script containing $'\x62\x64' close fm-example executes as `bd close` while the lint accepted it. Fix: the normalizer now tracks whether a single-quote context came from $' (ansi flag) and decodes ANSI-C escapes while building the command word: \xHH (1-2 hex), \uHHHH / \UHHHHHHHH (4/8 hex), \NNN (1-3 octal), \cX control characters, simple escapes (a/b/e/E/f/n/r/t/v decode to a placeholder that can never spell bd), NUL (value 0) truncates the word per bash C-string semantics, and unknown escapes drop the backslash per bash. Values outside printable ASCII decode to a placeholder so they can neither falsely match nor collide into `bd`. Regression coverage added to the existing lint-interface test test_rejects_direct_beads_cli_invocations in tests/fm-lint.test.sh: $'\x62\x64', $'\142\144' (octal), and b$'\x64' (split word), each written into a fixture and asserted rejected through the real fm-lint.sh executable. Verified regression property: with the fix stashed, the new hex case passes lint (reproduces the reported bypass); with the fix, it is rejected. Verification: all 29 tests in tests/fm-lint.test.sh pass, and the full CI-parity lint (CI=true bin/fm-lint.sh: ShellCheck 0.11.0 full analysis over the canonical set, backend-purity check, and actionlint workflow validation) exits 0. One iteration was needed because an awk comment containing a literal $'\x62\x64' sample broke shell-level quoting (bash -n / SC1001/SC2026); the comment was reworded without quotes. 2. PR must be raised via no-mistakes — NOT code-caused. The check failed with 'Pipeline attestation head_sha does not match the current PR head': the PR body attestation binds to 82b41c7 while the PR head is 1995cdb because a later pipeline push moved the head. This is exactly the stale-attestation condition the user intent describes; it clears when the outer pipeline re-runs 'git push no-mistakes' and re-binds the attestation to the new head (which now includes this Greptile fix). No code change can or should address it. Files changed: bin/fm-lint.sh (ANSI-C escape decoding in the backend-purity normalizer), tests/fm-lint.test.sh (three encoded-bd rejection cases) * fix(bin): preserve captain calls during teardown (kunchenguid#3595) * fix(bin): never close a captain call during cleanup A scout that held its own work item for the captain, which is what captain-hold-lifecycle prefers ("hold the work item the question gates"), was closed by bin/fm-teardown.sh's automatic backlog transition. The completion gate passed, cleanup ran, and the captain's question moved to Done with no recorded answer: the one thing the policy says must never happen. `tasks-axi done` closes a held row silently, and nothing in teardown asked whether the row was the captain's own call. bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when the task is still an open captain call, 1 when it is not, 2 when that cannot be established. It reads the row through the transition library's backend-aware probe, so it addresses the same backlog teardown does; the script's other commands now address the configured data directory the same way instead of FM_HOME, which also fixes captain holds in a home with a relocated data directory. Teardown asks `open` before any destructive step and refuses on 2. On 0 only the close changes: after cleanup and still under the task's own lock, the row gets one "Deliverable of the finished work" line at the end of its body and returns to Queued through `tasks-axi reopen`, keeping its hold, so it lands in Captain's Call instead of reading as work under way. --force does not lift this: it authorizes discarding unlanded work, never the captain's question. The deliverable goes into the body because `tasks-axi update --report` rewrites the title of a row that is not Done. The crash window reuses the pending-close record teardown already stages: a `mode=retain` line makes the existing replay record the deliverable and reopen instead of closing, with the same validator, stale-generation check, cleanup-incomplete marking, and non-blocking bootstrap lock as an ordinary close. A retained row the captain answered first simply retires the record. No parallel record type, recovery command, or second bootstrap loop is introduced. Regressions run the real executables: the captain-held scout survives cleanup queued, held, with its deliverable and on the board, only `answer` closes it, --force keeps it open, and an ordinary scout still closes with its report; an interrupted cleanup leaves the row untouched and the next session start retains it; a relocated backlog keeps the retention in its one configured file; and a ship row whose hold cannot be read refuses cleanup before anything destructive. Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np * no-mistakes(review): Serialize captain holds and fix backend-aware listing * no-mistakes(document): Update captain-call retention documentation * no-mistakes(document): Fix relocated captain-hold backlog diagnostics * fix(backlog): honor configured task adapters * no-mistakes(document): Update lifecycle backend documentation * fix(backlog): close adapter routing gaps * no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint * no-mistakes(lint): Fix empty local variable assignment * fix(backlog): close quoted path gaps * test(backlog): keep structural cases markdown-only * no-mistakes(review): Harden markdown lifecycle routing and close recovery * no-mistakes(review): fix tasks.toml hang, root authorization, lint quoting * no-mistakes(review): validate tasks config before exemption; fix lint quote gap * fix(backlog): address the markdown backlog as <data>/backlog.md Resolving the markdown backlog through a configured `[markdown] path` was scope this task never asked for. It is absent from main, which addresses `<data>/backlog.md` everywhere, and it came from an earlier review round rather than the task brief. Making it effective on the transition path alone put that path at odds with every other consumer of the same backlog - fm-captain-hold.sh, fm-session-start.sh, fm-fleet-snapshot.sh, fm-inbox.sh, fm-backlog-handoff.sh - which all still address `<data>/backlog.md`. In fm-captain-hold.sh the split was live: its reads had already moved to the shared gate while its writes had not, so the two could address different files. Address `<data>/backlog.md` from the shared gate, delete the unused resolver, and drop the two tests that pinned the withdrawn behaviour. What this task actually changes is unaffected: a configured non-markdown adapter is still addressed by its own root, without `--file`. * no-mistakes(review): restore home boundary guard and tighten purity lint * no-mistakes(review): authorize home boundary for every backlog adapter * no-mistakes(test): complete tasks-axi stubs in fm-gotmp teardown fixtures * no-mistakes(document): align backlog transition docs with adapter-neutral addressing * no-mistakes(review): label adapter data-dir authorization, drop dead row_probe local * no-mistakes(review): pin markdown backend at relocated-data addressing roots * no-mistakes(document): point lint-definition mention at fm-lint.sh header * no-mistakes(document): point mutate comment at adapter addressing owner * no-mistakes(review): Fix leftover-symlink refusal on non-markdown homes; hoist config check and lint/dedup cleanups * no-mistakes(document): Align fm-lint purity scope header with bin/backends --------- Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
jorguez96
added a commit
to jorguez96/firstmate
that referenced
this pull request
Sep 25, 2026
…ts resolved (#19) * fix(bin): support process events under symlinked homes (#3484) * fix(bin): resolve process-event state roots before validating them The process-event module validated the caller's spelling of a home's state root instead of the directory it operates on: it required the supplied path to equal its own lexical normalization, which rejects any path reached through a symlinked ancestor. On macOS both /tmp and $TMPDIR are symlinks, so an operator home under either could never claim a source. Reconcile still reported the runner started, while the detached runner died writing "cannot claim source" to the discarded stderr, and the source silently never fired. Resolve the state root to its physical directory once, then apply the existing private-directory validation to that resolved directory and derive every path, recorded claim identity, and later confinement check from it. This keeps the confinement contract for the directory actually operated on rather than only for callers that already spelled it physically, and removes the window where an ancestor symlink could be repointed between check and use. Homes already spelled physically behave identically. This was the single cause of both deterministic macOS failures in tests/fm-procevent.test.sh ("reconcile never claimed the registered source") and tests/fm-procevent-when.test.sh ("the winning concurrent arm did not produce an outcome"). The new case pins the behavior with an explicit symlinked-ancestor home, so it fails without the fix on any platform rather than only where the temp root happens to be a symlink. * fix(bin): pin the external capture staging boundary to its physical path The extension capture path pinned its registry staging boundary by comparing `pwd -P` against the caller-spelled registry directory, so a home reached through a symlinked ancestor still refused to start an extension-backed source after the state root itself resolved correctly. That left such a home half working: built-in sources ran while external ones failed. The staging preparer now prints the physical registry directory it validated, matching the inbox and reservation preparers beside it, and the start path pins on that returned path. The new end-to-end case drives the shipped file-signal package from a symlinked home spelling. * no-mistakes(review): Propagate canonical process-event state roots * no-mistakes(review): Propagate canonical state to process-event adapters * no-mistakes(document): Document physical process-event state roots * fix(pi): deliver captain outcomes as deterministic transcript entries (#3312) * fix(pi): persist captain outcomes visibly * no-mistakes(review): Recover captain outcomes after cold-start lock acquisition * no-mistakes(document): Document cold-start captain-outcome recovery * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes(review): Prove immediate Pi captain-outcome transcript delivery * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * fix(pi): process captain outcomes through a sequence-keyed turn PR #3312 made every captain-facing supervision outcome a durable, exact-once visible transcript entry with the read cursor advancing only after that entry exists. That is the display half of the delivery contract. Left alone it turns a probabilistic silent loss into a deterministic one: the captain sees an anchor line, and firstmate never acts, because nothing opens a turn and nothing records whether main ever processed the outcome. The 2026-08-31 timeline showed the two shapes this must survive on the previous hidden-turn path: seven delivered decision outcomes each answered by an empty assistant message (cursor advanced, no retry, unanswered for close to three hours), and two answered by an unrelated prior reply. Both happened because delivery advanced the cursor at enqueue and accepted whatever the next assistant message was. Add the processing half on top of the persistence half: - bin/fm-branch-outcome.sh keeps a processed marker separate from the read cursor (`unprocessed`, `mark-processed --through`, `processed-init`). It only advances through an explicit sequence-bound acknowledgement, never past the read cursor and never backwards; an absent marker reads as zero and `processed-init` migrates delivered history once so an upgraded home is not re-presented its past. - After the visible entry for a captain outcome exists, the extension hands every still-unprocessed captain row to main as one hidden, typed `fm-branch-process` request listing each `[seq N] task: summary`, opening exactly one main turn. Main closes it only by calling the new `fm_branch_processed` tool with the highest sequence listed. An unrelated, empty, or paraphrased answer leaves the sequence open, and the same request is presented again at the end of the next main run and at session start. The first two presentations of a sequence set open a turn of their own; after that the request rides the captain's next prompt so an ignored request cannot loop, and a session replacement resets that budget. Routine outcomes stay turn-free. - The regressions cover exactly those incident shapes against the real store scripts: an empty answer and an unrelated prior answer neither advance the marker nor stop re-presentation, the acknowledgement is refused beyond the read cursor and outside lock ownership, a partial acknowledgement keeps the newer sequence open, and #3312's own assertions now forbid an unkeyed turn rather than any turn. The store suite pins the marker's bounds and the migration; the real-SDK guard for appendEntry persistence and model exclusion is unchanged. Docs move the protocol from "no model turn" to "one sequence-keyed processing turn closed only by its acknowledgement", and the verification record carries the dated run against Pi 0.84.4. * no-mistakes(review): Harden outcome listing and sequence-bound acknowledgements * no-mistakes(review): Harden outcome state validation and request pacing * no-mistakes(review): Reject unsafe sidecars and unterminated outcome stores * no-mistakes(review): Validate canonical mark-read cursor state * no-mistakes(review): Guard cursor advancement against corrupt processed state * no-mistakes(review): Bind acknowledgements to active processing requests * no-mistakes(review): Reset pacing when processing sequence membership changes * no-mistakes(review): Enforce silent outcome invariants at storage boundary * no-mistakes(document): Document hardened captain outcome processing contracts --------- Co-authored-by: kunchenguid <kun@kunchenguid.com> * feat: add bounded concurrent Bearings ledger collection (#3481) * feat: bound Bearings remote ledger collection * no-mistakes(review): Clarify default remote-ledger collection behavior * no-mistakes(review): Detach reconcile delivery from watcher loop * no-mistakes(review): Enforce bounded snapshot and request captures * no-mistakes(review): Bound legacy summary capture before parsing * no-mistakes(review): Bound primary remote ledger captures * no-mistakes(document): Correct snapshot and reconcile documentation * no-mistakes(lint): Fix ShellCheck quoting in bounded collector * no-mistakes(ci): Fixed all three CI failures: updated the macOS Bearings assertion to 44 tests, made the home-summary test deterministic and aligned with default ledger consumption, and increased the asynchronous reconcile retirement wait for loaded CI. Verified both focused suites, all 44 Bearings tests, ShellCheck, actionlint, Bash parsing, and git diff checks * test: await reconcile request retirement * no-mistakes(review): Avoid empty reconcile queue process churn * no-mistakes(review): Read ledger summaries from immutable snapshots * no-mistakes(review): Reject multi-document home ledger streams * no-mistakes(review): Coalesce durable reconcile requests per target * no-mistakes(review): Unify reconcile keys and reject snapshot streams * no-mistakes(review): Key reconcile requests by stable target ID * no-mistakes(document): Document per-target reconcile request coalescing * no-mistakes(lint): Remove unused snapshot summary file variable * no-mistakes(ci): Adjusted the concurrent collector regression’s end-to-end timing ceiling to account for stock macOS process/jq overhead outside the three-second remote collection budget, while remaining below the 15-second serial-read floor. Verified with stock /bin/bash 3.2: all 44 Bearings tests pass; bash syntax and git diff checks pass * no-mistakes(ci): Fixed legacy summary validation to require exactly one top-level JSON document and added behavioral regression coverage. Stabilized CI by conditionally waiting longer for durable reconcile delivery and synchronously stopping the fm-on worker tree before fixture cleanup. Removed a redundant flaky healthy-path timing assertion; the wedged-reader test still proves concurrent bounded collection. Verified fm-bearings-snapshot, fm-secondmate-reconcile, and fm-on tests, plus project ShellCheck, bash syntax, and git diff checks * ci: rebalance portable serial test shards (#3489) * fix(ci): rebalance the portable serial shards on measured durations The "Behavior portable serial 3" shard ran 17-20 minutes against its 20-minute job cap and intermittently timed out seconds after a passing test, on branches and on main alike. Shards are packed longest-processing-time from per-script duration hints, and those hints were last measured on 2026-08-21 at 116 scripts. The lane has since grown to 139 scripts and from ~42 to ~63 minutes: 17 scripts had no hint at all and fell back to the 20 s default, and several existing hints were low by 2-5x (fm-watch-triage 142 s hinted vs 263 s measured, fm-public-followup 36 s vs 197 s). The partition therefore looked perfectly balanced in hint space, 734.6 s per shard, while really running 11.5, 13.6, 18.8 and 16.5 minutes. Script-count balance, which is what the tests asserted, stayed normal throughout and hid it. Refresh the hints from the timing artifacts of three green runs, taking the slowest measurement of each script so the balance holds on a slow runner, and split the lane across five shards instead of four. Replayed against those runs' real per-script durations the worst shard is now 12.54 minutes, 63% of the unchanged 20-minute cap, and the serial lane's wall clock drops from ~20 to ~12.5 minutes. Bound the drift that caused this rather than relying on the hints being refreshed by hand: the coverage guard now reports the unmeasured share as serial_unhinted= and refuses past PORTABLE_SERIAL_MAX_UNHINTED_PERCENT, which leaves room for newly added tests while making a stale table fail the guard instead of silently pushing one shard into its cap. No test changes what it asserts and no test stops running; only the partition across shards changes. * no-mistakes(document): Clarify conservative shard timing aggregate * fix(pi): fall back on incomplete supervision branch prompts (#3491) * fix(pi): fall back after settled branch errors * no-mistakes(review): Detect provider errors across prompt compaction * no-mistakes(review): Preserve in-flight branch state across selection changes * fix(pi): re-probe supervision branch after cooldown (#3497) * fix(pi): recover supervision branch after cooldown * no-mistakes(review): Defer branch recovery until prompt settlement * no-mistakes(document): Clarify supervision cooldown recovery contract * fix(bin): remove legacy remote snapshot reads (#3501) * refactor: remove legacy remote summary reads * no-mistakes(document): Document ledger-only snapshot reads * no-mistakes(ci): Fixed the snapshot test fixture so ledger refreshes use the same fake executable PATH as the snapshot consumer. This preserves observable endpoint freshness after removing legacy summary computation. Verified stock Bash parsing and all 44 Bearings tests pass under /bin/bash; git diff checks pass * no-mistakes(ci): Fixed the CI-only snapshot fixture failure by ensuring the bounded-ledger refresh uses its fake tmux backend. This removes host tmux availability as a source of nondeterminism. Verified all 44 Bearings tests pass, Bash syntax passes, and git diff checks are clean * no-mistakes(ci): Fixed CI nondeterminism in the Bearings fixture: all local ledger refreshes now use the fixture’s fake tmux backend when available, instead of depending on host tmux state. Verified stock /bin/bash syntax, git diff checks, and all 44 Bearings tests with a deliberately failing host tmux * fix(pi): preserve watcher continuity across session replacement (#3498) * fix(pi): rearm watcher after session replacement * no-mistakes(review): Queue actionable closes across Pi session replacement * no-mistakes(review): Stop replacement arm when handoff persistence fails * no-mistakes(review): Preserve actionable wakes through branch and late child races * no-mistakes(review): Surface late handoff failures without crashing Pi * no-mistakes(review): Coordinate replacement delivery settlement and unique handoff tokens * no-mistakes(review): Retry stale deliveries and release settled claims * no-mistakes(review): Distinguish branch settlement and retry handoff cleanup * no-mistakes(review): Deduplicate persistent handoff cleanup alerts * no-mistakes(review): Acknowledge watcher follow-ups only when consumed * no-mistakes(review): Persist idle follow-ups until agent consumption * no-mistakes(review): Preserve pending outcomes when handoff persistence fails * no-mistakes(review): Arm replacement before awaiting prior delivery settlement * no-mistakes(review): Adopt pending handoffs after lock reclamation * no-mistakes(review): Prevent stale generations from adopting replacement handoffs * no-mistakes(review): Scope replacement handoffs by watcher state * no-mistakes(document): Clarify replacement handoff documentation * no-mistakes(ci): Fixed the failing branch-extension tests to model the new settlement-promise contract. Failure cases now assert that delivery ownership returns to the watcher instead of expecting direct extension fallback. Verified the updated branch suite, Pi watcher suite, shell syntax, and diff checks * no-mistakes(review): Update branch settlement tests and preserve chunked outcomes * no-mistakes(document): Document watcher-owned replacement handoffs * no-mistakes(document): Verify replacement handoff documentation * test(pi): cover watcher-owned branch fallback * no-mistakes(document): Refresh watcher-owned fallback documentation * fix(bin): resurface task statuses missed by wake handling (#3495) * fix(bin): resurface terminal statuses lost after branch handling * test(watch): canonicalize process-event fixture homes * no-mistakes(review): Index branch outcomes by causal status position * no-mistakes(review): Recover outcome indexes and deduplicate resurfaced statuses * no-mistakes(review): Handle legacy ambiguity and oversized status diagnostics * no-mistakes(review): Keep unclassifiable oversized statuses silent * no-mistakes(document): Document lost-wake outcome backstop * no-mistakes(document): Update outcome backstop documentation * no-mistakes(ci): Fixed CI regressions in wake-drain: parseable reserved-key decisions can no longer bypass the durable decision-fold guard, and status output is prepared and receipt-committed before presentation to prevent repeated one-shot outcomes after later failures. Added a behavioral regression for receipt commit failure and retry. Targeted backstop, correlation-token, decision-cursor, open-decision, unread-status, syntax, and diff checks pass locally. Shard-4 failures appeared unrelated/flaky; the network-parallel test passed locally * no-mistakes(ci): Fixed the Greptile P1 data-loss issue by committing presentation receipts only after prepared output reaches stdout. Added behavioral coverage proving output failure leaves the backstop retryable and receipt failure may duplicate but never lose a presentation. Relevant wake-drain suites and syntax/diff checks pass. The shard-4 Pi extension failure is unrelated to this PR and did not warrant changes * no-mistakes(ci): Stabilized tests/fm-bootstrap-network-parallel.test.sh by replacing scheduler-sensitive equal-sleep timing with bounded synchronization between mocked fetch and remote probes. This preserves detection of real serialization while avoiding false failures under CI load. Verified with five consecutive test runs, bash syntax validation, ShellCheck, and git diff checks. The separate Pi stock-rendering failure reproduces locally but is unrelated environment/version drift * no-mistakes(ci): Fixed Behavior portable serial 4 by adding fm-classify-lib.sh and fm-timeout-lib.sh to the broken-root Pi test fixture; fm-branch-outcome.sh now depends on them. Verified the full Pi branch-extension suite with real-Pi checks skipped, the wake-drain outcome-backstop suite, Bash syntax, and git diff checks. Greptile findings are already addressed at HEAD; the no-mistakes attestation failure is external head-SHA state * fix(bin): collect follow-up results from remote work homes (#3503) * fix(bin): deliver typed terminal results from remote work homes A public commitment whose work is bound to a REMOTE secondmate home could never receive its typed terminal result. `fm-public-followup.sh brief` printed an emit command carrying this home's own absolute path and this checkout's own script path, neither of which exists on the machine the worker runs on, so the worker had nothing it could write to that the owning home would ever read - and `consume` kept finding nothing while the promise stayed open. The brief is now route-aware: for a remote work home it prints that route's own code root and home with `--stage-in`, so the typed event is staged in the home where the work actually runs, and the closing paragraph names the owning home as the one on the other machine instead of pointing at the path above it. The owning home collects those staged results over the same SSH route it reaches that secondmate on, because the transport only runs outbound: `consume` pulls them into its own inbox and reconciles them exactly as it reconciles a local report. Collection is non-destructive until the result is durably held, so a dropped connection cannot lose a terminal result, and a route that could not be reached is named in `consume`'s output with the promise left open rather than reported as an empty inbox. A local work home is untouched: the brief still prints `--home` with this home and this checkout's script, and the event still lands directly in this home's typed terminal-result inbox. This is the emit-side counterpart of the retire/clear fix in #3479 and reuses the remote-route resolution that landed with it. Reconciling a loop bound to a remote route now reaches that route, so the existing remote cases drive `consume` through the same faked transport their other steps already use. * no-mistakes(review): Fail loudly on unresolved routes and invalid staging homes * no-mistakes(review): Fail collection when remote outbox is unreadable * no-mistakes(review): Surface reassigned remote routes during empty collection * no-mistakes(review): Fail remote collection on invalid registrations * no-mistakes(review): Reject unsafe registration entries during remote collection * no-mistakes(review): Restore healthy empty remote collection behavior * no-mistakes(review): Skip remote collection for delivered registrations * no-mistakes(review): Skip delivered registrations before route validation * no-mistakes(document): Document remote follow-up collection semantics * fix(bin): exclude secondmates from home-summary validity (#3504) * fix(bin): exclude secondmates from home-summary child inventory kind=secondmate meta records never have backlog rows, so counting them in unowned_children or terminal_in_flight made a clean main home look invalid once earlier ledger checks passed. * no-mistakes(review): Cover terminal secondmate in-flight exclusion * no-mistakes(ci): Updated the stock macOS Bash CI snapshot expectation from 15 to 16 tests. Verified all 16 snapshot/fleet-view tests pass under Bash 3.2.57 and `git diff --check` succeeds * fix(bin): self-heal outcome indexes on first drain (#3509) * fix(bin): self-heal status-outcome indexes on every drain Missing ready markers were skipping the lost-wake backstop on non-Pi homes because only the Pi branch ran processed-init. Drain now rebuilds those indexes under the outcome lock and fails closed only on a real store fault. * no-mistakes(review): Guard held-lock initialization and fail marker writes * no-mistakes(document): Document cross-harness outcome-index self-healing * fix(bearings): keep active children underway during captain holds (#3505) * fix(bearings): keep active children underway beside a captain hold Project each readable home's active children into Underway independently of the home-level captain-decision classification so a hold no longer hides live work. * no-mistakes(review): Preserve Underway repos and disclose child truncation * no-mistakes(review): Fall back to task project for Underway repos * no-mistakes(ci): Updated the stock macOS Bash CI assertion from 44 to 45 Bearings tests, matching the newly added behavioral regression. Verified all 45 tests pass under /bin/bash, Bash syntax checks pass, and git diff validation is clean * fix(pi): settle watcher delivery on Pi accepting the follow-up (#3513) * fix(pi): settle watcher delivery on Pi accepting the follow-up A follow-up queued while main is streaming joins the running run without ever raising before_agent_start, so waiting on that event before clearing the successor pipeline (#3498) stalled every later actionable close: no successor started, no wake was delivered or offered to the branch, and the turn-end guard woke main to re-arm by hand after every close. The pipeline now settles once Pi accepts the follow-up. Consumption is observed at before_agent_start for an idle main and at the user message_start for a streaming main, and decides only what a replacement session (/new, /resume, /fork, reload) replays. An exhausted restoration delivers its typed failure without launching an arm past the retry bound, which the stall had hidden. The replacement-coordinator map is typed so the strict no-emit typecheck passes again. Tests: the doubles no longer raise before_agent_start for a streaming send, a portable regression drives two actionable closes while main streams and proves the successor chain plus consumption-scoped replay, and a credential-free real-SDK probe pins Pi's event contract for both the streaming and the idle follow-up. Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a * fix(pi): retry a verified successor that fails during wake delivery A verified successor can exit while the wake it was started for is still being delivered, most plausibly during a branch turn that holds the settlement for minutes. Its failure close arrived while the pipeline's single-flight guard was set, so the close handler skipped the retry, and the pipeline's end no longer launched an arm, which left the live generation with no watcher and no retry timer. The close handler now records that failure when the child had reported readiness and was not retired by the restoration itself, and the pipeline runs the ordinary bounded, lock-checked retry for it once the delivery settles. A restoration started for a later pending supersedes it, and an exhausted restoration still hands repair to main without a further arm. The regression holds a branch settlement open while the verified successor exits with a failure and proves one retry watcher starts after the settlement releases, none while it is held. Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a * fix(bin): bound repeat stale wakes for parked workers (#3532) * fix(bin): bound repeat stale wakes for a parked but live worker A worker parked on a declared wait - `paused:` for an external or pipeline wait, or a verified `captain-held` transfer - kept waking firstmate far inside FM_PAUSE_RESURFACE_SECS. Observed as five consecutive alarms on one captain-held worker and dozens across a day on a pipeline wait, and reported upstream as four wakes in 75 minutes against a 3600s window. pause_state_class deliberately answers `none` for a still-live agent even under a declared wait, so a worker genuinely waiting on a decision is never silenced. That classification is correct and is left alone; it routes every parked but live worker through surface_nonterminal_stale on first sight of each distinct stale hash, and an idle parked pane still churns its hash on a clock or a token counter without changing what is being waited on. Two places let that churn re-alarm: - surface_nonterminal_stale queued the wake BEFORE consulting whether a wait was declared, then wrote `.paused-resurfaced-<key>` - the very throttle that should have suppressed it. The throttle was never read on this path and was advanced by the wake it should have prevented. - The hash-change path cleared that throttle through clear_pause_tracking whenever the classification came back `none`, so each tick also bought the same declared wait a fresh window. Fixing only the first site changes nothing. Read the throttle before anything is queued and advance it only on a wake that really fires, and on the hash-change path reset only the per-hash bookkeeping while the declaration still stands, via a clear_stale_hash_tracking split so neither half of clear_pause_tracking is duplicated. The throttle is keyed to the declaration, not to the pane. First sight still wakes, so an inconclusive state is still inspected, and the window's end still re-surfaces once, so a forgotten wait cannot rot invisibly - noise traded for a bounded cadence, never for silence. The wake identity stays the plain `stale: <win>` the away-mode handoff depends on. Tests cover both observed forms and were confirmed to fail against three deliberate breaks: each site reverted on its own, and a re-surface that never fires again. * fix(document): Clarify declared-wait wake cadence documentation * fix(ci): Captain, fixed the stale-throttle inheritance: cadence markers now bind to the current wait declaration, so replacement paused and captain-held waits each emit their first plain `stale:` wake. Added behavioral coverage for both forms. Bite proof failed as expected when identity matching was removed, then passed after restoration. Full watcher triage suite, `bin/fm-lint.sh`, syntax checks, and diff checks pass. Changes remain uncommitted for the outer executor * fix(ci): Captain, fixed the confirmed Greptile finding. `resurface_absorbed` now applies a throttle only when its stored declaration scope matches the current wait, so replacement `paused:` and `captain-held` waits surface immediately without changing classification. Added executable coverage for both absorbed forms. Bite proof failed before the fix at the intended assertion; afterward the full watcher triage suite, `bin/fm-lint.sh`, shell syntax checks, and `git diff --check` passed * fix(bin): accept the away-mode daemon as the turn-end supervision owner (#3567) * fix(turnend): accept the away-mode daemon as the supervision owner While state/.afk exists the away-mode daemon owns supervision and runs bin/fm-watch.sh one-shot: the watcher exits on every wake and the daemon starts its replacement. The turn-end guard tested for a live watcher process holding the watch lock at that instant, so a turn boundary that landed in the hand-off blocked with "TURN WOULD END BLIND" while supervision was completely healthy, costing a full handling turn each time. Reproduced with the real daemon wrapping the real watcher and the real guard sampling the same home: 6 of 40 samples blocked, every one of them with the daemon alive and the beacon 2-3 seconds old, and a new watcher pid on each cycle. After the fix the same reproduction blocks 0 of 40, and killing the daemon and its watcher (away mode still on, beacon still fresh) blocks again. The guard now accepts a live, identity-matched daemon holding this home as proof of supervision while away mode is active. The identity match is the same discipline the watcher lock uses, so a recycled pid or a lock left by a killed daemon proves nothing. The fresh-beacon half of the predicate is unchanged: a daemon that stops restarting its watcher still blocks once the beacon passes grace, a home with no supervisor blocks exactly as before, and with away mode off the strict watcher predicate is untouched. The predicate reads only durable state, so it behaves identically for every primary harness and runtime backend. * no-mistakes(document): clarify away-mode daemon supervision proof and test coverage * no-mistakes(document): generalize stale turn-end predicate summary in architecture.md * fix(backlog): omit --file from row probes for non-markdown backends (#3582) * fix(backlog): omit markdown file for beads probes * no-mistakes(document): Narrow backlog addressing doc to mutations for backend-aware probes * no-mistakes(ci): Fixed the Greptile P2 review comment (the only failing check) on tests/fm-backlog-atomicity.test.sh. The comment correctly noted that an exported TASKS_AXI_BACKEND environment variable would inherit into the spawned scripts and, because fm_tasks_axi_backend gives it top precedence, override each test case's .tasks.toml backend fixture — making the backend-specific argv assertions fail for environmental reasons. Fix: unset TASKS_AXI_BACKEND in the test harness right after sourcing tests/lib.sh, with a comment explaining why, so every case deterministically exercises its declared backend (4 lines added; no production code touched). Verified: reproduced the leak before the fix (TASKS_AXI_BACKEND=beads made the markdown dispatch case fail with 'beads show failed', exactly the reported failure mode); after the fix the full suite passes (0 failures, exit 0) both with and without TASKS_AXI_BACKEND=beads exported. The added lines are shellcheck-clean (the only shellcheck note, SC1091 on the lib.sh source line, pre-exists this change) * fix(bin): classify progress updates on requested work as routine (#3589) The supervision branch's verdict rule escalated every outcome that answered a captain request, so "the work started" and "still working" notes reached the captain with nothing to look at. The rule now keeps a finished result of requested work captain-facing, even when healthy, and treats start or still-working updates that bring no new artifact, finding, or decision as routine. The captain list for review-ready PRs, ask-user findings, exhausted blockers, credentials, and destructive or security-sensitive cases is unchanged, as are the unsolicited-routine, silent-fleet-review, and doubt-chooses-captain rules. The fm_branch_report tool description and the two docs that restated the old unconditional rule now point at the prompt's "Verdict: routine or captain" section as the one owner instead of carrying a second copy. * fix(bin): preserve captain calls during teardown (#3595) * fix(bin): never close a captain call during cleanup A scout that held its own work item for the captain, which is what captain-hold-lifecycle prefers ("hold the work item the question gates"), was closed by bin/fm-teardown.sh's automatic backlog transition. The completion gate passed, cleanup ran, and the captain's question moved to Done with no recorded answer: the one thing the policy says must never happen. `tasks-axi done` closes a held row silently, and nothing in teardown asked whether the row was the captain's own call. bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when the task is still an open captain call, 1 when it is not, 2 when that cannot be established. It reads the row through the transition library's backend-aware probe, so it addresses the same backlog teardown does; the script's other commands now address the configured data directory the same way instead of FM_HOME, which also fixes captain holds in a home with a relocated data directory. Teardown asks `open` before any destructive step and refuses on 2. On 0 only the close changes: after cleanup and still under the task's own lock, the row gets one "Deliverable of the finished work" line at the end of its body and returns to Queued through `tasks-axi reopen`, keeping its hold, so it lands in Captain's Call instead of reading as work under way. --force does not lift this: it authorizes discarding unlanded work, never the captain's question. The deliverable goes into the body because `tasks-axi update --report` rewrites the title of a row that is not Done. The crash window reuses the pending-close record teardown already stages: a `mode=retain` line makes the existing replay record the deliverable and reopen instead of closing, with the same validator, stale-generation check, cleanup-incomplete marking, and non-blocking bootstrap lock as an ordinary close. A retained row the captain answered first simply retires the record. No parallel record type, recovery command, or second bootstrap loop is introduced. Regressions run the real executables: the captain-held scout survives cleanup queued, held, with its deliverable and on the board, only `answer` closes it, --force keeps it open, and an ordinary scout still closes with its report; an interrupted cleanup leaves the row untouched and the next session start retains it; a relocated backlog keeps the retention in its one configured file; and a ship row whose hold cannot be read refuses cleanup before anything destructive. Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np * no-mistakes(review): Serialize captain holds and fix backend-aware listing * no-mistakes(document): Update captain-call retention documentation * no-mistakes(document): Fix relocated captain-hold backlog diagnostics * fix(bin): deliver secondmate outcomes to the parent channel (#3592) * fix(bin): deliver every secondmate outcome on the parent channel from the recording scripts A secondmate's captain-facing outcomes could miss: the mate model addressed the captain in its own unread chat instead of appending to the parent channel, and a PR-ready report, a finding, a decision, a blocker, and a failure all depended on that one remembered append. Make delivery structural, so the parent channel never depends on the model: - bin/fm-parent-channel-lib.sh is the one owner of channel resolution and exact-line append-once; the merge outcome path and the inactive-outcome scan now publish through it instead of two private copies. - bin/fm-inactive-reconcile.sh gains a ledger-first path that runs on every watcher poll in a secondmate home: a direct child's whole terminal done or failed line is delivered at once with its note, recorded PR, mode, merge posture, and scout report pointer, keyed and receipted so it is delivered once, and the inactive path yields to it. `report <task-id>` runs the same delivery for a caller holding the child's meta lock. - bin/fm-pr-check.sh publishes the PR-ready line with the canonical URL at registration. - bin/fm-captain-hold.sh publishes a hold and its answer, keyed by task id and resolution-record count, with no new persisted state. - bin/fm-teardown.sh delivers the child's final line before removing its record and refuses, retaining every record, while the channel cannot be written. - The charter opens with the parent-channel rule and confines the mate's own appends to judgement; AGENTS.md carries the carve-out at the persona address rule and the escalation list. docs/secondmate-parent-channel.md records the design and its coverage, and docs/verification/secondmate-parent-channel.md records the live run with real tmux panes and both real watchers delivering every line with no model. Supersedes #3569. * no-mistakes(review): Fix parent outcome retries and reconciliation locking * no-mistakes(review): Prevent busy children from starving ledger delivery * no-mistakes(review): Correct ledger metadata and hold occurrence handling * no-mistakes(review): Disambiguate ledger outcomes and normalize hold reasons * no-mistakes(review): Close ledger races and preserve teardown records * no-mistakes(document): Correct parent-channel receipt and scanner documentation * no-mistakes(lint): Quote done arguments for ShellCheck compliance * no-mistakes(ci): Fixed both CI failures. Updated GOTMP teardown fixtures for the new final-outcome reporter and isolated them from host tmux state. Updated the PR security assertion to distinguish the accepted PR-ready line from duplicate merge outcomes. Verified with both failing test suites, bash syntax checks, and git diff checks * no-mistakes(ci): Fixed Greptile’s duplicate-delivery race in bin/fm-inactive-reconcile.sh. Ledger events now claim matching already-delivered inactive receipts using the prior status fingerprint, preventing duplicate parent reports while preserving later same-state completions. Added behavioral regression coverage. Verified inactive-reconcile tests, project lint, documentation audience checks, syntax, and diff checks. Teardown tests passed relevant cases before the documented pre-existing herdr-preflight-missing-adapter failure * fix(bin): sync remote second mates to primary commit (#3599) * fix(bin): sync remote second-mate homes to the parent primary commit Session start and remote launch pointed a remote second-mate home at whatever Firstmate copy its own host kept, so a home that had already advanced past that copy refused as a non-fast-forward and every other home stopped at the host's older commit while the primary ran ahead. The parent now resolves ITS primary default-branch commit with the existing helper and hands that commit to the host on both paths. Because a remote home is a standalone clone, the host imports that one commit before advancing - already present, else from that host's Firstmate copy without moving it, else from the home's own origin - and then runs the SAME ff_target guards a local home gets, so dirty, diverged, feature-branch, and unresolvable targets skip untouched and the ancestry rules keep one owner. An unimportable target now names /updatefirstmate instead of failing opaquely, and a host still running an older Firstmate copy is reported the same way rather than echoing a bare refusal. The host-local launch leg no longer re-runs its own secondmate sync, so the spawn it drives cannot re-target that host's copy after the parent has already converged the home. /updatefirstmate is unchanged: it still refreshes the remote code root from that host's origin and then syncs the home to that refreshed copy, which is what the sync call with no target commit means. * no-mistakes(document): Document primary-targeted remote secondmate synchronization * fix(bin): separate captain intent from firstmate specs (#3597) * fix(bin): split brief task into captain intent and firstmate spec Keep no-mistakes --intent as the captain's ask plus later captain words, not the build spec or worker tradeoffs. * fix(bin): stop task-subsection copies at the next heading Promotion was swallowing the scout Setup contract into Firstmate spec, and pre-subsection briefs lost their # Task body. * no-mistakes(review): Validate brief content and preserve nested specifications * no-mistakes(review): Scope placeholder validation to scaffold-only subsection bodies * no-mistakes(review): Ignore fenced subsection headings during brief validation * no-mistakes(review): Preserve captain intent across scout promotion * no-mistakes(review): Enforce safe intent boundaries for legacy promotions * no-mistakes(review): Allow marked legacy intent and reject empty promotions * no-mistakes(review): Scope task parsing and overlay legacy intent contracts * no-mistakes(review): Overlay current intent contract for all no-mistakes spawns * no-mistakes(review): Preserve later captain clarifications in intent overlays * no-mistakes(document): Document brief intent enforcement and ownership * no-mistakes(ci): Updated spawn-related test fixtures to use valid Captain intent and Firstmate spec subsections, corrected launch-path expectations to launch-brief.md, and resolved ShellCheck quoting findings. Verified with fm-lint.sh and 15 affected behavior tests, including real Herdr tests; all passed * no-mistakes(ci): Updated stale spawn/promotion fixtures in the Muse, Orca, secondmate-harness, and public-followup suites to provide valid Captain's intent and Firstmate spec subsections. Verified full Orca and secondmate-harness suites, targeted public-followup promotion behavior, Bash syntax, diff checks, and fm-lint * fix: start a fresh supervision branch for every main session (#3600) * fix(pi): start a new supervision branch conversation per main session The supervision branch reopened one recorded conversation forever, so every main session start reloaded the current generated prompt and then weeks of accumulated thread, where a superseded rule could still outweigh today's. The branch conversation is now scoped to one main session: the session generation owns the recorded conversation, so a cold start, /new, /resume, /fork, or a reload always builds a new one, while a rebuild inside one session (a model or effort change) still continues that session's own conversation. The dialog mirror re-anchors with it. Its durable cursor records what the previous branch conversation received, so a /resume or reload - which keeps main's own session file - would otherwise leave the new branch blind to dialog main itself still has. The reset is bounded by the current main session, and the cursor keeps advancing incrementally within it. The durable outcome store and its processed marker are untouched, so unacknowledged captain-facing outcomes still re-present on the new main session. * no-mistakes(document): Document fresh Pi supervision conversations * no-mistakes(ci): Fixed the flaky concurrent inbox failure. Lock acquisition now retries when a competing lock disappears between a failed claim and inspection. Added a behavioral regression covering that race. Verified the full inbox test four times, project lint, and git diff checks * feat: restart second mates after instruction updates (#3614) * feat(update): restart second mates whose instructions changed /updatefirstmate pulled new bytes onto disk and then asked each advanced second mate to re-read them. A running agent holds AGENTS.md and every loaded skill frozen from launch and no verified harness offers a reload, so that steer could not reach a loaded skill at all and left the mate holding two contradictory copies of its own job description. An eligible mate is now restarted instead, in the same home and endpoint, through the existing transactional relaunch. The restart is gated on the mate first writing down the open work it holds only in conversation - the open-record half of /stow, never its memory sweeps - so an unregistered captain call is flushed before the conversation is spent. Anything that leaves the reload unprovable falls back to the old re-read message and is reported as exactly that, never as a clean reload. Remote mates take the same path: fm-remote-secondmate-control.sh gains a relaunch verb whose host-local leg runs that same control plane, since the mate is an ordinary local secondmate from its host's point of view. The primary resolves the profile and passes it explicitly, because config/secondmate-harness is not inherited and the file on that host belongs to a different home. fm-update.sh now splits its advanced live mates into a restart set and a nudge residual, and both sets require a changed instruction surface, which also closes the over-nudge against the session-start sweep. Restart is stricter still: a bin/-only advance reloads itself on the next call, so it never costs a conversation. Colocated tests cover the gating, the persist-then-restart order, the task-subset persist request, each unsafe fallback, the remote hop, and the remote sync's new instruction-surface report. * no-mistakes(review): Fix restart correlation, concurrent waits, and lifecycle reporting * no-mistakes(review): Parallelize relaunches and classify replacement incarnations * no-mistakes(review): Gate restart actions on live agent state * no-mistakes(review): Handle failed restart workers without hanging * no-mistakes(review): Nudge legacy remotes and preserve persist recovery * no-mistakes(review): Document one-time secondmate restart rollout * no-mistakes(review): Honor arrived replies and refresh remote profiles * no-mistakes(review): Revert remote parent profile reconciliation * no-mistakes(review): Reset remote profile defaults and honor published results * no-mistakes(review): Preserve fallback nudges for unverifiable secondmates * no-mistakes(document): Document second-mate restart update flow * no-mistakes(lint): Fix ShellCheck warnings in restart scripts * perf: accelerate local validation with bounded concurrency (#3644) * perf(tests): route gate verification through the bounded concurrent runner Local validation was the pipeline's dominant cost: across 67 recorded no-mistakes agent sessions on this repo, 99.3% of command execution was `bash tests/*.test.sh`, run strictly one script at a time, and 2% of those calls were killed by an agent-guessed timeout and paid for twice. Three changes, each measured: - `.no-mistakes.yaml` pins `commands.test` to `bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner already owns changed-file selection, bounded concurrency, the refusal of unproven scripts, and a generous automatic per-script bound, so the gate's baseline is neither a serial chain nor a guessed timeout. It stays intent-targeted - the Test step still runs its evidence agent on top - and excludes the live-Herdr family the required Herdr lane owns. - `bin/fm-test-run.sh` gives a plain list of script paths the same bounded automatic scheduler and automatic bound that `--changed` gets. Naming several subjects is how a verification round asks for exactly those scripts. The curated selections are untouched: `--lane` still composes CI shards whose serial lane must stay serial, `--family` is what the required Herdr lane runs, and `--all` stays a deliberate complete regression. - `pr-forge` is admitted to the concurrent-safe family registry on two consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those, and records `secondmate` and `session-bootstrap` as refused with the exact script and reason each failed on, so the refusals are actionable rather than silent. Measured on this host, 0 failures on both sides: verification round, 4 scripts 448s chained -> 231s through the runner (-48%) pr-forge family 409.2s at 1 worker -> 237.9s at 4 (1.72x) watcher-wake-lock family 1311.1s at 1 worker -> 539.3s at 4 (2.43x) A fourth lever was implemented and then removed because the measurement refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made `fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s unchanged, back to back. Those sleeps are not overhead added to the clock - they are how a test waits for a subject moving on fm-watch.sh's own one-second cadence - so sampling less often only delays detection. It also broke `fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a settled condition. CONTRIBUTING.md records that result so the experiment is not repeated. * no-mistakes(review): Separate concurrent runs by isolation proof family * no-mistakes(review): Limit automatic timeouts to changed-file validation * no-mistakes(document): Clarify validation concurrency documentation * fix: copy PR URLs from durable records (#3648) * fix: copy PR URLs from records or abstain, never assemble them Supervision reported a plausible but dead PR link three times because its prompt demanded a full https:// URL at a moment when only a PR number was observable, so the model assembled an owner/repository from memory, and the PR check then accepted that URL and wrote it into the task record, after which the model kept defending its own tool-endorsed guess over the worker's real link. Three changes close that chain without any live forge lookup, so private forges are treated exactly like public ones: - bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy or abstain" section requires a URL to be copied verbatim from a durable record (the done: PR <url> status line, pr= metadata, or the backlog note), forbids assembling owner, repository, host, or number from memory, and has the branch report only the identifier it actually holds when no record names the URL yet, leaving the PR check unarmed until the worker's ready line arrives. AGENTS.md section 7 and 9 carry the same copy-or-abstain rule for main in place of the bare full-URL mandate. - Worker briefs (bin/fm-brief.sh, ship and scout rules) require the full https:// URL wherever a PR is mentioned - status line, terminal, or summary - never a bare "PR 108", so the link is in view as early as the number is. - bin/fm-pr-check.sh refuses, offline and before any side effect, a URL that the task's own done lines contradict, printing both spellings; a log naming no URL still records the argument as before. fm_pr_status_ready_urls in bin/fm-pr-lib.sh owns reading those lines. The refusal also reaches bin/fm-pr-merge.sh, so nothing merges under a contradicted URL. Tests cover the offline refusal with zero side effects, the recorded spelling being accepted, markdown-wrapped and punctuated URLs, working lines not counting, the merge wrapper propagation, a self-hosted merge request with no forge call, the prompt carrying the rule, and the brief carrying the worker rule. * no-mistakes(review): Remove stale PR URL enforcement * no-mistakes(ci): Removed backlog notes as an accepted PR identity source. PR URLs may now be copied only from the task’s `done: PR <url>` status or canonical `pr=` metadata; otherwise supervision reports only the known identifier and leaves PR checking unarmed. Updated related guidance/docs and verified with branch-supervision tests, brief tests, ShellCheck, and `git diff --check` * fix(bin): disable Claude feedback drafts for fleet launches (#3661) * fix(bin): disable Claude's feedback-draft flow for fleet-launched agents Scope --settings '{"feedbackDrafts":"off"}' to every Firstmate-launched Claude crewmate and secondmate, so /bug and /feedback never queue or submit a bug report on the captain's behalf. feedbackDrafts is the documented settings key (Claude Code changelog 2.1.247); the per-launch CLI flag never touches the captain's global settings.json. Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3 * no-mistakes(review): Prevent managed settings from re-enabling Claude feedback drafts * no-mistakes(document): Fix Claude feedback documentation formatting * fix(bin): layer both feedback-draft controls for defense in depth The prior --settings-only fix can be overridden by a managed Claude settings policy (feedbackDrafts precedence). Keep CLAUDE_CODE_SEND_FEEDBACK=0 alongside --settings '{"feedbackDrafts":"off"}': either control alone disables the SendFeedback tool, so a managed override of one still leaves the other in force. Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3 * no-mistakes(document): Document Claude feedback-draft suppression ownership * feat(tests): run three more validation families concurrently (#3662) * perf(tests): admit three more families to concurrent validation The three families that `docs/fm-test-isolation-proof.md` recorded as refused were not refused for concurrency. Each blocker was a test that decided a property by wall clock, or a script filed where it cannot run. Fixing those three things admits all three families and recovers 28.6 minutes of local validation with no assertion removed or weakened. - `tests/fm-backlog-handoff.test.sh` injected its pre-move crash by killing the handoff, sleeping a fixed second, then delegating the move to the real binary. Nothing ever killed the fake, so on a host slow enough for the case's next assertions to take longer than a second, the orphan woke and completed the very move the case requires left undone, and recovery then failed with `Task "pre-move-crash" not found in this backlog`. Watching the two backlogs during the injected crash showed exactly that, the item moving one second after the crash. All four crash injections in the file now go through a new `fm_fake_crash_injector` shim that signals the target and returns only once it is observably gone, and the pre-move fake never delegates the move at all. - `tests/fm-session-start.test.sh` proved the startup digest does not block on a slow current-state read by timing the whole digest against a fixed eight-second sleep, which a loaded host exceeds without the property being violated. It now holds that read open until the case releases it and asserts, the moment the digest returns, that the read has not finished. A digest that waited would wait indefinitely rather than for an interval a slow host can out-run, so the assertion is stronger than the bound it replaces. Its scan budget moves to the maximum, because the old value left two seconds of margin over the fixed sleep and measured the host rather than the deadline that `tests/fm-inactive-reconcile.test.sh` owns. - `fm-backend-herdr-focus-flash-e2e` was filed in the family map's catch-all, which put it in the portable serial lane, where Linux CI gate-skips it: that real-Herdr regression was running nowhere. It moves to `real-herdr-gated` and the required Herdr lane. `fm-claude-stop-autoarm-live-e2e` gate-skips on its opt-in variable and moves to `live-harness-optin`. The 28 remaining ungrouped scripts become an enumerated `standalone` family instead of admitting `unclassified` itself. `unclassified` is the family map's `*)` arm, so admitting it would silently grant concurrency to every test added afterwards, which is exactly the population with no proof. A new test still lands in `unclassified` and stays serial, and `tests/fm-test-run.test.sh` covers that split behaviorally. Each family passes two consecutive four-worker proofs with zero failures. On the production runner, `secondmate` goes 1233.1s to 453.4s, `session-bootstrap` 756.4s to 286.4s, and `standalone` 724.6s to 261.1s: 2.71x overall and 1713.2s recovered. The whole suite runs 177 scripts in 52.6 minutes of wall clock against 121 minutes of summed script time. * no-mistakes(document): Refresh concurrent validation and shard documentation * no-mistakes(ci): Fixed the real-Herdr focus-flash E2E race exposed by reclassification. Part C now starts its persistent child atomically via `pane run` and verifies stable child identity through Herdr’s public `process-info` interface, avoiding the racy send-text/send-keys sequence and platform-specific `ps` matching. Verified with bash syntax checking, ShellCheck, git diff checks, and the complete E2E test on Herdr 0.8.2 * feat: structure no-mistakes ask-user escalations (#3670) * feat(brief): structure no-mistakes ask-user escalation as event + snapshot file Crewmates escalating a no-mistakes ask-user gate now report one status event naming every finding id plus a snapshot file holding the gate's axi finding records verbatim (id, severity, file, line, description, authority), using the same shape even for a single finding. The status line never paraphrases. The format is defined once in fm-dod-lib.sh and rendered into both the scout and ship rule 6 in fm-brief.sh, so a promoted scout - whose rule 6 fm-promote.sh preserves unchanged - gets the identical contract as a freshly-spawned no-mistakes ship worker. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PpiWaDerbYavTLPPtEjQei * no-mistakes(review): Preserve ask-user escalation output contract * no-mistakes(review): Align escalation format test expectation * no-mistakes(review): Scope ask-user escalation instructions correctly * no-mistakes(review): Remove ask-user from generic decision rules --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> * fix(bin): require self-sufficient no-mistakes intent (#3671) * fix(bin): require a self-sufficient no-mistakes intent A no-mistakes worker's --intent is only as useful as the string it passes. PR #3604 shipped with an intent that was only "do 1, 2, 3, 7 from the report": the real contract lived in a private scout report and never reached --intent, so nobody holding that string plus the codebase could have derived the specification. This is pure instruction at the contract's one owner; no spawn-side or promotion-side check is added. - bin/fm-dod-lib.sh: the generated no-mistakes Definition of done now states that the --intent string must be self-sufficient (the string plus the codebase reconstructs roughly the same specification) and tells the worker to write the substance of any report, decision, or PR the captain's intent refers to into --intent rather than the pointer, while Firstmate build instructions and the worker's own decisions still stay out. The spawn-time overlay points back at that rule so its "supersedes" wording cannot cancel it, and the header's owner statement carries the rule. - AGENTS.md section 11 and bin/fm-brief.sh's header ask Firstmate to include the substance of referenced material when filling ## Captain's intent, and section 11 points at the owner of the rule. - tests/fm-brief.test.sh and tests/fm-task-delivery.test.sh assert the rendered brief and launch contract carry the rule. Claude-Session: https://claude.ai/code/session_01YMhEe42q7BAAoN6RxNuzim * no-mistakes(document): Replace incident-specific intent test commentary * fix: accelerate local Bearings snapshot composition (#3499) * Speed local fleet snapshot composition * no-mistakes(review): Stabilize task inventory during concurrent snapshot composition * no-mistakes(document): Document local snapshot observation concurrency * no-mistakes(ci): Fixed CI failures by making empty task manifests compatible with stock macOS Bash 3.2, snapshotting task metadata before concurrent observations to prevent generation drift, strengthening the behavioral race regression, and updating the stock-Bash Bearings test count to 45. Verified fleet snapshot tests (15), Bearings tests (45), workflow lint tests, project lint, Bash 3.2 parsing, and diff checks * no-mistakes(ci): Fixed the Linux CI failure caused by passing large backlog/task JSON through jq command-line arguments, which exceeded the per-argument size limit. Both inventory projections now stream large JSON inputs through stdin. Verified with fm-bearings-snapshot.test.sh (45 tests), fm-fleet-snapshot-view.test.sh (15 tests), Bash syntax, and git diff checks * no-mistakes(ci): Fixed concurrent task teardown during metadata capture: vanished metadata is now omitted while genuine copy failures remain fatal. Added a deterministic public Bearings regression test and updated CI’s expected test count. Verified with the full Bearings suite, workflow-lint suite, Bash syntax checks, and git diff checks * no-mistakes(ci): Fixed PR-caused CI and review issues: streamed large fleet JSON through jq stdin to avoid Linux argument limits, kept crew-state reads bound to captured metadata generations, and strengthened the behavioral race test. Bearings (46 tests), fleet snapshot (15 tests), crew-state, backend, lint, Bash syntax, and diff checks pass locally. Serial shard 5’s unrelated task-inbox segmentation fault appears infrastructural/flaky * no-mistakes(ci): Fixed endpoint-state generation crossing by validating captured spawn_gen before and after local endpoint probes, falling back to exact metadata identity for legacy tasks. Stale probe results now become unknown instead of false unhealthy state. Added a behavioral relaunch-race regression test. Verified the full Bearings snapshot suite, shellcheck, bash syntax, and git diff checks * fix(snapshot): keep live observations generation-coherent * no-mistakes(review): Keep secondmate observations generation-bound without copying reports * no-mistakes(document): Document generation-coherent snapshot observations * test(bearings): measure local read overlap instead of wall-clock budget The large-local-snapshot regression asserted that a whole snapshot composed in under five seconds. That bound measures how loaded the host is, not whether the per-task reads actually overlap, so it failed intermittently on a contended machine: one run in six on a box at load 16-20, landing exactly on the five second boundary. Time a serialized run and a concurrent run of the same workload instead and require the concurrent one to save at least two seconds. Both runs pay the same composition overhead, so the difference isolates the overlap this change delivers. Five one-second reads serialize into five seconds and overlap into about one, and re-serializing the reads collapses the saving to roughly zero, so the assertion still fails loudly if the concurrency regresses. Also bump the pinned Bearings test count to 48, since rebasing onto the current default branch picked up its captain-hold test. * no-mistakes(review): Restore JSON-derived decision flags * no-mistakes(review): Unify status-derived snapshot observations * no-mistakes(ci): Updated the stock macOS Bash CI check’s Bearings test count from 48 to 49. Verified the full Bearings suite passes and emits exactly 49 TAP successes; git diff checks pass * fix: prevent stale supervision wake loops (#3672) * fix(bin): stop the supervision branch's stale-ack and ghost-report loops Clean-slate implementation of the four authorized recommendations from the supervision-ghost-retrigger analysis (items 1, 2, 3, and 7), in their minimal form, superseding PR #3604: - fm_branch_report refuses a task the wake being handled never named. The extension fixes the reportable task set from the eligible rows before each prompt (signal and stale rows resolve to their tasks, a heartbeat allows any task with a live record, fleet is always allowed), so a report typed from memory about a task whose records teardown already removed is never stored or delivered. - An acknowledgement that consumes nothing says "nothing was acknowledged through N" and prints the exact --ack-through / --recovery-generation command for the current presented wake, instead of "re-run the drain", which re-fed the same stale acknowledgement in a loop. - bin/fm-guard.sh no longer tells the branch actor to drain queued wakes while it is handling them; it names the granted rows instead. - Teardown removes state/.<task>.branch-outcome-index for ordinary tasks and descendants; the index rebuild and the append-side index write both skip a task with neither a live record nor a status log, so the branch's report of a teardown it just performed is stored without recreating the index. No new locking, no spawn-generation binding, and no retired-task refusal: the branch can still report the outcome of a task it just tore down, and the teardown test now proves that path end to end. * fix(bin): narrow the branch report scope and guard silence to the minimal form Apply the four review decisions on the clean-slate branch: - A signal or stale prompt may report only the tasks its own rows resolve to; fleet is refused there too. A heartbeat review is not scoped by task at all, so the extension no longer tracks live task records and refuses nothing by task id during a fleet review. - The outcome-index rebuild no longer skips retired tasks; the append-side skip alone keeps a torn-down task's index from being recreated. - bin/fm-guard.sh keeps the queued-wakes warning silent for the branch actor instead of printing a replacement note. * no-mistakes(document): Align supervision docs with scoped wake handling * fix(bin): avoid fleet snapshot argument limits (#3677) * Fix fleet snapshot large JSON transport * no-mistakes(review): Captain: file-back fleet snapshot transport safely * no-mistakes(review): Captain: file-back parent summary aggregation * no-mistakes(ci): Rebased the PR's three commits onto f4d7875824ecc5e274b4bb896f10c1e1f207b7e4 and resolved the fleet snapshot conflict while preserving the base's task-observation lifecycle. Fixed Greptile's valid finding by recursively removing the private mktemp transport directory, so future transport files cannot cause cleanup to fail. Verified with tests/fm-home-summary-refresh.test.sh, bin/fm-lint.sh, git diff --check, and ancestry checks. All passed; the fix remains as an uncommitted worktree change for the outer executor * fix(bin): attribute active runs with unfetched pipeline heads (#3681) * fix(bin): rec…
sctru
added a commit
to sctru/firstmate
that referenced
this pull request
Sep 29, 2026
* feat(bin): add IMAP/SMTP mail plane with standing poll (#3765)
Opt-in IMAP/SMTP mail plane (fm-mail.sh / fm-mail-check.sh). Absent FM_MAIL_* stays off.
Speaking as Kun's firstmate: this is merged. Thank you @feilipu — really appreciate you taking the time on this.
* fix: launch remote Herdr through the user login shell (#4061)
* fix(remote): start the fm-remote Herdr agent through a login shell
Launchd was exec-ing herdr directly, so the Aqua agent inherited a background session without login-keychain access. Start it via /bin/zsh -lc exec so panes keep login env and can refresh OAuth tokens after reboot.
* fix(remote): start fm-remote Herdr via the account login shell
Resolve UserShell from Directory Services and invoke it with separate -l and -c so bash, fish, and zsh all get login-keychain access. Fall back to SHELL, then /bin/zsh, then /bin/sh without failing the render.
* no-mistakes(review): Fix launch-agent shell fallback resolution
* no-mistakes(review): Preserve and escape Directory Services shell paths
* no-mistakes(document): Document login-shell LaunchAgent behavior
* no-mistakes(ci): Updated the trusted fm-remote-doctor SHA-256 identity in bin/fm-remote-entrypoint.sh. Verified with tests/fm-on.test.sh, tests/fm-remote-doctor.test.sh, bash syntax checks, and git diff --check
* no-mistakes(ci): Resolved the login shell exactly once per doctor invocation and threaded it through plist rendering, installed/loaded contract validation, repair reporting, and post-repair checks. Added a regression test proving repeated repair remains healthy and performs no reload when a hypothetical second Directory Services lookup would differ. Updated the trusted doctor hash. Verified doctor, fm-on, remote-entrypoint, lint tests, ShellCheck, syntax, and diff checks
* no-mistakes(ci): Made Darwin shell resolution hermetic with executable injection and a 2-second Directory Services timeout. Updated tests to inject shells by default, isolate dscl-specific cases, parse plists semantically, and verify stalled dscl fallback. Updated the trusted doctor hash. Doctor, fm-on, entrypoint, syntax, hash, and diff checks pass
* no-mistakes(ci): Raised portable serial CI timeout from 20 to 30 minutes, refreshed the specified timing hints, added missing hints, and recomputed shard documentation. Verified coverage, runner behavior tests, workflow lint tests, shell syntax, requested timing maxima, and diff checks
* fix: distinguish landed deliveries from resolved captain calls (#3710)
* fix(bearings): keep captain-approved deliveries in Recently Landed
A closed task is never held: tasks-axi clears the held flag when a task
closes and keeps hold-kind and the hold reason as the record of the call
that was made. Recently Landed excluded every Done row whose hold-kind was
captain, so the marker it treated as "closed while still waiting on the
captain" was in fact the proof that the captain had approved the work. Every
merge routed through a captain decision disappeared from the list of what
shipped, including under --all-landed.
The selector now asks whether the closed row delivered something. Recently
Landed is merged PRs, completed scouts, and finished local-only merges, so a
row carrying one of those artifacts belongs there whoever approved it. A
captain question closes with an answer and no artifact of its own, and that
is what still stays out, so an answered question is never rendered as
shipped work.
The same rule was written twice - the bearings projection selects this
home's Done rows and the fleet snapshot selects each secondmate home's Done
rows into the roll-up the same section merges in - which is why one defect
hid deliveries in every home. Both now share bin/fm-landed-lib.sh.
* fix(review): Normalize landed evidence and exclude answered captain questions
* fix(review): Normalize captain delivery evidence across relocated data
* fix(review): Record authoritative delivery provenance with legacy fallback
* fix(review): Harden delivery provenance across forced and pruned completions
* fix(review): Replace premature merge closure with existing release contract
* fix(review): Document provenance-based Recently Landed selection
* fix(document): Align documentation with completion provenance
* fix(lint): Fix targeted ShellCheck warnings
* fix(ci): order the pinned tasks-axi install before its stock-Bash consumers
In `.github/workflows/ci.yml` the pinned tasks-axi install now precedes both
stock-Bash consumers, and the Bearings expectation is updated from 49 to 50
tests.
Verified with macOS Bash 3.2: snapshot 16/16, Bearings 50/50, public-followup
1/1. Full repository lint and all three workflow validations pass, and
`git diff --check` is clean.
* fix(review): Make completion provenance unambiguous
* fix(review): Make completion verdict authoritative over quoted provenance
* fix(review): Preserve retained artifacts through resumed captain closes
* fix(review): Unify completion provenance ordering across writer and reader
* fix(review): Preserve artifacts across failed captain closes
* fix(review): Refresh v1 assertions; provenance authority remains unresolved
* fix(review): Remove unreliable provenance while preserving landed deliveries
* fix(review): Reject stale home summaries visibly
* fix(review): Restore retained deliverable recording
* fix(review): Match landed artifacts and restore retention documentation
* fix(review): Disambiguate captain calls and restore landed artifact matching
* fix(review): Persist retained report and PR artifacts
* fix(review): Preserve staged artifacts before captain answers
* fix(review): Avoid wedging answers on unsupported report paths
* fix(review): Exclude unreleased captain-held pull requests
* fix(review): Exclude held local-only answers from landed
* fix(review): Preserve retained scout reports across snapshot rendering
* fix(document): Align landed lifecycle documentation with release semantics
* fix(review): Enforce landed artifact-kind ownership
* fix(review): Infer canonical task kinds in snapshots
* fix(review): Require captain-hold release before merges
* fix(review): Qualify merge lifecycle regression evidence
* fix(review): Serialize captain holds with merge operations
* fix(review): Document merge cleanup residuals honestly
* fix(test): Replace vacuous Bearings regression with behavioral cases
* fix(document): Align Bearings verification and merge lifecycle documentation
* fix(review): Serialize merges and exclude captain calls from landed
* fix(review): Harden merge identity and landed selection
* fix(document): Clarify landed selector compatibility filtering
* fix(bin): keep merge entrypoints usable on records without an incarnation
The merge identity guard refused any task record with no spawn_gen field.
That field identifies one exact incarnation, so comparing it across the wait
for the merge lock is what catches a task relaunched while the merge was
queued. Requiring it to be present is a different rule, and it refused every
record written before the field existed: a legacy task could no longer be
merged at all, and five behaviour suites refused before reaching the check
they were written to exercise.
The comparison only needs to notice a change. An absent field is now read as
an empty incarnation and compared like any other value, so a record that
gains, loses, or alters one is still refused, while a record that simply
predates the field merges. An ambiguous or unreadable field stays an error,
because a record that cannot name one incarnation cannot be compared. The
missing-record message each entrypoint had before the guard is restored, so
a genuinely absent record still says so in its own words.
The role partition now precedes reading the record. Refusing the supervision
branch is a statement about the actor, not about the task, so it cannot
depend on a record the wrong actor may not have.
A backlog file that does not exist meant "no longer an open captain call".
For a caller that asked to tell absence apart it now means absent, so a board
card whose home carries no backlog stays visible instead of being dropped as
resolved.
Fixture repositories pin their initial branch instead of inheriting
init.defaultBranch, which resolved to main on a developer machine and master
on a runner, so a fixture naming main failed only in CI.
* fix(review): read local-only note from body; surface pending-close failures
* fix(review): keep kindless local-only landings in Recently Landed
* fix(review): bind local-only note scan to the tasks-axi note line
* fix(review): Guard unavailable captain-hold authority records
* fix(document): Document unreadable authority predicate outcome
* fix(bin): read an absent backlog as absence, not an unreadable record
The merge gate refused every task whose home carries no backlog file. A
backlog that does not exist holds no captain call, so nothing can be held and
the merge is safe; only a backlog that exists and cannot be read may hide a
live hold. Those two states were collapsed into one refusal, which stopped
merges in any home that keeps no backlog.
The predicate now reports a missing backlog file as absence, alongside a row
the backlog does not carry. A record that exists but cannot be read still
leaves by the existing cannot-tell path, which both merge entrypoints already
refuse, so the restrictive direction is unchanged.
That leaves no way to reach the separate unavailable-record result, so the
result and the two branches that handled it are removed rather than left
describing an outcome that can no longer occur. The lifecycle documentation
loses the same claim.
Regressions cover both directions in each entrypoint: a home with a task
record and no backlog merges, and a backlog present but unreadable refuses
without reaching the forge.
* fix(review): Fail closed unreadable backend configuration
* fix(tests): pin the bare origin's initial branch in the remote seed fixture
The fixture created its bare origin with no initial branch, so that
repository's HEAD followed init.defaultBranch while the source repository
pushed the branch fm_git_init_commit pins. On a host that still defaults to
master the two disagreed: the bare origin's HEAD named a branch the push never
created, cloning it warned that the remote HEAD referred to a nonexistent ref
and checked out nothing, and the seed assertion for the cloned README failed.
A machine whose default is already main paired the two by accident and hid it,
which is why the fixture passed locally and failed on the runner.
Pinning the bare origin to the same branch removes the dependency on the
ambient default from both sides. Verified under both conditions: with
init.defaultBranch set to master, and set to main, the suite passes 26 of 26.
* fix(review): Fail closed unreadable user backend configuration
* fix(bin): republish the home summary as v1 and record two load-bearing rules
The published home-summary schema had moved to v3, which routed every
secondmate home still emitting the earlier version to the stale branch: their
landed rows, open decisions and holds all came back empty and their state read
as unknown until each home was updated. The payload never justified that. Its
field set, field order, truncations and the landed array construction are
byte-identical to v1, so only which rows the selector places in landed
differs, and a v1 consumer reads that the same way.
Republishing as v1 removes the rollout regression and, with it, the tolerance
machinery that existed only to soften the bump: the stale-schema predicate,
its two collection branches, the flag and its provenance branch, the omitted
surface that can no longer be reached, and the fixtures and assertions that
covered them.
Two rules that a scope review proposed removing are kept, each now carrying
the reason it exists, because both were measured to be load-bearing:
The artifact-kind ownership clause is what keeps an explicit scout that
recorded no report out of Recently Landed. Without it such a row has none of
the three artifacts, satisfies the compatibility fallback and renders as
shipped work with an empty artifact.
The kind fallback is needed because tasks-axi omits the kind metadata
entirely when a title begins with a canonical keyword. Without it a scout
titled "SCOUT ..." reports no kind, its recorded report stops counting as a
delivery, and it drops out of the section this selector exists to repair.
* fix(bin): move the scout guard note onto the rule and drop two dead pieces
The LOAD-BEARING note sat on an unreachable branch. Measured in both
directions: removing that branch together with the kind-is-not-scout guards
lets an explicit reportless scout into Recently Landed and fails
tests/fm-captain-hold-lifecycle.test.sh, while removing the branch alone
leaves that suite passing at 49 assertions. The guards carry the rule, so the
note now sits on them and the unreachable branch is gone. A note pointing a
later reader at the wrong line is the hazard this change corrects elsewhere.
summary_file_has_schema lost its only caller when the stale-schema machinery
was removed, so it goes with it.
* fix(review): Fix legacy report artifacts and canonical keyword boundaries
* fix(review): Update pinned Bearings test count to 56
* fix(document): Clarify landed summary compatibility documentation
* fix(review): Preserve unreadable backend configuration errors
* fix(review): Honor backend resolution errors at existing call sites
* fix(test): Stabilize remote collector tests under host load
* fix(document): Document backend resolution failure contracts
* fix(lint): Suppress intentional deferred probe expansion warnings
* fix(ci): Captain, quoted the two literal test IDs in tests/fm-backlog-atomicity.test.sh to fix SC2100 without changing behavior. Both warnings reproduced before the fix; the targeted fm-lint.sh run now passes with ShellCheck 0.11.0. Bash syntax and git diff --check also pass
* fix(ci): Fixed the resolver’s two configuration-parent checks to return 2 for inaccessible directories while preserving genuine absence. Added two behavioral tests; RED/GREEN and both requested mutation proofs confirmed. All 10 focused checks, targeted lint, syntax, and whitespace checks passed. Broader merge suite stopped after 10 passing cases under host load. Declined portable checks and merge-authority code remain unchanged
* fix(remote): keep fm-remote Herdr servers in the Aqua session (#4090)
* fix(remote): let the Aqua launch agent own the fm-remote Herdr session
A herdr server keeps the macOS audit session of whatever started it, and
only the Aqua login session (gui/<uid>) can read the login keychain
without a prompt. Herdr's SSH remote attach starts the fm-remote server
as its own child when it finds none, wins the socket at boot because sshd
accepts connections before the login session exists, and every claude
pane under that server then gets `security` exit 36, falls back to a stale
plaintext credentials file, and reports "Login expired". launchd's own job
lost the socket on every KeepAlive retry and the doctor still reported the
session ready because it only asked whether any server answered.
- Add bin/fm-remote-herdr-guard.sh, the launch agent's exec target: start
the server in the foreground when nothing owns the socket, exit 0 when an
Aqua-born server does, and otherwise stop the foreign server, wait for the
socket, and exec the server at once.
- Add bin/fm-remote-herdr-owner-lib.sh, the single owner of socket-owner
discovery (lsof; pgrep cannot see herdr's argv on macOS) and the birth
markers (SSH_*, XPC_SERVICE_NAME, FM_REMOTE_JOB_ACTIVE, sshd or
remote-client-bridge ancestry matched on argv[0] and whole arguments).
- Render the agent as the login shell exec'ing the guard with
KeepAlive={SuccessfulExit=false} and ThrottleInterval=10, check the loaded
job's successful-exit semaphore, and report a session served outside the
Aqua login session as fixable so --fix retakes it through launchd; the
reload waits for an Aqua-born owner rather than any running server.
- Correct the doctor and docs: the launch shell provides environment parity,
the launchd domain provides keychain access.
- Pin the guard's decision table and the doctor's verdicts against real
marker-carrying processes, and record the dated audit-session evidence.
* no-mistakes(review): Verify Aqua ownership through launchd domains
* no-mistakes(document): Document macOS lsof ownership requirement
* fix(bin): prefer a live no-mistakes run over a terminal one (#2881)
* fix(bin): prefer a live no-mistakes run over a terminal one
A worktree can bind to more than one recorded no-mistakes run at once.
The branch-and-code-identity rule in bin/fm-nm-run-lib.sh accepts both an
exact-equal commit and a worktree-is-an-ancestor match, but never stated
which wins when both bind, so the tie fell to whichever candidate the
caller reached first.
Observed on a live fleet: a crashed validation daemon left a FAILED run
at the worktree's own commit while the live run that replaced it
validated a descendant commit on the same branch. Bare `axi status`
answers with the most-recently-touched run - the corpse - and it bound by
the equal-commit rule, so every recomputation reported `failed` for a
task whose real run was healthy. The same label had also read `failed`
earlier while the work was genuinely stalled, so the signal was wrong in
both directions.
State the live-over-terminal policy in the matching rule's own contract,
where the equal-commit and ancestor rules already live, and add
fm_nm_run_status_class as the one classifier that decides liveness from a
recorded status word. fm-crew-state.sh applies it on both selection
paths: the runs listing now scans past a terminal row for a live one, and
a terminal `axi status` answer is provisional until the listing has been
asked whether this worktree also has a live run.
Same-liveness-class candidates keep the listing's newest-first
precedence, and a status word the classifier cannot place keeps the
caller's own ordering rather than displacing a known result, so a
single-run task and a task whose runs are all terminal are unchanged.
Regression coverage reproduces the proven case (terminal run at the
worktree's exact commit plus a live run descending from it) and its
runs-list twin; both fail under the old tie-break. Two companion cases
pin the no-widening half - two terminal rows still resolve newest-first,
and a terminal run with no live sibling keeps its full run-step detail -
and both pass before and after the change.
* no-mistakes(review): accept unfetched live sibling anchored at exact worktree head
* docs(bin): name both ledger reads behind the runs-limit setting
The FM_CREW_STATE_RUNS_LIMIT comment in bin/fm-crew-state.sh still described
the runs ledger as scanned only by the cross-branch fallback, but the
live-over-terminal fix also consults it as the live-sibling probe behind a
terminal axi status answer. Point the comment at docs/configuration.md as the
setting's owner instead of restating a second copy.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* fix(bin): repair process-event shutdown and tighten guard timing (#4009)
* fix(procevent): make the ordinary stop signal actually stop a runner
The owner guard that shipped in #3904 half-reaps. Against a poll child that
handles the ordinary stop signal and keeps waiting, the guard signals the group,
loses the runner leader to its own signal, then reads that success as a
leaderless group and exits without escalating. It destroys the only proof of
ownership that would have authorised the forced signal, so the survivor becomes
unreachable by retire, reconcile, sweep-home and the guard alike. A guard that
turns a leaking-but-identifiable generation into a permanently unreachable one
is worse than no guard at all.
Two defects, and they hid each other:
- The escalation re-derived ownership from the leader. `runner_group_signal`
now takes a `proved` mode, passed only by the escalation inside the stop that
already proved and signalled that exact generation moments earlier. A leader
dying to our own signal is the ordinary outcome, not fresh ambiguity.
- Every stop held the per-source lock across its wait while the runner's own
exit cleanup waited unboundedly for that same lock. That circular wait was
broken only by the forced signal, so the forced signal silently became the
normal path - and, by keeping the leader alive through the whole window, it
masked the escalation defect above. The runner's exit cleanup now refuses that
lock instead of waiting for it, which is what its existing `return 0` already
said it did.
Fixing the lock alone would have turned every stop of a signal-proof child into
a refusal that leaves it running, so both land together and the tests pin that.
Measured on macOS with a stand-in poll child that traps TERM, INT and HUP:
the guard left it running past 70s and now clears the group within the lease
plus one check; retiring a healthy runner fell from ~2.8s with a forced group
signal every time to ~0.6s on the ordinary signal alone.
Unchanged and stated deliberately: a leader lost to anything other than the
stop's own signal still leaves a group that retire, reconcile, sweep-home and
the guard all refuse, permanently - and that source stops listening without
saying so. Whether such a group may ever be signalled is an open decision and
is not answered here.
* fix(review): Fix proved escalation race and stop regression assertions
* fix(review): Preserve proved escalation through transient identity failures
* fix(review): Simplify proved escalation and correct guard timing documentation
* fix(document): Clarify process-event stop ownership and cleanup limits
* fix(document): Clarify process-event stop ownership and fixture comments
* revert(skills): restore the leaderless-ambiguity limit to the loaded skill
An automatic documentation step in this branch's validation edited
.agents/skills/process-event-sources/SKILL.md, which no instruction in this
change asked it to touch. That file is not documentation about the code: it is
the agent-loaded instruction surface, what an agent reads to know what it is
permitted to do.
The step deleted this line:
- leaderless PID/PGID-reuse ambiguity preserves the claim without signalling
or replacement, as owned by the operating contract in
[`docs/configuration.md`](../../../docs/configuration.md#process-to-event-sources-stateprocevent);
and folded it, with its neighbour, into a generic "registration and ownership
transitions, stop authority, and claim reclamation follow the operating
contract".
That deleted line states a PROHIBITION - that such a group is preserved WITHOUT
SIGNALLING - and it is the exact limit an open captain decision currently rests
on. Folded into a pointer, an agent reading the skill to learn what it may do
would have to chase a second document to discover it may not signal. A
prohibition that requires a second lookup is not a prohibition. The effect was
to weaken, in the instructions themselves, the boundary that keeps one home from
signalling another's process group - while the question of whether that boundary
should move at all is still open.
This is a deliberate revert, not an oversight, and it restores the file exactly
to its pre-branch state. The full statement also survives in
docs/configuration.md; that does not rescue it, because the agent handling a
process-event wake loads the skill and not the documentation.
* revert(procevent): restore the open-question marking beside the escalation
The same automatic documentation step that edited the loaded skill also removed
this from the comment above runner_group_signal:
A leaderless group nobody in this call ever proved remains refused too, for
every caller. That untouched refusal is what makes a crashed leader's group
permanent, and relaxing it is a separate open question, not something this
path assumes.
and replaced it with a pointer to docs/configuration.md.
This one fails differently from the skill deletion, which is why it is restored
separately. There, a prohibition was moved out of the reader's path, and a
missing prohibition gets violated. Here the prohibition survives in code - the
unproved path still refuses - and what was removed is the fact that the limit is
UNDECIDED. A prohibition that has quietly lost its "this is still open" reads as
settled design, and settled design gets relied on, extended, and eventually
relaxed by someone confident they understand why it is there. That question is
open right now.
The rule this branch's four instances produce, stated once here because this is
the point of decision: an unresolved question must be marked unresolved AT THE
POINT OF DECISION, not only where the contract is documented. A reader who does
not know something is open will treat it as closed, and that default is stronger
than any pointer overcomes.
The pointer added by that step is kept alongside; this restores what it replaced
rather than reverting it.
* docs(verification): restore the measured guard bound and its reason
The document step's rewrite of this record dropped the concrete figure while
keeping the surrounding measurements. What went missing was the bound itself -
lease plus two consecutive failed checks plus the stop's grace, roughly 630
seconds at the shipped 600-second lease and 15-second check - together with the
reason there are two checks rather than one: a single unreadable read must not
be enough to kill a live runner.
The mechanism survived elsewhere and the reason survived in
docs/configuration.md, so nothing was lost from the repository. The concreteness
was, and that is what this restores. A number recorded without why it is that
number is the one a later reader shortens; the reason is the whole safety
argument for the debounce, and the debounce is what stops the reaper killing a
live runner on one bad read.
* fix(document): Replace stale stop-authority summaries with owner pointers
* test(procevent): make the guard-bound case able to fail for its own reason
An automated reviewer observed that this case allowed sixty seconds for a bound
of roughly eight, so it could not go red for the reason it names: it would have
passed a guard that took fifty-five seconds. That is correct, and it is the same
family as the defect the case exists to defend against - a check that is green
because it cannot fail, rather than because the thing it guards is working.
The deadline is now derived from the bound itself - the lease, plus the two
consecutive failed checks the guard debounces on, plus the stop's own ordinary
and forced signal windows - rather than from a flat wall-clock number, and the
shortened lease and check the fixtures run under have a single definition so a
derived deadline cannot silently diverge from the settings the guard is given.
The doubling that remains is a load allowance and is documented as one; widening
it to make a slow guard pass would convert the assertion back into decoration.
Proven by mutation rather than by argument. Against the repaired case:
correct code ok
guard debounces on 20 misses instead of 2 not ok - "still holding
the group after 16s,
against a documented
bound of 8s"
proved escalation removed (the original defect) not ok - same
code restored ok
The previous sixty-second version passes every one of those mutations.
The reviewer's other claim, that the guard can survive past the announced bound
when an owner disappears immediately after a check, was measured and does not
hold against what this branch announces. Sweeping the phase deliberately at
0.0, 0.2, 0.4, 0.6 and 0.8 of a check interval gave 7.21s, 7.31s, 6.75s, 6.49s
and 6.31s, worst 7.31s, against the announced lease plus two consecutive failed
checks plus stop grace, which is up to 8s at those settings. The mechanism the
reviewer describes is real and is the announced mechanism; the bound it was
measured against is a phrasing this branch no longer carries.
* revert(scope): return the instruction surfaces to their base state
This delivery is being split. It carries the two proven process fixes alone; the
instruction text travels separately, through a run that removes the
documentation step rather than refusing it at its gate.
Two surfaces are therefore returned to exactly what the base branch has, so this
delivery neither adds to them nor removes from them:
.agents/skills/process-event-sources/SKILL.md - identical to base again. Three
bullets an automatic documentation step had folded into a pointer, including
that leaderless PID/PGID-reuse ambiguity preserves the claim WITHOUT
SIGNALLING and that there is one identity-matched owner per canonical source
across homes sharing one store.
The header comment block of bin/fm-procevent.sh, which is what the script
prints as its own help. Seven lines were removed from it: that a live owner is
never displaced, that only a claim whose stale owner and independently absent
process group prove its whole generation gone is reclaimed, that a crashed
leader or reused pid whose process group still has members cannot relax
ownership cleanup, and that reconcile signals only a live identity-matched
runner group and otherwise keeps the claim without starting a replacement.
The help output is now byte-identical to base.
Neither removal was requested by any instruction in this change, and both were
made to text that predates it. Returning them is scoping, not a third
restoration: nothing is being added to those files here.
* fix(ci): Captain, live CI revealed a fixture deadlock: it suspended the runner before startup released its lock. Added a public-list synchronization barrier in tests/fm-procevent.test.sh. Forced-delay reproduction detected the deadlock before the fix; all four cases passed afterward. Targeted lint, Bash syntax, and whitespace checks passed. Greptile’s watchdog requirement conflicts with the recorded R2 decision; runtime behavior and documentation remain unchanged. Full CI rerun belongs to the outer executor
* test(procevent): make the post-TERM cases report what they saw when they fail
On the failure path only, these cases now print what they actually saw: the
identity recorded at claim time, the identity readable at that moment, the size
of the signals file, the leader's state and wchan, every live member of the
runner's process group with its own state and wchan, the elapsed time since the
stop began, and what retire said. None of it runs when a case passes.
WHY THIS IS KEPT, stated accurately rather than by its original reason. It was
written to make an unexplained CI failure verifiable. That failure is now
explained - it was a fixture deadlock, diagnosed and repaired in the preceding
commit - so that justification has expired and is not the reason given here.
The reason it stays is smaller and independent of that failure: it is already
written, it is small, it sits in the file whose assertion this change reworked,
and an assertion that could not say why it failed cost most of a morning to
diagnose from the outside. The next failure will not be this one.
WHAT A PASSING RUN WOULD NOT MEAN: a pass is a sample of behaviour already
observed many times, not proof that anything is fixed. Only a failure carrying
the evidence above establishes a cause.
* fix(document): Clarify process-event fixture diagnostic rationale
* fix(ci): Captain, fixed two cleanup races in tests/fm-procevent.test.sh: removed premature child completion and waited for runner exit before retiring the restart fixture. Controlled Linux reproductions demonstrated failure before and success after. The full Linux process-event suite, six focused macOS checks, targeted ShellCheck, Bash syntax, and whitespace checks passed. Runtime behavior, guard debounce, and documentation remain unchanged. CI rerun belongs to the outer executor
* fix(procevent): bound owner-guard cleanup at one check interval, not two
A THIRD WAY, not a capitulation to the reviewer and not a refusal of it.
The automated reviewer's grievance was the LOOSENESS OF THE BOUND, not the
number of observations the guard makes before it acts. It asked for a single
read because that was the only route it could see to an acceptable bound. There
was another route, and this change takes it: the bound is reached and both reads
are kept.
TIGHTENED - the SPACING of the guard's two reads, not their number. The owner
watchdog now sleeps half the configured check interval and still requires two
consecutive failing reads, so the pair completes inside one check interval
instead of costing two. Worst-case detection falls from the lease term plus TWO
check intervals to the lease term plus ONE. At the shipped 600s lease and 15s
interval the stated bound falls from ~635s to ~620s.
PRESERVED - the second read. bin/fm-procevent.sh's two-consecutive-miss rule is
untouched. WHY IT PROTECTS: the guard's inputs are a lease read and a state-root
identity read, and either can fail transiently on a live, healthy home. Acting
on the first failure would let one isolated unreadable read kill a live service.
Requiring a second, independent read is what makes that impossible, and it is a
protection rather than padding. Nothing was traded away to reach the bound.
Both properties are now guarded by their own case, and each was proven by
MUTATION rather than asserted:
- putting a full interval back between the two reads fails the bound case:
"still running 17.0s after the last owner activity, against a documented
bound of 15s";
- acting on one failed read fails the new debounce case: "one unreadable lease
read ended a runner whose home was still alive" - while the bound case then
passes FASTER, 9.9s against 13.1s. The unsafe variant being the quicker one
is exactly why these are two cases: one elapsed-time case would have
registered the removal of the protection as an improvement.
MEASURED, sampling the phase between the guard's check clock and the lease clock
across eight runs per variant, on macOS (Darwin 25.5.0). Reaping an orphaned
listener whose home stopped refreshing its lease:
lease 2s / interval 1s: 4.41-5.29s before, 3.48-4.65s after
lease 2s / interval 4s: 7.69-8.12s before, 5.94-6.13s after
The 4s configuration is the informative one: the gap is about one check
interval, which is precisely the term that was removed.
A previously unstated term of the bound surfaced while measuring: the lease age
is compared in whole seconds, so a configured lease of N is honoured until that
age reads N+1. It is now part of the documented bound and of the regression's
derivation instead of being absorbed into a fudge factor.
The bound regression derives its deadline from the documented bound instead of a
flat number, and PINS the phase between the guard's check clock and the lease
clock rather than sampling it, because with a sampled phase a guard spending two
intervals passes about half the time on a lucky alignment. Its load slack is
additive and stays under half a check interval, so an extra whole interval
cannot hide inside it. The two flat deadlines that were there before (40s and
20s) and the doubling allowance on the derived one are gone; that looseness was
the reviewer's third complaint.
The stop's own grace is untouched: 2s for the ordinary signal, then 2s for the
forced one. It is a ceiling paid only by a group that outlives the signal it was
sent, not a delay every stop pays - a healthy runner's whole retire measures
0.40-0.66s on this host. The reviewer's literal "lease plus one tick" is
unreachable by any implementation, since signalling a process and giving it any
chance to exit takes non-zero time; detection now meets it and the stop runs
inside its own ceiling, and the contract says so rather than glossing it.
NECESSARY BUT NOT SUFFICIENT, and written BEFORE this head's integration runs
start rather than after they report. On the previous head, "Behavior portable
serial 1" and "Behavior portable serial 4" were both CANCELLED at the job
ceiling, independently of this finding. A new head triggers fresh runs, so those
two lanes MAY complete this time. IF THEY DO, THAT IS NOT EVIDENCE THE CEILING
DEFECT IS FIXED. It is one more sample of a lane that has been cut repeatedly
and sometimes is not; the shard-packing repair for it is open separately. Do not
reread a lucky pass here as a resolution.
Relatedly, and deliberately: the per-script duration hint in bin/fm-test-run.sh
was NOT updated even though the two new cases add ~19s of wall clock.
docs/fm-test-portable-shards.md says those hints are replaced wholesale from CI
timing artifacts of green runs, and that repair is the open request doing it; a
hand-edited estimate here would collide with it and silently repack the shards.
This suite runs in portable serial shard 3, which was green in the last run.
Verification: tests/fm-procevent.test.sh green, plus
tests/fm-captain-hold-lifecycle.test.sh, the test-coverage guard, and
bin/fm-lint.sh. The unrelated "reconcile stops a runner whose registration was
removed" case flaked in 4 of 7 local full runs; an isolated 20-trial
reproduction measured it at 13/20 unclean before this change and 11/20 after, so
it is issue 4080 and is not aggravated here.
* fix(procevent): repair our decimal-interval regression and enforce the timing phase
REPAIRED BEFORE PUBLICATION, AND IT WAS OURS. The half-interval arithmetic added
by the previous commit read a zero-prefixed interval as octal: 010 halved to 4
instead of 5, and 08 was not a number at all, so the owner guard died before
reporting ready and the runner failed closed and never listened. The validator
accepts those values and `[` compares them as decimal, so this broke a
configuration that worked before. Introduced by this delivery, found in review,
repaired here. Forcing base ten before the arithmetic is the whole runtime fix.
Proven by driving it rather than by reading the source: a new case starts a real
listener at 08 and at 010 and observes the guard's actual sleep argument - 4s and
5s. Removing the normalisation turns that case red with "a zero-prefixed decimal
interval (08) prevented the listener from starting".
THE TIMING PHASE IS NOW OBSERVED AND ENFORCED, NOT ASSUMED. The bound case
pinned its phase by CONSTRUCTION, from an assumed startup time, and enforced
nothing. Review was right that this is not enough: once startup reaches about two
seconds the expiry lands in a different part of the interval and the case
silently stops rejecting a two-interval guard while still reporting success. A
bound that cannot fail for the reason it names is the defect this whole delivery
exists to correct, so it must not ship inside the fix for it.
Now the lease is synchronised to the guard's own FIRST observed lease read,
every later real read is recorded, and the case REFUSES unless one recorded read
proves the required phase: it read the synchronised reference, it was still
fresh, and it began late enough that two further full intervals could not finish
before the deadline. An unestablished precondition refuses; it does not proceed
on trust. The derived deadline, the two-read debounce and the additive slack are
unchanged, and the slack invariant is now asserted rather than left to a comment.
Review also found the deadline was only ever checked while the group was still
alive, so a sampler descheduled past it would see the group gone and certify
success. The observed completion time is now checked too.
PROVEN BY MUTATION, each one run against this code:
- remove the decimal normalisation -> the interval case fails on 08;
- a full interval between the two reads -> "the guard exceeded its bound:
group still running 17.1s ... against a documented bound of 15s";
- a full interval WITH startup forced to ~2.5s, which is exactly the condition
the old construction pin could not survive -> still red, same message;
- the same ~2.5s startup with the correct guard -> still passes, 13.0s against
the 15s bound, so the delay alone does not break the case;
- phase evidence made unavailable -> "could not establish the required
pre-expiry guard-read phase", a refusal rather than a pass, even though the
group stopped quickly;
- act on one failed read -> the debounce case fails and the bound case passes
FASTER, 9.5s against 12.7s, which is why these remain separate cases.
Verification: full tests/fm-procevent.test.sh green, and bin/fm-lint.sh clean.
* fix(document): Correct process-event timing and debounce comments
* fix(backlog): route lifecycle transitions through configured adapters (#3417)
* fix(backlog): honor configured task adapters
* no-mistakes(review): Harden backend purity lint against prefixed Beads calls
* no-mistakes(document): Document configured backend lifecycle transitions
* fix(backlog): preserve markdown exemptions
* no-mistakes(review): Enforce backend purity for explicit lint paths
* no-mistakes(document): Update lifecycle backend documentation
* no-mistakes(lint): Remove redundant backend lint pattern
* fix(backlog): close adapter routing gaps
* no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint
* no-mistakes(document): Document environment-selected backlog adapters
* no-mistakes(lint): Fix empty local variable assignment
* fix(backlog): close quoted path gaps
* no-mistakes(review): Reject partially quoted direct Beads commands
* no-mistakes(document): Align lifecycle documentation with configured adapters
* test(backlog): keep structural cases markdown-only
* fix(backlog): honor configured task adapters
* no-mistakes(review): Harden backend purity lint against prefixed Beads calls
* no-mistakes(document): Document configured backend lifecycle transitions
* fix(backlog): preserve markdown exemptions
* no-mistakes(review): Enforce backend purity for explicit lint paths
* no-mistakes(document): Update lifecycle backend documentation
* no-mistakes(lint): Remove redundant backend lint pattern
* fix(backlog): close adapter routing gaps
* no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint
* no-mistakes(document): Document environment-selected backlog adapters
* no-mistakes(lint): Fix empty local variable assignment
* fix(backlog): close quoted path gaps
* no-mistakes(review): Reject partially quoted direct Beads commands
* no-mistakes(document): Align lifecycle documentation with configured adapters
* test(backlog): keep structural cases markdown-only
* no-mistakes(review): Harden markdown lifecycle routing and close recovery
* fix(lint): catch dollar-quoted beads commands
* no-mistakes(review): Drop unused authorized-data-dir param; canonicalize fm-lint ROOT with pwd -P
* no-mistakes(document): Document fm-lint backend-purity check in header and CONTRIBUTING
* no-mistakes(ci): Two failing checks, one code-caused and one attestation-only. 1. Greptile Review (P1: ANSI-C quoting bypasses lint) — REAL DEFECT, FIXED. The backend-purity normalizer in bin/fm-lint.sh (invokes_bd awk function) stripped $'...' quote delimiters without decoding ANSI-C escapes, so a core script containing $'\x62\x64' close fm-example executes as `bd close` while the lint accepted it. Fix: the normalizer now tracks whether a single-quote context came from $' (ansi flag) and decodes ANSI-C escapes while building the command word: \xHH (1-2 hex), \uHHHH / \UHHHHHHHH (4/8 hex), \NNN (1-3 octal), \cX control characters, simple escapes (a/b/e/E/f/n/r/t/v decode to a placeholder that can never spell bd), NUL (value 0) truncates the word per bash C-string semantics, and unknown escapes drop the backslash per bash. Values outside printable ASCII decode to a placeholder so they can neither falsely match nor collide into `bd`. Regression coverage added to the existing lint-interface test test_rejects_direct_beads_cli_invocations in tests/fm-lint.test.sh: $'\x62\x64', $'\142\144' (octal), and b$'\x64' (split word), each written into a fixture and asserted rejected through the real fm-lint.sh executable. Verified regression property: with the fix stashed, the new hex case passes lint (reproduces the reported bypass); with the fix, it is rejected. Verification: all 29 tests in tests/fm-lint.test.sh pass, and the full CI-parity lint (CI=true bin/fm-lint.sh: ShellCheck 0.11.0 full analysis over the canonical set, backend-purity check, and actionlint workflow validation) exits 0. One iteration was needed because an awk comment containing a literal $'\x62\x64' sample broke shell-level quoting (bash -n / SC1001/SC2026); the comment was reworded without quotes. 2. PR must be raised via no-mistakes — NOT code-caused. The check failed with 'Pipeline attestation head_sha does not match the current PR head': the PR body attestation binds to 82b41c7 while the PR head is 1995cdb because a later pipeline push moved the head. This is exactly the stale-attestation condition the user intent describes; it clears when the outer pipeline re-runs 'git push no-mistakes' and re-binds the attestation to the new head (which now includes this Greptile fix). No code change can or should address it. Files changed: bin/fm-lint.sh (ANSI-C escape decoding in the backend-purity normalizer), tests/fm-lint.test.sh (three encoded-bd rejection cases)
* no-mistakes(review): fix tasks.toml hang, root authorization, lint quoting
* no-mistakes(review): validate tasks config before exemption; fix lint quote gap
* fix(backlog): address the markdown backlog as <data>/backlog.md
Resolving the markdown backlog through a configured `[markdown] path` was
scope this task never asked for. It is absent from main, which addresses
`<data>/backlog.md` everywhere, and it came from an earlier review round
rather than the task brief.
Making it effective on the transition path alone put that path at odds
with every other consumer of the same backlog - fm-captain-hold.sh,
fm-session-start.sh, fm-fleet-snapshot.sh, fm-inbox.sh,
fm-backlog-handoff.sh - which all still address `<data>/backlog.md`. In
fm-captain-hold.sh the split was live: its reads had already moved to the
shared gate while its writes had not, so the two could address different
files.
Address `<data>/backlog.md` from the shared gate, delete the unused
resolver, and drop the two tests that pinned the withdrawn behaviour.
What this task actually changes is unaffected: a configured non-markdown
adapter is still addressed by its own root, without `--file`.
* fix(backlog): honor configured task adapters
* no-mistakes(review): Harden backend purity lint against prefixed Beads calls
* no-mistakes(document): Document configured backend lifecycle transitions
* fix(backlog): preserve markdown exemptions
* no-mistakes(review): Enforce backend purity for explicit lint paths
* no-mistakes(document): Update lifecycle backend documentation
* no-mistakes(lint): Remove redundant backend lint pattern
* fix(backlog): close adapter routing gaps
* no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint
* no-mistakes(document): Document environment-selected backlog adapters
* no-mistakes(lint): Fix empty local variable assignment
* fix(backlog): close quoted path gaps
* no-mistakes(review): Reject partially quoted direct Beads commands
* no-mistakes(document): Align lifecycle documentation with configured adapters
* test(backlog): keep structural cases markdown-only
* fix(lint): catch dollar-quoted beads commands
* no-mistakes(review): Drop unused authorized-data-dir param; canonicalize fm-lint ROOT with pwd -P
* no-mistakes(document): Document fm-lint backend-purity check in header and CONTRIBUTING
* no-mistakes(ci): Two failing checks, one code-caused and one attestation-only. 1. Greptile Review (P1: ANSI-C quoting bypasses lint) — REAL DEFECT, FIXED. The backend-purity normalizer in bin/fm-lint.sh (invokes_bd awk function) stripped $'...' quote delimiters without decoding ANSI-C escapes, so a core script containing $'\x62\x64' close fm-example executes as `bd close` while the lint accepted it. Fix: the normalizer now tracks whether a single-quote context came from $' (ansi flag) and decodes ANSI-C escapes while building the command word: \xHH (1-2 hex), \uHHHH / \UHHHHHHHH (4/8 hex), \NNN (1-3 octal), \cX control characters, simple escapes (a/b/e/E/f/n/r/t/v decode to a placeholder that can never spell bd), NUL (value 0) truncates the word per bash C-string semantics, and unknown escapes drop the backslash per bash. Values outside printable ASCII decode to a placeholder so they can neither falsely match nor collide into `bd`. Regression coverage added to the existing lint-interface test test_rejects_direct_beads_cli_invocations in tests/fm-lint.test.sh: $'\x62\x64', $'\142\144' (octal), and b$'\x64' (split word), each written into a fixture and asserted rejected through the real fm-lint.sh executable. Verified regression property: with the fix stashed, the new hex case passes lint (reproduces the reported bypass); with the fix, it is rejected. Verification: all 29 tests in tests/fm-lint.test.sh pass, and the full CI-parity lint (CI=true bin/fm-lint.sh: ShellCheck 0.11.0 full analysis over the canonical set, backend-purity check, and actionlint workflow validation) exits 0. One iteration was needed because an awk comment containing a literal $'\x62\x64' sample broke shell-level quoting (bash -n / SC1001/SC2026); the comment was reworded without quotes. 2. PR must be raised via no-mistakes — NOT code-caused. The check failed with 'Pipeline attestation head_sha does not match the current PR head': the PR body attestation binds to 82b41c7 while the PR head is 1995cdb because a later pipeline push moved the head. This is exactly the stale-attestation condition the user intent describes; it clears when the outer pipeline re-runs 'git push no-mistakes' and re-binds the attestation to the new head (which now includes this Greptile fix). No code change can or should address it. Files changed: bin/fm-lint.sh (ANSI-C escape decoding in the backend-purity normalizer), tests/fm-lint.test.sh (three encoded-bd rejection cases)
* fix(bin): preserve captain calls during teardown (#3595)
* fix(bin): never close a captain call during cleanup
A scout that held its own work item for the captain, which is what
captain-hold-lifecycle prefers ("hold the work item the question gates"),
was closed by bin/fm-teardown.sh's automatic backlog transition. The
completion gate passed, cleanup ran, and the captain's question moved to
Done with no recorded answer: the one thing the policy says must never
happen. `tasks-axi done` closes a held row silently, and nothing in
teardown asked whether the row was the captain's own call.
bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when
the task is still an open captain call, 1 when it is not, 2 when that
cannot be established. It reads the row through the transition library's
backend-aware probe, so it addresses the same backlog teardown does; the
script's other commands now address the configured data directory the
same way instead of FM_HOME, which also fixes captain holds in a home
with a relocated data directory.
Teardown asks `open` before any destructive step and refuses on 2. On 0
only the close changes: after cleanup and still under the task's own
lock, the row gets one "Deliverable of the finished work" line at the end
of its body and returns to Queued through `tasks-axi reopen`, keeping its
hold, so it lands in Captain's Call instead of reading as work under way.
--force does not lift this: it authorizes discarding unlanded work, never
the captain's question. The deliverable goes into the body because
`tasks-axi update --report` rewrites the title of a row that is not Done.
The crash window reuses the pending-close record teardown already stages:
a `mode=retain` line makes the existing replay record the deliverable and
reopen instead of closing, with the same validator, stale-generation
check, cleanup-incomplete marking, and non-blocking bootstrap lock as an
ordinary close. A retained row the captain answered first simply retires
the record. No parallel record type, recovery command, or second bootstrap
loop is introduced.
Regressions run the real executables: the captain-held scout survives
cleanup queued, held, with its deliverable and on the board, only
`answer` closes it, --force keeps it open, and an ordinary scout still
closes with its report; an interrupted cleanup leaves the row untouched
and the next session start retains it; a relocated backlog keeps the
retention in its one configured file; and a ship row whose hold cannot be
read refuses cleanup before anything destructive.
Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np
* no-mistakes(review): Serialize captain holds and fix backend-aware listing
* no-mistakes(document): Update captain-call retention documentation
* no-mistakes(document): Fix relocated captain-hold backlog diagnostics
* fix(backlog): honor configured task adapters
* no-mistakes(document): Update lifecycle backend documentation
* fix(backlog): close adapter routing gaps
* no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint
* no-mistakes(lint): Fix empty local variable assignment
* fix(backlog): close quoted path gaps
* test(backlog): keep structural cases markdown-only
* no-mistakes(review): Harden markdown lifecycle routing and close recovery
* no-mistakes(review): fix tasks.toml hang, root authorization, lint quoting
* no-mistakes(review): validate tasks config before exemption; fix lint quote gap
* fix(backlog): address the markdown backlog as <data>/backlog.md
Resolving the markdown backlog through a configured `[markdown] path` was
scope this task never asked for. It is absent from main, which addresses
`<data>/backlog.md` everywhere, and it came from an earlier review round
rather than the task brief.
Making it effective on the transition path alone put that path at odds
with every other consumer of the same backlog - fm-captain-hold.sh,
fm-session-start.sh, fm-fleet-snapshot.sh, fm-inbox.sh,
fm-backlog-handoff.sh - which all still address `<data>/backlog.md`. In
fm-captain-hold.sh the split was live: its reads had already moved to the
shared gate while its writes had not, so the two could address different
files.
Address `<data>/backlog.md` from the shared gate, delete the unused
resolver, and drop the two tests that pinned the withdrawn behaviour.
What this task actually changes is unaffected: a configured non-markdown
adapter is still addressed by its own root, without `--file`.
* no-mistakes(review): restore home boundary guard and tighten purity lint
* no-mistakes(review): authorize home boundary for every backlog adapter
* no-mistakes(test): complete tasks-axi stubs in fm-gotmp teardown fixtures
* no-mistakes(document): align backlog transition docs with adapter-neutral addressing
* no-mistakes(review): label adapter data-dir authorization, drop dead row_probe local
* no-mistakes(review): pin markdown backend at relocated-data addressing roots
* no-mistakes(document): point lint-definition mention at fm-lint.sh header
* no-mistakes(document): point mutate comment at adapter addressing owner
* no-mistakes(review): Fix leftover-symlink refusal on non-markdown homes; hoist config check and lint/dedup cleanups
* no-mistakes(document): Align fm-lint purity scope header with bin/backends
---------
Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
* fix(bin): suppress false Claude long-turn supervision alarms (#4119)
* fix(bin): stop false watcher-down alarms on long Claude turns
A healthy Stop auto-arm rewake or open claim already explains a mid-turn
beacon that has aged past grace, because turn-end will re-arm. Keep the
supervision-off banner for a missing, failed, or exhausted generation.
* no-mistakes(review): Bind Claude rewakes to active recovery generation
* no-mistakes(document): Document Claude long-turn supervision exception
* no-mistakes(lint): Fix empty ShellCheck assignment
* no-mistakes(ci): Fixed all reported CI failures: quoted the hyphenated recovery-delivery value to satisfy ShellCheck SC2100, and updated the session-lock auto-arm fixture to emit the recovery marker and watcher beacon now required for a valid rewake. Verified fm-session-lock-ancestry, fm-test-run, stale-banner, Claude auto-arm, targeted lint, ShellCheck, and workflow lint checks pass
* fix(herdr): recover gone and drifted worker endpoints (#4120)
* fix(herdr): classify a gone session's endpoint as recoverable
A task whose Herdr endpoint could not be read was classified `unreadable`,
which blocks recovery by design. The commonest reason that read fails is
that the recorded session's server is not running at all - a host reboot, a
server exit, a session never restored - and that is authoritative absence
for every pane in that session, not an ambiguous answer about one of them.
Tasks in that state had no sanctioned way back.
The recovery-grade read now settles an uninterpretable pane read with the
session server's own `.server.running` state: positively stopped reads
`missing`, while a running server, or a server state that cannot itself be
read, still reads `unreadable`. Resting the verdict on that field rather
than on the `server_not_running` error code is what keeps it working across
Herdr 0.8.x and 0.9.0, since the field is present on both and the code is
not.
Only that one boundary is widened. The husk classifier under it stays
strict, so duplicate prevention, rollback, and teardown - the paths that can
destroy something - keep refusing on exactly the reads they refused on
before.
Separately, a relaunch refused outright when the endpoint's shell had
drifted out of the recorded worktree. An agent's own exit routinely leaves
its shell somewhere else, so that refusal stranded tasks whose work was
sitting untouched on disk. The shell is now told once to return, and only a
shell that will not go refuses; the replacement still never starts outside
the copy holding the work.
Herdr 0.8.x is not installed on this host, so protocol-20 coverage is
structural plus the adapter fixture exercising both response shapes, and is
recorded as such rather than as a live result.
Fixes #4091.
* no-mistakes(review): Restrict drift recovery to Herdr endpoints
* no-mistakes(review): Correct Herdr recovery verification coverage
* no-mistakes(document): Document Herdr endpoint recovery boundaries
* fix(bin): prevent receiver wake failures from blocking remote handoffs (#4033)
* fix(bin): keep an escalated undelivered handoff wake retryable
A remote backlog handoff holds its outbox until the backlog receipt and
the receiver wake are both confirmed, and retries the wake under the same
pending-reply correlation on every resume. When that wake's remote
transport was lost, the correlation stayed undelivered in delivery_unknown
and the watcher's next pending-reply tick escalated it. Both the reuse
predicate and the known-undelivered reset refused an escalated record, so
the resume refused to resend the wake forever and every later handoff to
that mate jammed behind the outbox.
Treat an escalated record with no confirmed delivery as the undelivered
correlation it is: fm_pending_reply_corr_reusable accepts it for its own
task and fm_pending_reply_reset_known_undelivered returns it to
awaiting_report for the idempotent remote resend, while a delivered
record is still never reset and a missed-report escalation keeps its
meaning. The published delivery-unknown decision stays open until the
record resolves, so a repeat loss neither re-notifies nor strands it.
Reproduce the deadlock end to end in the remote handoff test (lost wake
transport, watcher escalation, resume) and pin the predicate contract in
the pending-reply suite; the fm-send fixture that pinned the refusal now
uses a genuinely stale delivered escalation.
* no-mistakes(review): Decouple durable outboxes from best-effort wake retries
* no-mistakes(review): Align handoff documentation with durable receipt release policy
* no-mistakes(review): Handle unrecordable wake state as dropped
* no-mistakes(review): Prevent stale wake markers blocking handoffs
* no-mistakes(review): Prevent stale delivered markers suppressing new wakes
* no-mistakes(document): Clarify retry escalation decision lifecycle
* no-mistakes(document): Document pending receiver wake retries
* fix: restrict captain address rule to user chat (#4075)
* docs: bound the mandatory captain address to the chat channel
AGENTS.md's opening address rule said "address the user as captain at
least once in every response" and never said what a response is. The
artefact exclusion two lines below governed only the optional nautical
seasoning, not the mandatory address. An agent that reads this file
without being the first mate - a pipeline corrector agent running inside
a copy of this repo - therefore read the obligation as applying
everywhere and the exclusion as applying only to flavour, and opened its
delivery message with "Captain,". That reading was correct.
Patch the existing owner rather than adding a rule elsewhere:
- bound the obligation to chat messages sent to the captain;
- state the artefact exclusion once, explicitly binding every agent that
reads this file whether or not it is the first mate, and naming commit
messages, PR and issue descriptions, briefs, code and comments;
- fold the seasoning under the same bound instead of carrying a second,
narrower copy of the exclusion.
The obligation itself is unchanged: the captain is still addressed in
every chat message.
AGENTS.md goes from 603 to 602 lines: the redundant "never send a
response with zero direct address" clause and the duplicated seasoning
exclusion pay for the new bound.
The two cross-references that paraphrased the unbounded wording
(bin/fm-parent-channel-lib.sh's header and
docs/secondmate-parent-channel.md's problem statement) now match the
owner; neither restates the rule.
* fix(review): Limit address exclusions to artifacts while preserving public replies
* fix(document): Consolidate captain address guidance
* fix(herdr): allow detached teardown of persisted-focused tabs (#4131)
* fix(herdr): close persisted-focused tabs when no live client is attached
The teardown active-tab guard treated Herdr's last-focused pointer as a live viewer, so detached sessions could not close panes on that tab.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(herdr): allow detached seeded-tab prune after live-client gate
Projection create still restored the persisted focused tab after a successful prune, so a detached last-focused seeded tab still quarantined the spawn.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(herdr): probe live client after seeded prune only when that tab was focused
The extra title-clear read after every prune shifted canned CLI fixtures and failed projection create.
Co-authored-by: Cursor <cursoragent@cursor.com>
* no-mistakes(review): Tighten Herdr active-tab close guard
* no-mistakes(review): Guard Herdr mutations with fresh target focus
* no-mistakes(document): Document Herdr live-viewer teardown guard
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(bin): escalate decision-owned wakes once as the decision (#4169)
* fix(bin): escalate decision-owned wakes once as the decision
The away-mode daemon treated a needs-decision: queued payload as an
unknown wake, so suppression markers never committed and the same open
decision re-escalated on every poll.
Classify that payload through the existing signal path so it escalates
once, labelled as the decision, and an unchanged repeat is suppressed
on the same terms as any other signal.
Fixes #4096
* no-mistakes(review): Escalate captain-held decision-owned rows once as the decision
* no-mistakes(review): Self-handle captain-held decision-owned rows instead of escalating them
* no-mistakes(document): Name away daemon as needs-decision payload reader
* test(herdr): cover agent exit-to-shell liveness (#4172)
* test(herdr): pin leftover-shell vs live-idle via agent get
Herdr 0.9.0 already distinguishes a Pi that exits to a surviving pane shell
from a sibling live idle occupant. Pin that pair through agent get and the
recovery classifier so a lagged pane-get status cannot silently reclaim the
leftover shell as alive.
Co-authored-by: Cursor <cursoragent@cursor.com>
* no-mistakes(document): Document Herdr leftover-shell liveness regression
* no-mistakes(ci): Fixed Lint failure SC2034 by replacing the unused wait-loop variable with `_`. Verified with the pinned project lint command, Bash syntax check, and git diff check
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(bin): rebalance portable parallel test lanes using CI timings (#4151)
* ci: rebalance the portable …
yelenplays
added a commit
to yelenplays/firstmate
that referenced
this pull request
Sep 29, 2026
* test: make timestamp fixtures portable across macOS and Linux (#4037)
* test(lib): set fixture mtimes through one portable epoch helper
On macOS the visible symptom was ONE red case in the turn-end guard suite. The
actual damage was TWO cases that had quietly stopped testing their subject. The
red one was the harmless half - people read one red case as one broken thing,
and here that intuition is wrong.
`touch -d @<epoch>` is a GNU extension; BSD touch rejects it outright and leaves
the file at its current mtime. So on macOS the three away-mode beacon cases
never aged their beacon at all.
The 400s case exists to pin that 400s is stale under the flat 300s default but
fresh under the poll-derived grace (660s at FM_POLL=600). Deleting the grace it
guards (FM_POLL=60, so max(300,120)=300) and re-running proves what it was
worth on this platform:
pre-fix input (beacon left at now): ok - passes with the feature DELETED
post-fix input (beacon 400s old): not ok - expected exit 0, got 2
It was green while measuring nothing, and could not have caught a regression in
the grace it names. Only the 700s case broke loudly.
`touch -t [[CC]YY]MMDDhhmm[.SS]` is POSIX and both platforms accept it, so the
only host-specific step left is formatting the epoch into that stamp, which
date(1) spells two incompatible ways. fm_touch_epoch in tests/lib.sh owns that
probe once and fails loudly rather than leaving an unset timestamp behind - the
failure mode that caused this. Verified on BSD touch/date here and on GNU
coreutils 9.7 in a container.
Three real sites, and one consistency change - not four fixes. The stale
destination lock in tests/fm-remote-backlog-handoff.test.sh was never a defect:
its `uname = Darwin` branch made the `touch -d` line unreachable on macOS, and
BSD touch accepts that space-separated form anyway. `touch -t` takes the date
directly on both platforms, so the branch goes rather than standing as a second
copy of the same platform assumption.
Known limit: the third case (away mode off) is only HALF recovered here. It now
receives the input its name claims, but it is still insensitive after this fix -
its verdict is identical with a 0s and a 400s beacon, because the fixture
records a daemon lock and no watcher lock, and with away mode off the daemon
lock proves nothing. Not fixed here; tracked separately, with the requirement
that any fix be shown to FAIL when the protection is removed.
FULL SUITE ON macOS: 189 scripts, four red, none of them this change. The
turn-end guard and remote-handoff suites are clean. Attribution was established
by running the four failures at the base commit and at this head on an idle
machine, because base-idle against head-under-load moves two variables at once:
script head/loaded base/idle head/idle verdict
fm-calm-pi-extension red red red pre-existing
fm-backlog-atomicity red red red pre-existing
fm-procevent red red red pre-existing
fm-startup-network red green green cause unestablished,
load-sensitive under
a full run
Reported, not fixed. fm-calm-pi-extension deserves its own note: it FAILS
because Chrome is absent instead of declaring the capability it needs and
standing aside, so its verdict is about the machine rather than its subject -
the same family as the defect above, with the red at least announcing itself.
Neighbouring class, reported not changed: `file_mode()` - a verbatim
`uname = Darwin ? stat -f %Lp : stat -c %a` - is copy-pasted across at least
five test scripts plus a `reread_mode` variant, and epoch-mtime reads are
open-coded as `date -r … || stat -c %Y` in three more; same one-owner shape as
the defect above. `git init` without `-b main` depends on the host's
init.defaultBranch in several scripts (branch-name case, tracked elsewhere).
timeout, sha256sum and sed -i uses are all correctly guarded where checked.
Observed while building the check rather than the fix: the first watcher I
wrote to wait for the suite matched its own command line, so it was waiting on
its own existence and could never fire. Same shape as the cases above -
machinery answering confidently about something other than its subject, by
including itself in the evidence it was meant to judge. The file sentinel it
was replaced with cannot be produced by the observer that reads it.
* fix(review): Pin fixture timestamps to UTC across DST transitions
* fix(document): Clarify shared fixture suite coverage
* feat(pi): accept native Codex ultra effort with progress-aware supervision (#4038)
* fix(pi): preserve native Codex effort and guarded supervision
* test(pi): identify native compatibility guard versions
* no-mistakes(review): share native-main follow rule between build and picker
* no-mistakes(document): document native progress marker and ultra effort owners
* no-mistakes(ci): Failing check "Behavior portable serial 1" was caused by this PR. The shard ran the default-on live guard tests/fm-pi-branch-responsiveness-live-e2e.test.sh, whose idle arm loads .pi/extensions/fm-branch-supervision.ts into a scratch project with a fixed list of copied libs. This PR added `import { registerFirstmateTool } from "./lib/fm-native-contract.ts"` to that extension and updated every other loading fixture's copy list, but missed this guard. Pi 0.85.1 therefore refused to load the extension ("Cannot find module './lib/fm-native-contract.ts'"), never drew its TUI, the test failed with "Pi 0.85.1 never drew its TUI in the idle arm", and the job hit its 20-minute cap. Fix (one line): added fm-native-contract to the lib copy loop in tests/fm-pi-branch-responsiveness-live-e2e.test.sh. Swept all other suites referencing fm-branch-supervision.ts / fm-primary-pi-watch.ts; the remaining ones without the new lib only hash, path-reference, or string-match the files and do not load them into Pi, so no further fixture changes are needed. Verification: reproduced the mechanism against the installed Pi 0.85.1 by building the lab copy with the old lib list (load error as above) and the fixed list (loads cleanly). tmux is not installed on this machine, so the live guard itself gate-skips locally ("skip: live: tmux absent") and could not be run end to end here; CI (which has tmux and Pi) will exercise it. shellcheck is clean on the edited file
---------
Co-authored-by: Talon Stark <talonstark@gmail.com>
* fix(herdr): bypass stale clients rejected by running servers (#4041)
* fix(herdr): step around a stale client the running server refuses
A remote host can carry a self-updated herdr in ~/.local/bin beside a
package-managed one, and the fixed remote-job PATH resolves ~/.local/bin
first. After the server upgraded to 0.9.0 (protocol 22) the stale 0.8.2
client (protocol 20) was answered with protocol_mismatch on every command,
which the read classifiers folded into `unreadable`: the live remote
secondmate read unknown, every doorbell into it failed, and both the spawn
and relaunch recovery paths refused, so the defect trapped itself.
The adapter's session-scoped CLI wrapper now recognizes that refusal, reads
status per session from each distinct herdr on PATH, adopts the first one
the running server reports compatible, retries once, and keeps it for the
process. The happy path makes no extra call and no other failure reselects.
An endpoint that still reads unreadable names the refused client, both
protocols, and the fix on stderr; the remote state read, fm-crew-state, and
the launch refusal carry that reason, and fm-remote-doctor reports the
selected client and rebinds the launch agent to it.
Regression coverage: fake two-client hosts in the herdr unit suite, the
doctor suite, the crew-state remote arm, and the real host-local control
script in the remote lifecycle e2e; the real-herdr smoke refreshes the
status shape the selection reads.
* no-mistakes(review): Reselect Herdr client after every protocol mismatch
* no-mistakes(review): Remove unrequired Herdr diagnostics and launch-agent rebinding
* no-mistakes(review): Scope cached Herdr clients to their selected session
* no-mistakes(review): Restrict herdr client selection to reactive CLI calls
* no-mistakes(document): Document session-scoped Herdr client reselection
* no-mistakes(document): Clarify Herdr client selection documentation
* no-mistakes(ci): Updated the trusted fm-remote-doctor.sh SHA-256 in bin/fm-remote-entrypoint.sh after the PR changed the doctor, restoring git-unavailable bootstrap authentication. Verified tests/fm-on.test.sh, tests/fm-backend-herdr.test.sh, bin/fm-lint.sh, and git diff --check all pass
* feat: add durable AFK posture lifecycle (#4048)
* feat(afk): record the away posture and its lifecycle (phase 1)
Away mode becomes a posture of the one supervision session, recorded in
state/.afk-contract by the new bin/fm-afk-contract.sh: the one owner of the
record schema, the mandate-clause grammar and compiler, refusal naming the
missing part, the read-back rendering, the entry announcement (hold-for-return
only, no phone channel), and the archive at return. This release records
clauses and does not execute them; the announcement and return brief say so.
bin/fm-afk-launch.sh gains propose and confirm, confirms the record before any
daemon launch, refuses to launch the daemon on Pi and pi-signed, and archives
the record last on stop. bin/fm-afk-return.sh snapshots supervisor health
before shutdown, renders the return brief (health, mandate, waiting on the
captain, could not fix, handled, cost) from the archived record, the outcome
store, the held set, and the status logs, and shrinks the blocker gate to what
the away session could not fix.
While the record exists the watcher and the daemon never recheck an item held
for the captain. Declared external waits get a four-hour default cadence and
honor `until <UTC ISO 8601>` on the paused line, in both postures, bounded by
FM_PAUSE_UNTIL_MAX_SECS.
The /afk skill, AGENTS.md's layout and away-mode stub, the session-start
digest, and the architecture, Pi branch, configuration, and scripts docs
describe the record. The Pi/Herdr e2e now proves the no-daemon posture on a
real Pi primary; its verification record carries the 2026-09-08 run.
* no-mistakes(review): Fix AFK confirmation, grammar, waits, and return gating
* no-mistakes(review): Harden AFK authority and posture lifecycle
* no-mistakes(review): Preserve AFK history and tighten authority grammar
* refactor(afk): record clause fields with no natural-language parser
By the captain's mandate the away-posture record keeps no static parser
that tries to understand natural language. A mandate clause is now given
as explicit fields (--action, --object, --when, optional --stop) that
bin/fm-afk-contract.sh records verbatim. The structural check asserts
only that the action, object, and precondition fields are present and
that the action is a listed verb; whether a precondition holds is the
supervision session's judgment at execution time in a later phase.
The never-set stays as a forbidden-concept safety scan: fields mentioning
credentials, passwords, logins, legal or financial acceptance, payments,
invoices, one-time codes, or an attended prompt are refused, matched at
token prefixes after punctuation normalization so compound and plural
spellings are caught. The red-check grammar, class-word rejection,
unconditional-word detection, clause-reference resolution, and condition
aliases are removed. --words-file keeps the captain's words verbatim,
trailing newline included.
The skill, docs, launcher help, and tests describe the field form.
* no-mistakes(review): Preserve AFK words and tighten safety refusals
* no-mistakes(review): Preserve clause bytes and honor declared waits
* no-mistakes(review): Harden deny-list and gate unreadable outcomes
* no-mistakes(review): Demote never-set scan and clarify authority
* no-mistakes(review): Gate return on unreadable held and status data
* no-mistakes(review): Validate posture archives and enforce Pi detection
* fix(afk): make the never-set a non-refusing flag and keep return fail-safe
Per the captain's decision the never-set scan is a coarse best-effort
flag, never a refusal and never the gate: a clause naming a listed
concept is still recorded with a flag the read-back, announcement, and
return brief show, and the scan matches listed terms exactly or with a
plain inflection at punctuation-delimited token boundaries, so unrelated
names such as ping-service or tokenize-worker are never flagged and
joined compounds remain a documented miss. Authoritative never-set and
forbidden-action enforcement is the supervision session's judgment at
execution time in phase 4.
A replacement copies the superseded record through a temporary name and
renames it atomically so a failed copy leaves no partial archive, the
record owner gains validate and flags subcommands, and the return keeps
catch-up gated when a superseded archive cannot be read.
* no-mistakes(review): Harden AFK record validation and return reconciliation
* no-mistakes(review): Harden AFK record validation and simplify commands
* no-mistakes(review): Harden mandate validation and retain missing records
* no-mistakes(review): Refuse blank explicit mandate stops
* no-mistakes(review): Recover restored posture epoch before return
* no-mistakes(review): Prevent return brief status symlink reads
* no-mistakes(document): Refresh AFK posture documentation
* no-mistakes(ci): Fixed both CI failures: updated lint telemetry for the new fourth source directive, quoted the hyphenated fixture value, and removed unreachable test cleanup. Verified with tests/fm-lint.test.sh, targeted CI-mode ShellCheck, bin/fm-lint.sh, bash syntax checks, and git diff checks
* no-mistakes(ci): Bound structured pause deadlines by FM_PAUSE_RESURFACE_SECS in watcher and daemon housekeeping, added distinct bounded-horizon reasons, regression coverage for near, passed, and wrong-year deadlines, and updated documentation. Verified targeted behavior tests, full daemon tests, ShellCheck source-following lint, syntax, and diff checks
* feat(bin): add IMAP/SMTP mail plane with standing poll (#3765)
Opt-in IMAP/SMTP mail plane (fm-mail.sh / fm-mail-check.sh). Absent FM_MAIL_* stays off.
Speaking as Kun's firstmate: this is merged. Thank you @feilipu — really appreciate you taking the time on this.
* fix: launch remote Herdr through the user login shell (#4061)
* fix(remote): start the fm-remote Herdr agent through a login shell
Launchd was exec-ing herdr directly, so the Aqua agent inherited a background session without login-keychain access. Start it via /bin/zsh -lc exec so panes keep login env and can refresh OAuth tokens after reboot.
* fix(remote): start fm-remote Herdr via the account login shell
Resolve UserShell from Directory Services and invoke it with separate -l and -c so bash, fish, and zsh all get login-keychain access. Fall back to SHELL, then /bin/zsh, then /bin/sh without failing the render.
* no-mistakes(review): Fix launch-agent shell fallback resolution
* no-mistakes(review): Preserve and escape Directory Services shell paths
* no-mistakes(document): Document login-shell LaunchAgent behavior
* no-mistakes(ci): Updated the trusted fm-remote-doctor SHA-256 identity in bin/fm-remote-entrypoint.sh. Verified with tests/fm-on.test.sh, tests/fm-remote-doctor.test.sh, bash syntax checks, and git diff --check
* no-mistakes(ci): Resolved the login shell exactly once per doctor invocation and threaded it through plist rendering, installed/loaded contract validation, repair reporting, and post-repair checks. Added a regression test proving repeated repair remains healthy and performs no reload when a hypothetical second Directory Services lookup would differ. Updated the trusted doctor hash. Verified doctor, fm-on, remote-entrypoint, lint tests, ShellCheck, syntax, and diff checks
* no-mistakes(ci): Made Darwin shell resolution hermetic with executable injection and a 2-second Directory Services timeout. Updated tests to inject shells by default, isolate dscl-specific cases, parse plists semantically, and verify stalled dscl fallback. Updated the trusted doctor hash. Doctor, fm-on, entrypoint, syntax, hash, and diff checks pass
* no-mistakes(ci): Raised portable serial CI timeout from 20 to 30 minutes, refreshed the specified timing hints, added missing hints, and recomputed shard documentation. Verified coverage, runner behavior tests, workflow lint tests, shell syntax, requested timing maxima, and diff checks
* fix: distinguish landed deliveries from resolved captain calls (#3710)
* fix(bearings): keep captain-approved deliveries in Recently Landed
A closed task is never held: tasks-axi clears the held flag when a task
closes and keeps hold-kind and the hold reason as the record of the call
that was made. Recently Landed excluded every Done row whose hold-kind was
captain, so the marker it treated as "closed while still waiting on the
captain" was in fact the proof that the captain had approved the work. Every
merge routed through a captain decision disappeared from the list of what
shipped, including under --all-landed.
The selector now asks whether the closed row delivered something. Recently
Landed is merged PRs, completed scouts, and finished local-only merges, so a
row carrying one of those artifacts belongs there whoever approved it. A
captain question closes with an answer and no artifact of its own, and that
is what still stays out, so an answered question is never rendered as
shipped work.
The same rule was written twice - the bearings projection selects this
home's Done rows and the fleet snapshot selects each secondmate home's Done
rows into the roll-up the same section merges in - which is why one defect
hid deliveries in every home. Both now share bin/fm-landed-lib.sh.
* fix(review): Normalize landed evidence and exclude answered captain questions
* fix(review): Normalize captain delivery evidence across relocated data
* fix(review): Record authoritative delivery provenance with legacy fallback
* fix(review): Harden delivery provenance across forced and pruned completions
* fix(review): Replace premature merge closure with existing release contract
* fix(review): Document provenance-based Recently Landed selection
* fix(document): Align documentation with completion provenance
* fix(lint): Fix targeted ShellCheck warnings
* fix(ci): order the pinned tasks-axi install before its stock-Bash consumers
In `.github/workflows/ci.yml` the pinned tasks-axi install now precedes both
stock-Bash consumers, and the Bearings expectation is updated from 49 to 50
tests.
Verified with macOS Bash 3.2: snapshot 16/16, Bearings 50/50, public-followup
1/1. Full repository lint and all three workflow validations pass, and
`git diff --check` is clean.
* fix(review): Make completion provenance unambiguous
* fix(review): Make completion verdict authoritative over quoted provenance
* fix(review): Preserve retained artifacts through resumed captain closes
* fix(review): Unify completion provenance ordering across writer and reader
* fix(review): Preserve artifacts across failed captain closes
* fix(review): Refresh v1 assertions; provenance authority remains unresolved
* fix(review): Remove unreliable provenance while preserving landed deliveries
* fix(review): Reject stale home summaries visibly
* fix(review): Restore retained deliverable recording
* fix(review): Match landed artifacts and restore retention documentation
* fix(review): Disambiguate captain calls and restore landed artifact matching
* fix(review): Persist retained report and PR artifacts
* fix(review): Preserve staged artifacts before captain answers
* fix(review): Avoid wedging answers on unsupported report paths
* fix(review): Exclude unreleased captain-held pull requests
* fix(review): Exclude held local-only answers from landed
* fix(review): Preserve retained scout reports across snapshot rendering
* fix(document): Align landed lifecycle documentation with release semantics
* fix(review): Enforce landed artifact-kind ownership
* fix(review): Infer canonical task kinds in snapshots
* fix(review): Require captain-hold release before merges
* fix(review): Qualify merge lifecycle regression evidence
* fix(review): Serialize captain holds with merge operations
* fix(review): Document merge cleanup residuals honestly
* fix(test): Replace vacuous Bearings regression with behavioral cases
* fix(document): Align Bearings verification and merge lifecycle documentation
* fix(review): Serialize merges and exclude captain calls from landed
* fix(review): Harden merge identity and landed selection
* fix(document): Clarify landed selector compatibility filtering
* fix(bin): keep merge entrypoints usable on records without an incarnation
The merge identity guard refused any task record with no spawn_gen field.
That field identifies one exact incarnation, so comparing it across the wait
for the merge lock is what catches a task relaunched while the merge was
queued. Requiring it to be present is a different rule, and it refused every
record written before the field existed: a legacy task could no longer be
merged at all, and five behaviour suites refused before reaching the check
they were written to exercise.
The comparison only needs to notice a change. An absent field is now read as
an empty incarnation and compared like any other value, so a record that
gains, loses, or alters one is still refused, while a record that simply
predates the field merges. An ambiguous or unreadable field stays an error,
because a record that cannot name one incarnation cannot be compared. The
missing-record message each entrypoint had before the guard is restored, so
a genuinely absent record still says so in its own words.
The role partition now precedes reading the record. Refusing the supervision
branch is a statement about the actor, not about the task, so it cannot
depend on a record the wrong actor may not have.
A backlog file that does not exist meant "no longer an open captain call".
For a caller that asked to tell absence apart it now means absent, so a board
card whose home carries no backlog stays visible instead of being dropped as
resolved.
Fixture repositories pin their initial branch instead of inheriting
init.defaultBranch, which resolved to main on a developer machine and master
on a runner, so a fixture naming main failed only in CI.
* fix(review): read local-only note from body; surface pending-close failures
* fix(review): keep kindless local-only landings in Recently Landed
* fix(review): bind local-only note scan to the tasks-axi note line
* fix(review): Guard unavailable captain-hold authority records
* fix(document): Document unreadable authority predicate outcome
* fix(bin): read an absent backlog as absence, not an unreadable record
The merge gate refused every task whose home carries no backlog file. A
backlog that does not exist holds no captain call, so nothing can be held and
the merge is safe; only a backlog that exists and cannot be read may hide a
live hold. Those two states were collapsed into one refusal, which stopped
merges in any home that keeps no backlog.
The predicate now reports a missing backlog file as absence, alongside a row
the backlog does not carry. A record that exists but cannot be read still
leaves by the existing cannot-tell path, which both merge entrypoints already
refuse, so the restrictive direction is unchanged.
That leaves no way to reach the separate unavailable-record result, so the
result and the two branches that handled it are removed rather than left
describing an outcome that can no longer occur. The lifecycle documentation
loses the same claim.
Regressions cover both directions in each entrypoint: a home with a task
record and no backlog merges, and a backlog present but unreadable refuses
without reaching the forge.
* fix(review): Fail closed unreadable backend configuration
* fix(tests): pin the bare origin's initial branch in the remote seed fixture
The fixture created its bare origin with no initial branch, so that
repository's HEAD followed init.defaultBranch while the source repository
pushed the branch fm_git_init_commit pins. On a host that still defaults to
master the two disagreed: the bare origin's HEAD named a branch the push never
created, cloning it warned that the remote HEAD referred to a nonexistent ref
and checked out nothing, and the seed assertion for the cloned README failed.
A machine whose default is already main paired the two by accident and hid it,
which is why the fixture passed locally and failed on the runner.
Pinning the bare origin to the same branch removes the dependency on the
ambient default from both sides. Verified under both conditions: with
init.defaultBranch set to master, and set to main, the suite passes 26 of 26.
* fix(review): Fail closed unreadable user backend configuration
* fix(bin): republish the home summary as v1 and record two load-bearing rules
The published home-summary schema had moved to v3, which routed every
secondmate home still emitting the earlier version to the stale branch: their
landed rows, open decisions and holds all came back empty and their state read
as unknown until each home was updated. The payload never justified that. Its
field set, field order, truncations and the landed array construction are
byte-identical to v1, so only which rows the selector places in landed
differs, and a v1 consumer reads that the same way.
Republishing as v1 removes the rollout regression and, with it, the tolerance
machinery that existed only to soften the bump: the stale-schema predicate,
its two collection branches, the flag and its provenance branch, the omitted
surface that can no longer be reached, and the fixtures and assertions that
covered them.
Two rules that a scope review proposed removing are kept, each now carrying
the reason it exists, because both were measured to be load-bearing:
The artifact-kind ownership clause is what keeps an explicit scout that
recorded no report out of Recently Landed. Without it such a row has none of
the three artifacts, satisfies the compatibility fallback and renders as
shipped work with an empty artifact.
The kind fallback is needed because tasks-axi omits the kind metadata
entirely when a title begins with a canonical keyword. Without it a scout
titled "SCOUT ..." reports no kind, its recorded report stops counting as a
delivery, and it drops out of the section this selector exists to repair.
* fix(bin): move the scout guard note onto the rule and drop two dead pieces
The LOAD-BEARING note sat on an unreachable branch. Measured in both
directions: removing that branch together with the kind-is-not-scout guards
lets an explicit reportless scout into Recently Landed and fails
tests/fm-captain-hold-lifecycle.test.sh, while removing the branch alone
leaves that suite passing at 49 assertions. The guards carry the rule, so the
note now sits on them and the unreachable branch is gone. A note pointing a
later reader at the wrong line is the hazard this change corrects elsewhere.
summary_file_has_schema lost its only caller when the stale-schema machinery
was removed, so it goes with it.
* fix(review): Fix legacy report artifacts and canonical keyword boundaries
* fix(review): Update pinned Bearings test count to 56
* fix(document): Clarify landed summary compatibility documentation
* fix(review): Preserve unreadable backend configuration errors
* fix(review): Honor backend resolution errors at existing call sites
* fix(test): Stabilize remote collector tests under host load
* fix(document): Document backend resolution failure contracts
* fix(lint): Suppress intentional deferred probe expansion warnings
* fix(ci): Captain, quoted the two literal test IDs in tests/fm-backlog-atomicity.test.sh to fix SC2100 without changing behavior. Both warnings reproduced before the fix; the targeted fm-lint.sh run now passes with ShellCheck 0.11.0. Bash syntax and git diff --check also pass
* fix(ci): Fixed the resolver’s two configuration-parent checks to return 2 for inaccessible directories while preserving genuine absence. Added two behavioral tests; RED/GREEN and both requested mutation proofs confirmed. All 10 focused checks, targeted lint, syntax, and whitespace checks passed. Broader merge suite stopped after 10 passing cases under host load. Declined portable checks and merge-authority code remain unchanged
* fix(remote): keep fm-remote Herdr servers in the Aqua session (#4090)
* fix(remote): let the Aqua launch agent own the fm-remote Herdr session
A herdr server keeps the macOS audit session of whatever started it, and
only the Aqua login session (gui/<uid>) can read the login keychain
without a prompt. Herdr's SSH remote attach starts the fm-remote server
as its own child when it finds none, wins the socket at boot because sshd
accepts connections before the login session exists, and every claude
pane under that server then gets `security` exit 36, falls back to a stale
plaintext credentials file, and reports "Login expired". launchd's own job
lost the socket on every KeepAlive retry and the doctor still reported the
session ready because it only asked whether any server answered.
- Add bin/fm-remote-herdr-guard.sh, the launch agent's exec target: start
the server in the foreground when nothing owns the socket, exit 0 when an
Aqua-born server does, and otherwise stop the foreign server, wait for the
socket, and exec the server at once.
- Add bin/fm-remote-herdr-owner-lib.sh, the single owner of socket-owner
discovery (lsof; pgrep cannot see herdr's argv on macOS) and the birth
markers (SSH_*, XPC_SERVICE_NAME, FM_REMOTE_JOB_ACTIVE, sshd or
remote-client-bridge ancestry matched on argv[0] and whole arguments).
- Render the agent as the login shell exec'ing the guard with
KeepAlive={SuccessfulExit=false} and ThrottleInterval=10, check the loaded
job's successful-exit semaphore, and report a session served outside the
Aqua login session as fixable so --fix retakes it through launchd; the
reload waits for an Aqua-born owner rather than any running server.
- Correct the doctor and docs: the launch shell provides environment parity,
the launchd domain provides keychain access.
- Pin the guard's decision table and the doctor's verdicts against real
marker-carrying processes, and record the dated audit-session evidence.
* no-mistakes(review): Verify Aqua ownership through launchd domains
* no-mistakes(document): Document macOS lsof ownership requirement
* fix(bin): prefer a live no-mistakes run over a terminal one (#2881)
* fix(bin): prefer a live no-mistakes run over a terminal one
A worktree can bind to more than one recorded no-mistakes run at once.
The branch-and-code-identity rule in bin/fm-nm-run-lib.sh accepts both an
exact-equal commit and a worktree-is-an-ancestor match, but never stated
which wins when both bind, so the tie fell to whichever candidate the
caller reached first.
Observed on a live fleet: a crashed validation daemon left a FAILED run
at the worktree's own commit while the live run that replaced it
validated a descendant commit on the same branch. Bare `axi status`
answers with the most-recently-touched run - the corpse - and it bound by
the equal-commit rule, so every recomputation reported `failed` for a
task whose real run was healthy. The same label had also read `failed`
earlier while the work was genuinely stalled, so the signal was wrong in
both directions.
State the live-over-terminal policy in the matching rule's own contract,
where the equal-commit and ancestor rules already live, and add
fm_nm_run_status_class as the one classifier that decides liveness from a
recorded status word. fm-crew-state.sh applies it on both selection
paths: the runs listing now scans past a terminal row for a live one, and
a terminal `axi status` answer is provisional until the listing has been
asked whether this worktree also has a live run.
Same-liveness-class candidates keep the listing's newest-first
precedence, and a status word the classifier cannot place keeps the
caller's own ordering rather than displacing a known result, so a
single-run task and a task whose runs are all terminal are unchanged.
Regression coverage reproduces the proven case (terminal run at the
worktree's exact commit plus a live run descending from it) and its
runs-list twin; both fail under the old tie-break. Two companion cases
pin the no-widening half - two terminal rows still resolve newest-first,
and a terminal run with no live sibling keeps its full run-step detail -
and both pass before and after the change.
* no-mistakes(review): accept unfetched live sibling anchored at exact worktree head
* docs(bin): name both ledger reads behind the runs-limit setting
The FM_CREW_STATE_RUNS_LIMIT comment in bin/fm-crew-state.sh still described
the runs ledger as scanned only by the cross-branch fallback, but the
live-over-terminal fix also consults it as the live-sibling probe behind a
terminal axi status answer. Point the comment at docs/configuration.md as the
setting's owner instead of restating a second copy.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* fix(bin): repair process-event shutdown and tighten guard timing (#4009)
* fix(procevent): make the ordinary stop signal actually stop a runner
The owner guard that shipped in #3904 half-reaps. Against a poll child that
handles the ordinary stop signal and keeps waiting, the guard signals the group,
loses the runner leader to its own signal, then reads that success as a
leaderless group and exits without escalating. It destroys the only proof of
ownership that would have authorised the forced signal, so the survivor becomes
unreachable by retire, reconcile, sweep-home and the guard alike. A guard that
turns a leaking-but-identifiable generation into a permanently unreachable one
is worse than no guard at all.
Two defects, and they hid each other:
- The escalation re-derived ownership from the leader. `runner_group_signal`
now takes a `proved` mode, passed only by the escalation inside the stop that
already proved and signalled that exact generation moments earlier. A leader
dying to our own signal is the ordinary outcome, not fresh ambiguity.
- Every stop held the per-source lock across its wait while the runner's own
exit cleanup waited unboundedly for that same lock. That circular wait was
broken only by the forced signal, so the forced signal silently became the
normal path - and, by keeping the leader alive through the whole window, it
masked the escalation defect above. The runner's exit cleanup now refuses that
lock instead of waiting for it, which is what its existing `return 0` already
said it did.
Fixing the lock alone would have turned every stop of a signal-proof child into
a refusal that leaves it running, so both land together and the tests pin that.
Measured on macOS with a stand-in poll child that traps TERM, INT and HUP:
the guard left it running past 70s and now clears the group within the lease
plus one check; retiring a healthy runner fell from ~2.8s with a forced group
signal every time to ~0.6s on the ordinary signal alone.
Unchanged and stated deliberately: a leader lost to anything other than the
stop's own signal still leaves a group that retire, reconcile, sweep-home and
the guard all refuse, permanently - and that source stops listening without
saying so. Whether such a group may ever be signalled is an open decision and
is not answered here.
* fix(review): Fix proved escalation race and stop regression assertions
* fix(review): Preserve proved escalation through transient identity failures
* fix(review): Simplify proved escalation and correct guard timing documentation
* fix(document): Clarify process-event stop ownership and cleanup limits
* fix(document): Clarify process-event stop ownership and fixture comments
* revert(skills): restore the leaderless-ambiguity limit to the loaded skill
An automatic documentation step in this branch's validation edited
.agents/skills/process-event-sources/SKILL.md, which no instruction in this
change asked it to touch. That file is not documentation about the code: it is
the agent-loaded instruction surface, what an agent reads to know what it is
permitted to do.
The step deleted this line:
- leaderless PID/PGID-reuse ambiguity preserves the claim without signalling
or replacement, as owned by the operating contract in
[`docs/configuration.md`](../../../docs/configuration.md#process-to-event-sources-stateprocevent);
and folded it, with its neighbour, into a generic "registration and ownership
transitions, stop authority, and claim reclamation follow the operating
contract".
That deleted line states a PROHIBITION - that such a group is preserved WITHOUT
SIGNALLING - and it is the exact limit an open captain decision currently rests
on. Folded into a pointer, an agent reading the skill to learn what it may do
would have to chase a second document to discover it may not signal. A
prohibition that requires a second lookup is not a prohibition. The effect was
to weaken, in the instructions themselves, the boundary that keeps one home from
signalling another's process group - while the question of whether that boundary
should move at all is still open.
This is a deliberate revert, not an oversight, and it restores the file exactly
to its pre-branch state. The full statement also survives in
docs/configuration.md; that does not rescue it, because the agent handling a
process-event wake loads the skill and not the documentation.
* revert(procevent): restore the open-question marking beside the escalation
The same automatic documentation step that edited the loaded skill also removed
this from the comment above runner_group_signal:
A leaderless group nobody in this call ever proved remains refused too, for
every caller. That untouched refusal is what makes a crashed leader's group
permanent, and relaxing it is a separate open question, not something this
path assumes.
and replaced it with a pointer to docs/configuration.md.
This one fails differently from the skill deletion, which is why it is restored
separately. There, a prohibition was moved out of the reader's path, and a
missing prohibition gets violated. Here the prohibition survives in code - the
unproved path still refuses - and what was removed is the fact that the limit is
UNDECIDED. A prohibition that has quietly lost its "this is still open" reads as
settled design, and settled design gets relied on, extended, and eventually
relaxed by someone confident they understand why it is there. That question is
open right now.
The rule this branch's four instances produce, stated once here because this is
the point of decision: an unresolved question must be marked unresolved AT THE
POINT OF DECISION, not only where the contract is documented. A reader who does
not know something is open will treat it as closed, and that default is stronger
than any pointer overcomes.
The pointer added by that step is kept alongside; this restores what it replaced
rather than reverting it.
* docs(verification): restore the measured guard bound and its reason
The document step's rewrite of this record dropped the concrete figure while
keeping the surrounding measurements. What went missing was the bound itself -
lease plus two consecutive failed checks plus the stop's grace, roughly 630
seconds at the shipped 600-second lease and 15-second check - together with the
reason there are two checks rather than one: a single unreadable read must not
be enough to kill a live runner.
The mechanism survived elsewhere and the reason survived in
docs/configuration.md, so nothing was lost from the repository. The concreteness
was, and that is what this restores. A number recorded without why it is that
number is the one a later reader shortens; the reason is the whole safety
argument for the debounce, and the debounce is what stops the reaper killing a
live runner on one bad read.
* fix(document): Replace stale stop-authority summaries with owner pointers
* test(procevent): make the guard-bound case able to fail for its own reason
An automated reviewer observed that this case allowed sixty seconds for a bound
of roughly eight, so it could not go red for the reason it names: it would have
passed a guard that took fifty-five seconds. That is correct, and it is the same
family as the defect the case exists to defend against - a check that is green
because it cannot fail, rather than because the thing it guards is working.
The deadline is now derived from the bound itself - the lease, plus the two
consecutive failed checks the guard debounces on, plus the stop's own ordinary
and forced signal windows - rather than from a flat wall-clock number, and the
shortened lease and check the fixtures run under have a single definition so a
derived deadline cannot silently diverge from the settings the guard is given.
The doubling that remains is a load allowance and is documented as one; widening
it to make a slow guard pass would convert the assertion back into decoration.
Proven by mutation rather than by argument. Against the repaired case:
correct code ok
guard debounces on 20 misses instead of 2 not ok - "still holding
the group after 16s,
against a documented
bound of 8s"
proved escalation removed (the original defect) not ok - same
code restored ok
The previous sixty-second version passes every one of those mutations.
The reviewer's other claim, that the guard can survive past the announced bound
when an owner disappears immediately after a check, was measured and does not
hold against what this branch announces. Sweeping the phase deliberately at
0.0, 0.2, 0.4, 0.6 and 0.8 of a check interval gave 7.21s, 7.31s, 6.75s, 6.49s
and 6.31s, worst 7.31s, against the announced lease plus two consecutive failed
checks plus stop grace, which is up to 8s at those settings. The mechanism the
reviewer describes is real and is the announced mechanism; the bound it was
measured against is a phrasing this branch no longer carries.
* revert(scope): return the instruction surfaces to their base state
This delivery is being split. It carries the two proven process fixes alone; the
instruction text travels separately, through a run that removes the
documentation step rather than refusing it at its gate.
Two surfaces are therefore returned to exactly what the base branch has, so this
delivery neither adds to them nor removes from them:
.agents/skills/process-event-sources/SKILL.md - identical to base again. Three
bullets an automatic documentation step had folded into a pointer, including
that leaderless PID/PGID-reuse ambiguity preserves the claim WITHOUT
SIGNALLING and that there is one identity-matched owner per canonical source
across homes sharing one store.
The header comment block of bin/fm-procevent.sh, which is what the script
prints as its own help. Seven lines were removed from it: that a live owner is
never displaced, that only a claim whose stale owner and independently absent
process group prove its whole generation gone is reclaimed, that a crashed
leader or reused pid whose process group still has members cannot relax
ownership cleanup, and that reconcile signals only a live identity-matched
runner group and otherwise keeps the claim without starting a replacement.
The help output is now byte-identical to base.
Neither removal was requested by any instruction in this change, and both were
made to text that predates it. Returning them is scoping, not a third
restoration: nothing is being added to those files here.
* fix(ci): Captain, live CI revealed a fixture deadlock: it suspended the runner before startup released its lock. Added a public-list synchronization barrier in tests/fm-procevent.test.sh. Forced-delay reproduction detected the deadlock before the fix; all four cases passed afterward. Targeted lint, Bash syntax, and whitespace checks passed. Greptile’s watchdog requirement conflicts with the recorded R2 decision; runtime behavior and documentation remain unchanged. Full CI rerun belongs to the outer executor
* test(procevent): make the post-TERM cases report what they saw when they fail
On the failure path only, these cases now print what they actually saw: the
identity recorded at claim time, the identity readable at that moment, the size
of the signals file, the leader's state and wchan, every live member of the
runner's process group with its own state and wchan, the elapsed time since the
stop began, and what retire said. None of it runs when a case passes.
WHY THIS IS KEPT, stated accurately rather than by its original reason. It was
written to make an unexplained CI failure verifiable. That failure is now
explained - it was a fixture deadlock, diagnosed and repaired in the preceding
commit - so that justification has expired and is not the reason given here.
The reason it stays is smaller and independent of that failure: it is already
written, it is small, it sits in the file whose assertion this change reworked,
and an assertion that could not say why it failed cost most of a morning to
diagnose from the outside. The next failure will not be this one.
WHAT A PASSING RUN WOULD NOT MEAN: a pass is a sample of behaviour already
observed many times, not proof that anything is fixed. Only a failure carrying
the evidence above establishes a cause.
* fix(document): Clarify process-event fixture diagnostic rationale
* fix(ci): Captain, fixed two cleanup races in tests/fm-procevent.test.sh: removed premature child completion and waited for runner exit before retiring the restart fixture. Controlled Linux reproductions demonstrated failure before and success after. The full Linux process-event suite, six focused macOS checks, targeted ShellCheck, Bash syntax, and whitespace checks passed. Runtime behavior, guard debounce, and documentation remain unchanged. CI rerun belongs to the outer executor
* fix(procevent): bound owner-guard cleanup at one check interval, not two
A THIRD WAY, not a capitulation to the reviewer and not a refusal of it.
The automated reviewer's grievance was the LOOSENESS OF THE BOUND, not the
number of observations the guard makes before it acts. It asked for a single
read because that was the only route it could see to an acceptable bound. There
was another route, and this change takes it: the bound is reached and both reads
are kept.
TIGHTENED - the SPACING of the guard's two reads, not their number. The owner
watchdog now sleeps half the configured check interval and still requires two
consecutive failing reads, so the pair completes inside one check interval
instead of costing two. Worst-case detection falls from the lease term plus TWO
check intervals to the lease term plus ONE. At the shipped 600s lease and 15s
interval the stated bound falls from ~635s to ~620s.
PRESERVED - the second read. bin/fm-procevent.sh's two-consecutive-miss rule is
untouched. WHY IT PROTECTS: the guard's inputs are a lease read and a state-root
identity read, and either can fail transiently on a live, healthy home. Acting
on the first failure would let one isolated unreadable read kill a live service.
Requiring a second, independent read is what makes that impossible, and it is a
protection rather than padding. Nothing was traded away to reach the bound.
Both properties are now guarded by their own case, and each was proven by
MUTATION rather than asserted:
- putting a full interval back between the two reads fails the bound case:
"still running 17.0s after the last owner activity, against a documented
bound of 15s";
- acting on one failed read fails the new debounce case: "one unreadable lease
read ended a runner whose home was still alive" - while the bound case then
passes FASTER, 9.9s against 13.1s. The unsafe variant being the quicker one
is exactly why these are two cases: one elapsed-time case would have
registered the removal of the protection as an improvement.
MEASURED, sampling the phase between the guard's check clock and the lease clock
across eight runs per variant, on macOS (Darwin 25.5.0). Reaping an orphaned
listener whose home stopped refreshing its lease:
lease 2s / interval 1s: 4.41-5.29s before, 3.48-4.65s after
lease 2s / interval 4s: 7.69-8.12s before, 5.94-6.13s after
The 4s configuration is the informative one: the gap is about one check
interval, which is precisely the term that was removed.
A previously unstated term of the bound surfaced while measuring: the lease age
is compared in whole seconds, so a configured lease of N is honoured until that
age reads N+1. It is now part of the documented bound and of the regression's
derivation instead of being absorbed into a fudge factor.
The bound regression derives its deadline from the documented bound instead of a
flat number, and PINS the phase between the guard's check clock and the lease
clock rather than sampling it, because with a sampled phase a guard spending two
intervals passes about half the time on a lucky alignment. Its load slack is
additive and stays under half a check interval, so an extra whole interval
cannot hide inside it. The two flat deadlines that were there before (40s and
20s) and the doubling allowance on the derived one are gone; that looseness was
the reviewer's third complaint.
The stop's own grace is untouched: 2s for the ordinary signal, then 2s for the
forced one. It is a ceiling paid only by a group that outlives the signal it was
sent, not a delay every stop pays - a healthy runner's whole retire measures
0.40-0.66s on this host. The reviewer's literal "lease plus one tick" is
unreachable by any implementation, since signalling a process and giving it any
chance to exit takes non-zero time; detection now meets it and the stop runs
inside its own ceiling, and the contract says so rather than glossing it.
NECESSARY BUT NOT SUFFICIENT, and written BEFORE this head's integration runs
start rather than after they report. On the previous head, "Behavior portable
serial 1" and "Behavior portable serial 4" were both CANCELLED at the job
ceiling, independently of this finding. A new head triggers fresh runs, so those
two lanes MAY complete this time. IF THEY DO, THAT IS NOT EVIDENCE THE CEILING
DEFECT IS FIXED. It is one more sample of a lane that has been cut repeatedly
and sometimes is not; the shard-packing repair for it is open separately. Do not
reread a lucky pass here as a resolution.
Relatedly, and deliberately: the per-script duration hint in bin/fm-test-run.sh
was NOT updated even though the two new cases add ~19s of wall clock.
docs/fm-test-portable-shards.md says those hints are replaced wholesale from CI
timing artifacts of green runs, and that repair is the open request doing it; a
hand-edited estimate here would collide with it and silently repack the shards.
This suite runs in portable serial shard 3, which was green in the last run.
Verification: tests/fm-procevent.test.sh green, plus
tests/fm-captain-hold-lifecycle.test.sh, the test-coverage guard, and
bin/fm-lint.sh. The unrelated "reconcile stops a runner whose registration was
removed" case flaked in 4 of 7 local full runs; an isolated 20-trial
reproduction measured it at 13/20 unclean before this change and 11/20 after, so
it is issue 4080 and is not aggravated here.
* fix(procevent): repair our decimal-interval regression and enforce the timing phase
REPAIRED BEFORE PUBLICATION, AND IT WAS OURS. The half-interval arithmetic added
by the previous commit read a zero-prefixed interval as octal: 010 halved to 4
instead of 5, and 08 was not a number at all, so the owner guard died before
reporting ready and the runner failed closed and never listened. The validator
accepts those values and `[` compares them as decimal, so this broke a
configuration that worked before. Introduced by this delivery, found in review,
repaired here. Forcing base ten before the arithmetic is the whole runtime fix.
Proven by driving it rather than by reading the source: a new case starts a real
listener at 08 and at 010 and observes the guard's actual sleep argument - 4s and
5s. Removing the normalisation turns that case red with "a zero-prefixed decimal
interval (08) prevented the listener from starting".
THE TIMING PHASE IS NOW OBSERVED AND ENFORCED, NOT ASSUMED. The bound case
pinned its phase by CONSTRUCTION, from an assumed startup time, and enforced
nothing. Review was right that this is not enough: once startup reaches about two
seconds the expiry lands in a different part of the interval and the case
silently stops rejecting a two-interval guard while still reporting success. A
bound that cannot fail for the reason it names is the defect this whole delivery
exists to correct, so it must not ship inside the fix for it.
Now the lease is synchronised to the guard's own FIRST observed lease read,
every later real read is recorded, and the case REFUSES unless one recorded read
proves the required phase: it read the synchronised reference, it was still
fresh, and it began late enough that two further full intervals could not finish
before the deadline. An unestablished precondition refuses; it does not proceed
on trust. The derived deadline, the two-read debounce and the additive slack are
unchanged, and the slack invariant is now asserted rather than left to a comment.
Review also found the deadline was only ever checked while the group was still
alive, so a sampler descheduled past it would see the group gone and certify
success. The observed completion time is now checked too.
PROVEN BY MUTATION, each one run against this code:
- remove the decimal normalisation -> the interval case fails on 08;
- a full interval between the two reads -> "the guard exceeded its bound:
group still running 17.1s ... against a documented bound of 15s";
- a full interval WITH startup forced to ~2.5s, which is exactly the condition
the old construction pin could not survive -> still red, same message;
- the same ~2.5s startup with the correct guard -> still passes, 13.0s against
the 15s bound, so the delay alone does not break the case;
- phase evidence made unavailable -> "could not establish the required
pre-expiry guard-read phase", a refusal rather than a pass, even though the
group stopped quickly;
- act on one failed read -> the debounce case fails and the bound case passes
FASTER, 9.5s against 12.7s, which is why these remain separate cases.
Verification: full tests/fm-procevent.test.sh green, and bin/fm-lint.sh clean.
* fix(document): Correct process-event timing and debounce comments
* fix(backlog): route lifecycle transitions through configured adapters (#3417)
* fix(backlog): honor configured task adapters
* no-mistakes(review): Harden backend purity lint against prefixed Beads calls
* no-mistakes(document): Document configured backend lifecycle transitions
* fix(backlog): preserve markdown exemptions
* no-mistakes(review): Enforce backend purity for explicit lint paths
* no-mistakes(document): Update lifecycle backend documentation
* no-mistakes(lint): Remove redundant backend lint pattern
* fix(backlog): close adapter routing gaps
* no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint
* no-mistakes(document): Document environment-selected backlog adapters
* no-mistakes(lint): Fix empty local variable assignment
* fix(backlog): close quoted path gaps
* no-mistakes(review): Reject partially quoted direct Beads commands
* no-mistakes(document): Align lifecycle documentation with configured adapters
* test(backlog): keep structural cases markdown-only
* fix(backlog): honor configured task adapters
* no-mistakes(review): Harden backend purity lint against prefixed Beads calls
* no-mistakes(document): Document configured backend lifecycle transitions
* fix(backlog): preserve markdown exemptions
* no-mistakes(review): Enforce backend purity for explicit lint paths
* no-mistakes(document): Update lifecycle backend documentation
* no-mistakes(lint): Remove redundant backend lint pattern
* fix(backlog): close adapter routing gaps
* no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint
* no-mistakes(document): Document environment-selected backlog adapters
* no-mistakes(lint): Fix empty local variable assignment
* fix(backlog): close quoted path gaps
* no-mistakes(review): Reject partially quoted direct Beads commands
* no-mistakes(document): Align lifecycle documentation with configured adapters
* test(backlog): keep structural cases markdown-only
* no-mistakes(review): Harden markdown lifecycle routing and close recovery
* fix(lint): catch dollar-quoted beads commands
* no-mistakes(review): Drop unused authorized-data-dir param; canonicalize fm-lint ROOT with pwd -P
* no-mistakes(document): Document fm-lint backend-purity check in header and CONTRIBUTING
* no-mistakes(ci): Two failing checks, one code-caused and one attestation-only. 1. Greptile Review (P1: ANSI-C quoting bypasses lint) — REAL DEFECT, FIXED. The backend-purity normalizer in bin/fm-lint.sh (invokes_bd awk function) stripped $'...' quote delimiters without decoding ANSI-C escapes, so a core script containing $'\x62\x64' close fm-example executes as `bd close` while the lint accepted it. Fix: the normalizer now tracks whether a single-quote context came from $' (ansi flag) and decodes ANSI-C escapes while building the command word: \xHH (1-2 hex), \uHHHH / \UHHHHHHHH (4/8 hex), \NNN (1-3 octal), \cX control characters, simple escapes (a/b/e/E/f/n/r/t/v decode to a placeholder that can never spell bd), NUL (value 0) truncates the word per bash C-string semantics, and unknown escapes drop the backslash per bash. Values outside printable ASCII decode to a placeholder so they can neither falsely match nor collide into `bd`. Regression coverage added to the existing lint-interface test test_rejects_direct_beads_cli_invocations in tests/fm-lint.test.sh: $'\x62\x64', $'\142\144' (octal), and b$'\x64' (split word), each written into a fixture and asserted rejected through the real fm-lint.sh executable. Verified regression property: with the fix stashed, the new hex case passes lint (reproduces the reported bypass); with the fix, it is rejected. Verification: all 29 tests in tests/fm-lint.test.sh pass, and the full CI-parity lint (CI=true bin/fm-lint.sh: ShellCheck 0.11.0 full analysis over the canonical set, backend-purity check, and actionlint workflow validation) exits 0. One iteration was needed because an awk comment containing a literal $'\x62\x64' sample broke shell-level quoting (bash -n / SC1001/SC2026); the comment was reworded without quotes. 2. PR must be raised via no-mistakes — NOT code-caused. The check failed with 'Pipeline attestation head_sha does not match the current PR head': the PR body attestation binds to 82b41c7 while the PR head is 1995cdb because a later pipeline push moved the head. This is exactly the stale-attestation condition the user intent describes; it clears when the outer pipeline re-runs 'git push no-mistakes' and re-binds the attestation to the new head (which now includes this Greptile fix). No code change can or should address it. Files changed: bin/fm-lint.sh (ANSI-C escape decoding in the backend-purity normalizer), tests/fm-lint.test.sh (three encoded-bd rejection cases)
* no-mistakes(review): fix tasks.toml hang, root authorization, lint quoting
* no-mistakes(review): validate tasks config before exemption; fix lint quote gap
* fix(backlog): address the markdown backlog as <data>/backlog.md
Resolving the markdown backlog through a configured `[markdown] path` was
scope this task never asked for. It is absent from main, which addresses
`<data>/backlog.md` everywhere, and it came from an earlier review round
rather than the task brief.
Making it effective on the transition path alone put that path at odds
with every other consumer of the same backlog - fm-captain-hold.sh,
fm-session-start.sh, fm-fleet-snapshot.sh, fm-inbox.sh,
fm-backlog-handoff.sh - which all still address `<data>/backlog.md`. In
fm-captain-hold.sh the split was live: its reads had already moved to the
shared gate while its writes had not, so the two could address different
files.
Address `<data>/backlog.md` from the shared gate, delete the unused
resolver, and drop the two tests that pinned the withdrawn behaviour.
What this task actually changes is unaffected: a configured non-markdown
adapter is still addressed by its own root, without `--file`.
* fix(backlog): honor configured task adapters
* no-mistakes(review): Harden backend purity lint against prefixed Beads calls
* no-mistakes(document): Document configured backend lifecycle transitions
* fix(backlog): preserve markdown exemptions
* no-mistakes(review): Enforce backend purity for explicit lint paths
* no-mistakes(document): Update lifecycle backend documentation
* no-mistakes(lint): Remove redundant backend lint pattern
* fix(backlog): close adapter routing gaps
* no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint
* no-mistakes(document): Document environment-selected backlog adapters
* no-mistakes(lint): Fix empty local variable assignment
* fix(backlog): close quoted path gaps
* no-mistakes(review): Reject partially quoted direct Beads commands
* no-mistakes(document): Align lifecycle documentation with configured adapters
* test(backlog): keep structural cases markdown-only
* fix(lint): catch dollar-quoted beads commands
* no-mistakes(review): Drop unused authorized-data-dir param; canonicalize fm-lint ROOT with pwd -P
* no-mistakes(document): Document fm-lint backend-purity check in header and CONTRIBUTING
* no-mistakes(ci): Two failing checks, one code-caused and one attestation-only. 1. Greptile Review (P1: ANSI-C quoting bypasses lint) — REAL DEFECT, FIXED. The backend-purity normalizer in bin/fm-lint.sh (invokes_bd awk function) stripped $'...' quote delimiters without decoding ANSI-C escapes, so a core script containing $'\x62\x64' close fm-example executes as `bd close` while the lint accepted it. Fix: the normalizer now tracks whether a single-quote context came from $' (ansi flag) and decodes ANSI-C escapes while building the command word: \xHH (1-2 hex), \uHHHH / \UHHHHHHHH (4/8 hex), \NNN (1-3 octal), \cX control characters, simple escapes (a/b/e/E/f/n/r/t/v decode to a placeholder that can never spell bd), NUL (value 0) truncates the word per bash C-string semantics, and unknown escapes drop the backslash per bash. Values outside printable ASCII decode to a placeholder so they can neither falsely match nor collide into `bd`. Regression coverage added to the existing lint-interface test test_rejects_direct_beads_cli_invocations in tests/fm-lint.test.sh: $'\x62\x64', $'\142\144' (octal), and b$'\x64' (split word), each written into a fixture and asserted rejected through the real fm-lint.sh executable. Verified regression property: with the fix stashed, the new hex case passes lint (reproduces the reported bypass); with the fix, it is rejected. Verification: all 29 tests in tests/fm-lint.test.sh pass, and the full CI-parity lint (CI=true bin/fm-lint.sh: ShellCheck 0.11.0 full analysis over the canonical set, backend-purity check, and actionlint workflow validation) exits 0. One iteration was needed because an awk comment containing a literal $'\x62\x64' sample broke shell-level quoting (bash -n / SC1001/SC2026); the comment was reworded without quotes. 2. PR must be raised via no-mistakes — NOT code-caused. The check failed with 'Pipeline attestation head_sha does not match the current PR head': the PR body attestation binds to 82b41c7 while the PR head is 1995cdb because a later pipeline push moved the head. This is exactly the stale-attestation condition the user intent describes; it clears when the outer pipeline re-runs 'git push no-mistakes' and re-binds the attestation to the new head (which now includes this Greptile fix). No code change can or should address it. Files changed: bin/fm-lint.sh (ANSI-C escape decoding in the backend-purity normalizer), tests/fm-lint.test.sh (three encoded-bd rejection cases)
* fix(bin): preserve captain calls during teardown (#3595)
* fix(bin): never close a captain call during cleanup
A scout that held its own work item for the captain, which is what
captain-hold-lifecycle prefers ("hold the work item the question gates"),
was closed by bin/fm-teardown.sh's automatic backlog transition. The
completion gate passed, cleanup ran, and the captain's question moved to
Done with no recorded answer: the one thing the policy says must never
happen. …
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Captain's intent: Yeah similar to the other PR, have the crewmate do a clean slate minimal sufficient PR to supersede it. Meaning: do not land #3502. Implement the smallest sufficient fix from a clean origin/main, then supersede 3502 with the new PR.
Context: the bug is that bin/fm-teardown.sh closed a captain-held scout task with no recorded captain answer. The captain-hold-lifecycle policy prefers holding the very work item a question gates, and says a captain-held task must only ever close through fm-captain-hold.sh answer; but teardown's automatic backlog transition ran tasks-axi done on that row, which silently closes a held row and drops its hold. PR 3502 fixed this with a parallel retention-record family (a state/.backlog-retain record, a recover-retain command, a second bootstrap replay loop, and a lock handoff from teardown to a child process); an independent audit found that machinery mostly pipeline-induced and asked for the smallest sufficient fix from a clean main.
Deliberate decisions in this change, all intended: (1) fm-captain-hold.sh gains a read-only open predicate with exit 0 (open captain call), 1 (not), 2 (cannot tell), and any read error is a 2 so a mechanical closer never treats 'cannot tell' as permission to close. (2) Teardown asks open before any destructive step and refuses on 2; on 0 only the close changes: after cleanup and still under the task's own meta lock, teardown appends one 'Deliverable of the finished work: ...' line to the task body and runs tasks-axi reopen, so the row returns to Queued with its hold intact and appears in Bearings' Captain's Call; --force deliberately does not lift this because it authorizes discarding unlanded work, never the captain's question. (3) The deliverable is written into the body rather than through tasks-axi update --report/--pr, because on a row that is not Done that flag rewrites the task title (verified with the installed tasks-axi). (4) The crash window reuses the existing pending-close record with a new optional mode=retain line, so the existing validator, stale-generation check, cleanup-incomplete marking, bootstrap replay loop, and its non-blocking try-lock all apply unchanged; replay in retain mode records the deliverable and reopens instead of closing, and a row the captain already answered and closed simply retires the record. No parallel record type, no recover command, no second bootstrap loop, no lock handoff: that omission is the point of the change, not an oversight. (5) fm-captain-hold.sh now addresses the configured data directory (FM_DATA_OVERRIDE, relocated data dirs) the same way bin/fm-backlog-transition-lib.sh does for every command, and reads rows through the library's backend-aware show helper; before this it ran tasks-axi from FM_HOME and ignored a relocated data directory, which would have made open return 'not found' and closed the call. (6) The ordinary non-captain-held teardown path is behaviorally unchanged. (7) Tests exercise the real executables only: the captain-held scout case including --force, ordinary scout, and answer-only closure; an interrupted cleanup leaving the row untouched and the next session start retaining it; a relocated backlog; and a ship-kind row whose hold cannot be read refusing before anything destructive. Race tests with paused fake binaries were deliberately not added. (8) tests/fm-teardown.test.sh has one case (herdr-preflight-missing-adapter) that fails identically on unmodified origin/main on this Mac where a real herdr is installed; it is pre-existing and out of scope. Constraints: bash 3.2 safe, bin/fm-lint.sh clean, one sentence per line and plain dashes in tracked Markdown, no agent co-author trailer, single-owner contracts with docs/captain-hold-lifecycle.md and the script headers updated rather than duplicated.
What Changed
openpredicate for captain-held tasks that honors configured and relocated data directories.--force, and expand lifecycle documentation and executable regression coverage.Risk Assessment
✅ Low: The per-task control lock now serializes teardown and replay with captain hold/answer transitions, and backend-aware listing resolves the previously identified relocated non-Markdown read risk without introducing a substantiated regression.
Testing
The focused Bash 3.2 lifecycle suite passed, covering held and forced teardown, ordinary teardown, answer-only closure, interrupted-cleanup replay, relocated backlog handling, and fail-closed read errors; a real CLI transcript additionally demonstrated that teardown preserves the queued captain-held row and deliverable until the captain answer closes it.
Evidence: Captain-held teardown and answer end-to-end CLI transcript
Source: Captain-held teardown and answer end-to-end CLI transcript
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
🔧 **Review** - 2 issues found → auto-fixed ✅
bin/fm-teardown.sh:324- Theopenresult is a TOCTOU snapshot rather than a durable invariant. After it returns 1, a concurrentfm-captain-hold.sh holdcan establish a captain call before teardown runstasks-axi done, reproducing the forbidden answerless close. Conversely, after it returns 0, a concurrentanswercan close the row beforefm_backlog_retainreopens it; replay has the same probe-to-transition race. Serialize hold/answer and teardown/replay with the existing per-task lifecycle/control lock across the predicate and final transition.bin/fm-captain-hold.sh:241- The relocated-data helper says it owns mutations, butopen_task_idsalso uses it, causingtasks-axi list --file .... Non-Markdown backends must be read from their configured addressing root without--file(the immediately preceding backend fix establishes this contract), sodivergedcan silently see no Beads rows. Route list reads through a backend-aware helper while retaining--filefor mutations.🔧 Fix: Serialize captain holds and fix backend-aware listing
✅ Re-checked - no issues remain.
✅ **Test** - passed
✅ No issues found.
bash tests/fm-captain-hold-lifecycle.test.shIn an isolated home using real executables, ranfm-teardown.sh evidence-retain, inspected the row withtasks-axi show --full, confirmed it remained in Bearings' Captain's Call, answered it throughfm-captain-hold.sh answer, and inspected the final closed row.🔧 **Document** - 1 issue found → auto-fixed ✅
bin/fm-captain-hold.sh:384- Three failure messages still name FM_HOME/data/backlog.md even when FM_DATA_OVERRIDE selects a relocated backlog; executable-message edits were outside this documentation-only phase.🔧 Fix: Fix relocated captain-hold backlog diagnostics
✅ Re-checked - no issues remain.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.